SemIf (formerly OpenJev): Integration and API Guide


Verdict

SemIf is currently a local Python research package and JSONL command-line scorer, not a hosted service or a ready-made HTTP API. Integrate it through semif-score for batch or subprocess use, or call its Python scorer functions inside a long-lived worker; if an application needs HTTP, place a small service around that worker rather than expecting a built-in /v1/systemone endpoint.[1][4]

The project was formerly named OpenJev, but it is independent of TypeSafe and does not reproduce TypeSafe's undisclosed Jev model or training.[2] This distinction also matters because other, newer projects use the OpenJev name and expose different APIs.


5-minute overview

You send three things: the state holding the evidence, one question, and 2-16 explicitly described options.[5] SemIf labels those options with answer-token letters, then reads the model's last-position logits for exactly those letters and applies softmax over them.[6] Nothing is written back: there is no decoding loop, no prose to parse, and no JSON to repair or retry.

That is where the speed comes from. Scoring is a forward pass over the prompt instead of token-by-token generation, and shared mode prefills one long state once and reuses that prefix across every independent question asked about it.[8]

Decision quality comes from the bounded, explicit option set — the model chooses among alternatives you wrote down, including an explicit insufficient-evidence option when you supply one — and from validating that set on your own workload. It does not come from typed output alone.

The returned values are conditional scores over the options you supplied, not deployment confidence. They are relative to that option set and must be calibrated on representative labeled data before any action threshold is attached to them.[6][11]


1. What SemIf Does

SemIf turns one item of state plus one runtime-defined question and 2-16 described options into a probability distribution over those options.[5] It reads the language model's next-token logits for letter slots (A through P) and applies softmax; it does not generate an answer sentence or JSON response.[6]

Typical uses:

SemIf provides four execution modes:

ModePurposeState reuseSupported backends
directScore each row independentlyNoneTorch, MLX, llama.cpp
serialProcess rows in order and reuse the immediately repeated stateOne retained prefix cacheTorch, MLX, llama.cpp
sharedEvaluate a batch whose rows all have exactly the same stateOne prefill, multiple question suffixesTorch, MLX, llama.cpp
rerankerAlternative native reranker baselineBackend-specificTorch/CUDA only

The CLI rejects MLX or llama.cpp with reranker; the Torch reranker is CUDA-only.[4]


2. Important Naming Boundary

Use these identifiers to avoid integrating the wrong project:

NameMeaning in this guide
SemIfgithub.com/TheoLeeCJ/SemIf, package semif-phase1, CLI semif-score
Former OpenJevThe previous name of the same TheoLeeCJ/SemIf project
TypeSafe JevA separate hosted commercial model and API
openjev/openjevA different open-weights project that now uses the OpenJev name
api.openjev.shA separate hosted API, not the SemIf CLI documented here

Do not point a SemIf client at an OpenJev or TypeSafe endpoint and assume semantic or calibration equivalence. SemIf's own output explicitly labels its probabilities as conditional option scores and uncalibrated decision confidence.[6]


3. Requirements

The researched package declares Python 3.10 or newer and pins Torch 2.10.0, Transformers 5.17.0, Accelerate 1.12.0, Safetensors 0.8.0, Hugging Face Hub 1.31.0, Tokenizers 0.23.2, NumPy 2.2.6, SentencePiece 0.2.1, and Protobuf 7.36.1.[3]

Choose one runtime:

RuntimeRecommended useNotes
Torch + CUDAMain GPU pathExpose exactly one CUDA GPU; Qwen3.5-4B BF16 must fit in memory.[2][5]
MLXApple SiliconText scoring through Metal; direct, serial, and shared modes.[9][10]
Torch + MPSApple Silicon fallbackSupported, but some Qwen3.5 kernels use slower fallback paths.[9]
llama.cppCPU/local GGUFRequires the llamacpp extra and a local GGUF checkpoint.[2][4]
Torch + CPUFull-precision referenceExplicitly pass --device cpu --dtype float32; expect it to be slow.[2][4]

The default Qwen3.5-4B MLX source weights are about 9 GB on disk and require additional execution memory.[10]


4. Installation

Clone the repository because SemIf is not presented as a stable PyPI SDK:

git clone https://github.com/TheoLeeCJ/SemIf.git
cd SemIf
git checkout 23cf1f39fc9534fe81437200959b6dfc7106e45a
python3 -m venv .venv
source .venv/bin/activate

Install the runtime you need:

# Torch: CUDA, MPS, or explicit CPU
pip install -e '.[test]'

# Apple Silicon MLX
pip install -e '.[test,mlx]'

# llama.cpp with a local GGUF
pip install -e '.[test,llamacpp]'

For a production integration, pin both the SemIf commit and model revision. Remote model revisions must be immutable 40-character commit IDs; local model directories require a nonempty revision or manifest label.[5]

Optional: put Hugging Face downloads on a large drive before installation or first use:

export HF_HOME=/path/to/large-drive/huggingface

5. Input API: JSONL Decision Rows

semif-score reads newline-delimited JSON. Each line is one independent decision:

{
  "id": "route-1",
  "state": "Customer cannot access an account after a password reset.",
  "question": "Which queue should handle this request?",
  "options": [
    {"id": "access", "description": "Account access support."},
    {"id": "billing", "description": "Billing support."},
    {"id": "insufficient", "description": "There is not enough evidence to decide."}
  ]
}

Field reference

FieldTypeRequiredConstraints and meaning
idstringYesNonempty decision identifier; maps output back to application code. Must be unique within a shared batch.
statestring, object, or arrayYesNonempty, finite JSON-compatible evidence and context. Objects and arrays retain structured JSON in direct modes.
questionstringYesNonempty criterion to apply to the state. It must be self-contained.
optionsarrayYes2-16 entries. Order is mapped to answer-token letters, so preserve it when comparing runs.
options[].idstringYesApplication-facing option ID; unique within the row.
options[].descriptionstringYesMeaning the model judges. Descriptions may be empty under the current validator, but meaningful descriptions are required for useful behavior.

These constraints come directly from validate_row in the researched source.[5]

Design rules

  1. Put facts and evidence in state; put the single judgment in question.
  2. Keep each question narrow and independently answerable.
  3. Describe every option explicitly; IDs are for application code, not model semantics.
  4. Include an insufficient, unknown, or none option when the evidence may not support any substantive answer.
  5. Do not pass a URL and expect browsing. Fetch external content in your application and place the relevant content in state.
  6. Keep policy and execution in code. SemIf should choose among allowed alternatives, not invent operations.

6. CLI API

The installed executable is semif-score.[3]

Full command shape

semif-score \
  --mode direct|serial|shared|reranker \
  --backend torch|mlx|llamacpp \
  --model MODEL_OR_LOCAL_PATH \
  --revision IMMUTABLE_REVISION_OR_LOCAL_LABEL \
  --input INPUT.jsonl \
  --output NEW_OUTPUT.jsonl \
  [--max-tokens 4096] \
  [--device auto|cuda|mps|cpu] \
  [--dtype bfloat16|float16|float32] \
  [--mlx-bits 4|8] \
  [--mlx-cache-limit-mib N] \
  [--gguf /path/to/model.gguf] \
  [--llama-threads N]

The output path must not already exist. SemIf creates it exclusively and refuses to overwrite it. Input that exceeds --max-tokens fails rather than being truncated.[4][6]

CUDA example

CUDA_VISIBLE_DEVICES=0 semif-score \
  --mode direct \
  --model Qwen/Qwen3.5-4B \
  --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
  --input decisions.jsonl \
  --output results.jsonl

Apple Silicon MLX example

semif-score \
  --backend mlx \
  --mode direct \
  --model Qwen/Qwen3.5-4B \
  --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
  --input decisions.jsonl \
  --output results-mlx.jsonl

Use --mlx-bits 8 or --mlx-bits 4 for in-memory affine quantization. Quantization changes probabilities and must be evaluated separately; it does not create another model artifact.[10]

Apple Silicon MPS example

semif-score \
  --mode direct \
  --device mps \
  --model Qwen/Qwen3.5-4B \
  --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
  --input decisions.jsonl \
  --output results-mps.jsonl

CPU llama.cpp example

First obtain a compatible local GGUF, then run:

semif-score \
  --backend llamacpp \
  --mode direct \
  --model Qwen/Qwen3.5-4B \
  --revision local-gguf-manifest-v1 \
  --gguf /models/Qwen_Qwen3.5-4B-Q4_K_M.gguf \
  --llama-threads 8 \
  --input decisions.jsonl \
  --output results-cpu.jsonl

The --model value remains relevant because prompt construction uses the reference tokenizer; --gguf supplies the local scoring weights.[2][4]


7. Output API

A direct-mode output line has this logical shape:[6]

{
  "id": "route-1",
  "option_ids": ["access", "billing", "insufficient"],
  "probabilities": [0.82, 0.06, 0.12],
  "option_logits": [8.1, 5.5, 6.2],
  "input_tokens": 137,
  "forward_seconds": 0.24,
  "total_seconds": 0.25,
  "prompt_sha256": "...",
  "prompt_version": "direct-options-v1",
  "model": {
    "source": "Qwen/Qwen3.5-4B",
    "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
    "dtype": "bfloat16",
    "device": "...",
    "torch_version": "...",
    "transformers_version": "..."
  },
  "readout": "native full-vocabulary last-position logits restricted to declared answer slots",
  "probability_status": "conditional option score; uncalibrated as decision confidence"
}

The numeric values above are illustrative; the field shape is source-derived.

Output field reference

FieldMeaning
idInput row identifier.
option_idsIDs in the exact order corresponding to probabilities and option_logits.
probabilitiesSoftmax over only the declared answer-slot logits; values sum approximately to 1.
option_logitsRaw selected next-token logits. Preserve these if you plan to apply temperature calibration later.
input_tokensPrompt token count.
forward_secondsTimed model-forward work for the selected execution path.
total_secondsEnd-to-end row scoring time inside the scorer.
prompt_sha256Hash of the rendered prompt, useful for reproducibility and cache checks.
prompt_versionPrompt contract identifier; currently direct-options-v1.
modelModel source, revision, runtime, device, precision, and possibly serving configuration.
readoutDescription of the scoring path.
probability_statusExplicit warning about interpretation.

Serial output adds fields such as cache_hit, prefix_tokens, cache/prefix timing, answer token IDs, and vocabulary diagnostics.[7] Shared output adds the same shared_timing object to every output row, including batch size, prefix length, prefill time, replication time, and suffix-forward time.[4][8]

Safely selecting an answer

Never assume the highest probability is at the same array index across different row schemas. Pair the arrays from the same output object:

best_index = max(range(len(result["probabilities"])),
                 key=result["probabilities"].__getitem__)
best_option_id = result["option_ids"][best_index]
best_probability = result["probabilities"][best_index]

Treat best_probability as a relative score among the supplied options until calibrated on your workload.


8. Integration Pattern A: Batch Files

Use the CLI directly when decisions can be prepared and consumed in batches.

import json
import subprocess
import tempfile
from pathlib import Path

MODEL = "Qwen/Qwen3.5-4B"
REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"


def score_batch(rows: list[dict]) -> list[dict]:
    with tempfile.TemporaryDirectory() as directory:
        root = Path(directory)
        source = root / "input.jsonl"
        destination = root / "output.jsonl"
        source.write_text(
            "".join(json.dumps(row, ensure_ascii=False) + "\n" for row in rows),
            encoding="utf-8",
        )
        subprocess.run(
            [
                "semif-score",
                "--mode", "direct",
                "--model", MODEL,
                "--revision", REVISION,
                "--input", str(source),
                "--output", str(destination),
            ],
            check=True,
        )
        return [
            json.loads(line)
            for line in destination.read_text(encoding="utf-8").splitlines()
            if line.strip()
        ]

This is simple but inefficient for request-per-decision serving because every process invocation reloads the model. Use it for offline pipelines, cron jobs, experiments, and reproducible evaluation—not a latency-sensitive HTTP route.


9. Integration Pattern B: Long-Lived Python Worker

For an application service, load the model once and reuse it. The Python functions are implementation-level APIs, not a separately versioned public SDK, so pin the repository commit and place them behind your own adapter.

Direct scorer

from semif_phase1.core import load_causal_model, validate_row
from semif_phase1.direct import score

MODEL = "Qwen/Qwen3.5-4B"
REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"

model, tokenizer, metadata = load_causal_model(
    MODEL,
    REVISION,
    device="cuda",       # or "mps" / "cpu"
    dtype="bfloat16",    # use float32 for explicit CPU reference work
)


def decide(row: dict) -> dict:
    validate_row(row)
    return score(model, tokenizer, row, metadata, max_tokens=4096)

Serial prefix reuse

Use serial mode when consecutive requests often share the same exact state:

from semif_phase1.serial import SerialPrefixScorer

scorer = SerialPrefixScorer(model, tokenizer, metadata, max_tokens=4096)
result_a = scorer.score(row_a)
result_b = scorer.score(row_b)  # cache hit only if state matches

The scorer retains one state's native prefix cache. A changed state invalidates it.[7]

Shared-state fan-out

Use shared scoring when many independent questions all see the exact same state:

from semif_phase1.shared import score_shared

results, timing = score_shared(
    model,
    tokenizer,
    rows_with_identical_state,
    metadata,
    max_tokens=4096,
)

All rows must have exactly equal states and unique IDs. Questions are independent; one question cannot consume another question's result in the same call.[8]

Concurrency

Do not call one mutable model/cache object concurrently without an explicit queue or lock. The llama.cpp backend owns one stateful scoring context, and serial/shared paths manipulate prefix caches.[2][7][8] A safe starting architecture is one worker queue per loaded model/device, with bounded input and application-level timeouts.


10. Integration Pattern C: Your Own HTTP API

SemIf does not ship an HTTP server at the researched commit.[1][3][4] If you need one, expose a narrow adapter around the long-lived worker.

Recommended contract:

POST /v1/decisions
Content-Type: application/json

{
  "mode": "direct",
  "decisions": [<SemIf decision row>, ...]
}

Recommended response:

200 OK

{
  "results": [<SemIf output row>, ...],
  "model_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
}

Adapter requirements:

  1. Validate every row with validate_row before enqueueing it.
  2. Allow only configured model and revision values; do not accept arbitrary Hugging Face repositories from public callers.
  3. Cap body size, decision count, and prompt tokens.
  4. Serialize or bound GPU work to avoid out-of-memory failures.
  5. Return 422 for invalid rows, 429 for queue saturation, and 503 for unavailable workers.
  6. Do not expose raw traceback, filesystem paths, model cache paths, or unrestricted local model loading.
  7. Preserve option_ids, probabilities, option_logits, model metadata, and prompt hash in internal logs, subject to data-retention policy.
  8. Authenticate the endpoint and terminate TLS before exposing it outside localhost.

This wrapper contract is a recommended integration design, not a SemIf-provided API.


11. Choosing Between direct, serial, and shared

Are all questions known now and based on the exact same state?
  Yes -> shared
  No  -> Will consecutive decisions often use the exact same state?
           Yes -> serial
           No  -> direct

Use a second scoring call when:

Shared mode is not a chain-of-thought or dependency mechanism; it is independent fan-out over a common prefix.[8]


12. Confidence, Calibration, and Thresholds

Raw SemIf probabilities are conditional on the options supplied. They are not universally calibrated confidence values.[6][11]

Consequences:

SemIf includes an offline scalar temperature-calibration script. Given labeled results containing option_logits, it applies:

calibrated_probabilities = softmax(option_logits / T)

Because division by a positive temperature is monotonic, calibration changes confidence but not the selected argmax.[11]

Example application of an already selected temperature:

python benchmarks/calibrate.py \
  --predictions predictions.jsonl \
  --temperature 1.23 \
  --calibrated-out calibrated.jsonl

Fit and validate T on representative, labeled data from the actual deployment workload. Then choose thresholds based on consequences—for example, auto-route only above the measured review threshold and send other cases to a human or a stronger model.


13. Error Handling

SemIf is a local program, so errors appear as Python exceptions or a nonzero subprocess exit—not HTTP statuses.

FailureLikely causeHandling
Missing row fieldsNo id, state, question, or optionsReject before inference.
Invalid stateEmpty, unsupported type, NaN, infinity, or non-JSON dataNormalize to finite JSON-compatible content.
Invalid optionsFewer than 2, more than 16, duplicate IDs, or malformed entriesFix the schema.
Token limit exceededRendered prompt is larger than --max-tokensReduce state intentionally; SemIf will not truncate it.
Remote revision rejectedRevision is not a 40-character commit hashResolve and pin an immutable model commit.
CUDA device errorZero or multiple visible GPUsSet CUDA_VISIBLE_DEVICES to exactly one GPU.
Existing output fileCLI is create-onlyChoose a new output path or remove an abandoned file deliberately.
Unsupported mode/backendReranker selected with MLX, llama.cpp, MPS, or CPUUse direct/serial/shared or Torch/CUDA reranker.
Out of memoryModel, batch, or suffix fan-out is too largeReduce batch size, use serial mode, quantize where supported, or use a smaller model.

For subprocess integration, capture stderr, set a timeout, check the exit status, and treat partial output as incomplete. The CLI flushes rows as it writes, so an interrupted run can leave a partial file.[4]


14. Security and Privacy

Local execution means decision state does not need to be sent to TypeSafe or another hosted inference API, but model downloads still contact Hugging Face unless weights are already local.[2][10]

Production checklist:

SemIf's repository code is MIT-licensed, while upstream models retain their own licenses.[2]


15. Testing an Integration

Schema smoke test

PYTHONPATH=src python3 - <<'PY'
import json
from pathlib import Path
from semif_phase1.core import validate_row

rows = [
    json.loads(line)
    for line in Path("examples/decisions.jsonl").read_text().splitlines()
    if line.strip()
]
for row in rows:
    validate_row(row)
print(f"validated {len(rows)} rows")
PY

Repository tests

After installing the selected extras:

pytest -q

Application tests

Maintain a labeled suite that covers:

  1. normal cases for every option;
  2. insufficient-evidence cases;
  3. ambiguous cases near review thresholds;
  4. reordered options;
  5. missing or adversarial evidence;
  6. long inputs near the token limit;
  7. repeated-state serial behavior;
  8. shared batches with mixed states, which must fail;
  9. model or runtime upgrades compared against the pinned baseline;
  10. calibrated-threshold behavior, not only argmax accuracy.

Do not silently promote a new SemIf commit, Transformers version, model revision, quantization, prompt version, or backend. Compare choices, probability drift, calibration, latency, and memory before deployment.


16. Observability

Keep these output fields where policy permits:

The state itself may contain sensitive data. Store it only when necessary and under an explicit retention policy.


17. Deployment Checklist


18. Known Limitations


Sources