TypeSafe Jev — First-Hands Evaluation Report

Verdict

Jev is a decision-only model — send state + typed questions, get calibrated probabilities instead of generated text — and it does what it claims on first hands: ~0.9s end-to-end per call from Vietnam, ~$0.00004 for a 9-question fan-out (vendor price list, not metered on our side), perfect calibration on a 10-fact toy probe (MAE 0.018), clean typed errors (400/422), and fan-out batching that cut wall time 8.65x and tokens 4.29x vs sequential calls with no answer drift. It is not a chat model and cannot replace one; it is a fast structured-decision primitive to sit beside an LLM. The TypeSafe agent skill is installed and its two warnings (fan out, don't invent request fields) were both validated by the experiments.

What Jev is

Measured results (live, 2026-09-19)

#ExperimentResult
E1Fan-out: 1 call, 9 mixed questions (Noul/Choice/Score)200 OK, 0.98s, 799 in / 232 out tok; sensible answers (urgent noul 0.99, department=billing conf 1.0, churn score 2.0 conf 1.0)
E2Same 9 questions as 9 sequential calls8.478s total (mean 0.942s/call), 4428 tokens total → fan-out is 8.65x faster wall, 4.29x fewer tokens, answers equivalent
E3Calibration probe: 10 labeled yes/no facts, one callMAE 0.018, Brier 0.001; all p within 0.05 of label. TOY probe (n=10, general facts) — says nothing about domain calibration
E4Confidence behaviorClear question conf 1.0; ambiguous question still chose singapore conf 0.88 (wrong-ish: Quito is closer to the equator — real hallucination case caught, see findings); contextless score ("how happy is the writer?" with weather-only state) returned score 1.0 conf 0.99 — model fills defaults confidently when state lacks the dimension
E5Structured JSON state + backticked path refs (ticket.messages[0].text, order.charges)All correct: refund_requested 0.99, policy_supports 0.99, n_charges=two conf 1.0
E6Error shapesUnknown model → 400 {"detail":{"error_type":"api_usage_error",...}}; empty questions → 422 FastAPI validation (min_length 1); bad question type → 400 "Invalid request."
E7jev-preview aliasResolves to the SAME jev-1.13.0 as jev-latest today
E8Response headersserver: istio-envoy, x-typesafe-request-id: req_... — no rate-limit headers on non-throttled responses

Findings

  1. Fan-out is the idiomatic pattern and it delivers. One call with all independent questions was 8.65x faster and 4.29x cheaper than sequential single-question calls on identical state (measured). The docs' "speculative fan-out" advice is real: batch every independent question, including ones only relevant for some inputs, and let code pick which answers to consume.
  2. Calibration claim held on the toy probe (MAE 0.018, Brier 0.001, n=10). This is the product's core bet; a real calibration harness on YOUR domain data is the first serious experiment to build (see Next steps).
  3. Confidence ≠ probability, and both have blind spots. Confidence is derived from distribution shape. E4 caught two instructive cases: the model was wrong-but-confident on a geography comparison (0.88 on singapore over quito), and confidently defaulted (1.0) on a question whose state lacked the needed dimension. Calibrated ≠ correct per-call; never gate destructive actions on a single call.
  4. Noul has no confidence field — the probability IS the signal. Choice/Score carry confidence. Design question sets accordingly.
  5. Errors are typed and cheap to handle. 400 = api_usage_error with message, 422 = FastAPI validation detail. Empty questions is a 422, not "no-op" — fail fast in client code.
  6. Latency from Vietnam is ~0.9s/call end-to-end (vendor claims 70-500ms measured US West). Still 3-300x faster than LLM-roundtrip workflows. Plan network latency into UX budgets.
  7. Version pinning matters: jev-latest moved under us before (docs: jev-1.12 → 1.13). If tuning thresholds, pin the versioned ID (jev-1.13.0) — the response's model field always reports what answered, so log it.
  8. Token metering includes output tokens for structured answers (232 out on E1) even though pricing says output is "free" — usage is reported, cost is input-only.

Where it fits the user's stack

Next steps (proposed, not started)

  1. Real calibration harness on user-domain data (Covidence abstracts with known verdicts) — reliability bins + ECE, JSON + markdown report. This is the experiment that matters.
  2. A small experiment CLI in a ~/Projects/jev-experiments sibling repo (key in gitignored .env) — built via superpowers plan + omp -p --model anthropic/claude-opus-5:xhigh per standing workflow, when user says go.
  3. If the harness holds up: wire a Jev gate into the Covidence full-text-review workflow as a first-pass sorter with confidence-gated escalation.

Caveats

References