Decision AI that never writes a word, and why you should care
When the job is a decision, route this ticket, approve this refund, escalate this alert, you don’t want prose. You want a typed label and a confidence number your code can branch on: “refund=approve” and “confidence=0.92”. Decision AI models return those machine-native answers (choices, scores, probabilities) instead of freeform text. That changes how you build automation, agents, and realtime systems.
That shift accelerated in 2026 with commercial launches and open-weight releases. The category promises cheaper, faster, programmatic judgments, but vendor claims and experimental reproductions vary. Know what decision models do well, where they fail, and which operational checks matter before you let one start calling shots.
What a decision model returns
At the core a decision model consumes a compact “state”, a string or small structured payload describing context, and returns one of a few typed primitives that code understands directly:
- Choice: pick one option from a predefined list (with a probability distribution across options). TypeSafe states Jev supports up to 255 options.
- Score: place the input into an ordered rubric (for example, severity buckets), returned as a distribution over bins and a confidence.
- Binary / Bernoulli: a 0-1 probability that a statement is true. (TypeSafe refers to this primitive as “Noul” in its docs.)
These primitives are machine-native: your code branches on a label plus a numeric confidence, not on parsed natural language. That removes a brittle parsing layer, but it does not remove the need for validation, schema enforcement, and audit logs upstream and downstream.
TypeSafe says that because Jev “never generates strings, it cannot return a type error.”
That phrasing is a TypeSafe engineering claim about API behavior. Treat it as a design guarantee at the interface level, not a mathematical proof against every downstream schema or integration bug.
How Jev works, and what those vendor terms mean
TypeSafe markets Jev as a “System One model, ” an explicit reference to Kahneman’s idea of fast, intuitive judgments. The message is clear: immediate, programmatic answers, not long-form generation.
Two vendor-named technical ideas drive Jev’s positioning:
- Parallel sampling: multiple questions are evaluated in parallel and in isolation against the same state. Vendors say this keeps per-question latency nearly constant as you add checks.
- RLCD (Reinforcement Learning for Calibrated Decisions): a training approach TypeSafe describes as optimized for probability calibration, so a declared confidence score correlates with empirical accuracy. TypeSafe contrasts RLCD with RLHF, which targets human-preference alignment rather than calibrated probabilities.
Both are vendor-named design choices. The high-level consequence is clear: these models aim to return fast, well-calibrated probabilities rather than human-style prose. The precise training recipes and reproducibility details remain company-level claims and should be validated for your domain.
The competitive landscape, quickly
The category split into two pragmatic camps: managed, API-first decision engines and compact, open-weight decision encoders you can run locally.
- TypeSafe Jev (API), positioned as a managed, low-latency, calibrated decision service (TypeSafe publishes demo timing and pricing examples on typesafe.ai; validate against live workloads).
- Fastino Labs, published GLiNER2.5‑Decide (an open-weight model) on September 24, 2026 and followed with GLiDE later that month. Fastino’s post includes p50 latency numbers on GPUs and an internal “Fast Decisions” benchmark.
- Community reproductions, projects like JevK5, OpenJev, kev and Laya let teams experiment locally or air‑gap deployments; reproductions are not identical to proprietary hosted models and can differ in calibration and behavior.
- Tooling integrations, OpenRouter, Langfuse, Arize and CI tools such as Buddy are adding decision-model hooks for routing, gating and evaluation.
Open weights matter for auditability and on‑prem needs. API-first offerings matter for quick integration and a managed calibration story. Choose based on governance, latency, and reproducibility requirements.
A brief note on benchmarks and vendor numbers
Vendors publish promising numbers, demo costs, median latencies, workflow eval tables and adoption signals, but these are vendor-reported and often use different datasets, baselines, or community reproductions for comparisons. Treat them as directional signals to shortlist options, not definitive, apples-to-apples proof. Always run your own benchmark with your inputs, labels and latency constraints.
Where decision models excel
- Agent control flow: pick the next tool or subagent, gate expensive calls, or decide which microservice to invoke. Example: auto-route a refund to fraud investigation if refund_reason ∈ {“suspected_fraud”, “charge_dup”} else auto_approve.
- High-volume triage and classification: ticket routing, intent classification, spam and abuse triage at scale where speed and cost matter more than narrative explanation.
- Guardrails and verified cascades: draft text with a cheap LLM, run a decision model as a fast verifier, and escalate only low-confidence cases to humans or to stronger models.
- Reranking and automated evaluation: search reranking and CI/eval pipelines where structured judgments and confidence numbers simplify automation.
- Realtime and edge checks: latency-sensitive policy decisions at CDN or edge points, an arXiv-cited use case reported median-latency wins for a decision-model approach at the edge (see vendor-cited results; validate for your stack).
Where they break, concrete limits
- When you need explanations or narrative: decision models give labels and probabilities, not human-readable reasoning. If regulators, auditors, or customers require textual justifications, pair the decision model with a generative LLM or keep humans in the loop.
- Exact arithmetic, counting, date math: vendors warn about “jaggedness” (non‑smooth, error-prone behavior on arithmetic, exact counts, date calculations). Don’t use decision models as precise calculators.
- High‑stakes, legally consequential decisions: hiring, credit denial, parole-like decisions demand auditability, traceable data, and legal review, probabilities alone don’t satisfy regulatory or fairness requirements.
Calibration: the practical currency
Calibration is not a marketing buzzword, it is the property that declared confidence matches empirical accuracy. Ask vendors for these metrics on representative data:
- Expected Calibration Error (ECE), how far predicted confidences deviate from observed accuracy across bins.
- Brier score, a mean-squared error metric on probability predictions.
- Reliability diagrams, visual plots showing calibration across confidence ranges.
Two operational rules of thumb:
- Measure calibration on your domain data, not only vendor-provided benchmarks.
- Stress-test calibration under domain shift, edge-case language, nonstandard formats, or foreign languages can miscalibrate quickly.
Operational checklist before production (prioritized)
- Logging & observability, high priority: log inputs, outputs, declared confidences, timestamps, and model version. Make these logs queryable for audits and debugging.
- Fallbacks & escalation, high priority: define what happens on low confidence (human review, retry with a stronger model, or safe default). Implement immediate human-in-the-loop for the first production batches.
- Calibration pilot, medium priority: run a focused pilot on representative workloads to measure ECE/Brier and tune thresholds before full rollout.
- Schema enforcement & validation, medium priority: ensure your integration enforces allowed labels and handles unexpected or missing outputs.
- Governance & legal gates, high priority for regulated use: decide whether you need open weights for audit, and involve legal and compliance early if decisions materially affect people.
- Drift monitoring, ongoing: watch calibration drift, label distribution shifts, and model version changes. Automate alerts when thresholds cross.
Picking between Jev, GLiDE, GLiNER2.5‑Decide and open models
- Managed API (TypeSafe Jev), choose this if you want a turnkey integration, vendor-published calibration narrative and ecosystem integrations. TypeSafe’s demo (typesafe.ai) shows a tiny example completing in 0.114 s for $0.000081; treat demo numbers as illustrative and validate on your payloads.
- Open-weight encoders (GLiNER2.5‑Decide and friends), choose open weights when you must run air‑gapped, reproduce benchmarks exactly, or perform model-level audits. Fastino published GLiNER2.5‑Decide on September 24, 2026 and released the weights; their post reports p50 latencies (GPU/CPU) for a 15-label schema.
- Community reproductions (JevK5, OpenJev, kev, Laya), useful for experimentation and local validation, but they are not identical to proprietary hosted models; results can differ in accuracy and calibration.
- Hybrid pattern, a common pragmatic pattern is draft-and-check: use a cheaper generative model for text then validate with a decision model and escalate low-confidence items.
Risks you cannot outsource to a vendor
- Hidden bias: probabilities are comfortingly precise but do not guarantee fairness. Run counterfactual and subgroup analyses and retain human oversight for sensitive classes.
- Vendor lock-in: API contracts, label