AI News Feed
Market watch
Products & Applications

Decision AI Models Enter Mainstream as TypeSafe's Jev Draws Fastino and Open-Source Rivals

TypeSafe AI's Jev returns typed choices, scores and probabilities instead of generated text, pushing decision AI models into the mainstream. Fastino Labs and open-source developers have followed, while early evals show competitive accuracy at much lower cost and latency.

Decision AI models are a new class of model that returns a decision, not a paragraph. A user sends text with typed questions, and the model returns choices, scores or yes/no probabilities that code can branch on directly. TypeSafe AI launched Jev after two years in stealth and calls it a 'System One model,' after Daniel Kahneman's fast, intuitive System 1 thinking. Fastino Labs later shipped two rival models, and open-source developers published several Jev-style reproductions, according to the report.

Jev accepts a 'state' — a string, array or set of name-value pairs — and one or more questions. According to TypeSafe's documentation, it supports three primitives. Choice picks one option from a list, with probabilities and confidence, and TypeSafe says Jev supports up to 255 options. Score rates the state on a rubric of ordered levels, with probabilities and confidence. Noul, short for Bernoulli, returns a 0 to 1 probability that a statement is true. Every question is evaluated in parallel and in isolation against the same state, and adding questions barely changes response time. Because Jev never generates strings, TypeSafe says it cannot return a type error. Under the hood, TypeSafe describes a parallel sampler and a training method called Reinforcement Learning for Calibrated Decisions, or RLCD. While RLHF optimizes for human preference, RLCD optimizes for calibrated probabilities, where higher confidence should mean higher accuracy. Jev costs $0.042 per million input tokens, with output free, and OpenRouter lists a 32K context window. TypeSafe reports end-to-end responses in 70 to 500 milliseconds.

TypeSafe built workflow evals across four tasks: security incidents, agent trace observability, invoice processing and customer service. Reference labels came from averaging GPT-6 Astra and Claude Fable 5.1 at high thinking. Jev scored 67.8% mean accuracy at $0.0004 per case and 0.4 seconds. Claude Sonnet 5, run on the same workflow, also scored 67.8%, at $0.1174 per case and 78.1 seconds. The best comparison model, OpenAI 'sol,' scored 74.1% at $0.0836 per case and 23.3 seconds. Jev matched Sonnet 5 on accuracy at a fraction of the cost and latency, but still trailed the top frontier configuration by 6.3 points. Per task, Jev scored 76.0% on customer service but only 61.8% on invoice processing.

The rule of thumb is simple. If code needs a bounded answer it will branch on, a decision model is a candidate. If a human needs to read the output, use an LLM. For agent control flow, Vercel lists choosing the next tool or subagent as a primary use. A single Choice question can replace a fragile JSON-parsing step when an agent must continue, retry, ask the user or stop. Fastino lists model routing by destination, complexity or escalation level, sending easy requests to a cheap model and hard ones to a frontier model.

Classification and triage are also central. Support ticket routing, email triage and intent detection make up much of Fastino's 17-dataset benchmark. Simon Willison calls spam detection, label suggestions and prioritization natural classification fits. TypeSafe's simplest published security workflow decides whether to close an alert, pass it to an analyst or contain it.

For verification and safety, TypeSafe pitches Jev for scoring prompts, reasoning traces and outputs. GLiNER2.5-Decide can decode 'safety' and 'harm type' together so the answers never contradict. OpenRouter's Jev guide describes drafting with a cheap model, checking with Jev, and escalating only on failure.

For evaluation and observability, Arize and Langfuse now run Jev evaluators on traces. Buddy lets pipelines score, classify or gate runs with a Jev action.

For search and data processing, Jev can score 100 BM25 candidates for relevance in one parallel call. TypeSafe pitches map-reduce over large datasets to turn bulk data into features at low cost. Fastino lists choosing which context to keep before an LLM call.

In real-time applications, TypeSafe demoed Jev playing Doom and Wikiracing. Sub-second decisions make AI usable inside interactive flows. Vercel also reports that Jev became the fastest-adopted model in AI Gateway history.

There are clear cases not to use a decision model. It is not suited to generating text, summaries or explanations, or to tasks requiring exact arithmetic, counting or date math. TypeSafe's Jev 1.13 jaggedness guide flags all three. It also warns about decisions that affect people's livelihoods, such as hiring, and Willison warns that hidden bias is hard to inspect.