In mid-September 2026 TypeSafe AI released Jev, and within a week LangChain had published two guides on using it inside agents. Most developers have heard the name. Fewer know why you’d call a separate „decision model” when GPT or Claude can already answer „is this ticket urgent?”. This post explains it.
What is Jev?
Jev is what TypeSafe calls a System One model: a model that makes fast, structured decisions your software can use directly. You give it some state (a ticket, a JSON payload, a chat history) and a set of typed questions. It gives back typed answers with calibrated probabilities. It never writes prose.
The name comes from Daniel Kahneman’s split between fast, intuitive „System 1” thinking and slow, deliberate „System 2” thinking. LLMs act as System 2: they reason step by step, in text. Jev handles the quick judgement calls, like „urgent or not?”, „billing or technical?” or „safe to run this command?”.
Why Jev isn’t an LLM
An LLM’s output is text. Even when you ask it for JSON, you’re parsing generated tokens and hoping they match your schema. Jev’s output is the answer itself: a number for each question. It supports three question types:
- Noul: a yes/no question. Returns the probability that the answer is „yes”.
- Choice: pick one of several options. Returns a probability for each option.
- Score: rate something on ordered levels (low / medium / high). Returns a continuous score and a distribution.
System One vs. traditional LLMs
| Jev (System One) | LLM (System Two) | |
|---|---|---|
| Output | Typed answers + probabilities | Free-form text |
| Best at | Classification, routing, yes/no checks | Reasoning, writing, code, summaries |
| Answer set | Known in advance | Open-ended |
| Speed & cost* | Up to 200× faster, 400× cheaper | Baseline |
| Consistency | Scores barely move across repeated runs | Varies between runs |
How Jev makes decisions
Every question in a request is evaluated in parallel, so asking five questions takes about as long as asking one. You only pay for the tokens of the extra questions. Because each answer is a calibrated probability, you choose the threshold. For example: act automatically above 0.9, ask a human between 0.6 and 0.9, and ignore anything lower.
A simple example: classifying a support ticket
With the langchain-typesafe package and a TYPESAFE_API_KEY environment variable set, a single call answers three questions about one ticket:
from langchain_typesafe import Noul, TypeSafeClassifier
classifier = TypeSafeClassifier()
ticket = "I was charged twice this month and the invoice page returns a 500."
response = classifier.invoke({
"state": ticket,
"questions": {
"urgent": Noul(instructions="Does this need attention right now?"),
"is_billing": Noul(instructions="Is this about payments or invoices?"),
"needs_human": Noul(instructions="Does resolving this require a human decision?"),
},
})
urgent = response.nouls["urgent"].noul # e.g. 0.94
is_billing = response.nouls["is_billing"].noul # e.g. 0.88
needs_human = response.nouls["needs_human"].noul # e.g. 0.12Your code makes the decision with ordinary if statements, so the logic is easy to test and to log:
if needs_human > 0.7 or urgent > 0.9:
escalate_to_on_call(ticket)
elif is_billing > 0.8:
route_to("billing", ticket)
else:
draft_reply_with_llm(ticket) # only now do we pay for an LLMJev inside an AI agent
Many agents today use a frontier model for every step, including simple ones like „which tool should I call?”. Jev reverses that. The graph (for example in LangGraph) sets the control flow, Jev makes the small judgement calls, and the LLM is only called when you need open-ended reasoning or generated text. LangChain calls this „cheap by default, frontier on exception”.

Model routing with Jev
Not every request needs your most expensive model. LangChain’s experimental router middleware asks Jev which tier fits each request, then calls that model:
from langchain_typesafe.experimental.middleware import (
ModelChoice,
ModelRouterMiddleware,
)
router = ModelRouterMiddleware(
choices={
"fast": ModelChoice(model="openai:luna", criteria="Direct lookups"),
"powerful": ModelChoice(model="openai:sol", criteria="Complex decisions"),
},
)Using Jev as a safety/guardrail layer
Before an agent runs a shell command, issues a refund or sends an email, Jev can check whether the proposed tool call looks risky, and your code blocks it if it does. The check adds very little latency, so you can run it on every call:
from langchain.agents import create_agent
from langchain_typesafe.experimental.middleware import AutoModeMiddleware
guardrail = AutoModeMiddleware(tools=["bash"])
agent = create_agent("openai:gpt-5.6-luna", middleware=[guardrail])Both middlewares are in an experimental namespace. Pin your versions and expect the API to change.
Jev vs. an LLM: when should you use each?
Use Jev when you already know the possible answers: routing, triage, tagging, intent detection, „is this safe?”, „is this done?”, and quality checks on an LLM’s output.
Use an LLM when the output is text or code, when the task needs several reasoning steps, or when it means reconciling several documents. Jev doesn’t generate strings, and early reports show it trails frontier models on cross-document tasks such as invoice matching.
Is Jev actually useful or just another new AI model?
The headline figures (200× faster, 400× cheaper) are TypeSafe’s own benchmarks, so treat them with care. LangChain’s independent tests are more modest but still meaningful: the classification step in a document-review graph ran 5–6× faster than with Sonnet, and a Stagehand browser agent’s decision step fell from 1.97 s to 0.46 s.

It also has limits. TypeSafe’s own documentation notes that Jev is unreliable at arithmetic and date comparison, and that it answers the literal question you wrote rather than the one you meant. Keep the state you send short and relevant, and don’t let Jev alone make irreversible, high-cost decisions such as money transfers or eligibility checks.
Conclusion: code controls the workflow, AI makes narrow decisions
Jev doesn’t replace GPT or Claude. It replaces using them for small yes/no and multiple-choice decisions. The architecture that works well is simple: deterministic code owns the control flow, a System One model makes the narrow decisions quickly and cheaply, and an LLM is called only when the task needs real reasoning or writing. For teams running agents in production, that means lower cost, lower latency and more predictable behaviour.
Building an AI agent and want to cut latency or cost? Talk to the Fireup team.








