Jev: State In, Typed Decisions Out
Diogo Almeida posted on X about a new model, Jev, from a company called TypeSafe. TypeSafe calls it its first “System One” model, a term borrowed from Daniel Kahneman’s split between fast, intuitive System 1 thinking and slow, deliberate System 2 reasoning. A normal LLM writes out its reasoning in text. Jev skips straight to a typed answer.
What it actually does
Ask a normal LLM “given this incident, what should we do?” and it answers in text:
<think>Let's look at the incident closely...</think>
Fetch up metric for service A.
Send the same situation to Jev as structured state, along with a set of typed questions, and it returns a typed, probabilistic answer to each one:
{"check_up_metric": 0.9, "page_engineer": 0.1}
TypeSafe’s docs describe System One models plainly: built “to make fast, structured decisions that software can use directly.” They “do not write replies, produce code, or generate explanations of their reasoning.”
The three primitives
Every question you send to Jev is one of three types, called primitives in the docs:
| Primitive | Question shape | Returns |
|---|---|---|
| Choice | Pick one option from a list (up to 255) | choice, probabilities per option, confidence |
| Score | Rate the state on an ordered rubric (2-10 levels) | score, legend, probabilities per level, confidence |
| Noul | Is this statement true? | noul, a probability from 0 to 1 |
Score’s number isn’t a discrete pick — it’s the probability-weighted mean of the level positions. For a 3-level scale (0, 1, 2) with probabilities 0.0, 0.7, and 0.3:
$$ \text{score} = 0 \times 0.0 + 1 \times 0.7 + 2 \times 0.3 = 1.3 $$
Noul has no separate confidence field. The probability itself carries that: 0.92 says “very likely yes,” 0.5 says “no idea.”
Every question in a request runs against the same state, independently of the others. The docs put it directly: “One question’s answer is not hidden context for another. You can add or remove questions without changing the others’ results.”
Calling it
The API is one endpoint, POST https://api.typesafe.ai/v1/systemone, authenticated with a bearer token. A request is state plus a map of questions:
{
"state": "Help! My payouts have been failing for 3 days.",
"model": "jev-latest",
"questions": {
"is_urgent": {
"type": "noul",
"instructions": "Does this convey urgency?"
},
"department": {
"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {
"billing": "Payments, invoicing, refunds",
"technical": "Bugs, outages, integrations",
"sales": "Pricing, upgrades, new accounts"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Frustrated", "Very angry"]
}
}
}
The response answers every question in one shot:
{
"model": "jev-latest",
"answers": {
"is_urgent": { "type": "noul", "noul": 0.92 },
"department": {
"type": "choice",
"choice": "technical",
"probabilities": { "billing": 0.08, "technical": 0.85, "sales": 0.07 },
"confidence": 0.82
},
"frustration": {
"type": "score",
"score": 1.6,
"legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" },
"probabilities": { "0": 0.05, "1": 0.3, "2": 0.65 },
"confidence": 0.78
}
},
"usage": { "input_tokens": 312, "output_tokens": 48 }
}
A request holds roughly 32,000 tokens of state and questions combined, about 150,000 English characters. The current model is jev-1.13.0, with jev-latest and jev-preview both pointing at it for now. Pricing is $0.042 per million input tokens; output tokens are free. There’s a Python SDK too:
from typesafe_sdk import TypeSafeClient
client = TypeSafeClient()
response = client.system_one(state=ticket, questions={...})
It can’t produce text
Choice, Score, and Noul all output a probability distribution over a fixed schema you define upfront. None of them output free text. That constraint is what makes the speed and the “can’t hallucinate” claim possible.
I saw someone work around this anyway: define a Choice over a single-character alphabet, “which letter comes next,” and ask that question over and over, appending each answer to the state before the next call. Run enough rounds, and you’ve spelled out a sentence one character at a time.
It works, but it’s an abuse of the design. Each character costs a full API call, so generating a paragraph this way costs far more than a normal model just writing the paragraph. Jev is built to pick one of a few known options fast, not to write open-ended text.
Speed and cost
TypeSafe’s own numbers, from the launch post: end-to-end response time of 70-500ms, against 3 to 329 seconds for frontier LLMs on the same tasks, which they put at 40x-200x faster. For full workflow evaluations they report 193.6x faster and 444.6x cheaper, while noting that figure sits at “the higher end of real world gains.”
RLCD: trained for calibration, not just correctness
TypeSafe trains Jev with what it calls Reinforcement Learning for Calibrated Decisions (RLCD), which it lines up against two more familiar objectives:
- RLHF (Reinforcement Learning from Human Feedback) optimizes for text a human rater prefers.
- RLVR (Reinforcement Learning with Verifiable Rewards) optimizes for an answer that can be checked as correct.
- RLCD optimizes for calibrated probabilities: “answers with epistemically honest probabilities,” in TypeSafe’s phrasing.
Calibrated means the number tracks reality: if Jev outputs 0.8 across many similar predictions, about 80% of them should turn out correct. The docs describe the training this way: “their probabilities are optimized against outcomes to reflect uncertainty.” TypeSafe hasn’t published the RLCD algorithm itself.
The confidence field in a Choice or Score answer is a separate, simpler thing — a single number computed from the shape of the probabilities you already got back. A distribution stacked on one option gives a high confidence; a flat spread gives a low one. It’s there so you can threshold on certainty without computing that yourself, not a second model output.
“Can’t hallucinate”
TypeSafe’s launch post states plainly that Jev “can’t hallucinate,” then adds a specific caveat: “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.”
That’s a claim about the output’s shape, not its content. If the schema is:
Choice(["refund", "rebook", "support"])
Jev cannot return "give_customer_a_free_spaceship". No malformed JSON, no invented enum value, no stray prose, no parser failure — the guarantee is exact and unconditional.
What it doesn’t guarantee is that the chosen option is the right one. Jev can still assign page_engineer a probability of 0.9 when paging the engineer is the wrong call. Whether that happens depends on the training data, same as any model.