Jev: an AI model that only makes decisions

Oct 1, 2026

Most of the AI plumbing I write doesn't need an essay. It needs an answer. Is this listing a duplicate? Which queue does this support message go to? Should this PR get flagged for a human? For two years the standard move has been to ask a chat model, beg it for JSON, parse the JSON, and hope.

On September 15, TypeSafe AI came out of stealth with a model built for exactly that gap. It's called Jev, and the pitch is simple: it never generates text. You hand it some text or JSON and a typed question, and it hands back a choice and the probabilities. That's the whole product.

What it actually returns

A request is a state (the thing to judge) plus one or more questions. Every question is one of three shapes:

  • Choice: pick one option from a fixed list. You get the pick, a probability for every option, and a confidence value.
  • Score: rate on ordered levels you describe. You get a probability-weighted score and the distribution across levels.
  • Noul: a yes/no proposition. You get the probability that the answer is yes.

You can ask several questions in one call, and the call looks roughly like this:

{
  "state": "Hi, I was charged twice for my order last week and want the extra one refunded.",
  "questions": {
    "intent": {
      "type": "choice",
      "instructions": "What does the customer want?",
      "criteria": ["refund", "order status", "cancel", "other"]
    },
    "urgent": {
      "type": "noul",
      "instructions": "Does this need a reply today?"
    }
  }
}

No prose comes back, which means no output tokens to pay for. Pricing is $0.042 per million input tokens, with a 32,000-token context window. TypeSafe's own telemetry puts production latency at a P50 of 0.21 seconds and a P95 of 0.34 seconds. For comparison, their passage-checking demo clocks Jev at 0.35 seconds against 8.83 seconds for Fable 5.1, which is a vendor number, but the shape of it matches what everyone else is measuring.

The founders call this a "System One model," after Kahneman's fast, intuitive mode of thinking. A chat model reasons its way to an answer one token at a time. Jev scores all the options in a single pass. The CEO, Diogo Almeida, co-wrote the InstructGPT paper, so this isn't a side project; the company raised a $40 million seed from DCVC on day one.

Does it beat a real LLM?

On Banking77, a 3,080-message intent classification set, Jev scored 81.1 percent against 76.4 percent for a Qwen baseline at the same median latency, about 245 milliseconds. The more interesting number is calibration: when Jev reported a confidence of 1.00, it was right 97.1 percent of the time. That's the feature. A model that tells you when to trust it is worth more than a model that's two points more accurate.

Except the calibration has a soft middle. In the 0.7 to 0.9 confidence band, accuracy dropped to 53 percent. And when the Towards Data Science reviewer fed it 30 messages that fit none of the categories, Jev picked a listed category every single time, with confidence of at least 0.99. It cannot say "none of these." You have to give it that option yourself, and even then it's your job to check.

It's also text only, no images. No arithmetic, no date math, no tool calls, no multi-step anything. And since it reads whatever you put in state, it's as open to prompt injection as anything else.

The two-week gap

Fourteen days after Jev shipped, OpenAI announced a Decisions API at DevDay that does the same thing by constraining GPT-6 Luna to a fixed answer set. OpenAI claims 150 milliseconds against 1.6 seconds for the unconstrained model, at $0.10 per million input tokens, and it takes images. It isn't broadly available yet. So the category is real; the biggest lab in the world just agreed.

Where I'd use it

Routing and gating, not judgment. "Which of these five buckets" and "is this safe to run" are Jev-shaped problems. "Is this listing actually the same property as that one" is not, because the honest answer is often "neither," and Jev will confidently pick one anyway.

The pattern I'd reach for: Jev decides, with a "none of the above" option always on the list, and anything under 0.95 confidence goes to a slower model or a person. Cheap, fast, and honest about the one thing it can't do.

Sources