Jev: the model that can't talk

Jev: the model that can't talk

September 22, 2026


A model with nothing to say

Everything in AI changes forever roughly every eight days, so we have learned to discount the phrase. Still, something interesting landed this month: a new model, built by an ex-OpenAI researcher's startup after a long stretch in stealth, that cannot write code, cannot draft your emails, and will never open a reply with "you're absolutely right."

It cannot talk at all. That is the feature.

The pitch is that large language models have one structural flaw for a large class of jobs: they will not shut up. Ask a frontier model for a simple true or false and it may well discover a third option, reason about it for four thousand tokens, and bill you for the privilege. The response here was to delete language from the language model and ship what is essentially a classifier with a typed interface.

Typed questions, typed answers

You send it a question and some unstructured context, same as any model. The difference is that the question is strongly typed and must return one of three shapes:

  • A choice, one option from a set you define.

  • A score, a number in a range you define.

  • A bool, plain yes or no.

Because there is no free text anywhere in the output path, schema conformance is not a thing you validate and retry. It is structurally guaranteed. If you have ever written a parser to catch a model wrapping its JSON in a friendly paragraph, you can see the appeal immediately. The company name (Typesafe AI) is doing the explaining, and the mental model is closer to TypeScript than to chat.

System 1, not system 2

The framing borrowed here is the fast-and-slow split popularised by Thinking, Fast and Slow. Reasoning models are system 2: deliberate, expensive, and capable of burning forty thousand tokens naming a variable. This is system 1: gut instinct, no deliberation, answer or nothing.

The self-reported numbers are startling. Roughly 200 times faster, several hundred times cheaper, no output token cost, and no hallucinations in the usual sense because there is no text to hallucinate in. Treat those figures as vendor benchmarks until someone independent runs them, but the shape of the claim is plausible: you are not paying for generation you were going to throw away.

Fast and cheap enough changes what you build. Real-time moderation, NPC behaviour in games, per-keystroke decisions: all things nobody sane routes through a reasoning model today.

Type-safe is not the same as correct

This is the caveat worth holding onto. A guaranteed shape tells you nothing about the content inside it, and the model is not deterministic: same question, same context, different answer is entirely possible.

A type error is impossible. A wrong answer is not.

What it does return alongside the answer is a calibrated confidence value, trained through a reinforcement learning method aimed at calibration rather than at pleasing human raters. That last distinction matters more than it sounds. Chat models are tuned on human preference, humans reward confidence, and that is a large part of why we ended up with models that are wrong in a very self-assured tone. A number that actually means what it says is a real improvement, if it holds up.

The sceptics have a point

The architecture has not been published. There is a paper "possibly coming," which is not a citation. Meanwhile the objections are specific rather than vibes-based:

  • Researchers were building zero-shot classifiers over a decade ago, and the launch materials credit none of that lineage.

  • At least one researcher says a paper published a year ago describes the same approach.

  • Someone has already shipped an open-source reimplementation that reproduces the whole interface by reading option probabilities off a frozen Qwen 4B in a single forward pass. No new training, runs on a consumer GPU, with a WebGPU demo in the browser.

If a single forward pass over an open 4B model reproduces the product, the interesting thing was never the weights. It was noticing that most teams are using a conversation to do the work of a function call.

What this means for teams building with AI

We have been routing decisions like these badly across the industry, and we include ourselves in that. The lesson is not "adopt this model." It is that most agentic systems make a lot of small, boring, high-volume judgements, and pointing a reasoning model at every one of them is how budgets and latency quietly die.

  • Sort your calls by kind. Anything answerable as a choice, a score or a bool is a classification problem wearing a chat interface.

  • You can do this today. Constrained decoding and reading option probabilities off a small open model are established techniques. You do not need a new vendor to stop paying for tokens you discard.

  • Calibration beats confidence. A score you can threshold on, and route to a human below, is worth more than an eloquent answer.

  • Keep the human where the stakes are. Fast gut-instinct decisions are exactly the ones worth sampling and reviewing, precisely because nobody reads them individually.

The product may or may not survive contact with independent benchmarks. The design principle behind it, that the interface to a model should be a type and not a conversation, looks right regardless of who ends up owning it.