Design & ProductARTICLE

Jev: the decision model that trades text generation for instant responses in production

In an episode of the podcast How I AI, educator John Lindquist shows eight real use cases for Jev, a decision model from TypeSafe AI that trades text generation for near-instant responses and cost low enough for ideas that used to be impractical.

In the September 30, 2026 episode of the podcast How I AI, hosted by Claire Vo, educator John Lindquist (creator of egghead.io and now leading mega.dev) showed eight real use cases for a model called Jev, from TypeSafe AI. The episode's promise is right there in the title: it's the fastest and cheapest model he's ever used. The interesting angle for anyone building product isn't speed itself, but the architectural proposal behind it: Jev wasn't built to converse, it was built to decide.

The fastest, cheapest model I've ever used.

John Lindquist, creator of egghead.io and founder of mega.dev

Decision, not conversation

TypeSafe AI calls Jev a "System One model," a direct reference to Daniel Kahneman's distinction between fast, intuitive thinking (system 1) and slow, deliberate thinking (system 2). A generative LLM like the ones running behind ChatGPT or Claude generates text token by token, with room for chain-of-thought reasoning before the response. Jev doesn't generate free-form text: it takes an input and returns a structured decision within a finite set of options, something closer to an extremely fast classifier than a chatbot.

That difference in contract is what Lindquist calls a "decision engine, not a chatbot." In practice, this changes the kind of question it makes sense to ask the model: instead of "write a response to this," the question becomes "which of these actions corresponds to this." It's a narrower slice of the problem, but one that opens room for uses that would be too costly or too slow with a full generative model.

Eight demos, a pattern that repeats

Lindquist ran live demonstrations throughout the episode, and most follow the same pattern: take a classification or routing task that today is handled with code rules or a heavy LLM, and solve it with a single call (or a short sequence of calls) to Jev. Among the examples shown:

  • A voice-driven task app that classifies and executes commands in real time, with no perceptible pause between speech and action
  • Deduplication of messy records, merging duplicate entries based on confidence scores
  • Translation of free text into function names, tested with a shopping cart that interprets everyday phrases as commands
  • A multi-level router, where a single text input navigates the user to deep points within an app
  • A Wikipedia route mapper, following the informal "path to philosophy" challenge by clicking through successive links
  • Coordination among multiple agents with collision avoidance
  • A real-time presentation coach

The common thread is that none of these cases called for creative text generation. They all called for a quick decision among known options, and that's exactly where, according to Lindquist, Jev's architecture makes up for what it gives up in flexibility.

Chess as an informal benchmark

The most direct point of comparison in the episode is a chess match between Jev and what the episode's notes describe as "a low-reasoning LLM." The source doesn't disclose exact latency or cost-per-call figures for this test, but the point being demonstrated is qualitative: in a task where every move is, at bottom, a choice among a finite set of valid moves, a decision model like Jev responds much faster than a generative model spending tokens on reasoning before deciding its move.

It's worth noting that the episode doesn't name GPT-4o or Claude as the specific opponent in this match, nor does it publish a benchmark table. Lindquist's point is more about the category of problem than about which specific generative model loses. For anyone deciding on a production stack, the practical lesson is: if the task is choosing among known options, it's worth measuring whether a decision model solves it before setting aside a token budget for a full reasoning model.

Where the decision model doesn't work

The episode's own script sets aside a block for Jev's limits, and that's just as important as the success stories. Lindquist is clear that when a task requires original text generation, long-form synthesis, or open-ended reasoning without pre-defined options, the decision model simply isn't the right tool. That's where he still turns to a full generative model: Anthropic's Opus 5.5 shows up in the episode precisely in the context of iteratively building the demos, not as part of the production decision pipeline.

That combination is the most useful takeaway for anyone architecting AI systems: Jev doesn't replace a generative LLM, it complements one. A pattern that emerges from the demos is using a decision model as an input router or classifier, and only escalating to a full generative model when the task genuinely calls for free-form generation.

How this reaches your stack

The episode cites two access routes to models that are already part of daily life for anyone working with multiple providers: Vercel AI Gateway and OpenRouter. Neither is exclusive to Jev, but both make it simpler to test a decision model side by side with the generative models you already use, without rewriting your entire integration layer.

The question left for anyone deciding AI architecture in production isn't "does Jev replace GPT-4o or Claude in my stack," because the episode's material doesn't support that direct comparison. The more honest question, and the one closer to what Lindquist actually demonstrated, is a different one:

How many of my calls to a generative LLM today
are, at bottom, a decision among known options
disguised as a free-text prompt?

If the answer is "several," it's worth measuring latency and cost by swapping those calls for a dedicated decision model before assuming the generative model is the only piece available.

Translated from the Brazilian Portuguese original · Read the original

View profile →