Jev: How to use it for better, faster, and cheaper user decisions in UX
A judgment call, wired in like any other function.
A judgment call, wired in like any other function.

Typed answers at keystroke speed let the interface act, suggest, or ask. How UX designers can put a decision model to work today in their applications.

On September 15, 2026, a San Francisco lab called TypeSafe AI shipped Jev, a classifier model. You hand it context, questions and possible fixed choices; it hands back a singe choice, a score, or a yes-or-no, each with a probability.

Absolutely brilliant.

I’m using Jev to power by newsletter scoring. It’s cost me all of 35 cents for over 5000 articles.
I’m using Jev to power by newsletter scoring. It’s cost me all of 35 cents for over 5000 articles.

Where do I sign up?

Can’t really argue about up and to the right.
Can’t really argue about up and to the right.

Everyone else thought the same.: Within 24 hours 13% of paid teams on Vercel’s AI Gateway had used it, and TechCrunch reported demand briefly knocked it offline.

What BERT Classification looks like.

Here’s the reality: Jev is an old concept with new shoes, and it’s going to change the landscape of AI. The open-source ones have been free for a decade:

  • Sort text with fastText. Facebook has shipped it since 2016.
  • Fine-tune BERT. Google’s encoder, since 2018.
  • Route with ModernBERT. Answer.AI built it in 2024 for the jobs Jev pitches.

What changed is the quality of the judgment, the speed, and the cost per call, making it incredibly accessible. That matters to designers because judgment is now cheap and fast enough to sit inside a control instead of behind a text box, so the confidence threshold becomes a design decision.

Here is what Jev is, how a call works, what it costs, why the economics change the brief, and where to put typed decisions in your product.

A frontier badge pinned on a 1958 idea.
A frontier badge pinned on a 1958 idea.

Call it a classifier

Strip the launch language away and the shape is familiar: an input goes in, one of a fixed set of labels comes out, usually with odds attached.

Frank Rosenblatt built one in 1958 and called it the perceptron. Forty years later, Mehran Sahami, Susan Dumais, David Heckerman, and Eric Horvitz published a Bayesian junk-mail filter at Microsoft Research that scored each message against a threshold. The same pattern approves your credit card and picks the ad you see.

If you have run a card sort you already know the shape.

https://medium.com/media/b0c397d68a47d19b6dc41aa4a0da3ac1/href

As Donna Spencer lays it out in Card Sorting, a closed sort hands participants fixed categories and asks where each card goes, while an open sort lets them invent the categories.

For this, a chat model runs the open sort and Jev runs the closed one. That is a powerful approach.

TypeSafe trained that shape at frontier scale on synthetic data, and its launch post benchmarks the result against chat models, not against a fine-tuned open classifier, so that gap is unmeasured. The founder, Diogo Almeida, helped build the instruction-following methods behind ChatGPT at OpenAI, and left because superhuman chat produced very little automation; software does not consume prose.

The launch post compresses the pitch to one line: “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” The novelty is that it’s a frontier model’s judgment and a spam filter’s price tag.

“System One” borrows from Daniel Kahneman’s Thinking, Fast and Slow, the fast, intuitive mode that pattern-matches instead of deliberating.

Jev is the gut check.

One state in, every answer out at once.
One state in, every answer out at once.

Run one call

It really is this simple, yet powerful.

All laid out in a nice flow chart.

The Inputs

  • A state in the form of content: a support ticket, a form in progress, a tool call an agent wants to make
  • A set of questions
  • A set of fixed options.

The outputs

  • A choice you named
  • A score returns a value on a scale you define
  • The Noul, TypeSafe’s name for a yes-or-no, returns the confidence of yes on a zero to one scale

The documentation allows dozens of questions per request, answered in one parallel pass instead of token by token, which is where the speed comes from: 70 to 500 milliseconds end to end, by TypeSafe’s own West Coast measurements.

Every answer carries two numbers: a probability per option and a confidence score for the question. A Noul that returns 0.08 is not unsure; it thinks the answer is no. A Choice that spreads 0.4, 0.35, and 0.25 across three options is telling you the state does not settle the matter.

That puts Jev closer to the deterministic end of the scale than the chat models, and for software that is the incredible news in a probabilistic world.

The answers are still probabilities, but the output shape is fixed in advance, a type error is impossible, and TypeSafe says similar states return similar answers.

A control that behaves the same way twice is something a designer can build on; a conversation that drifts is not. The second case lands on design. Jev cannot hallucinate, since it can only pick from the options you gave it, but it can pick wrong, and it tells you how sure it was.

Armin Ronacher, chief technology officer at Earendil, told TechCrunch the model “delegates the hallucination problem a little bit to the user.” That user has to decide whether 50% is a coin toss to ignore and 95% is a result to act on.

It’s still human in the loop in a really good way.

Jev moves it from the model to the person who sets the threshold. That person should be a designer and the threshold belongs in the specification next to the error states.

Metered by the billion, not the million.
Metered by the billion, not the million.

Count the cost

TypeSafe lists Jev at $0.042 per million input tokens, with output free, or $42 per billion. Its comparison table puts frontier chat models between $0.20 and $10 per million on input, with output costing about five times that.

At it’s latest cost, that puts it 250 times cheaper than Claude Fable 5.1.

Read that again, please — 250 times cheaper.

Other headline multiples of 194 times faster and 445 times cheaper come from TypeSafe’s own workflow evaluations, and the launch post says so, adding that it cannot yet show the pricing is unsubsidized. Treat those as the ceiling.

Jev’s Evaluation on cost.
Jev’s Evaluation on cost.
Jev’s Evaluation on time.
Jev’s Evaluation on time.

Per TechCrunch’s reporting, Bryo AI tested Jev against Gemini on classifying business emails and found Gemini slightly more accurate but up to 20 times more expensive. Vercel swapped a ChatGPT-based safety classifier for Jev and saw results five to 18 times faster, with better accuracy.

At these prices, the question stops being whether you can afford a decision and becomes which decisions you have been faking with rules.

TypeSafe’s own Doom demo makes it concrete: a bot querying the model 10 times a second costs about $7 an hour. A million decisions a day at 500 tokens each works out to about $21, below the line where anyone asks for a budget.

Cheaper coal, more engines. Cheaper judgment, more decisions.
Cheaper coal, more engines. Cheaper judgment, more decisions.

Expect more decisions

TypeSafe named the model after William Stanley Jevons, the economist who observed in 1865 that more efficient steam engines did not cut Britain’s coal consumption, they raised it, because cheaper power found jobs nobody had bothered to power before because of the cost. Yes, they named it after Jevons paradox.

TypeSafe expects intelligence to follow the same curve.

The gateway number is about developers, not end users. Teams reached for it to route, score, and verify, jobs they had been doing with regular expressions and hand-tuned rules.

For designers, the latency number matters more than the price. The Doherty Threshold comes from IBM research in 1982: once a system answers in under about 400 milliseconds, people stop waiting and start thinking with it.

The results:

  • End-to-end chat answers have never cleared it; the benchmark TypeSafe cites puts frontier models at three seconds to more than five minutes.
  • Jev clears the line on most calls, by the vendor’s own numbers.

For the first time, frontier judgment is fast enough to live inside an interaction instead of beside it.

That is the user-experience change. Judgment arrives at the speed of a keystroke, so the interface reacts mid-task instead of after the user stops and asks.

  • Rank suggestions by intent. Order search suggestions by what the query means, not what it starts with.
  • Validate before the error. Score a field as it is filled and nudge before submit; Luke Wroblewski measured in 2009 that inline validation made people faster and less error-prone.
  • Triage on arrival. Sort tickets, inbox items, and review queues by urgency the moment they land.
  • Check actions before they run. Hold a risky bulk edit or payment for one-tap confirmation; wave it through when routine.
  • Flag content before it posts. Catch a comment while it is being typed, not an hour after it went live.

When intelligence is slow and expensive, it gets its own screen. When it is fast and nearly free, it disappears into the controls you already have, and the designer gets the decision back.

Five places a decision already sits in your interface.
Five places a decision already sits in your interface.

How to apply it to UX

Here are five patterns. Each is a decision your product makes today with a rule, a default, or a human.

  • Sort the pile. Score each article for fit with my newsletter topics, then read from the top instead of the whole pile.
  • Route a request. Someone types “I need next Friday off.” A Choice question sends them to the leave form, not a page of search results.
  • Approve the easy ones. An expense under policy with a clean receipt goes straight through, a borderline one gets a one-tap confirm, and a messy one goes to a person.
  • Check the bot’s answer. Before a support reply shows, ask a yes-or-no question: does this answer what was asked? If not, hand off to a person.
  • Pick the empty state. A Choice question shows a brand-new account “Import your data” and a solo user “Invite your team.”

The ranking problems data science owned are priced at interaction patterns a designer can prototype in hours. I did.

One caution: Automation bias is the tendency to defer to a machine recommendation even when it is wrong, and a confident-looking percentage invites it.

Show confidence where it changes what the user should do, and hide it where it would only decorate.

Pick one decision

Jev will not be the only one of these.

Ronacher expects competitors now that the utility is obvious, community reproductions on open-weight models appeared within days, and TypeSafe promises more modalities soon.

The architecture is undisclosed and the price may be subsidized. None of that changes the design question, which is older than the company.

Classifiers have sat inside software since its inceptions; what kept them out of the interface was that good ones were expensive to build and slow to run, so the judgment lived in a batch job or a chat window and the designer worked around it.

That constraint just went away for a whole class of decisions, and a lot of current design patterns assume it still holds.

Choose one decision and run it.

Not a feature, not an assistant, just one place where your product currently guesses with a rule and could judge with a probability instead. Sketch the three states, set the gate, and measure the result against what you run now. If it wins, you have learned a new material. If it loses, you have learned where the old rule was right. Either way, you did not ship yet another chat box.

Action Items

  • Inventory the judgment calls your product fakes. List every place a rule, a default setting, or a support queue stands in for a decision: triage, sort order, what to show next. That is the backlog.
  • Design the three confidence states before touching the model. Sketch the confidence gate for one decision, including the copy for the ask state. The two thresholds are interface design.
  • Run one decision through a gateway for a week. Jev is on Vercel’s and other gateways without the waitlist. Pick your lowest-stakes routing decision, log the probabilities against what people did, and read the distribution before you trust a threshold.
  • Measure against the rule you run today. Vercel compared Jev with its previous classifier, not with nothing; your baseline is the regular expression, the dropdown, or the human. If the model does not beat it on speed and accuracy, keep the rule.

Jev Resources


Jev: How to use it for better, faster, and cheaper user decisions in UX was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.

Need help?

Don't hesitate to reach out to us regarding a project, custom development, or any general inquiries.
We're here to assist you.

Get in touch