← Reports

Jev: ~50× faster, over 100× cheaper. But is it any good?

A new kind of AI model went viral this month. We put it on two of our own tasks against Claude, using real formalization errors from Prover projects.

What's Jev?

Background

Jev is the first model from TypeSafe AI, a US start-up whose cofounder Diogo Almeida co-invented RLHF, the training method behind ChatGPT. It opened to the public on 15 September 2026, and demand was high enough that new sign-ups were paused a few days later.

You give Jev some text and a few questions. Each question is a yes/no, a choice between options you name, or a score on a scale you define. Jev returns an answer and a probability for each question. It never writes a sentence.

How it differs from a chat model

Chat model (Claude) requirement+ question Claude The arrow is reversed, so the verdict is inconsistent. written word by word, then your code reads the label out of the text Jev requirement+ 3 questions Jev verdict → inconsistent92% contradiction? → yes63% condition missing? → no71% all answers in one pass, about 0.5 s
Claude writes, and a label has to be read back out of its text. Jev answers every question at once, each with a probability. Values are an example.

What Jev gives up is the writing: it cannot explain an answer, and it cannot think a problem through before it answers.

How it differs from a BERT classifier

BERT + classification head "The IXL shalllock the route…" BERT 768 numbers 1234 5? fixed table, unnamed columns 0.91 scores forced to add up to 1 new class → relabel + retrain Jev "The IXL shalllock the route…" requirement: an obligation design: describes a solution info: background + definition: defines a term options written in plain English Jev 85% new option answered at once
BERT's classes are unnamed columns in a fixed table, learned from labeled examples. Jev reads the options as text, so a new option works in the next request. Values are an example.

Why BERT is overconfident

Every training example says "100% this class", even the ambiguous ones. Training keeps rewarding a push toward 100% and never rewards an honest 70%.

Why Jev is not

It is trained against the confidence itself. Saying 90% and being right 70% of the time is penalized; saying 70% and being right 70% is full marks.

What's special

  • Confidence you can code against. If "90% sure" really means right nine times in ten, you can write if confidence > 0.9: act, else: ask someone.
  • Price and speed. $0.042 per million input tokens, output free, and about half a second per request.
  • No training data. The task is described in plain English at request time.
  • Always a valid answer. It can only return one of the options you defined.

Experiment in Prover context

Test subjects

Jev

Picks an answer, no thinking step

Time per case
0.5 s
150 calls
$0.06
Claude Sonnet 5

Thinks first, then answers. What Prover Labs uses today

Time per case
27 s
150 calls
$9.78
Claude Haiku 4.5

Claude's small, fast model, thinking off

Time per case
9 s
150 calls
$1.28

The tasks are two apps on Prover Labs: the Formal-Req Description Checker and the Req Quality Checker. Each model ran every test case three times. Time and cost were measured on the 50 real cases in scenario 1.

Scenario 1: Does the formal requirement match its description?

App: Formal-Req Description Checker

  • 50 real cases. 18 carry an error an engineer actually made and later fixed; 32 are correct. They come from Prover's GSS training package and the Shift2Rail level-crossing project, with every name rewritten.
  • 10 made-up cases. The app's own test set, each with one obvious edit.
JevSonnet 5Haiku 4.5

Real cases: right verdict

Sonnet wins
Sonnet
83%
Jev
70%
Haiku
51%
Most important

Real errors caught

Sonnet wins
Sonnet
89%
Jev
39%
Haiku
83%

False alarms, lower is better

Jev wins
Sonnet
15%
Jev
6%
Haiku
47%

Made-up cases: right verdict

Jev wins
Sonnet
87%
Jev
100%
Haiku
90%

Catching real errors is what this checker is for: a missed error ends up in the safety case, a false alarm only costs an engineer a second look. On that measure Sonnet caught nine in ten and Jev four in ten. Jev was perfect on the made-up errors, which are one visible edit each. Real errors hide in structure, like a bracket that turns "some gate" into "every gate".

Scenario 2: Is the requirement well written?

App: Req Quality Checker

  • 12 cases from the app's own test set, each judged on five criteria: unambiguous, verifiable, understandable, singular, complete.

Criterion verdicts right

Sonnet wins
Sonnet
83%
Jev
72%
Haiku
78%

Requirements fully right

Sonnet wins
Sonnet
78%
Jev
58%
Haiku
64%
9

Warn or fail? Jev saw the problem but picked the milder verdict. Where that line falls is Prover's own rubric.

9

Flaws in clean text. It flagged a well-formed requirement. Sonnet did the same on two others.

6

Something missing. "The route shall be released" never says by whom. Jev passed it.

Jev came last here, even behind Haiku. Judging quality means applying a severity scale and noticing what is absent, which takes more than one look.

Pros & cons

Sonnet Reads the input→ Thinks it through in writing→ Writes an answer
Jev Reads the input→ no thinking step→ Picks an option, with a probability

Jev reads like a large language model and knows as much: on an exam-style knowledge test it scores 80%, level with Claude Opus 5. What it lacks is the step in between. Sonnet thought for about 2,000 words per real case before answering.

Good at

  • Judgements you can make at a glance
  • Comparing two texts when the difference is visible
  • Sorting and routing at high volume, for almost nothing
  • Knowing when it is sure

Not good at

  • Anything that needs steps: tracing a quantifier's scope, applying a rubric
  • Noticing what is missing
  • Explaining itself: it cannot write a reason or a fix
  • Long inputs: 64k tokens per request

How can we use it in Prover tasks?

Trust it when it is sure

Share of Jev's answers that were right, keeping only answers at or above each confidence level. The surer it is, the more often it is right. On one-glance questions its surest answers were all right; on real formalization errors 6 of its 45 surest answers were still wrong.
each item Jev confidence≥ 0.9? yes no act on it automatically send to Sonnet or an engineer
The pattern: act on sure answers, pass on the rest. Use it where a wrong answer is cheap, not for safety checks.

Suggested tasks

Sort a customer's requirement documents

"The interlocking shall lock every point in the route before…" → requirement 94%constraintdesigninfo

Hundreds of statements per specification, sorted before formalization starts.

Find the object type a requirement is about

"A signal may show proceed only if its overlap is free." → signal 88%routepointtrack section…60 types

Pulls up similar requirements that are already formalized. About 1,000 labeled examples to test on.

Check that a knowledge article answers the question

"How do I export SEIF files in batch?" + one retrieved article → answers it 91%does not

Drops the retrieved articles that miss the point. 854 question–article pairs already judged by five people.

Tag project documents on upload

"Interlocking test specification, version 3.2" → test specificationinterlockingsuperseded? no

Tags from the fixed list, plus a flag when a newer version exists.

The percentages are illustrations. None of these tasks has been tested yet, and any customer text needs approval before it goes to a US service.

Test cases rewritten from non-customer sources; originals never left Prover. Jev reached via Vercel AI Gateway (US).