What's Jev?
Background
Jev is the first model from TypeSafe AI, a US start-up whose cofounder Diogo Almeida co-invented RLHF, the training method behind ChatGPT. It opened to the public on 15 September 2026, and demand was high enough that new sign-ups were paused a few days later.
You give Jev some text and a few questions. Each question is a yes/no, a choice between options you name, or a score on a scale you define. Jev returns an answer and a probability for each question. It never writes a sentence.
How it differs from a chat model
What Jev gives up is the writing: it cannot explain an answer, and it cannot think a problem through before it answers.
How it differs from a BERT classifier
Why BERT is overconfident
Every training example says "100% this class", even the ambiguous ones. Training keeps rewarding a push toward 100% and never rewards an honest 70%.
Why Jev is not
It is trained against the confidence itself. Saying 90% and being right 70% of the time is penalized; saying 70% and being right 70% is full marks.
What's special
- Confidence you can code against. If "90% sure" really means right nine times in ten, you can write
if confidence > 0.9: act, else: ask someone. - Price and speed. $0.042 per million input tokens, output free, and about half a second per request.
- No training data. The task is described in plain English at request time.
- Always a valid answer. It can only return one of the options you defined.
Experiment in Prover context
Test subjects
Picks an answer, no thinking step
- Time per case
- 0.5 s
- 150 calls
- $0.06
Thinks first, then answers. What Prover Labs uses today
- Time per case
- 27 s
- 150 calls
- $9.78
Claude's small, fast model, thinking off
- Time per case
- 9 s
- 150 calls
- $1.28
The tasks are two apps on Prover Labs: the Formal-Req Description Checker and the Req Quality Checker. Each model ran every test case three times. Time and cost were measured on the 50 real cases in scenario 1.
Scenario 1: Does the formal requirement match its description?
App: Formal-Req Description Checker
- 50 real cases. 18 carry an error an engineer actually made and later fixed; 32 are correct. They come from Prover's GSS training package and the Shift2Rail level-crossing project, with every name rewritten.
- 10 made-up cases. The app's own test set, each with one obvious edit.
Real cases: right verdict
Sonnet winsReal errors caught
Sonnet winsFalse alarms, lower is better
Jev winsMade-up cases: right verdict
Jev winsCatching real errors is what this checker is for: a missed error ends up in the safety case, a false alarm only costs an engineer a second look. On that measure Sonnet caught nine in ten and Jev four in ten. Jev was perfect on the made-up errors, which are one visible edit each. Real errors hide in structure, like a bracket that turns "some gate" into "every gate".
Scenario 2: Is the requirement well written?
App: Req Quality Checker
- 12 cases from the app's own test set, each judged on five criteria: unambiguous, verifiable, understandable, singular, complete.
Criterion verdicts right
Sonnet winsRequirements fully right
Sonnet winsWarn or fail? Jev saw the problem but picked the milder verdict. Where that line falls is Prover's own rubric.
Flaws in clean text. It flagged a well-formed requirement. Sonnet did the same on two others.
Something missing. "The route shall be released" never says by whom. Jev passed it.
Jev came last here, even behind Haiku. Judging quality means applying a severity scale and noticing what is absent, which takes more than one look.
Pros & cons
Jev reads like a large language model and knows as much: on an exam-style knowledge test it scores 80%, level with Claude Opus 5. What it lacks is the step in between. Sonnet thought for about 2,000 words per real case before answering.
Good at
- Judgements you can make at a glance
- Comparing two texts when the difference is visible
- Sorting and routing at high volume, for almost nothing
- Knowing when it is sure
Not good at
- Anything that needs steps: tracing a quantifier's scope, applying a rubric
- Noticing what is missing
- Explaining itself: it cannot write a reason or a fix
- Long inputs: 64k tokens per request
How can we use it in Prover tasks?
Trust it when it is sure
Suggested tasks
Sort a customer's requirement documents
Hundreds of statements per specification, sorted before formalization starts.
Find the object type a requirement is about
Pulls up similar requirements that are already formalized. About 1,000 labeled examples to test on.
Check that a knowledge article answers the question
Drops the retrieved articles that miss the point. 854 question–article pairs already judged by five people.
Tag project documents on upload
Tags from the fixed list, plus a flag when a newer version exists.
The percentages are illustrations. None of these tasks has been tested yet, and any customer text needs approval before it goes to a US service.