
Jev Explained: The AI Model That Never Writes a Word
Yoni Fraimorice
Most AI models today are built to talk. You ask a question, and they write an answer.
Jev, the first model from San Francisco startup TypeSafe AI, never writes a word. You give it some text and a list of questions. It returns typed answers, such as a label, a score, or a probability, that your code can use right away.
TypeSafe released Jev in early access on September 15, together with a $40 million seed round led by DCVC. The company was co-founded by Diogo Almeida, who worked on RLHF, InstructGPT, and ChatGPT during about four years at OpenAI. His explanation for starting the project, as he told TechCrunch: "We have lightning in a bottle, and yet it is not useful."
This guide explains what makes Jev different, why it costs so little, where it helps, and how to start using it.
The problem Jev is trying to solve
Much of the AI in production software is not a chat. It is a quick decision inside a program:
- Which team should handle this support ticket?
- Is this message spam?
- Is this shell command safe for an agent to run?
- Does this document answer the user's question?
Today, developers often send these questions to a large language model and write "return JSON" in the prompt. The model writes text, token by token. Then your code parses that text, checks the shape, and retries when the model returns something unexpected.
That works, but it is slow, expensive, and fragile. You pay for a writer when you only needed a judge.
What makes Jev different
A Jev request has two parts:
- State – the input text. It can be a string, a JSON object, or an array of text values.
- Questions – one or more typed questions about that state.
There are only three question types, which TypeSafe calls primitives:
| Primitive | Use it for | Example answer |
|---|---|---|
| Choice | Pick one option from a fixed list | "technical", plus a probability for every option |
| Score | Rate something on ordered levels | 1.0 on a scale of "calm" to "very angry" |
| Noul | A yes/no statement | 0.95, the probability that the answer is yes |
Because you define every possible answer in advance, Jev cannot return a value outside your schema. No broken JSON, no invented category, no retry loop.
TypeSafe calls Jev a "System One" model, after Daniel Kahneman's idea of fast, intuitive thinking. An LLM that reasons step by step is closer to System Two: slow and careful.
Trained for honest probabilities, not for pleasing people
The other big difference is training. Chatbots are tuned with RLHF, which rewards answers that people prefer. TypeSafe argues that this can reward confident-sounding mistakes.
Jev uses a method TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions). The goal is calibration: across many answers, things Jev rates at 80% should be true about 80% of the time. That turns the probability into something your code can trust and act on.
TypeSafe has not published Jev's architecture, weights, or a technical paper. It says Jev is transformer-based and trained on synthetic data. Some outside observers think it may be built on top of an open-weight LLM.
Why is Jev so cheap?
The price is $0.042 per million input tokens, and output tokens are free. For comparison, reading a 10,000-token document costs about $0.0004.
Three design choices explain most of the saving:
1. No text generation. In an LLM, writing the answer is usually the slow part. Each new token needs another pass through the model, one after the other. Jev returns a handful of numbers instead of hundreds of words, and it does not write any "reasoning" text first.
2. Read once, ask many. Jev reads the state once and evaluates all questions against it in parallel. In TypeSafe's own test with 13 questions about a long article, one batched call was 12.2× cheaper and 10× faster than 13 separate calls, with the same answers.
3. A narrow job. Jev only has to choose between answers you already wrote. It does not need to write code, poems, or long explanations. TypeSafe hasn't said how big the model is, but a narrow task generally allows a smaller, cheaper model.
TypeSafe reports response times of 70–500 ms and claims Jev is 40–200× faster and 40–400× cheaper than frontier LLMs on comparable tasks. Treat these as marketing numbers. TypeSafe itself notes that its own team built the test workflows, that bias is possible, and that the results are probably at the high end. Always benchmark on your own data.
How to use Jev
You can try it in the Playground without code. For real use, install the Python SDK (pip install typesafe-sdk). This example, adapted from the quick start, routes a support ticket:
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient() # reads TYPESAFE_API_KEY
ticket = "My Stripe integration has failed for 3 days. I'm losing sales. Please help ASAP."
result = client.system_one(
state=ticket,
questions={
"department": Choice(
instructions="Which team should handle this",
criteria={
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions",
},
),
"frustration": Score(
instructions="How frustrated the customer appears",
criteria=["Calm", "Frustrated but civil", "Very angry"],
),
"is_urgent": Noul(instructions="The message conveys urgency"),
},
)
dept = result.answers["department"]
print(dept.choice, dept.confidence) # e.g. "technical" 0.78Every Choice and Score answer includes a confidence value from 0 to 1. That is the most useful part of the design: the answer tells you what, and confidence tells you whether to act.
if dept.confidence >= 0.9:
assign_ticket(dept.choice)
elif dept.confidence >= 0.6:
assign_ticket(dept.choice, needs_review=True)
else:
send_to_human(ticket)Pick thresholds based on risk. Tagging a ticket can use a low bar. Deleting data or refunding money needs a much higher one, or a person.
Where Jev helps, and where it doesn't
Good fits:
- Ticket routing, content moderation, and spam detection
- Guardrails that screen messages going into and out of an LLM
- Choosing which tool or skill an AI agent should use
- Checking whether a citation actually supports a claim
- Filtering RAG passages before they reach a more expensive model
A common pattern is a cascade: Jev handles the easy, high-volume cases, and only uncertain ones go to a large LLM or a human.
Jev is not a replacement for the LLM in your chatbot or coding agent. TypeSafe's known-limitations page is refreshingly honest about other gaps:
- It reads literally. It answers the question you wrote, not the one you meant.
- It is weak at math, counting, and date comparison. Do those in code.
- Probabilities don't always add up. A Noul and its opposite can sum to more than 1.
- Text only, English first. Other languages work, but less well. The limit is 64k tokens per request.
The security angle: "can't hallucinate" is not "can't be fooled"
A fixed output schema removes a whole class of bugs. It does not make the answer correct.
TypeSafe states that Jev does not treat the state as hostile by default. Text written to steer the model, such as an injected instruction, can move the answer. So if you use Jev as a guardrail for user or web content, an attacker may try to write content that argues for its own "safe" label.
Practical rules:
- Treat Jev's output as a signal, not a permission. Keep real access checks in code.
- Use higher thresholds for risky actions, and send low-confidence cases to a person.
- Test with adversarial examples before you ship.
- Pin a model version, such as
jev-1.13.0. Thejev-latestalias can change, and your tuned thresholds may shift with it.
The bottom line
Jev is a bet that most future AI will talk to software, not to people. Instead of making one model do everything, it makes one kind of decision fast, cheap, and with an honest measure of uncertainty.
The model is named after economist William Stanley Jevons. The Jevons paradox says that when a resource becomes cheaper to use, people often use much more of it. If AI decisions become almost free, we may soon see them in every function call, and good confidence thresholds will matter more than ever.
Sources
- Wikipedia: Jev (AI model)
- TypeSafe docs: System One
- TypeSafe docs: Quick start
- TypeSafe docs: Models and pricing
- TypeSafe docs: AI primer and RLCD
- TypeSafe docs: Confidence
- TypeSafe docs: Parallel questions cookbook
- TypeSafe docs: Jev 1.13 known limitations
- TechCrunch: A new kind of AI model from a ChatGPT inventor
Diagrams created for this article. Hero image: New Romney signal box lever frame, photo by John K Thorne, public domain.