A scattered cloud of grey and orange dots labelled a choice and a probability funnels into a structured grid where one clean orange route is compiled to a single governed output.
A model returns a choice and a probability. A semantic layer compiles one governed path to the answer.

Jev AI vs LLMs: A Model That Cannot Hallucinate Still Needs a Semantic Layer

A former OpenAI researcher raised $40 million for an AI model that never writes a sentence. Jev returns a choice and a probability. It still cannot tell you what "net revenue" means at your company.

Here is the short version. Three kinds of system, and what each one does with a business question.

Question LLM (GPT, Claude, Gemini) Jev (System One model) Semantic execution layer (Colrows)
What does it output? Free text, code, SQL A choice, a score or a probability, from options you supply Governed SQL, compiled from business definitions
How does it fail? A plausible but wrong answer A valid but wrong option An explicit error when it cannot resolve the question
Who defines "net revenue"? The model guesses from table and column names Your code must decide. Jev is not built for it A versioned business definition
Math and fiscal dates? Written as SQL. Logic can differ from run to run Its maker says to do math and date logic in code Computed in code. Same input, same result

The rest of this post gives the evidence behind each cell. We cite Jev's own documentation where we can. It makes the argument for us.

What Jev does, in plain words

TypeSafe AI released Jev on 15 September 2026. The company raised a $40 million seed round led by DCVC. Its founder, Diogo Almeida, is a former OpenAI researcher. TypeSafe calls Jev a "System One" model. It is built for fast judgments, not for long reasoning.

An LLM such as GPT, Claude or Gemini writes an answer one word at a time. It can write anything. Jev works like a multiple-choice grader. You send it a block of text (TypeSafe calls this the "state") and a list of typed questions. You supply the options. Jev scores every option and returns structured data with a probability. It answers all questions in one pass.

Picture two exam takers. One writes an essay and may state a false fact. The other fills in bubbles on a sheet you printed. The second one cannot write something outside the sheet. It can still fill in the wrong bubble.

How Jev is trained, and what is still unknown

TypeSafe trains Jev with a method it calls RLCD, short for reinforcement learning for calibrated decisions. The goal is honest probabilities. When Jev says 70%, it should be right about 70% of the time. Chat models are tuned mainly on human preference. That rewards answers people like.

Almeida told TechCrunch that the training data is synthetic. TypeSafe has not published a research paper, a parameter count or a model card. TechCrunch reports that outside observers suspect an open-weight LLM underneath. TypeSafe has not confirmed that. Treat the architecture as undisclosed.

What the evidence says about Jev so far

Early evidence is thin. Some of it comes from TypeSafe itself. Read each number with its source in mind.

  • Vendor claim, speed and cost. TypeSafe reports up to 193.6x faster and 444.6x cheaper than frontier LLMs. TypeSafe designed those tests. It calls the figures the high end of real-world gains.
  • Independent test, speed and cost. Every, a media company, ran its own test. Jev was about 25 times faster and about 580 times cheaper. Jev caught 6 of 7 planted defects. The LLM caught all 7.
  • Vendor claim, quality. On TypeSafe's workflow tests, Jev agreed with reference labels 67.8% of the time. The labels come from two frontier models. The best LLM tested, GPT-5.6 Sol, agreed 74.1% of the time. This measures agreement with other models, not truth.
  • Weakest result. On invoice processing, the most finance-like task, Jev agreed 61.8% of the time. GPT-5.6 Sol agreed 79.1% of the time.

Jev is fast and cheap. That is real. For a job such as routing a support ticket, a small quality gap can be a good trade. That is a different job from answering a finance question.

What "cannot hallucinate" really means

Jev cannot return an answer outside the options you gave it. That removes one class of error: broken formats and invented fields. It does not make the chosen option correct. A valid answer can still be wrong.

Calibration is the stronger selling point. If Jev returns 52%, your code can treat the result as a coin toss and send it to a person. Armin Ronacher, CTO of Earendil, made this point to TechCrunch. Nikhil Mudholkar, CTO of Bryo AI, said Jev is "the only one that hands back a real probability."

That only works if the probabilities hold on your data. At the time of writing, we found no public calibration curves and no independent calibration test. Ask for one before you automate a decision with it.

Read Jev's limits page: it describes a semantic layer

TypeSafe publishes a page of known weak spots for Jev 1.13. It was last reviewed on 17 September 2026. The page is honest. It reads like a requirements list for the layer that sits next to the model.

  • Math. "Jev is not a calculator. We strongly recommend implementing any mathematical logic in code."
  • Dates. Asking which of two dates comes first, how far apart they are, or whether one falls inside a window "is unreliable." The page names quarters, settlement windows and accrual periods.
  • Noisy input. "Accuracy falls as the state grows with content unrelated to the decision."
  • Indirection. A question about a property of a property "costs accuracy."
  • Writing. Jev "is not trained to generate text."

Now take a plain board question: "What was net revenue last quarter by region, excluding refunds?" The answer needs a fiscal calendar, a refund rule, a join between orders and regions, and a sum. Jev's own page tells you to do the date logic and the sum in code. It also tells you to filter the data before Jev sees it.

TypeSafe built a fast judgment tool and documents its limits openly. The code that does the math and holds the definitions still has to come from somewhere.

LLMs reach the same wall from the other side

Larger LLMs do not remove the gap. The data shows it.

  • Spider 2.0. This benchmark uses 632 tasks from real enterprise databases. An o1-preview code agent solved 21.3% of them. The same kind of agent scored 91.2% on the older Spider 1.0 and 73.0% on BIRD (arXiv 2411.07763).
  • Business context. In a data.world benchmark on an insurance schema, GPT-4 answered 16% of questions correctly when it queried SQL directly. It answered 54% when the same questions ran over a knowledge graph that carried the business meaning (arXiv 2311.07509). That test used a 2023 model.
  • Repeatability. Thinking Machines ran one prompt 1,000 times at temperature 0 on Qwen3-235B. It got 80 different outputs. The cause is server load, which changes how the math runs. Ask the same question twice and you may get a different query.

We go deeper on that enterprise accuracy drop in why text-to-SQL accuracy drops on real schemas.

What fixes it: put meaning in a layer, then compile

dbt Labs published a 2026 benchmark that compares text-to-SQL with a semantic layer. It used 11 questions, run 20 times each. dbt Labs sells a semantic layer, so read it as vendor research. The method is public and the result is candid.

  • On questions the semantic layer covered, accuracy was 98.2% (Sonnet 4.6) and 100% (GPT-5.3 Codex).
  • Text-to-SQL on the same modeled project scored 90.0% and 84.1%.
  • On questions the semantic layer could not cover, it scored 0.0%. It returned an error and no number.

dbt Labs states the lesson well: with text-to-SQL, "failure looks like a plausible but incorrect answer." With the semantic layer, "failure looks like an error message." A board pack can survive an error message. It cannot survive a silent revenue error.

The 0.0% shows a second problem. A semantic layer that people model by hand answers only what someone modeled. Every new table, renamed column or changed definition adds work. Agents ask new questions all day. The backlog grows faster than a team can clear it.

Gartner points the same way. In May 2026 it predicted that organizations that prioritize semantics in AI-ready data will raise agentic AI accuracy by up to 80% and cut costs by up to 60% by 2027. That is a forecast, not a measurement. The direction is clear.

Where Colrows and Jev each fit

A three-stage pipeline: an Interpreter (LLM or Jev) reads the question and scores the intent; the Semantic Execution Layer (Colrows) resolves business meaning, applies access policy and compiles governed SQL; the Data Warehouse runs governed SQL only. A callout notes the layer returns an explicit error instead of guessing.
Change the model on the left. The meaning in the middle stays the same.

Use each part for what it does well. A language model, whether an LLM or a model like Jev, reads the question and scores the intent. A fast classifier could route a request in well under a second. The semantic execution layer then resolves what the request means in your business and compiles it into SQL.

This is an architecture pattern. Colrows has not announced a Jev integration. The point is simple: whichever model you place on the left, the middle layer decides what the numbers mean.

Colrows is an autonomous semantic layer for enterprise AI. It builds and updates its model by reading your databases, Confluence pages, data catalogs and user feedback. People do not hand-maintain it. For each request, it compiles governed SQL. Role, row and column rules apply before the query runs. When it cannot resolve a join or a definition, it returns an error instead of a guess. The semantic layer buyer's guide covers what to check before you put one between your models and your warehouse.

We have not published a benchmark that matches dbt Labs' 11-question test. We will not quote one that does not exist. Run the test on your own schema and judge the result.

Replace an LLM with Jev and you change the model. The context problem stays. Business definitions, join paths and access rules still need one governed home. Fix the context, not the model.

Our earlier post on Gemini 4 Argon and your warehouse makes the same case for a frontier LLM. Jev is a different kind of model. The answer is the same.

What this means for your next AI decision

Model types will keep changing. LLMs, System One models like Jev and others will each win some jobs. Each one still needs the same things from your data estate: one definition of each metric, one set of access rules, and the same answer on every run.

Build that layer once. Then swap models as the market moves, and check each swap against the same governed definitions.

Frequently asked questions

What is Jev AI, and how is it different from an LLM?

Jev is a System One model from TypeSafe AI, released on 15 September 2026. An LLM writes an answer one word at a time and can produce any text. Jev works like a multiple-choice grader: you send it a block of text and a list of typed questions with the options you supply, and it returns a chosen option and a probability. It does not generate free text.

Can Jev AI hallucinate?

Jev cannot return an answer outside the options you give it, so it cannot invent fields or break the output format the way a chat model can. That removes one class of error. It does not make the chosen option correct. A valid answer can still be the wrong answer.

Does Jev AI need a semantic layer?

Yes, for data questions. TypeSafe's own limits page for Jev recommends doing math and date logic in code and filtering the input before Jev sees it. A plain question such as net revenue last quarter by region needs a fiscal calendar, a refund rule, a join and a sum. The definitions and the math have to live somewhere. That is what a semantic layer holds.

Is Jev accurate enough for finance questions?

On TypeSafe's own invoice-processing test, the most finance-like task, Jev agreed with the reference labels 61.8% of the time, against 79.1% for GPT-5.6 Sol. Jev is fast and cheap, which is a good trade for jobs like routing a support ticket. A board-level finance number is a different job, and no public calibration test on enterprise data exists yet.

How do Colrows and Jev work together?

As a pattern, not a product integration. A language model, whether an LLM or a model like Jev, reads the question and scores the intent. The semantic execution layer then resolves what the request means in your business and compiles governed SQL, with role, row and column rules applied before the query runs. Whichever model sits on the left, the middle layer decides what the numbers mean.

Sources

  1. TypeSafe AI, "Introducing System One Models and Jev," and model documentation.
  2. TypeSafe AI, "Jev 1.13 jagged edges," last reviewed 17 September 2026.
  3. TypeSafe AI, "Workflow evals." Vendor-designed tests.
  4. TechCrunch, "A new kind of AI model from a ChatGPT inventor is thrilling developers," 18 September 2026.
  5. Business Wire, "TypeSafe AI emerges from stealth with $40M."
  6. Every, "Mini-Vibe Check: TypeSafe's Jev," 15 September 2026.
  7. Spider 2.0, arXiv 2411.07763.
  8. Sequeda, Allemang and Jacob (data.world), arXiv 2311.07509.
  9. Thinking Machines, "Defeating Nondeterminism in LLM Inference."
  10. dbt Labs, "Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update," 7 April 2026. Vendor research.
  11. Gartner, "Lack of Semantics Causes Inaccurate AI Agents and Wasted Spending," 11 May 2026.

Ready to stop guessing?

Book a technical architecture review. Bring one schema and one hard question. See the governed SQL, the policy checks, and what the layer does when it cannot answer.