SPIRITT logoSPIRITTTypeSafe AIJev

Jev gives software a decision, not another paragraph

TypeSafe AI's first System One model accepts text state plus typed questions, then returns constrained values, probabilities, and confidence in parallel. The launch promises LLM-level judgment for narrow decisions at a fraction of the latency and cost. Jev is accessible through SPIRITT AI Gateway.

Jev builders and evaluations across X

28 source posts: adaptive forms, real-time games, browser automation, model routing, app testing, Mac troubleshooting, safety monitors, and published evaluations. Builders describe their own setups; in SPIRITT, just ask your workspace to use Jev.

A model built to sit inside the if statement

Jev is deliberately narrower than an LLM. Its bet is that automation needs constrained judgment more often than it needs generated prose.

What TypeSafe launched, and where Jev fits

Jev is TypeSafe AI's first System One model. Developers send unstructured text state plus Choice, Score, or Noul questions whose output spaces are declared in advance. Jev returns typed answers, probability distributions, and confidence values rather than free-form text.

TypeSafe reports 193.6x faster and 444.6x cheaper results in a four-workflow evaluation, separate from its short recorded demo. Jev averages 67.8% agreement with Astra/Fable-generated reference answers there, while GPT-5.6 Sol scores 74.1%. The vendor calls the speed and cost ratios the high end of real-world gains, not universal guarantees.

Current documentation lists Jev 1.13 at a TypeSafe direct list price of $0.042 per million input tokens with free output. Context is 64K tokens across state and all questions, with a separate 32K limit for state plus the longest question. Input is text only and Choice supports up to 255 options. Typed output still can be wrong or manipulated by adversarial input; keep exact arithmetic and date comparisons in code.

Jev is accessible through SPIRITT AI Gateway. Open a SPIRITT Workspace and ask the agent to use Jev; the gateway handles model access while the workspace supplies the code, files, tools, tests, and surrounding workflow.

Typed decisionsParallel questions64K total / 32K per-question contextText onlyTypeSafe list: $0.042 / 1M inputAvailable in SPIRITT

Mean workflow agreement

GPT-5.6 Sol
74.1%
Claude Opus 5
73.1%
GPT-5.6 Terra
67.9%
Jev
67.8%
Claude Sonnet 5
67.8%
GPT-5.6 Luna
66.8%
DeepSeek V4 Pro
65.5%
DeepSeek V4 Flash
64.4%
Claude Haiku 4.5
53.6%

TypeSafe's four published workflows, equally weighted. Reference answers are GPT-6 Astra plus Fable 5.1 at high thinking, not verified human labels. Jev's reported mean is 67.8%, about $0.0004/case and 0.4 seconds/case (rounded).

Published Jev benchmarks, demos, and real-world evaluations

The full four-workflow vendor comparison, all nine N8 accuracy conditions, fx and classifier latency/accuracy, calibration, grading costs, and production routing. Each source is scoped separately; none is an overall leaderboard.

TypeSafe vendor demonstration

Recorded demo latency (ms)

Jev
114
GPT-5.6 Terra
8566

TypeSafe's short, deliberately simplified launch demo; Terra used default reasoning. About 75.1x faster here, distinct from the 193.6x four-workflow headline.

TypeSafe vendor demonstration

Cost per 1,000 demo runs

Jev
$0.08
GPT-5.6 Terra
$13.88

Scaled from $0.000081 and $0.013880 per recorded run, about 171.4x lower cost. These are direct-provider reference costs, not a SPIRITT billing quote.

TypeSafe vendor workflow evaluation

Security incidents agreement

Jev
61.7%
GPT-5.6 Terra
51.2%
GPT-5.6 Sol
62.5%
Claude Opus 5
66.2%

evals.typesafe.ai: all models run the same coded workflow at default reasoning. Reference labels are the average of Astra and Fable 5.1 at high thinking, not human-verified ground truth.

TypeSafe vendor workflow evaluation

Agent trace observability agreement

Jev
71.6%
GPT-5.6 Terra
73%
GPT-5.6 Sol
76.6%
Claude Opus 5
75.2%

evals.typesafe.ai: all models run the same coded workflow at default reasoning. Reference labels are the average of Astra and Fable 5.1 at high thinking, not human-verified ground truth.

TypeSafe vendor workflow evaluation

Invoice processing agreement

Jev
61.8%
GPT-5.6 Terra
74.7%
GPT-5.6 Sol
79.1%
Claude Opus 5
78.4%

evals.typesafe.ai: all models run the same coded workflow at default reasoning. Reference labels are the average of Astra and Fable 5.1 at high thinking, not human-verified ground truth.

TypeSafe vendor workflow evaluation

Customer service agreement

Jev
76%
GPT-5.6 Terra
72.7%
GPT-5.6 Sol
78.3%
Claude Opus 5
72.4%

evals.typesafe.ai: all models run the same coded workflow at default reasoning. Reference labels are the average of Astra and Fable 5.1 at high thinking, not human-verified ground truth.

TypeSafe vendor headline

Workflow speed headline (relative)

Jev
193.6
Compared workflow
1

The vendor attributes 193.6x to its separate four-workflow evaluation and calls these gains the higher end of real-world results. Not the 114 ms demo or a universal guarantee.

TypeSafe vendor headline

Workflow cost efficiency (relative)

Jev
444.6
Compared workflow
1

TypeSafe's separate 444.6x workflow-cost claim. Results depend on workflow and comparator; the short demo instead implies about 171.4x.

TypeSafe published reference price

Input list price per billion tokens

Jev
$42.00
Launch LLM floor
$200.00

TypeSafe lists $0.042 per million input tokens and free output. The comparator is only the $0.20/M lower bound in TypeSafe's generic launch table, not an equal-quality benchmark or SPIRITT price.

Independent: N8 Programs

MMLU

Jev
90.91%
Terra, no reasoning
87.89%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

GPQA Diamond

Jev
68.18%
Terra, no reasoning
56.06%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

ARC-Easy

Jev
99.33%
Terra, no reasoning
98.82%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

ARC-Challenge

Jev
97.61%
Terra, no reasoning
96.76%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

WinoGrande

Jev
91.63%
Terra, no reasoning
78.69%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

HellaSwag

Jev
94.86%
Terra, no reasoning
95.32%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

GSM8K: four choices

Jev
79.68%
Terra, no reasoning
87.72%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

GSM8K: ten choices

Jev
54.59%
Terra, no reasoning
69.83%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

Chess: four legal moves

Jev
49.45%
Terra, no reasoning
50.25%

N8 Programs, September 16 numeric result table. Terra reasoning=none and letter-only output; explicit a/b/c/d instruction for chess. This is an external experiment, not an official TypeSafe benchmark or a full-reasoning comparison.

Independent: N8 Programs

Expected calibration error (pp)

Lowest benchmark
0.3
Reported average
1.7
Highest benchmark
7.0

N8 reports 1.74 percentage-point expected calibration error across the benchmark suite, with a 0.26 to 6.96 range. The reliability chart covers 33,735 predictions and uses ten equal-width bins. Lower is better; individual answers may still be wrong.

Independent: Vercel fx

fx safety accuracy

Jev
98.6%
GPT-5.6 Luna
96.7%

Vercel fx safety evaluation: 70 labeled cases, three runs each (210 decisions per classifier, not 210 independent cases). Original Pranit chart, September 16. Correct decisions: 207/210 versus 203/210. Jev had zero fallbacks and identical answers on all three runs of each case. Not a general safety guarantee.

Independent: Vercel fx

fx median latency (ms)

Jev
312
GPT-5.6 Luna
1458

Vercel fx safety evaluation: 70 labeled cases, three runs each (210 decisions per classifier, not 210 independent cases). Original Pranit chart, September 16. About 4.7x faster at the median.

Independent: Vercel fx

fx p95 latency (ms)

Jev
374
GPT-5.6 Luna
6583

Vercel fx safety evaluation: 70 labeled cases, three runs each (210 decisions per classifier, not 210 independent cases). Original Pranit chart, September 16. About 17.6x faster at p95, not a 5 to 18x p95 range.

Independent: Malte Ubl

Classifier accuracy

Jev
100%
Gemini 2.5 Flash Lite
98.2%

Malte Ubl classifier evaluation: 167 prompts run three times (501 decisions each). Original screenshot says not deployed yet. 501/501 versus 492/501 correct; saturation of this evaluation does not prove universal accuracy.

Independent: Malte Ubl

Classifier median latency (ms)

Jev
111
Gemini 2.5 Flash Lite
703

Malte Ubl classifier evaluation: 167 prompts run three times (501 decisions each). Original screenshot says not deployed yet. About 6.3x faster at the median.

Independent: Malte Ubl

Classifier p95 latency (ms)

Jev
217
Gemini 2.5 Flash Lite
1074

Malte Ubl classifier evaluation: 167 prompts run three times (501 decisions each). Original screenshot says not deployed yet. About 4.9x faster at p95.

Independent: Good Start Labs

Estimated grading cost / million answers

Jev
$160.00
Fable 5.1
$33000.00

Good Start's published study: 6,003 rubric checks across 1,203 financial-research answers, 91.5% agreement with Fable 5.1. Estimated costs use direct-provider rates. Agreement is not accuracy.

Builder report: 25 evaluated tasks

Firstmate routing cost, indexed

With Jev
29
Previous workflow
100

Kun Chen reports 71% lower end-to-end routing cost. The same 25 decisions matched Fable, which is agreement, not independently proven correctness.

Builder report: 25 evaluated tasks

Firstmate routing time, indexed

With Jev
10
Previous workflow
100

Kun Chen reports 90% lower end-to-end wall time and roughly 200 ms for the Jev decision. Indexed to 100 for the prior system; includes the surrounding agent invocation.

Independent: Good Start Labs

Jev agreement with other graders

With Fable 5.1
91.5%
With Astra
90.6%
With Gemini 3.8
90.6%
With DeepSeek V4.1
91.5%
With GPT-5.6 Luna
86.5%

6,003 rubric checks across 1,203 financial-research answers. Share of matching verdicts, not accuracy. Separate initial study: 96.0% agreement with Sonnet 4.5 on 741 checks/112 answers. A separate 10,500-call game-task run had zero unusable gradings, not zero wrong verdicts.

Official sources: TypeSafe AI's launch, homepage, workflow evaluation notes, System One documentation, and current model reference. Independent signals include the Vercel fx safety evaluation, Malte Ubl's classifier evaluation, Firstmate's 25-task production-routing report, Good Start Labs' grading sample, and N8 Programs' MMLU/GPQA comparison. These use different harnesses and are not one comparable leaderboard. Reproduce calibration, latency, and cost on your own decisions before autonomous production use.

How It Works

Ask the workspace agent to use Jev through SPIRITT AI Gateway

01

Open a workspace

Open a workspace and land in a fully equipped cloud computer: browser, files, terminal, integrations, and memory. No local setup. No thin chat box pretending to be an agent.

Open a SPIRITT workspace for a Jev integration
02

Just Ask to use Jev

Tell your workspace agent to use Jev for the decision, classifier, score, or routing step you need. SPIRITT AI Gateway handles model access, so you can stay focused on the outcome instead of wiring a provider API.

A bright isometric Jev decision workflow inside SPIRITT AI Gateway
03

Build or automate

Use the workspace codebase, tools, browser, tests, and durable context to wrap narrow Jev decisions inside a larger workflow. Keep deterministic code in control and route low-confidence cases to a person or reasoning model.

Build an agent workflow around Jev in SPIRITT

Build the workflow around the decision.

Ask your SPIRITT Workspace to use Jev, then give the agent the code, files, tools, tests, and constraints needed to turn that decision into a working system.

Questions

Jev FAQ

01What is Jev?+
TypeSafe AI's first System One model. It evaluates text state against developer-defined Choice, Score, and Noul questions, returning typed answers, probabilities, and confidence rather than generated prose.
02Can Jev chat, write code, or explain its reasoning?+
No. TypeSafe explicitly says System One models do not write replies, produce code, or generate reasoning explanations. Jev is for narrow decisions inside software; pair it with deterministic code or a generative model for broader work.
03How fast and inexpensive is Jev?+
TypeSafe's recorded demo took 114 ms versus 8,566 ms for GPT-5.6 Terra with default reasoning, about 75.1x faster. Separately, TypeSafe reports 193.6x faster and 444.6x cheaper in its four-workflow evaluation. These are vendor-controlled results, not universal guarantees. Its $0.042/M input price is a direct-provider reference price, not a SPIRITT billing quote.
04What are Jev's current technical limits?+
Jev accepts text only. The 64K budget covers state plus all questions; state plus the longest individual question is limited to 32K. Choice supports up to 255 options. Calibration is measured across predictions, not a guarantee for one answer. A typed decision can be incorrect or affected by prompt injection, so consequential actions still need checks and approval.
05Where can developers access Jev?+
Jev is accessible in SPIRITT through SPIRITT AI Gateway. TypeSafe also offers its own SDKs and HTTP endpoint, while other gateways have announced distribution. Check current limits and terms before production use.
06How do I use Jev in SPIRITT?+
Open a SPIRITT Workspace and ask the agent to use Jev. Jev is accessible through SPIRITT AI Gateway, so the workspace can route the relevant decision or classification step without you wiring provider credentials.
Buy from builders who use what they sellBuilt usingSPIRITT