Skip to content
Trending storyDeveloping

TypeSafe launches Jev, a structured-decision model 100x faster and cheaper than frontier LLMs

7 reports5 sourcesupdated 9 hours ago

What happened

AI digest

On September 16, 2026, TypeSafe AI, founded by former OpenAI employee and RLHF co-inventor Diogo Almeida, released Jev, a "System 1" model. Jev does not generate text; it outputs type-safe structured values with calibrated probabilities, focused on classification, judgment, routing and scoring. It is trained with RLCD (Reinforcement Learning for Calibrated Decisions) and pitched on parallel sampling, no hallucinations and calibratable output. Early reports put input cost at $0.042 per million tokens with free output, and end-to-end latency of 70-500 ms — 40-200x faster and 40-400x cheaper than GPT-5.6 Terra. A later Latent Space report said it is over 100x faster and over 200x cheaper than small frontier models, but did not disclose specific latency, per-call price or parameter count. On September 18, a hands-on test called it cost-effective for classification tasks, again without disclosing parameter count, training data sources or latency figures. On September 19, TechCrunch reported Jev is still Transformer-based, with a training objective of executing code and calling APIs rather than generating human language, and that developer feedback was positive. The same day, "Yage Research" said community tests using open-source small models that read logits directly came within under 4 percentage points in judgment quality, cut latency from 178 ms to 71 ms and were free, and noted that selling judgment as a separate product is not new. On September 22, Almeida said on a podcast that mainstream large models over-rely on autoregressive chat tuning, and that Jev aims to "disappear into the background" like a regular expression; its launch video drew about 40 million views. On September 30, a report discussed combining Jev with Apache Arrow.

Written by AI from the coverage · updated 4 hours ago

Developments

3 developments

Coverage

Follow the reports to see the story from different sides.

Sep 30
  1. Hacker News front page
    What if Jev spoke Arrow?

    TypeSafe AI 推出新模型 Jev,可将自然语言和应用状态转化为带类型决策,返回选项、分数和概率并以 JSON 输出。其采用并行采样器与名为 Reinforcement Learning for Calibrated Decisions 的训练方法,在决策工作流中相比通用 LLM 有显著速度和成本优势。

Sep 22
  1. Latent Space
    Jev: System One models for Prod, not God — with Diogo Almeida, CEO, TypeSafe AI

    Diogo Almeida, CEO of TypeSafe AI, explains Jev's origin in a podcast. He argues that mainstream LLMs (like ChatGPT) have gone down the wrong path by over-relying on autoregressive chat tuning, dropping all other modes. Jev is designed as a 'System One' model: fast, reliable, embeddable into workflows, aiming to 'disappear into the background' like regex. The launch video got ~40M views, surpassing GPT-4o's 22M. Almeida also criticizes all three RLHF branches as wrong north stars. The post does not disclose Jev's architecture, parameter count, or pricing.

Sep 19
  1. Computing Life · Share · YagePick
    Jev is a classification-only API, but open-source alternatives are faster, deterministic, and free

    TypeSafe's Jev outputs probability distributions instead of text, aiming to decouple judgment from generation. Community benchmarks show open-source models reading logits directly match Jev's quality within 4 percentage points, while cutting latency from 178ms to 71ms and offering deterministic outputs. This classification-as-a-service idea has cycled through four prior waves since 2017—Perspective API, OpenAI's /classifications, Cohere Classify, and GLiNER2—all stalling due to missing demand or infrastructure. Jev's timing works because agent architectures now require frequent cheap judgments, frontier base models enable high-quality distillation, and distribution partners like Vercel onboarded it within 72 hours. The tech itself isn't a must-buy; the timing is the real story.

  2. TechCrunch · AIPick
    A ChatGPT inventor built Jev, a model that runs code instead of chatting, and developers are excited

    Diogo Almeida, a former OpenAI researcher who co-invented RLHF, built a model called Jev that optimizes for code execution rather than human language. It's still a transformer, but TypeSafe AI designed it to run programs and call APIs directly, acting more like an automation agent. Developer reception has been enthusiastic. The post doesn't disclose benchmark scores, pricing, or whether weights will be open.

Sep 18
  1. AI HOT (Curated Pool)
    TypeSafe AI Launches Jev, a Model for High-Frequency Decisions; Test Shows Cost-Effective Classification

    TypeSafe AI released Jev, a model designed for high-frequency decisions. The author's test shows it offers strong cost-effectiveness in classification tasks. The post does not disclose model size, training data, or latency figures.

Sep 16
  1. Latent SpacePick
    TypeSafe launches Jev: a “System One” model that only decides, classifies, routes, and scores — >100x faster, >200x cheaper than small frontier LLMs

    TypeSafe released Jev, a model that skips chat, code, and reasoning to focus on classification, routing, and scoring. Trained with RLCD, it promises parallel sampling, no hallucination, and calibrated outputs. The team claims >100x speed and >200x cost savings over small frontier LLMs. The post doesn’t disclose exact latency, per-call pricing, or parameter count. I’d treat the speed/cost claims as directional until the evals and pricing pages fill in the details.

  2. Hacker News front pagePick
    TypeSafe launches Jev, a structured-decision model that’s 40–400× cheaper and 20–200× faster than frontier LLMs

    TypeSafe founder Diogo Almeida (ex-OpenAI) announced System One models and the first public model Jev. Jev doesn’t generate strings—it outputs type-safe structured values with calibrated probabilities, making hallucinations and type errors mathematically impossible. Input costs $0.042/MTok, output is free; end-to-end latency is 70–500ms, 40–200× faster than GPT-5.6 Terra. The training method, RLCD, optimizes for calibrated decisions rather than human preference. A side-by-side demo with GPT-5.6 Terra shows only one disagreement—on churn likelihood—which the author says is genuinely ambiguous. I’d hold off on full enthusiasm: the post doesn’t provide independent third-party benchmarks, and long-term pricing sustainability isn’t proven yet.

Heat over time

Heat now 8·Comparable peak 9(Sep 30 06:00)·Comparable change over 24 hours –

02.557.510Sep 3006:00Sep 3007:00Sep 3008:00Sep 3009:00

The trend only compares accounts observed without gaps, so its range may be smaller than the current heat. Hover or tap the chart for each hour; the left and right arrow keys step through it.