The modern AI industry has long advanced by scaling generative models. We have become accustomed to using large language models (LLMs) such as GPT-6.1 Sol, Claude Opus 5.5, Claude Fable, or Gemini 3.8 Flash for every task: from writing essays to parsing order statuses. However, when building reliable software, text generation becomes a bottleneck: models spend seconds unfolding tokens step by step, hallucinate invalid JSON, require endless regular expressions, and are too costly for microservice architectures.
Jev from the TypeSafe AI lab is the world's first model of a new class known as System One AI. It does not generate text at all. Jev is built exclusively for instant, typed, structural decisions inside software systems: classification, routing, rubric scoring, hypothesis checking, and agent firewalls, with latency around 100 ms and a cost of $0.042 per 1 million tokens.
In this guide, as practitioners, we will dissect Jev's technology down to the smallest details: the fundamental difference between System One and System Two, the three key primitives, the mathematics of calibrated confidence, architectural patterns, ready-to-use TypeScript and Python pipelines, and ten breakthrough real-world case studies with actual cost receipts and video demonstrations.
Practical toolkit from Yuriy (@yuriisams):
For practical use of Jev in Claude Code, download the author's archive with ready-made UserPromptSubmit hooks for automatic model routing (jev-router) and dynamic skill selection (jev-skills):
Download the full JEV-Guide.zip archive (28 KB) →
TypeSafe Jev speed demonstration: instant claim verification with 69–83 ms latency and calibrated confidence1. Paradigm Shift: System One vs. Classical LLMs (System Two)
The concept's name comes from cognitive psychology and the foundational work by Nobel laureate Daniel Kahneman, Thinking, Fast and Slow.
Kahneman divides human thinking into two systems:
- System One: Fast, intuitive, automatic, nearly instantaneous reaction to familiar patterns (for example, recognizing a conversational partner's facial expression, detecting danger on the road, or distinguishing spam from important mail in a fraction of a second).
- System Two: Slow, deliberate, analytical work that requires focused effort and sequential steps (solving an equation, writing an architecture document, or logically proving a theorem).
Why Generative LLMs Break Software Pipelines
Most decisions in backend services are not philosophical essays, but concrete discrete choices:
- “Is this transaction suspicious?” (
trueorfalse). - “Which department should this ticket be routed to: billing, tech support, or sales?” (
billing | tech | sales). - “How irritated is the customer on a scale from 1 to 5?” (
1..5).
When an engineer forces a heavy generative model such as GPT-6.1 Sol, Claude Opus 5.5, or Claude Fable to answer these questions, a fundamental mismatch in the nature of the tool arises:
- Autoregressive latency: The model is required to generate tokens sequentially, one after another. Even a short answer takes 1.5–3 seconds due to network waits and generation quanta.
- JSON Mode fragility: Generative models often add comments, omit quotes, or wrap the response in
```json ... ```, forcing developers to write complex sanitizers and retry mechanisms. - Economic inefficiency: You pay the full cost of output tokens, which in commercial APIs cost 3–5 times more than input tokens.
- Attention degradation (Context Rot): If you ask the model several sequential questions in a single chat, the growing context dilutes attention and increases the probability of errors.
Jev Principle: Jev has no autoregressive text generation mechanism for the user. The model reads the input state (state), passes it through a transformer, and computes normalized probability distributions over predefined categories in a single forward pass. The result is returned immediately as a strictly typed data structure.
2. Economics, Speed, and Benchmarks: $0.042 per 1M Tokens and 100 ms
Jev’s economics change the rules of the game for high-load services. Instead of paying for each output word, developers pay only for input tokens.
Cost and Characteristic Comparison Table
| Characteristic | TypeSafe Jev 1.13 | OpenAI GPT-6.1 Sol | Anthropic Claude Opus 5.5 | Google Gemini 3.8 Flash |
|---|---|---|---|---|
| Architecture Class | System One (Decision) | System Two (Generative) | System Two (Generative) | System Two (Generative) |
| Input Token Cost (1M) | $0.042 | $3.00 | $5.00 | $0.10 |
| Output Token Cost (1M) | $0.00 (Free!) | $12.00 | $25.00 | $0.40 |
| Typical Latency (P50) | ~90–120 ms | 1200–2400 ms | 1400–2800 ms | 400–700 ms |
| Throughput Limits | 100K tok/s, 40 RPS | Depends on Tier | Depends on Tier | Depends on Tier |
| Request Context Window | 64,000 tokens | 256,000 tokens | 500,000 tokens | 2,000,000 tokens |
| Response Parsing Requirement | None (Native types) | JSON.parse / Regex | JSON.parse / Regex | JSON.parse / Regex |
Output Tokens Are Free: Because Jev does not generate free text, TypeSafe bills only the input context (state + question). Processing 1 billion tokens costs only $42. For comparison, the same volume on GPT-6.1 Sol or Claude Opus 5.5 would cost at least $3,000 – $7,500.
Real-World Benchmark: MotherDuck (SQL Text Classification)
The team behind the analytical cloud database MotherDuck integrated the prompt_jev() function directly into the SQL dialect for classifying text records at industrial scale.
Results from testing on an array of 100,000 rows:
- Classic frontier LLM: execution time — 32 minutes, compute cost — $37.00.
- TypeSafe Jev: execution time — 40 seconds, compute cost — $0.50.
- Summary: Jev completed the task 48x faster and 74x cheaper, demonstrating full parity in classification quality.
3. Request Anatomy: Structured State and Precise Dot-Path Addressing
A request to Jev consists of two fundamental components:
state(context): Task context as plain text, a string, or a nested JSON document (event log, transaction, user profile, code, page content).questions(question set): A dictionary of questions that the model must evaluate simultaneously against the provided state.
Context Isolation Principle and Preventing Context Rot
A classic mistake when working with large models is sending the entire history of previous dialogs and system prompts. Jev enforces strict context hygiene: pass into state only the data required for making specific decisions. The model receives the state once and computes all questions in parallel.
Precise Addressing via Dot-and-Index Paths
When your state is a complex structured JSON object, you can refer to specific fields and array elements directly in the question text using backticks:
The Jev model is specifically optimized for navigating the JSON tree. When it sees `transaction.amount_usd`, its attention layer focuses precisely on the specified key, eliminating ambiguity and preventing misinterpretation.
4. Three AI primitives: detailed breakdown of Choice, Score, and Noul
TypeSafe built the system around three minimal, complementary primitives. Each one is designed for a specific mathematical type of decision. Each subsection below includes a separate interactive demo where you can select a ready-made prompt and test the primitive in action.
4.1. The Choice primitive: selecting from a discrete list of categories
Choice is used when the answer must be a single category from a fixed, unordered list of options.
What Jev returns for Choice:
choice: Key of the selected category (for example,"technical").probabilities: Full probability distribution across all options:{"billing": 0.04, "technical": 0.92, "sales": 0.02, "other": 0.02}.confidence: A number from0.0to1.0that measures how clearly the leading option dominates the other options.
Golden rule for Choice: Always include an other or none_of_the_above category. If the input does not match any of the target options, the model will not be forced to artificially inflate the probability of an inappropriate category; instead, it will select other or signal low confidence.
4.2. The Score primitive: rating on an ordinal scale or rubric
Score is designed to evaluate properties located on a continuous or ordinal spectrum: bug severity, customer stress level, candidate resume quality, and code complexity.
You provide an ordered array of textual criteria (levels):
What Jev returns for Score:
score: Numeric value in the range from0toN-1. Important: the value can be fractional (for example,2.37) if the situation falls between the second and third rubric levels!probabilities: Probability distribution across each scale level.confidence: Degree of certainty of the rating.
4.3. Noul Primitive: Calibrated Probability of Statement Truth
The term Noul comes from the idea of a binary judgment. It is a question that can be answered “Yes” or “No.” Instead of returning a simple boolean value, Jev returns a calibrated probability that the statement is true.
What Jev returns for Noul:
noul: A floating-point number from0.0to1.0.- A value of
0.98means a confident “Yes.” - A value of
0.02means a confident “No.” - A value of
0.50indicates complete model uncertainty.
- A value of
Nouldoes not have a separateconfidencefield because thenoulvalue itself is the mathematical probability estimate.
4.4. Parallel Batching
All three primitive types can be combined in a single request in any quantity.
In the official TypeSafe cookbook, an experiment with a batch of 13 questions against GDPR text showed that combining all checks into a single request was 12.2 times cheaper and 10.0 times faster than sending 13 separate requests, while producing identical results.
5. Confidence vs. Probability: Mathematics and Confidence-Gated Routing
One of the primary problems with classic LLMs is their tendency toward hallucination and overconfidence. When a generative model (even at the level of GPT-6.1 Sol or Claude Opus 5.5) encounters a borderline or under-specified context, it attempts to invent convincing, plausible text. Jev is designed with calibrated-decision reinforcement learning algorithms (RLCD), enabling the model to honestly signal: “I am not confident.”
Mathematical Difference
- Probability (
probability): Answers the question “What is the chance that option X is correct?” This is a measure of aleatoric uncertainty within the given options. - Confidence (
confidence): Answers the question “How clearly does one option dominate all others?” This is a measure of the model’s epistemic certainty in its choice.
For $K$ options in a Choice query, the TypeSafe normalized confidence calculation formula is:
$$\text{Confidence} = \max\left(0, \min\left(1, \frac{K \cdot p_{\max} - 1}{K - 1}\right)\right)$$
where:
- $K$ — total number of available categories.
- $p_{\max}$ — highest probability among all categories.
Example interpretation for 3 options ($K = 3$):
-
If probabilities are distributed as
[0.90, 0.06, 0.04], then $p_{\max} = 0.90$.$\text{Confidence} = \frac{3 \cdot 0.90 - 1}{2} = \frac{1.7}{2} = 0.85$ (High confidence).
-
If probabilities are uniform
[0.34, 0.33, 0.33], then $p_{\max} = 0.34$.$\text{Confidence} \approx \frac{3 \cdot 0.34 - 1}{2} = \frac{0.02}{2} = 0.01$ (The model has no clear leader; complete uncertainty).
Three-Level Routing Template (Confidence-Gated Routing)
Thanks to the confidence metric, engineers can build reliable automation systems with three safety loops:
6. Production-Grade Architectural Patterns (Production Patterns)
Experience from hundreds of Jev projects has crystallized five key architectural patterns for designing intelligent systems.
6.1. Speculative Fan-Out
In conventional code, developers first evaluate a condition and then call the next function. In the Jev world, questions are computed in parallel, and adding new questions barely affects response time.
Pattern: In the first request, send not only the required questions but also speculative questions whose answers are needed only in particular code branches.
6.2. Composite Scoring
Instead of asking the model to abstractly “score a lead from 1 to 100,” break the assessment into atomic, objective factors. Combine them into a final score using a deterministic mathematical formula in your code:
$$\text{LeadScore} = 0.40 \cdot \text{BudgetConfirmed} + 0.35 \cdot \text{DecisionMakerRole} + 0.25 \cdot \text{Urgency}$$
If the company’s priorities change, you can adjust the weights in your code instead of retraining the model.
6.3. Structured Data Extraction (SDE) Cascade
When you need to extract complex data from unstructured text, build a two-stage pipeline:
- Stage 1: A fast parser or regular expressions identifies candidate entities (dates, amounts, links, email addresses).
- Stage 2 (Jev): A series of
ChoiceorNoulquestions verifies and selects the candidates that match the context. - Stage 3 (Only for 2–3% of collisions): If Jev returns
confidence < 0.5, the request is escalated to a heavy deep-reasoning model (GPT-6.1 Sol or Claude Opus 5.5).
This approach reduces total cloud LLM spend by 90–95%.
6.4. Tool-Call Firewall
AI agents with access to consoles or tool invocation (MCP, bash, SQL) pose a major risk of irreversible data deletion or execution of malicious code.
Using Jev, you can create a low-latency firewall: every command generated by the agent is sent to Jev with 5–7 security questions before execution:
- “Does this command modify system files?”
- “Are sensitive environment variables being sent to external hosts?”
- “Does the action align with the user’s original request?”
The 100 ms latency is imperceptible to the agent while reliably protecting the system from dangerous operations.
7. Practical Coding: Production-Ready TypeScript and Python Pipelines
Below are production-ready examples for building a customer support ticket intake and financial risk-scoring service using the official TypeSafe SDKs for TypeScript and Python.
7.1. Production Pipelines: Customer Support Ticket Processing
typescriptimport { TypeSafeClient, choice, score, noul } from "@typesafe-ai/sdk"; // 1. Ініціалізація клієнта (бере TYPESAFE_API_KEY зі змінних середовища) const client = new TypeSafeClient(); interface SupportState { ticketId: string; userEmail: string; accountAgeDays: number; messageText: string; attachedLogs?: string; } export async function processCustomerMessage(ticket: SupportState) { try { // 2. Виклик System One з паралельними питаннями const response = await client.systemOne({ state: { ticket_id: ticket.ticketId, user_tier: ticket.accountAgeDays > 365 ? "vip" : "standard", content: ticket.messageText, logs: ticket.attachedLogs ?? "Немає логів" }, questions: { // Категоризація запиту (Choice) topic: choice("Яка основна тема звернення користувача?", { billing: "Проблеми з оплатою, картками, підпискою, запит на повернення грошей", bug_report: "Повідомлення про збій у додатку, помилку в інтерфейсі або API", feature_request: "Побажання щодо покращення функціоналу, нові інструменти", account: "Проблеми зі входом, зміна пошти або скидання пароля", other: "Питання, які не підпадають під жодну з попередніх категорій" }), // Оцінка за шкалою роздратування (Score) frustration_level: score("Наскільки користувач роздратований у `content`?", [ "Спокійний: діловий, нейтральний тон, виклад фактів", "Стурбований: відчувається легке невдоволення або нетерпіння", "Розлючений: агресія, погрози піти до конкурентів, скарги", "Екстремальний: ненормативна лексика, caps lock, вимога негайного дзвінка керівництва" ]), // Перевірка на терміновість (Noul) is_urgent: noul("Чи вказує користувач у `content`, що його продакшен зупинений або проблема критична для бізнесу?"), // Перевірка на наявність витоку секретів у тексті (Noul) contains_leaked_secrets: noul("Чи містить `content` або `logs` приватні API-ключі, токени доступу чи паролі?") } }); const { topic, frustration_level, is_urgent, contains_leaked_secrets } = response.answers; // 3. Детермінована логіка маршрутизації console.log(`[Ticket ${ticket.ticketId}] Тема: ${topic.choice} (Впевненість: ${topic.confidence.toFixed(2)})`); console.log(`[Ticket ${ticket.ticketId}] Рівень стресу: ${frustration_level.score.toFixed(2)}/3.00`); // Безпековий контур if (contains_leaked_secrets.noul > 0.85) { console.warn(`[SECURITY ALERT] Виявлено можливий витік ключів у тікеті ${ticket.ticketId}. Автоматичне маскування!`); } // Маршрутизація на основі впевненості if (topic.confidence < 0.50) { return { status: "manual_triage", reason: "Model uncertain about topic" }; } if (is_urgent.noul > 0.80 || frustration_level.score > 2.0) { return { status: "escalated_p1", department: topic.choice, priority: "CRITICAL", confidence: topic.confidence }; } return { status: "routed", department: topic.choice, priority: "NORMAL", confidence: topic.confidence }; } catch (error) { console.error("Помилка під час виклику Jev API:", error); throw error; } }
8. Breakdown of the 10 Best Global Use Cases with Video Demonstrations (Receipts)
The engineer community on shipwithjev.com, jevbest.com, and jevable.com has demonstrated dozens of revolutionary Jev applications. Below is a detailed breakdown of 10 leading global cases: on the left is an interactive player showing a real demonstration or telemetry feed, and on the right is a structured analysis of the problem, the Jev-based architectural solution, and the verified speed and cost receipt.
8.1. Browser Use + Jev: Autonomous Airfare Search Agent
8.2. Toolgate: Firewall for MCP Tool Calls and Claude Code
8.3. Astra + Jev in Minecraft: Real-Time System One + Two Agent
8.4. Jev Driver: Autonomous Vehicle Control in the Browser
8.5. 2048Jev: Jev Plays 2048 in Real Time
8.6. Semantic Jev: Natural-Language Queries Through SQL
8.7. MotherDuck: prompt_jev() Classification Directly in SQL
8.8. Jev Swap: Find LLM Calls That Should Be Replaced with Jev
8.9. ElevenLabs: Real-Time Scam Caller Detection
8.10. Softlint: AI Linter in CI for Semantic Code Rules
9. Integration into Agentic Pipelines: Claude Code, Codex, and MCP Toolgate
Modern autonomous coding agents (Claude Code, OpenAI Codex, Antigravity, Cursor) face three critical bottlenecks: context bloat from dozens of connected tools, irrational use of ultra-expensive flagship models for trivial tasks, and the risk of uncontrolled destructive actions in the terminal.
Jev System One acts as an ultra-fast reflexive layer (L0/L1) for agentic systems: it makes discrete decisions in ~80–90 ms at a cost under $0.0003, optimizing the agent's entire workflow loop.
9.1. Dynamic Model Tiering for Different Task Types (Model Tier Routing)
In classic agentic pipelines, developers either hard-pin a single model (for example, Claude 3.7 Sonnet or GPT-4.5) for all subtasks, or invoke a heavy LLM to analyze the request, adding 2–4 seconds of latency and unnecessary costs at every step.
Jev classifies intent and task complexity in ~85 ms, routing the request to the appropriate model tier:
- Fast Tier (quick micro-tasks): Code formatting, writing simple unit tests, generating validators and documentation. Routed to Gemini 2.5 Flash or GPT-4o-mini ($0.05–$0.15 per 1M tokens).
- Balanced Tier (standard coding): Feature implementation, function refactoring, API integration, and fixing medium-complexity bugs. Routed to Claude 3.5 Sonnet or DeepSeek V3 ($3.00 per 1M tokens).
- Deep Reasoning Tier (critical architecture): Deep analysis of race conditions, database design, comprehensive cryptographic audits. Routed to Claude 3.7 Sonnet Thinking or OpenAI o3-mini ($12.00–$15.00 per 1M tokens).
Result: For a typical agent with 10–20 steps per task, this cuts average execution cost by 60–85% and response time by 40–50%.
9.2. Ranking and On-Demand Skill Loading (Dynamic Skill Selection)
A modern developer may have 30–80 agent skills (skills/*), plugins, and tools in their environment. If the full specifications and instructions for all skills are passed into the agent's system prompt:
- 15,000 to 35,000 tokens are consumed on each dialogue iteration.
- The model begins confusing similar tools (Tool Hallucination / Overload).
- The agent's first-response latency increases to 5–10 seconds.
Jev implements an On-Demand Skill Ingestion architecture. The agent keeps only lightweight one-line descriptions of available skills in memory, and at each step Jev ranks them in ~80 ms and selects the top 1–2 most needed:
Why this is faster: Instead of sending 30k tokens, the agent sends only the user request (~200 tokens) to Jev. Once it receives the name of the required skill, the agent loads the corresponding SKILL.md file immediately before execution. This keeps the context window clean for the project code.
9.3. Security firewall toolgate for MCP and terminal
The open-source project toolgate adds a validation middleware layer for any Model Context Protocol (MCP) calls. Each time an autonomous agent initiates a terminal command (bash, npm run, git reset) or overwrites files, Jev concurrently computes 7 Noul destructive-risk probabilities:
- Dangerous file deletion: Probability of destructive actions (
rm -rf, wiping directories outside the repository). - Sensitive data leakage: Attempt to expose environment variables (
.env,AWS_SECRET_ACCESS_KEY, private SSH keys). - Privilege escalation: Use of
sudo, modification of system files in/etc/or~/.ssh/. - Git history destruction:
git push --forcecalls or branch resets without confirmation. - Network anomalies: Unauthorized socket opening or sending data to external IPs.
- Scope drift: Agent attempts to modify files unrelated to the assigned task.
- Operational cost: Launching heavy cloud scripts or deploying without user consent.
If the aggregate risk index exceeds the 0.85 threshold, the action is immediately blocked, and the operator receives an alert with a precise description of the threat.
9.4. Comparison matrix: Classic agent vs. Jev System One agent
| Operational parameter | Classic approach (Full Context / Heavy LLM) | Agent with Jev System One integration | Project benefit |
|---|---|---|---|
| Task-specific model selection | Fixed flagship model or heavy LLM router (2–4 s, ~$0.02) | Jev System One Choice router (~85 ms, $0.0003) | -98% latency, -98.5% cost |
| Skill selection and loading | All 40–80 skills in the prompt (25,000+ tokens per step) | Ranking in 80 ms and loading 1–2 skills on demand | Up to 90% context savings |
| Tool security control | Simple regex rules or no protection at all | 7 parallel Noul checks before each tool call | Protection against unauthorized actions |
| Agent startup speed | 4–9 seconds waiting for the first token | 600–900 ms until step execution begins | 5–7× faster startup |
| Average cost of 100 agent steps | ~$4.50 – $8.00 | ~$0.85 – $1.40 | 75–85% cost savings |
9.5. Installing the official skill for agents
TypeSafe provides ready-made integration plugins and skills for popular development environments:
bashclaude plugin marketplace add typesafe-ai/skills
9.6. Ready-made Claude Code toolkit from Yuriy (@yuriisams)
Developer and practitioner Yuriy (@yuriisams) created a ready-to-use automation based on Jev directly for the Claude Code terminal agent. These are two autonomous tools that integrate as UserPromptSubmit hooks and trigger automatically on every user message while keeping full control in the developer's hands.
By default, both tools are disabled: until activated, Claude Code runs in standard mode and does not send data anywhere. If the Jev service is unavailable or the model confidence is low, Claude transparently continues normal execution.
| Tool in archive | Purpose | How it works under the hood | Location |
|---|---|---|---|
jev-router.zip | Model router | Jev classifies the task (tiny, everyday, large, hardest), after which Claude either responds directly or invokes Haiku, Sonnet, or Opus. | ~/.jev-router/router.py |
jev-skills.zip | Skills picker | Jev reviews the skills in ~/.claude/skills/; when confidence is $\ge 60%$, it recommends the appropriate skill to Claude instead of guessing. | ~/.jev-skills/hook.py |
Jev-README.md | Guide | Full step-by-step instructions in Ukrainian and configuration setup. | Archive root |
Step-by-step installation on the workstation
tabs === Quick download (Curl)
=== 1. Configure jev-router
=== 2. Configure jev-skills
:::
Registering hooks in ~/.claude/settings.json
In the ~/.claude/settings.json configuration file, add the hooks to the hooks.UserPromptSubmit array (without removing any other existing hooks):
Model instructions: Add the contents of the claude-md-snippet.md file from the archive to your global ~/.claude/CLAUDE.md file. This teaches Claude itself how to respond correctly to the router and skills-picker control tags.
Control and quick terminal commands
After restarting Claude Code, you can enable or disable the tools at any time:
-
Model router control:
jev router on/jev router off/jev router status. -
Skills picker control:
jev skills on/jev skills off/jev skills status. -
One-off check without enabling hooks:
bashpython3 ~/.jev-skills/picker.py "налаштувати nginx для reverse proxy з ssl"
Download the ready-made JEV-Guide.zip archive from Yuri (@yuriisams) →
10. Pitfalls, Jev 1.13 Limitations, and Implementation Checklist
Jev is a powerful tool, but it is not a universal silver bullet. Understanding its limits helps avoid critical mistakes during the design phase.
Known Jagged Edges in the current jev-1.13.0 release
- Context limits: The maximum request size is 64k tokens, but the
stateplus the longest question must not exceed 32k tokens. As you approach the 32k boundary, model accuracy begins to gradually degrade. - Text only: Jev does not support direct ingestion of images, video, or audio. All media data must be transcribed (for example, with Whisper) or converted into structured text before being passed to
state. - Language specifics: The model was trained primarily on an English-language data corpus. It can process Ukrainian, Polish, or Spanish, but the highest accuracy is achieved with this pattern:
Tip
Tip for localized projects: Pass the user's local text (for example, in Ukrainian) into the
statefield, but write the instructions (instructions) and criteria (criteria) for questions in English. Jev maps English rules to Ukrainian context very well.
Production readiness checklist
- The article body and service contain no unnecessary system prompts; only clean
stateis passed. - All Choice questions include a default category (
otherornone_of_the_above). - Related questions are grouped into a single parallel
client.systemOnecall. - Three-tier routing based on
confidenceis implemented (Tier 1 / Tier 2 / Tier 3). - Confidence thresholds are differentiated: higher for destructive operations ($>0.90$) and moderate for read-only operations ($>0.60$).
- Automatic retry with exponential backoff is configured to handle possible HTTP 429 rate limits.
- Question constants and threshold values are extracted into a separate configuration file.
- Input state size is validated (does not exceed 32k tokens per question).
- A fallback route to a classic reasoning LLM is provided for anomalously low confidence.
-
response.modelandanswers.*.confidencevalues are logged for later analysis of decision distribution.
11. Local Open-Weight Alternatives: Laya, GLiNER2.5-Decide, and CLM-8B
Although TypeSafe's cloud Jev offers extremely affordable pricing ($0.042 per 1M tokens), many enterprise systems require full control over data (on-premises, GDPR, HIPAA, banking secrecy) or zero dependency on third-party APIs and network latency.
The open-source community quickly adopted the System One paradigm and released open-weight decision models (Open-Weight Decision Models) that can be run locally on your own server, GPU, or even an Apple Silicon laptop.
11.1. Laya by Convai Innovations: a direct open-source Jev equivalent
Laya is the first direct open-source analog of Jev, built around the same philosophy: a non-autoregressive decision model that never generates free text and instead evaluates typed questions in a single forward transformer pass (~33 ms on GPU).
The model is trained using reinforcement learning based on strictly correct evaluation rules (RLCD — Reinforcement Learning for Calibrated Decisions), which ensures mathematically honest probability calibration.
- Stack and architecture: a ModernBERT-large encoder (421M parameters) for English and an mmBERT-base encoder (322M parameters) for 100+ languages worldwide.
- Primitive support: native support for
choice,score, andnoulwith the same request and response format as TypeSafe Jev. - Document context: supports up to 1024 tokens by default and up to 8192 tokens in the
laya-multilingualversion (max_len=8192). - Compatible
laya-serveserver: includes a built-in proxy that implements thePOST /v1/systemoneendpoint. You can replace cloud Jev in your existing applications simply by changingbaseURLtohttp://localhost:8000. - Speed and cost: approximately 33 ms on GPU (6–8 times faster than a network call to cloud Jev) and $0 cost under the Apache 2.0 license.
pythonfrom laya import Router # Preload моделей у пам'ять для миттєвої маршрутизації (<35 мс) router = Router(preload=True) state = "Користувач скаржиться на подвійне списання коштів за підписку і вимагає повернення." questions = { "department": { "type": "choice", "instructions": "Which department should handle this request?", "criteria": { "billing": "invoices, payments, refunds", "tech_support": "bugs, outages, system errors", "other": "everything else" } }, "churn_risk": { "type": "noul", "instructions": "Does the user threaten to cancel or express high churn risk?" } } result = router.predict(state, questions) print("Відділ:", result["answers"]["department"]["choice"]) print("Ризик відтоку:", result["answers"]["churn_risk"]["noul"])
Multilingual support: The Router automatically detects the language of the text (including Ukrainian) and routes non-English requests to the laya-multilingual checkpoint, enabling high-quality classification across 45+ languages without manually switching weights.
Open repository on Hugging Face: convaiinnovations/laya →
11.2. GLiNER2.5-Decide from Fastino: 340M schema-oriented classifier
GLiNER2.5-Decide is a specialized model from the Fastino lab, designed for operational classification, safety filters, and task routing without requiring prompt engineering or parsing output tokens.
The model is based on the DeBERTa-v3-large architecture (340M parameters) and outperforms the commercial JevK5 on the fast-decisions benchmark (60.2% accuracy vs 57.6% for Jev).
- Dynamic label set: The list of categories is passed directly in the function call at runtime without retraining (Zero-Shot).
- Multi-Label support: It can return multiple labels simultaneously (for example, identifying several aspects of feedback or customer issues) based on the
cls_threshold. - Labels with descriptions: If a label name is ambiguous, it can be paired with a detailed description; the model considers it during decision-making.
- Ordinal scales and QA: It supports numeric urgency levels (for example, from
"0"to"5"), tone scoring, and binary questions (yes/no) about the provided text snippet. - Minimal hardware requirements: With 340M parameters, the model can operate with millisecond-level latency even on a standard CPU.
Open repository on Hugging Face: fastino/GLiNER2.5-Decide →
11.3. CLM-v0.1-8B by Contrastive-LM: contrastive action scoring built on Qwen3
CLM-v0.1-8B (Contrastive Language Model) is a development by researchers at Stanford and NVIDIA (September 2026), built for lightning-fast action selection inside agentic pipelines (Computer Use, Tool Calling, and step verification).
Instead of slow step-by-step text generation, the model works on a contrastive principle: it projects the current state (state) and a list of possible actions or tool calls into a shared vector space, then ranks them by cosine similarity.
- Stack and architecture: a frozen Qwen3-8B encoder combined with lightweight trained contrastive projection heads (~20M parameters).
- Agent acceleration: delivers up to 9× lower latency compared with generative LLMs when selecting the required tool from a large function list.
- Action Caching: because state and actions are encoded separately, vector embeddings for static tools or system functions can be computed once and kept in memory.
- Local hardware support: open weights under the Apache 2.0 license and official Apple MLX optimizations allow deploying the model on workstations and Macs with unified memory.
Open the Hugging Face repository: Contrastive-LM/CLM-v0.1-8B →
11.4. Summary table: Jev versus open System One alternatives
| Model | Developer / Organization | Architecture and size | Latency (P50) | Primary focus and features | Drop-in compatibility with Jev | License |
|---|---|---|---|---|---|---|
| TypeSafe Jev 1.13 | TypeSafe AI | Proprietary in-house architecture | ~90–120 ms | Cloud System One model, 3 primitives, calibrated confidence | Official API | Commercial ($0.042/1M) |
| Laya | Convai Innovations | ModernBERT-large (421M) / mmBERT (322M) | ~33 ms (GPU) | 1:1 support for Choice/Score/Noul, 100+ languages, RLCD calibration | Yes (laya-serve) | Apache 2.0 (Open Source) |
| GLiNER2.5-Decide | Fastino AI | DeBERTa-v3-large (340M) | ~15–40 ms | Multi-label classification, label descriptions, excellent CPU performance | No (custom SDK gliner2) | Apache 2.0 (Open Source) |
| CLM-v0.1-8B | Contrastive-LM (Stanford / NVIDIA) | Frozen Qwen3-8B + Heads (~8B) | ~50–80 ms | Contrastive action scoring for agents, Action Caching, Apple MLX | No (contrastive router) | Apache 2.0 (Open Source) |
11.5. How to Choose a Local Model for Your Stack
Practical rule for choosing a local System 1 engine:
- Choose Laya (
convaiinnovations/laya) if you have already designed your pipeline around TypeSafe Jev primitives (choice,score,noul), need multilingual processing (including Ukrainian-language text), or want to migrate an existing production workload to your own server without rewriting code usinglaya-serve. - Choose GLiNER2.5-Decide (
fastino/GLiNER2.5-Decide) if your task is fast message classification, multi-label categories (multiple tags at once), working with complex descriptive labels, or if you need to deploy the service on resource-constrained servers without dedicated GPUs. - Choose CLM-v0.1-8B (
Contrastive-LM/CLM-v0.1-8B) if you are building an autonomous AI agent with a rich toolset (MCP / Tool Calling) and need maximum action-verification speed and error protection through tool-embedding caching.