Laya: the open-source 421M-param "decision model" that wants to replace the LLM in your agents
🔎 A LLM to say "billing"? We have a problem
While the big labs fight over the crown of best general-purpose model — GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.7 — another, quieter race is accelerating: the one for models that decide. Every AI agent in production makes thousands of these per hour. Routing a ticket, scoring a request, validating a tool call. And today, almost all of these micro-decisions go through a giant LLM that generates text… to say "billing".
On September 21, 2026, TypeSafe AI opened its Jev API to the public, a "System 1 model" dedicated to these typed decisions. A few days later, the community answered with open weights: Laya, 421M parameters, Apache 2.0 license, released by Convai Innovations. The repo literally caught fire — 24,091 stars in 7 days, according to buildwithneej's recap "5 GitHub repos that caught fire" (September 2026).
This isn't yet another Llama clone. It's the emergence of a new category of models — decision models — and, frankly, it was about time someone took care of the least glamorous and most heavily exercised layer of your agents.
Key Takeaways
- Laya is not an LLM. It's a non-autoregressive decision model that returns typed decisions — choice, score, noul (yes/no) — with calibrated probabilities, in a single pass (~35 ms on GPU).
- Under the hood: ModernBERT-large (395M of the 421M params) + a decision head trained from scratch, with a built-in act-versus-escalate gate. Apache 2.0, full safetensors weights.
- Against Jev 1.13.0 (TypeSafe AI's closed API, opened on September 21, 2026, $0.042/M input tokens): ~7x lower latency, ECE calibration of 0.081 vs 0.246, zero cost when self-hosted. But Jev dominates large option sets (0.870 vs 0.425 on Banking77).
- The real issue: your agent router doesn't need a 400B LLM. A local-first cascade matches Jev's accuracy 1.8x faster, according to the independent lab yibie/laya-jev-lab.
Recommended Tools
| Tool | Primary use | Price (September 2026) | Best for |
|---|---|---|---|
| Laya — GitHub SDK | Open-weight decision model (choice/score/noul) | Free — Apache 2.0 | Self-hosting your decision layer |
| ONNX Port — receptron/laya | ONNX execution, including in the browser | Free | Demos, edge, serverless inference |
| Jev 1.13.0 — TypeSafe AI | Hosted decision API | $0.042/M input tokens (flowtivity.ai benchmark, verify on TypeSafe's site) | Classification with very large option sets |
| Hostinger | VPS for your Laya + local LLM pipeline | Depends on plan (check hostinger.com) | Deploying a local-first cascade |
Honest note: the jevmodel.org site is independent of TypeSafe — in fact, it's the one that had to publish a clarifying statement, evidence included, that Laya is neither Jev nor a TypeSafe release.
What is a "decision model"? The System 1 of your agents
A decision model doesn't generate text: it returns a typed decision, with a calibrated probability, in a single forward pass. That's the whole difference.
Borrow Kahneman's framework. A frontier LLM is your System 2: it thinks, breaks things down, generates token by token. Excellent for multi-step reasoning — GPT-5.5 scores 98.2 on agentic benchmarks — but grotesque for deciding "does this ticket go to billing?". Every routing triggers text generation: several hundred milliseconds, per-token cost, and fragile output parsing.
Laya is the System 1: the reflex. The SDK asks typed questions about a text and gets back structured answers:
- choice: "billing", with a confidence of 0.94;
- score: 1.84 out of 2.0;
- noul: a yes/no with a calibrated probability, for example 0.892.
Everything is computed in a single pass — about 35 ms on GPU, all questions in one batch. No generation, so no hallucination in the classical sense: a probability you can threshold, audit, and monitor.
The usage is, in fact, strictly bounded. The PyPI page for laya 0.3.21 makes it clear: structured usage only, no open Q&A, no text generation. This model doesn't want to write your blog. It wants to decide whether your agent should trigger the refund.
Vishal Mysore's analysis on Medium sums it up well: Laya returns calibrated probabilities instead of generated text. Routing, scoring, guardrails — the thousands of micro-decisions that separate a reliable agent from a PowerPoint demo.
Under the hood: lean, and that's a compliment
Laya doesn't invent a new architecture: it's a proven encoder plus a decision head trained from scratch. And that's precisely why it works.
According to flowtivity.ai's technical benchmark (September 2026), the English checkpoint is built on ModernBERT-large, which provides 395M of the 421M parameters. The remaining ~26M carry the decision head, two transformer layers, an option scorer, and an act-versus-escalate gate — act, or escalate to a bigger model. The cascade lives in the model, not in your code.
On the checkpoint side, the Reddit thread r/LocalLLaMA lists three: a 421M English one on ModernBERT-large, a 322M multilingual one on mmBERT-base — faster — and a typed-decisions variant. The positioning is unapologetic throughout: System 1 model, non-autoregressive.
The SDK fits in three lines. laya.load("convaiinnovations/laya") loads the root English checkpoint; the multilingual subfolder covers 100+ languages; the typed-decisions subfolder enables typed choice/score/noul responses.
Two details that signal real rigor. A laya_shortlist guardrail for choices with more than 20 options — the authors know exactly where their model breaks. And MCP tests included in the package: CI without weights, plus a local e2e test with real weights and a stdio handshake. This is a far cry from the showcase repo thrown together overnight to ride a hype wave.
To install it, one line is enough:
pip install laya
After that, it's all about architecture. And that's the next matchup we're interested in.
Laya vs Jev: what the numbers really say
On short decisions, Laya wins on latency, cost, and calibration. Jev keeps the advantage on large option sets. The flowtivity.ai table is crystal clear:
| Metric | Laya (open-weight) | Jev 1.13.0 (TypeSafe API) |
|---|---|---|
| Latency, 1 question | 32.8 ms (T4, self-measured) | 236–276 ms (third-party measurements) |
| 10 batched questions | 72.3 ms total (7.2 ms/question) | ~1,500 ms in series |
| Typed-decisions accuracy | 0.766 fine-tuned / 0.362 zero-shot | 0.727 (published) |
| Calibration (ECE) | 0.081 after temperature refit | 0.246 (reported) |
| Languages > 3x random | 45 of 51 evaluated | — |
| Cost | $0 self-hosted | $0.042/M input tokens |
| Weights | full safetensors, Apache 2.0 | No access (API waitlist) |
| Banking77 (77 labels) | 0.425 | 0.870 |
Three readings of the table
One, calibration. ECE measures the gap between stated confidence and actual accuracy — the lower, the better. 0.081 vs 0.246 is the difference between a model whose 0.9s you can trust and a model you have to be wary of. For agent guardrails, it's the metric that matters most: a badly calibrated "95% sure" decision is an agent derailing in silence.
Two, the 0.362 zero-shot. Don't skip this line: without fine-tuning, Laya is frankly mediocre. Its 0.766 fine-tuned beats Jev's published 0.727, but you'll have to train on your own data. Laya isn't a model you consume, it's a model you adapt.
Three, Banking77. On this 77-label classification benchmark, Jev scores 0.870 against 0.425 — its acknowledged strong suit. If your use case is very high-cardinality classification, the current answer is not Laya.
And where does the independent lab fit into all this?
yibie/laya-jev-lab tested 40 Chinese support-ticket classification cases: a local-first cascade matches Jev's accuracy while running 1.8x faster. The authors themselves flag the limitations — small samples (10 to 24 for some sub-analyses), a single task domain. Read these as trends, not point estimates. That honesty is refreshing; it's also rare.
Why your agent router doesn't need a 400B LLM
Because deciding "billing or technical?" is a reflex, not reasoning. And you don't pay a lawyer to open a door.
Do the math on the frontier-LLM reflex. A model like GPT-5.5 — 98.2 on agentic benchmarks, see our pick of the best LLMs for AI agents — is a scalpel for multi-step planning. Using it to route tickets means burning hundreds of milliseconds and fractions of a cent per decision, thousands of times per hour, to produce a word that a 421M-param encoder spits out in 33 ms, for free.
The worst part? This overworked LLM generates its response token by token, with no calibrated probability. You parse text and you pray. At the scale of an agent fleet, that's structural waste — and a risk.
The three-tier cascade
The winning architecture — Laya has built it in natively with its act-versus-escalate gate:
- Laya decides. Most cases are obvious: it makes the call in ~35 ms, with calibrated confidence to back it up.
- The gate escalates. Ambiguous cases go to a real LLM — Claude Opus 4.7, Gemini 3 Pro Deep Think, depending on your autonomous agent stack.
- The LLM executes. Generation, reasoning, everything it's actually made for.
This is exactly the setup validated by the yibie lab: Jev-level accuracy, 1.8× the speed, rock-bottom bill. The real question was never "LLM or no LLM" — it's "who decides what". And to pick who deserves your escalations, our monthly comparison of the best LLMs is made for exactly that.
Where Laya fits into an agent stack (and with what)
Laya takes the decision layer — upstream of the LLM, never in its place. In a modern agent pipeline, it reads in four stages: ingestion, decision, orchestration, execution.
Upstream, a crawler like Crawl4AI, the #1 open-source crawler on GitHub for feeding your agents and RAG pipelines collects the raw data. Laya sorts it: what intent, what priority score, does it need escalating? Orchestration then coordinates the tools — a framework like Vercel eve, the open-source framework that wants to do for AI agents what Next.js did for the web, or the runtime GoogleAx v0.3.0, Google's agent orchestrator that went open source. And only the subset of requests that deserve it ever reaches the LLM.
On the runtime side, optimizations stack up without stepping on each other: Life-Harness, which boosts LLM agents by 88.5% without retraining acts on execution, Laya on decision. Two complementary layers, not competing ones.
Two integration points worth knowing about. The PyPI package ships with MCP tests — everything you need to plug Laya into MCP tool ecosystems without any fiddling. And the receptron/laya port pushes ONNX all the way to the browser: your decision layer can run client-side, with no server. For edge cases, demos, or infrastructure savings, it's an angle to watch very closely.
The limitations, no corporate speak
Laya won't replace your agents' LLM — it takes over the reflex layer. The title is provocative, the reality is more nuanced, and — a rarity worth saluting — the limitations are documented by the authors themselves.
No generation, period. The PyPI page is explicit: structured usage only, no open-ended Q&A. If you were looking for a small model that writes, check out our favorite local LLMs instead.
Cardinality, its Achilles' heel. 0.425 on Banking77 versus 0.870 for Jev. Beyond 20 options, the laya_shortlist guardrail becomes mandatory — and even then, the gap remains.
Zero-shot is weak. 0.362 is the number to keep in mind. Without fine-tuning on your data, Laya underperforms even Jev's published 0.727. Budget for the annotation effort.
The measurements aren't apples-to-apples. Laya's 32.8 ms are self-measured on a T4; Jev's 236–276 ms come from third-party measurements. flowtivity.ai owns up to it — that's the price of an "honest" benchmark — but validate on your own workload.
The multilingual side has its gray areas. 45 out of 51 languages above 3x random: very good. But six fall below the bar, and the multilingual checkpoint (322M, mmBERT-base) prioritizes speed. For French, test on your own data before production.
The ecosystem still confuses everything. It took an independent site to remind people that Laya is neither Jev nor a TypeSafe release. When the community needs a clarifier, it means we're at the very beginning: excitement guaranteed, stability not yet.
Getting Started with Laya: the minimal path
One install, one checkpoint, typed questions. Nothing more.
pip install laya
The SDK loads the weights in three lines: laya.load("convaiinnovations/laya") for the root English checkpoint, the multilingual subfolder for the 100+ languages, the typed-decisions subfolder for choice/score/noul responses. All your questions go out in a single pass — count on ~35 ms on a T4-class GPU, and even less on the 322M multilingual checkpoint. To see it in action before writing a single line, Vishal Mysore's live demo on Medium walks you through it in pictures.
For infrastructure, two paths. First, the local machine: Laya coexists frictionlessly with a full local stack — our guide to installing a local LLM covers Ollama and LM Studio, and our open source AI agents with Ollama shows how to set the whole thing up end to end. Next, the VPS: a 421M-param checkpoint doesn't call for a GPU farm, and Hostinger hosting is enough for a starter local-first cascade.
My advice: a single use case to start — ticket routing or tool guardrails. Fine-tune the typed-decisions variant, measure your own ECE, compare. The published numbers are a compass, not a destination.
❌ Common Mistakes
Four pitfalls keep coming up in discussions around Laya. Here they are, along with the fix.
Mistake 1: Treating Laya as a mini-LLM
Expecting generated text, summarization, open-ended Q&A — then concluding that "it doesn't work." Laya returns typed decisions, nothing else. The solution: keep an LLM for generation and hand Laya the routing, scoring, and guardrails. It's a cascade, not a substitution.
Mistake 2: Confusing Laya and Jev
The confusion grew so widespread that an independent site had to publish a clarification: Laya is an Apache 2.0 open-weight model from Convai Innovations, meant to run on your own infrastructure; Jev is TypeSafe AI's hosted API, opened on September 21, 2026. Same job — choice, score, noul — different products.
Mistake 3: Comparing latencies without reading the table footnotes
32.8 ms vs 236 ms makes for a punchy headline. But Laya's numbers are self-reported on a T4, Jev's come from third parties, and the hardware differs. Before you decide, benchmark on your own workload — that's what the yibie lab did, and its local-first cascade changes the conclusion.
Mistake 4: Running Laya zero-shot on 50 options
0.362 zero-shot, 0.425 on Banking77: large option sets are the model's documented weak point. Solution: reduce the cardinality, enable laya_shortlist beyond 20 options, fine-tune — or just go with Jev for this specific case. Its 0.870 on Banking77 exists for a reason.
❓ Frequently Asked Questions
Short answers, no unnecessary jargon.
Can Laya really replace an LLM in an agent?
For deciding, yes; for generating, no. Laya excels at routing, scoring, and guardrails — choice, score, noul in ~35 ms with calibrated probabilities. For writing, multi-step reasoning, or open-ended Q&A, keep an LLM. The winning architecture is the cascade: Laya decides, the gate escalates, the LLM executes.
How much does Laya cost?
The model is Apache 2.0, with full safetensors weights: $0 when self-hosted. Only compute and any fine-tuning