Open source, self-hosted. There is no hosted API. You download the weights and run them on your own hardware. How open-source models are listed.
What GLiNER2.5-Decide is
GLiNER2.5-Decide is Fastino Labs’ System One model: an encoder-based decision model rather than a generative one. It evaluates a set of user-defined typed questions and rules against a piece of text in one pass, then jointly decodes the answers, returning probability distributions and confidence scores. Beyond typed Choice, Score and Noul questions, Fastino describes it returning character-level spans, relations, and structured records, and applying implications, exclusions, cardinality limits, and ordinal bounds across related decisions in one call.
The model card is direct about scope: “This release is not a general-purpose model. It does not reason, explain, or answer open questions.”
How it is built
The Hugging Face model card lists GLiNER2.5-Decide as a 340M-parameter model fine-tuned from Fastino’s own gliner2-large-v1, a DeBERTa-v3-large encoder. The Hugging Face API’s safetensors metadata for the published weights reports 486,444,053 parameters, which does not match the 340M figure. Neither source explains the gap.
It loads through the gliner2 Python package (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and accepts label sets at runtime, so a new task needs no retraining or prompt template.
What it’s good at, per Fastino’s own benchmark
Fastino evaluated the model on “Fast Decisions,” an internal suite of 17 datasets covering intent routing, triage, sentiment and content understanding. This is Fastino’s own benchmark, not an independently run one. The blog post and company X post give GLiNER2.5-Decide 60.1% average accuracy, ahead of JevK5 (57.5%), SemIf (56.4%) and Laya (46.6%), leading on 9 of 17 datasets. The Hugging Face card’s own results table gives slightly different figures for the same run: 60.2% and JevK5 at 57.6%. Fastino highlights support intent classification at 75.3%, 18.6 points ahead of the next model, and banking intent at 64.3%, 8.6 points ahead.
The GitHub repository behind it, fastino-ai/GLiNER2, had 2,169 stars and was last pushed on 24 September 2026, the day of release.
What an independent benchmark shows
The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, is a broader test than Fastino’s. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, GLiNER2.5-Decide scores 11.21 against Jev’s 57.91. That is the best result among the models under 500M parameters it tested, but far behind the larger open models. JevK5, which Fastino’s suite ranks below GLiNER2.5-Decide, scores 38.81 here. Its expected calibration error, the average gap between stated confidence and actual accuracy, is 0.088 against Jev’s 0.074.
Access and running it
Weights are Apache 2.0 on Hugging Face with no waitlist. Fastino says the model is “efficient enough to run locally on consumer-grade CPUs or in air-gapped environments,” and co-founder George Maloney gives CPU latency around 167ms and GPU latency of 38 to 47ms for short requests. Fastino also runs hosted inference and fine-tuning at http://agent.fastino.ai, for calling from inside a coding agent; no pricing for that service appeared in the sources checked.
What it’s not for
It does not generate text, reason through multi-step problems, or answer open-ended questions; the model card says so outright. No context-window limit, maximum Choice count, or per-call question cap is documented anywhere checked, so those limits are unknown rather than unlimited.
Specifications
| Question types | ChoiceScoreNoul |
| Max Choice options | Not documented |
| Score levels | Not documented |
| Questions per call | Not documented |
| Total context | Not documented |
| State budget | Not documented |
| Rate limit | None for self-hosted use. Fastino's hosted inference and fine-tuning API at http://agent.fastino.ai did not have published rate limits in the sources checked. |
| Endpoint | http://agent.fastino.ai |
| SDKs | Python: gliner2 |
Neither the blog post, the Hugging Face model card, nor the GitHub repo README documented a fixed context window, a maximum number of Choice options, or a cap on questions per call as of this check.
Versions
- fastino/GLiNER2.5-Decide, 24 Sep 2026, 340M-parameter encoder fine-tuned from fastino/gliner2-large-v1, a DeBERTa-v3-large base per the Hugging Face model card. The Hugging Face API's safetensors metadata reports 486,444,053 total parameters for the published weights, which does not match the 340M figure stated in the blog post and model card; this discrepancy is unresolved as of this check. Apache 2.0 license. Release notes
Use cases
What people use GLiNER2.5-Decide for, one page per pattern.
Support inbox triage with System One models
Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.
Intent and model routing with System One models
One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.
Agent routing and skill selection with System One models
An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess.
Structured extraction with System One models
Jev is not trained to generate text, so it cannot write a value out for you. Extraction works the other way round: a regex or a parser finds candidate spans, a choice question picks the one the question asks for, and code copies that span unchanged.
Examples built with GLiNER2.5-Decide
The most-starred and most-viewed entries in the directory. Browse all examples.

GLiNER2.5-Decide
Open-weights 340M-parameter decision model from Fastino Labs. It jointly decodes typed classification questions and, per Fastino's own internal benchmark, leads 9 of 17 datasets against JevK5, SemIf and Laya.
2.2k starsCPU latency 167 msGPU latency 38-47 ms

Decision Index: Jev against 50+ open decision models
A leaderboard that runs Jev and 54 open System One models and clones through the same 120,000-question suite on one RTX PRO 6000, scoring accuracy and calibration. Jev leads the 0.2 edition at 51.67 on a chance-corrected scale, with AutoJev-27B close behind at 50.94.