Skip to content
System One

GLiNER2.5-Decide

GLiNER2.5-Decide is Fastino Labs' open-weights System One model. A 340M-parameter encoder answers typed classification questions, extracts spans and relations, and enforces cross-decision rules, and Fastino shipped it on 24 September 2026 under Apache 2.0.

Updated

Open source, self-hosted. There is no hosted API. You download the weights and run them on your own hardware. How open-source models are listed.

What GLiNER2.5-Decide is

GLiNER2.5-Decide is Fastino Labs’ System One model: an encoder-based decision model rather than a generative one. It evaluates a set of user-defined typed questions and rules against a piece of text in one pass, then jointly decodes the answers, returning probability distributions and confidence scores. Beyond typed Choice, Score and Noul questions, Fastino describes it returning character-level spans, relations, and structured records, and applying implications, exclusions, cardinality limits, and ordinal bounds across related decisions in one call.

The model card is direct about scope: “This release is not a general-purpose model. It does not reason, explain, or answer open questions.”

How it is built

The Hugging Face model card lists GLiNER2.5-Decide as a 340M-parameter model fine-tuned from Fastino’s own gliner2-large-v1, a DeBERTa-v3-large encoder. The Hugging Face API’s safetensors metadata for the published weights reports 486,444,053 parameters, which does not match the 340M figure. Neither source explains the gap.

It loads through the gliner2 Python package (AutoExtractor.from_pretrained("fastino/GLiNER2.5-Decide")) and accepts label sets at runtime, so a new task needs no retraining or prompt template.

What it’s good at, per Fastino’s own benchmark

Fastino evaluated the model on “Fast Decisions,” an internal suite of 17 datasets covering intent routing, triage, sentiment and content understanding. This is Fastino’s own benchmark, not an independently run one. The blog post and company X post give GLiNER2.5-Decide 60.1% average accuracy, ahead of JevK5 (57.5%), SemIf (56.4%) and Laya (46.6%), leading on 9 of 17 datasets. The Hugging Face card’s own results table gives slightly different figures for the same run: 60.2% and JevK5 at 57.6%. Fastino highlights support intent classification at 75.3%, 18.6 points ahead of the next model, and banking intent at 64.3%, 8.6 points ahead.

The GitHub repository behind it, fastino-ai/GLiNER2, had 2,169 stars and was last pushed on 24 September 2026, the day of release.

What an independent benchmark shows

The Decision Index 0.2.1 by multimodalart, updated 28 September 2026, is a broader test than Fastino’s. On a chance-corrected score where 0 is random guessing and 100 is perfect, averaged over 38 benchmarks in five weighted areas, GLiNER2.5-Decide scores 11.21 against Jev’s 57.91. That is the best result among the models under 500M parameters it tested, but far behind the larger open models. JevK5, which Fastino’s suite ranks below GLiNER2.5-Decide, scores 38.81 here. Its expected calibration error, the average gap between stated confidence and actual accuracy, is 0.088 against Jev’s 0.074.

Access and running it

Weights are Apache 2.0 on Hugging Face with no waitlist. Fastino says the model is “efficient enough to run locally on consumer-grade CPUs or in air-gapped environments,” and co-founder George Maloney gives CPU latency around 167ms and GPU latency of 38 to 47ms for short requests. Fastino also runs hosted inference and fine-tuning at http://agent.fastino.ai, for calling from inside a coding agent; no pricing for that service appeared in the sources checked.

What it’s not for

It does not generate text, reason through multi-step problems, or answer open-ended questions; the model card says so outright. No context-window limit, maximum Choice count, or per-call question cap is documented anywhere checked, so those limits are unknown rather than unlimited.

Specifications

Question typesChoiceScoreNoul
Max Choice optionsNot documented
Score levelsNot documented
Questions per callNot documented
Total contextNot documented
State budgetNot documented
Rate limitNone for self-hosted use. Fastino's hosted inference and fine-tuning API at http://agent.fastino.ai did not have published rate limits in the sources checked.
Endpointhttp://agent.fastino.ai
SDKsPython: gliner2

Neither the blog post, the Hugging Face model card, nor the GitHub repo README documented a fixed context window, a maximum number of Choice options, or a cap on questions per call as of this check.

Versions

  • fastino/GLiNER2.5-Decide, 24 Sep 2026, 340M-parameter encoder fine-tuned from fastino/gliner2-large-v1, a DeBERTa-v3-large base per the Hugging Face model card. The Hugging Face API's safetensors metadata reports 486,444,053 total parameters for the published weights, which does not match the 340M figure stated in the blog post and model card; this discrepancy is unresolved as of this check. Apache 2.0 license. Release notes

Use cases

What people use GLiNER2.5-Decide for, one page per pattern.

Workflow controlStarter

Support inbox triage with System One models

Send a support ticket to Jev once with every question attached. Category comes back as a selected label, severity and frustration as numbers on scales you wrote, refund intent as a probability. Your code reads those values and decides what happens to the ticket.

ChoiceScoreNoul
Workflow controlIntermediate

Intent and model routing with System One models

One Jev call reads an incoming request and returns its intent as a label plus a difficulty rating on a scale you wrote. Your router reads both numbers and picks the handler: deterministic code, a cheap model, an expensive one, or a human queue.

ChoiceScore
Real-time and agentsIntermediate

Agent routing and skill selection with System One models

An agent choosing from a long skill roster reads one truncated line per entry and often loads the wrong thing. Jev ranks every entry in one request and separately answers whether any skill applies at all, so the agent gets a short hint instead of a guess.

ChoiceNoul
Data and operationsStarter

Structured extraction with System One models

Jev is not trained to generate text, so it cannot write a value out for you. Extraction works the other way round: a regex or a parser finds candidate spans, a choice question picks the one the question asks for, and code copies that span unchanged.

ChoiceNoul

Examples built with GLiNER2.5-Decide

The most-starred and most-viewed entries in the directory. Browse all examples.