News analysis · Published
Fastino GLiNER2.5-Decide: Shadow-Test Typed AI Decisions
By the ELYMENT AI editorial team · Free to read
Fastino released GLiNER2.5-Decide on 24 September 2026 as a 340-million-parameter, open-weight model that converts text and a user-defined schema into typed answers, probabilities and confidence data. It can run locally, including on CPUs, and is licensed under Apache 2.0. That makes it an interesting component for operational classification and triage, but not an autonomous authority. Businesses should first run it in shadow mode beside existing rules or reviewers, then measure error costs, abstentions and handoffs for each decision class.

What Fastino released
GLiNER2.5-Decide is a compact English-language classifier built for schema-defined decisions. Instead of generating an open-ended response, an operator supplies text plus questions, allowed answer types and optional constraints. The model returns structured answers with probabilities, confidence and constraint metadata. Fastino says it can also extract spans and relations or apply rule checks, while full and LoRA fine-tuning are supported.
The release is designed for local deployment on CPUs or GPUs and for air-gapped environments. Fastino describes it as a 340-million-parameter model, and its Hugging Face model card lists an Apache 2.0 licence. Those attributes can make privacy review, predictable output validation and unit economics easier than routing every classification task through a general-purpose model. They do not establish that the model is accurate enough for a particular workflow.
Treat the benchmark as a hypothesis, not approval
Fastino reports an average exact-match accuracy of 60.1% across 5,100 test examples drawn from 17 datasets, with the model leading the comparison on nine datasets. The vendor also reports median latency of 38.3 milliseconds on an NVIDIA V100 and 167.3 milliseconds on a 48-vCPU Intel Xeon. These are useful engineering signals, but the company labels the exercise an internal benchmark rather than an independent evaluation.
An average also hides the cost distribution. A false positive in a low-risk support queue is not equivalent to an incorrect fraud hold, safety escalation or customer eligibility decision. Confidence is likewise a model output, not permission to act. Teams need class-level precision and recall, a calibrated abstention policy and a human or deterministic fallback before using the result to change real-world state.
Run a shadow-decision acceptance test
Start with one bounded decision that already has a rule, reviewer or recorded outcome. Freeze the schema and context, replay historical cases, then dual-run live traffic without allowing the model to act. The trial should produce an evidence pack that a workflow owner can approve or reject.
- Define the action boundary: specify which outputs may inform, recommend, queue or execute, and which always require human approval.
- Measure by class and consequence: record false positives, false negatives, abstentions, confidence calibration, latency and cost for each decision type.
- Stress the inputs: test missing fields, conflicting instructions, long text, prompt injection, changed labels and out-of-distribution examples.
- Design the handoff: set minimum confidence and constraint checks, route uncertain cases to a named owner and preserve the source text with the decision.
- Prove rollback: version the model, schema and operating context together, log every result and keep the previous rule path available.
Where typed decision models fit
Good candidates are high-volume, narrow and reversible: document routing, lead classification, support triage, policy tagging and moderation pre-screening. Poor first candidates are irreversible or legally consequential actions with weak review paths. The model card also says GLiNER2.5-Decide is not a general-purpose language model and does not provide open-ended reasoning or explanations.
This release is best understood as a decision component, not a complete workflow. Pair the acceptance test with ELYMENT AI’s context release manifest so the schema, instructions and knowledge inputs are reproducible. Use its AI agent approval workflow for consequential actions, and apply the model-routing framework when deciding whether a small local model, a frontier model or deterministic code should handle each stage. ELYMENT AI can help turn that evidence into governed automation without confusing confidence with authority.
Sources
- Fastino, GLiNER2.5-Decide: The First Open-Weight Decision Model (24 September 2026) - Launch announcement, architecture overview and vendor-reported benchmark and latency results.
- Fastino, GLiNER2.5-Decide model card (Updated 25 September 2026) - Model scope, licensing, deployment guidance, structured output format and stated limitations.
- GLiNER2: Efficient Generalist and Specialist Encoder for Schema-Driven Information Extraction (25 July 2025) - Research paper describing the underlying schema-driven information-extraction architecture and compact deployment approach.
Continue learning
Frequently asked questions
What is GLiNER2.5-Decide?
It is Fastino’s 340-million-parameter open-weight model for turning text and a user-defined schema into typed classification and extraction results.
Can GLiNER2.5-Decide run locally?
Yes. Fastino and the model card describe CPU and GPU deployment, including local and air-gapped use, under an Apache 2.0 licence.
Should businesses let the model make decisions automatically?
Not by default. Run it in shadow mode, measure errors and confidence by class, define abstention and human handoff rules, then approve only bounded actions.