Strands Decider 2B: A Small, Open-Source Decision Model
By Siam Sukkhee Trading Co., Ltd — 2026-10-07 — Hacker News
AWS released a new model on October 1st. It's built to never write a single sentence.
That's the opening move in what looks like an entirely different approach to agentic AI. Most of what we talk about—LLMs, chatbots, reasoning models—they're all built to generate text. They can say anything. They're flexible. But that flexibility comes at a cost: latency, compute overhead, and the sheer probabilistic risk of generating nonsense when you actually need a yes or no.
Enter decision models. Strands Decider 2B is AWS's contribution to this emerging category. It's a 2-billion-parameter model that does one job. It looks at a set of options you give it, picks one, scores its own confidence, and returns the result in roughly 115 milliseconds on a local GPU. No text generation. No hallucinations. No faffing about.
The architecture is clean. They started with Qwen 3.5-2B—a standard language model from Alibaba—and surgically removed the head that generates words. In its place: a pointer head of just over a million parameters that scores each option you present against a designated answer position. That's it. A rank-16 LoRA adapter handles fine-tuning. The result is version 19 after rounds of iteration; earlier versions using a different head design performed significantly worse.
On benchmarks, it ranks third among 33 models in its size class for accuracy and calibration combined. For latency on a local RTX 3090, median decision time lands around 115 milliseconds on small tasks. On Apple silicon—an M3 MacBook—it's about 153 milliseconds. Not spectacular, but genuinely fast enough for the kinds of routing decisions agents need to make dozens of times per workflow.
Why 2B parameters, specifically? Because it's small enough to run on hardware developers already own, but large enough to do real work. The model passes every easy task on JevBench without stumbling. Those problems map directly to the rote decisions routing systems need: Is this input about the coffee machine? English or Zulu? Should this action be gated? Simple binary or multi-choice calls.
Actually, that's not quite right—it's not just the rote stuff anymore. The early adopter community is already finding wilder applications: hybrid agents that use LLMs for the hard problems and decision models for the easy ones, reducing both cost and latency. Model routing. Tool selection. Memory management. Guardrails. Game-playing. Maze navigation. The innovation velocity is genuinely astonishing.
The whole thing is open source. Weights on Hugging Face. Training data. Training scripts. Repository on GitHub. That openness matters because it signals something: AWS isn't treating this as a proprietary moat. They're treating it as a problem to solve collaboratively.
The catch is real. Decision models fail at anything requiring complex reasoning or text output. They can't code. They can't summarize documents. They can't play twenty questions with you. They're specialists. But for a system that needs to make a thousand small decisions per day with minimal latency and maximum confidence calibration, they might be exactly what you need.
The model sits inside Strands Labs, which is where AWS puts experimental agentic work that isn't yet ready for the production Strands Harness SDK. That positioning matters. This isn't positioned as the future of all AI. It's positioned as a tool for a specific problem set: fast, reliable, local decision-making inside agent workflows.
Whether it actually reshapes how agents are built will depend on what developers do with it next.
Source: "On a single Nvidia RTX 3090, it returns a decision in a median of about 115 milliseconds." — SQ Magazine
Tags: zinc trading Southeast Asia