Beam: Reflection's 501B Open-Weight Model
By Siam Sukkhee Trading Co., Ltd — 2026-10-06 — Hacker News
There's a new model everyone's talking about. Beam, from a startup called Reflection, just launched and it's doing something that's been harder than it looks in AI right now: delivering top-tier performance without top-tier electricity bills.
This is worth paying attention to.
Beam is a sparse Mixture-of-Experts model. That phrase does actual work — it means 501 billion parameters total, but only 23 billion fire up for any single token. The architecture is built specifically for coding, reasoning, and agentic tasks. Reflection trained it on 23.8 trillion tokens pulled from web data and proprietary licensed datasets, which is a proper amount of material to work with.
The reinforcement learning phase is where it gets interesting. They ran 100 million rollouts across 10.5 thousand NVIDIA GB300 GPUs over four weeks. This is genuinely a different scale of investment in making the reasoning pipeline work.
Why does that matter? Because it shows up in the outputs.
On coding benchmarks like SWE-Bench Verified, Beam scores 80.9. On Terminal Bench v2.1, 80.1. On AIME 2026 mathematical reasoning, 97.8. These aren't dominant numbers across every test — Kimi K3 still outperforms it on raw DeepSWE v1.1 scoring, and there are gaps where other models pull ahead. But here's the actual story: Beam delivers reasoning capabilities that match models like GLM-5.2 while using roughly three to four times less inference compute per query.
Less compute means lower cost. Lower cost means it actually gets used.
The model can be prompted to control its own reasoning effort. Lower settings produce shorter responses. Higher settings let it work through harder problems. Users get to pick the tradeoff. Early in reinforcement learning, the system learned to solve tasks more efficiently with fewer tokens. Later it discovered that longer reasoning chains could unlock further performance gains. The policy learned that inefficiency and capability aren't the same thing.
There's a generalization story here too, actually that's not quite right — there's a transfer learning story. During training on reasoning, software engineering, and terminal tasks, the model started performing well on web browsing despite never being trained on browsing tasks. It learned to synthesize broader agentic capabilities that moved between domains. Given web access, it organically figured out how to search, query other language models, and use OCR APIs to read documents. None of that was explicitly taught.
The engineering that enables this to run at scale gets overlooked. Reflection built new algorithms to handle policy staleness — the problem that arises when you're training on interactions generated hours or even days earlier by older versions of the model. They developed stable asynchronous policy gradients that can train productively even when the oldest samples in a batch are 107 weight versions behind the current policy. This matters because running 100 million rollouts at scale means you can't just throw away old data and pretend it doesn't exist.
That's another detail worth noting. The weights are coming out later this month without royalties, usage caps, or restrictions on derivative work. It's the kind of open release that forces other labs to respond.
None of this means Beam is superior to every model in every context. DeepSeek V4.1 Flash still leads on some benchmarks. Qwen and Kimi models have their own strengths. But what Reflection built is a model that trades some raw peak performance for efficiency, and then packages it openly. That's a different competitive move than people have been making.
If the efficiency claims hold — and early benchmarks suggest they do — then companies can run coding agents on hardware budgets they've already accepted as inevitable. That changes the unit economics of deploying models at scale.
Source: "Reflection AI claims Beam matches the performance of leading Chinese open models on advanced reasoning benchmarks at dramatically lower costs." — TechCrunch
Tags: metals trading Thailand 2026