Models

Amazon open-sources Strands Decider 2B as decision models multiply

AWS released Strands Decider 2B, an open-source decision model built on Qwen3.5-2B that topped Jevbench. It ships as dozens of Jev-style clones flood the market.

Amazon releases its own Jev clone as decision models flood the web
Amazon releases its own Jev clone as decision models flood the webAI-generated
By Elena Vasquez5 min read

Updated

Why it matters

  • AWS released Strands Decider 2B, a fully open-source decision model built on the Qwen3.5-2B LLM torso, available now and small enough to run locally.
  • Amazon distinguished engineer Marc Brooker's homebrew take on TypeSafe's Jev briefly topped the Jevbench ranking for its size before AWS productized it via Strands Labs.
  • TypeSafe CEO Diogo Almeida says the wave of clones is 'more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful.'

Amazon Web Services has released Strands Decider 2B, a fully open-source decision model built on the "torso" of an LLM — Qwen3.5-2B — that sorts between pre-decided options and returns a calibrated confidence score instead of generating text. The release landed the same week OpenAI announced a similar offering, and it arrives as dozens of research teams push out their own takes on the format TypeSafe pioneered with Jev.

The model is available now and small enough to run locally. That detail matters as much as the release itself: the entire premise of decision models is that much of the intelligence the market is currently buying — large, expensive, general-purpose LLMs — is overkill for many of the steps inside automated workflows.

The project started as a homebrew experiment by Marc Brooker, a distinguished engineer at Amazon. After seeing Jev, TypeSafe's decision model, he built his own take. It worked well enough to briefly claim the top spot on Jevbench, the ranking for models of its size. Amazon's engineers then cleaned it up and shipped it through Strands Labs, an internal organization developing new tools and protocols for deploying AI agents.

From customer conversations to a product

Brooker says the need came directly out of conversations with AWS customers. Their agentic workflows did not always require the capability, or the cost, of a fully-featured LLM at every step.

"What originally piqued my interest in this class of models was that they make a perfect decider for a workflow step — 'what is the next thing for me to do here, based on where I am?'" Brooker told TechCrunch. He framed the value in concrete engineering terms: "a workflow step that can be structured in a way that is more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and is] lower latency, potentially lower cost."

Those three properties — reliability, latency, cost — are the pitch. A decider that picks from a closed set of options and reports how confident it is can be dropped into an agent pipeline where a frontier model would be slower, more expensive, and less predictable. For companies wiring AI into production systems, the confidence score is the differentiator: it turns a probabilistic component into something a pipeline can gate, retry, or escalate on.

The Jevons logic behind the name

The lineage of the category is worth understanding. TypeSafe named its model Jev after the economist William Stanley Jevons, whose theory holds that the falling cost of something — like computer intelligence — can actually increase demand for it. Decision models are a direct bet on that dynamic: if routing a workflow step costs fractions of a cent instead of cents, and runs in milliseconds instead of seconds, the number of places intelligence gets embedded should grow.

That bet explains the flood of entrants. Since TypeSafe debuted the idea, researchers have produced dozens of similar models. The wide interest is a signal of real demand. It also raises an obvious question: how much of this wave will produce anything durable?

Brooker points to the central technical tension. The challenge is optimizing the model's fast decision-making without degrading the intelligence underneath.

"There is a very careful balance to be found where you want to push its performance on accuracy and calibration on these kinds of tasks, without degrading its performance on understanding different languages, on having the kind of knowledge it has, which is what makes it general purpose and interesting and useful," he told TechCrunch.

Strip too much from the base model and you get a fast classifier that fails on anything outside its lane. Strip too little and you give back the latency and cost advantages that justified the category in the first place.

Why the frontier labs may not own this market

Brooker does not expect the frontier labs to dominate the space. His reasoning is economic: these are smaller markets, and the cost to build something interesting is in the hundreds or thousands of dollars — not the hundreds of millions that frontier training runs require. That low barrier cuts both ways. It explains why so many clones have appeared so quickly, and why few of them may survive.

TypeSafe's own leadership is blunt about the quality of the current field. CEO and founder Diogo Almeida says his company is keeping its head down and improving future models, and he is not impressed by what has shipped so far.

"I get that people think it's a gold rush, but they might be underestimating the difficulty of making the models actually smart," Almeida told TechCrunch. He said he does not see real competition for his company emerging yet.

"The current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful," he said.

The stakes

For AWS, the release is a cheap, low-risk entry into a category adjacent to its core AI business: if agentic workflows become the dominant pattern for enterprise AI, the components that route, decide, and gate those workflows become infrastructure — and infrastructure is where AWS lives. An open-source, locally runnable model also fits customer demand for control over where and how inference happens.

For the broader market, Strands Decider 2B is a concrete test of the Jevons hypothesis at scale. The model is free, small, and fast. If decision models like it get embedded across thousands of pipelines, the total demand for AI inference grows even as the per-call price collapses — which is precisely the dynamic Amazon, TypeSafe, OpenAI, and the dozens of clone builders are now racing to position themselves for.

Original: strandsagents.com

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

166 articles

Related articles

  1. OpenAI's Decisions API chases TypeSafe's Jev in fast agent control
  2. OpenAI Rewrites Its Model Spec Using Public Input From 1,000 People
  3. Amazon Puts $50 Billion Into OpenAI in Sweeping Cloud Deal
  4. OpenAI Launches o3 and o4-mini, Its Smartest Models Yet
  5. OpenAI Brings GPT-5.5, Codex, and Managed Agents to AWS Bedrock

« Previous article