9 min read B2B FinOps

Open Source LLMs Just Got a Cheaper Job: Deciding

Amazon's Strands Decider 2B shows open source LLMs can replace frontier calls on agent routing steps. What that means for your token cost.

Open Source LLMs Just Got a Cheaper Job: Deciding

Searches for “open source llm” are up 196% in three months. That is not a fluke of curiosity. It tracks a real shift in how the people running agent fleets think about money: the frontier model is no longer the default answer for every step in a pipeline.

On October 1, 2026, Amazon Web Services released Strands Decider 2B, an open source decision model built on the Qwen3.5-2B framework. It does not write prose or reason through a hard proof. It answers a narrower question: given a set of already-decided options, which one comes next, and how confident is the model in that choice. It runs locally, it is small enough to sit on a laptop, and it shipped the same week OpenAI announced a comparable offering. TypeSafe’s Jev model kicked off the category, and dozens of research replicas have followed.

For heavy AI users paying hundreds of dollars a month, the interesting part is not the model itself. It is the job it takes away from the expensive models you are currently paying for.

Why open source LLMs are moving down the stack

Every agentic workflow is a chain of decisions. Read the ticket, classify it, pick a tool, check whether the result is good enough, decide whether to retry, decide whether to stop. Only a fraction of those steps need the reasoning depth of a frontier model. The rest are routing and classification, and routing does not require a $50-per-million-token output.

Amazon distinguished engineer Marc Brooker built the first version of Strands Decider after seeing Jev, and his framing is the whole argument: AWS customers’ agentic workflows “did not always require the capability or cost of a fully featured LLM all the time.” The decider gives a workflow step that is more reliable, thanks to calibrated confidence scores and a closed domain of answers, with lower latency and potentially lower cost.

That is the shape of the shift. The open source LLM is not replacing your frontier model at the center of the pipeline. It is replacing the frontier model at the edges, where the task is a multiple-choice question rather than an open-ended one.

The token cost math on a routing step

Take a support triage agent that handles 20,000 tickets a month. Say it makes four classification calls per ticket on a frontier model billed at $10 per million input tokens and $50 per million output tokens, with roughly 1,500 input tokens and 200 output tokens per call.

That is 80,000 calls a month: 120 million input tokens and 16 million output tokens. On frontier pricing, that is $1,200 in input and $800 in output, or about $2,000 a month for the routing layer alone. The same routing logic on a local decider running on hardware you already own costs the electricity to run it, plus the engineering time to maintain it.

Even if only half those steps genuinely need a classifier rather than a reasoner, moving them off the frontier model takes roughly $1,000 a month out of the bill for one workflow. Multiply that across the classification, tool-selection, and retry-check steps in every agent you run, and the savings compound fast.

The catch is accuracy. Brooker names it directly: the challenge is finding the balance where you push performance on accuracy and calibration “without degrading its performance on understanding different languages, on having the kind of knowledge it has, which is what makes it general purpose.” A decider that routes to the wrong branch saves you tokens and costs you a broken workflow. The savings only count if the cheaper model clears your accuracy bar.

A skill worth knowing right now: model routing

The broader lesson from the decision-model wave is that the highest-leverage cost skill of 2026 is not prompt tuning. It is model routing: knowing which step of your pipeline deserves which model.

The pattern you want to build looks like this. A frontier model handles the open-ended reasoning step, the one that genuinely benefits from depth. A mid-tier model handles generation where quality still matters. A small open source model, running locally or on cheap inference, handles the high-volume, low-ambiguity decisions: is this spam, which tool is next, does this result pass, should this loop continue.

This is also where cross-provider discipline pays. If your entire pipeline runs through a single provider, you have no routing lever, and you pay frontier prices for classification work. The moment you can send a step to a local Qwen-based decider, or to a cheap tier from OpenAI, Anthropic or Google, you have a dial you can turn. The decision-model release from Amazon and the parallel moves from OpenAI are notable precisely because they make the cheap tier good enough to trust with real work.

What to do this week

Three concrete moves, in order of payoff.

First, instrument one agent and label every step by task type. Most teams discover that a third to a half of their calls are routing decisions, not reasoning. You cannot route what you have not measured.

Second, pick the cheapest step and try a small model on it. Start with a classification or tool-selection step where the answer space is closed and you can compute an accuracy number. Run it in shadow mode against your current model, compare agreement, and only cut over when the cheaper model matches within a threshold you set in advance.

Third, watch the cost per completed task, not the cost per token. A model that is 10x cheaper per token but fails twice as often and triggers a retry is not cheaper. The decision-model category is a saving only when accuracy holds.

The bigger picture

The open source LLM is being asked to do a narrower, less glamorous job than “replace GPT-6.” It is being asked to make the small decisions that currently cost the most because there are so many of them. That is a healthier direction for the ecosystem than another benchmark race, and for heavy users it is the most direct cost lever to appear in months.

The frontier models are not going anywhere. They are just going to be doing less of the work you were paying them for, one routing step at a time.

Amazon’s Strands Decider is small enough to run locally and fully open sourced, which means the barrier to testing the routing pattern on your own workflow is now your afternoon, not your procurement budget. Given that “open source llm” search interest has nearly tripled in a quarter, plenty of teams are already running the experiment. The ones that come out ahead will be the ones that measured accuracy before they measured the savings.