How many tasks does it take to trust a cheaper model?
Bharath Bhat9 min read
Say you want to know whether routine codebase exploration — "where is X handled?", "how does Y work?" — can move from Opus 5.5 to DeepSeek v4.1 Flash. In this post, we discuss what it takes to run such a study.
- The method: replay real tasks from your own sessions on both models, and grade each against outcomes like merged PRs or reference answers.
- What it takes: 30 tasks and about $55 pins the gap to ±6 points (assuming a 0–100 point scoring scale).
- The payoff: if DeepSeek's quality on routine exploration holds up within what you'll accept, a routing rule sends that work to DeepSeek, at 2.5% of what it cost on Opus.
This follows on from what it would take to match model intelligence to the task.
Why this is hard
Most platform teams already suspect that a cheaper model would do fine on a lot of their work. What they lack is evidence that holds up when a team says "we need Opus for this." Public benchmarks are someone else's repositories. Budget caps and switching models off work, but they create friction, and they say nothing about which work actually needed the expensive model.
The evidence that would settle it already exists: every coding agent session is a record of a real task, on your code, with what the engineer actually accepted at the end. Replay those tasks on a candidate model, grade the result against what shipped, and you have a comparison on your own workload.
The hard part is making that comparison trustworthy without making it expensive. Replay three tasks and the number means nothing; replay every task and you've spent a fortune proving a point. So: how many tasks do you need? Which ones? Is a replay a fair test of the model? Is the resulting number meaningful enough to act on? And what does it cost?
How many tasks
The decision you're trying to make is: "How much worse is a cheaper model like DeepSeek going to be on our workload, if at all?" You probably have a number in mind below which the switch doesn't make sense. Call it the margin. To decide, you need to measure two things:
- The gap: the mean difference in score between your current model and the candidate, on the same tasks.
- The 95% confidence interval on that gap: how far the true gap could be from the one you measured.
The gap is whatever it turns out to be. The interval is what depends on how many tasks you replay, so that's what sets the number of tasks needed — call it k.
Measuring the gap
We replay the current model too, on the same frozen tasks, and look at the per-task difference:
dᵢ = scoreᵢ(candidate) − scoreᵢ(baseline)
Two important details:
Replay the baseline; don't compare against 100%. A replay of Opus doesn't score perfectly either. The sandbox, the brief and the grader all add noise, and that noise has to be on both sides of the comparison.
Compare task by task. The same tasks are hard for every model. Comparing on the same tasks cancels much of that shared difficulty, so the gap is measured more tightly than two separate averages would allow.
How wide the interval is
The precision of the gap comes down to one formula:
half-width h = 1.645 · s_d / √k
s_d is the standard deviation of the per-task differences. We use 1.645, a one-sided 95% bound, because the only question is whether the candidate could be worse: "at most X points worse than Opus." Scores are on a 0–100 scale, so the gap is in points, not percent.
s_d isn't known before you run anything. We plan with s_d = 20 points, which gives ±6 points at 30 tasks.
The interval moves with the result
Thirty tasks buy you a width, not a verdict. What you can say after the run is:
With 95% confidence, the candidate is at most (observed gap + half-width) points worse than the baseline.
How that plays out depends on where the candidate lands. With 30 explore tasks, so ±6:
| Observed gap vs. Opus | What you can say (95%) | Margin 10 pts | Margin 5 pts |
|---|---|---|---|
| 0 | at most ~6 worse | passes | not yet |
| −2 | at most ~8 worse | passes | not yet |
| −6 | at most ~12 worse | not yet | likely fails |
| −20 | at most ~26 worse | clearly fails | clearly fails |
The margin is a business decision, not a statistical one. Once it's set, the table shows the pattern that drives cost. Clear wins and clear losses are cheap to prove. Close calls are expensive. Settling a candidate 2 points behind against a 5-point margin takes about 120 tasks, not 30.
That points to the right way to run this: start with a first batch, look at where the gap lands, and only extend the suite when the answer is still open. Because tasks are frozen and the baseline is reused, extending is incremental.
Which tasks
From sessions to tasks
A coding agent session is rarely one task. It might start with exploring the codebase, move to a plan, then an implementation, then a review of the result. We split each session into contiguous blocks of turns, one per task, and tag each block with:
- Task type — explore, plan, implement, review, debug, testing, doc, email, and so on.
- Complexity — trivial, routine, substantial or open-ended.
Both come from a calibrated LLM classifier that reads each user message plus a few turns of prior context. The taxonomy is configurable. In our own team's sessions this produced ~1400 task blocks across ~400 sessions. This first step also gives us an audit of the current model mix, allowing us to answer questions like "How much are we spending on review tasks today, and which models are we using for them?"
Picking k tasks
The next step then is to run evals on a particular slice you care about. Say explore tasks, and you know from above that you need k tasks to get an interval in a range that you are comfortable with. Before a task can go into that suite, it has to pass a few checks:
- It can run in a sandbox. The work happened in a git repository we can check out at the base commit, a Docker image builds at that commit, and any data the session pulled from external systems can be staged as files.
- There is a reference to grade against. The merged diff for code tasks, or the final answer for analysis tasks.
- The prompt doesn't give away the answer.
- The prompt and the reference together form a well-specified task that can actually be completed.
From the tasks that pass, we pick k spread across complexity, repository and team. A rule will likely end up as type × complexity ("routine exploration to DeepSeek, substantial stays on Opus"), so each bucket needs enough tasks to stand on its own.
The chosen tasks are then frozen: the brief, the reference and the sandbox inputs are stored once and never change, so every run of a task sees identical inputs, and a model added next month is compared on exactly the same work.
Replaying a task
Each replay runs in a fresh sandbox. We use Harbor to orchestrate and Modal to execute, across Claude Code, Codex and OpenCode.
- The repo is checked out at the task's base commit with
git archive. No git history ships into the sandbox, so the shipped solution is unreachable. - Prior context is supplied. Most tasks start mid-session. For harnesses that can resume a conversation, the earlier turns are loaded as real history. For the rest, and for cross-harness comparisons, a model writes a short "context so far" summary instead.
- The instruction is one self-contained brief distilled from the user's messages for that task. It front-loads every decision the user made along the way, so the agent doesn't stall waiting for input that won't come. The instruction is written by a model that looks at the session transcript, and must survive multiple rounds of verification for information leakage and completeness.
- External tools are mocked out, in a sense. If the original session queried an observability tool or an MCP server that the sandbox can't reach, the results it got are handed to the agent as given inputs.
- Skills come with the repo. Anything checked into the repository, including agent skills, is present at the base commit.
This is a modified one-shot, not a simulated user. It is the biggest simplification in the pipeline, and we come back to it under limitations.
Grading a replay
Each task type is graded by a small set of weighted checks written specifically for that task type. We rely on deterministic checks when possible to act as gates, and then on a panel of LLM judges, either scoring against a rubric or checking claims against a reference.
Implement.
- Gates: the patch is non-empty, and the tests that passed at the base commit still pass (pass-to-pass).
- Scored: a rubric judge compares the candidate's diff with the merged PR (weight 3). New tests must fail on the base commit (weight 1).
- Recorded but not scored: whether the right files were touched.
Explore.
- Run the same task through a frontier model that acts as the oracle (one time) to establish ground truth claims, each with its file or symbol citations, and freeze that as the answer key.
- Measure precision and recall for the replay answer from the candidate model against this reference.
- Citations are checked against the real repository, so a fluent fabrication can't pass.
Review.
- Collect all the review comments made on the PR being reviewed in the task.
- For each, check if the comments were addressed:
- In a follow-up commit.
- Acknowledged as important, but perhaps punted to a follow-up PR.
- Measure replay review comments against this golden set. Calculate precision and recall.
Plan and writing. A rubric judge compares the candidate with the recorded deliverable. For writing, substance and form are scored separately.
What it costs
Before the run. A replay's cost can be estimated from the original session: its token counts, the candidate's prices, and a correction for how many more or fewer tokens that model tends to use. There's one extra factor we didn't expect. A single-shot brief is not the original interactive session, so a replay can cost anything from a twentieth of the original (long, back-and-forth planning sessions compress well) to about double (exploration replays wander a bit more). That factor is learned per task type from past replays. Until there's enough history, an honest estimate is a range.
The judge is most of the bill for cheap models. A panel of three judges costs roughly $0.20 to $2 per task, depending on the task type and the length of the transcripts. DeepSeek v4.1 Flash costs about $0.02 per exploration task. So the price of evaluating a cheap model is mostly the price of grading it, which makes judge efficiency, such as skipping the third judge when the first two agree, the main lever on evaluation cost.
A worked example. Take the 30 exploration tasks from the top of this post. Replaying them on Opus 5.5 costs about $0.90 per task, plus about $0.47 to grade, so roughly $40. That baseline is paid once: every later candidate is compared against the same frozen replays. DeepSeek v4.1 Flash costs about $0.02 per task plus the same grading, so roughly $15. That's $55 for the first comparison, and about $15 for each new model after that. The 2.5% comes from pricing: exploration is about 96% cache reads, which DeepSeek charges $0.006 per million tokens for, against $0.20 for Opus.
Look back, then look forward. Once a candidate clears the margin, we report two savings numbers:
- Look-back: what the tasks that match the rule cost over the last month, repriced at the candidate's rates. The model-mix audit already has those token counts.
- Forward: once a routing rule is live, we show what the routed traffic actually cost, versus what it would have cost on the default model.
Then confirm on live traffic. Replay gives you confidence before anything changes. Routing a slice of real traffic for the rule to the candidate, and watching outcomes such as PRs merged, rework and follow-up turns, confirms it on live work and catches any drift over time.
Limitations
- Single-shot, not a simulated user. Interactive sessions are compressed into one brief. That's fair for comparing models, but it isn't the session the engineer actually had, and it works better for some task types than others. We do plan to explore simulated users in the Harbor framework in a future iteration.
- Judges are noisy. Panels and claim-level grading narrow the noise, and the intervals above include it, but they don't remove it.
- External tool calls are not true mocks currently. Data from tool calls made to external systems is captured as files on the sandbox. This makes the task somewhat easier on replay, as the agent can just read a file rather than having to figure out an exact query. And if the replay asks a question the original session didn't, there's no captured answer for it.
- Complexity labels rate the ask, not the work. "Quick fix" requests that turn into afternoons get labelled routine.
- Tasks from the same session aren't independent. A suite drawn heavily from a few sessions has a wider real interval than the formula suggests.
If you'd like to see these numbers on your own sessions, reach out at founders@tuneloop.io.