zokie.BlogChangelogGet Zokie
← All posts

Small model or frontier model? Picking per task to cut AI coding costs

30 September 2026 · 3 min read

Short answer: size each task before you start. Small tasks, such as a version bump, a typo or a one-file fix, go to a small, cheap model. Normal tasks, like a feature in a few files, go to a standard model. Deep tasks, such as schema changes, money logic, security or architecture, go to the strongest frontier model, ideally for planning, with a standard model building from that plan. Using the best model for everything wastes money; using a small model for hard work wastes more, in retries and bugs.

Why it matters

Frontier models cost several times more per token than small ones, and coding agents use a lot of tokens: they read files, run commands and loop until tests pass. If most of your tasks are small, and for most teams they are, running them all on the top model is paying premium prices for work a cheaper model does just as well.

The opposite mistake costs more. A small model on a hard task produces plausible code that is subtly wrong, and you pay for the retries, the review time and sometimes the incident.

A simple sizing rule

SizeSignsModelExamples
SmallOne file, clear change, low riskSmall / fastBump a dependency, fix a typo, rename a variable, update copy
NormalA few files, a known patternStandardAdd an endpoint like the others, a new UI screen, a bug with a clear repro
DeepMany files, unclear design, or high riskFrontier to plan, standard to buildSchema migration, payments, auth, concurrency, a new module

Risk beats size. A three-line change to how refunds are calculated is a deep task. A 300-line change that adds similar pages is normal.

Plan big, build smaller

For deep tasks, split the work: let the frontier model write the plan (the files to touch, the approach, the risks and how to test it), then let a standard model carry it out. The expensive thinking happens once; the long tail of edits and test runs happens on the cheaper model. Check the plan yourself before the build starts.

Signs you picked wrong

  • Too small: the agent loops on failing tests, invents APIs, or asks the same question twice. Stop and move up a size.
  • Too big: a trivial change took minutes and a long explanation. Next time, go smaller.

Doing it automatically

Deciding by hand for every task gets old. You can automate it:

  • Write the rule into your prompts or scripts, for example a task template with a "size" field that picks the model.
  • Use a tool that routes for you. In Zokie, Auto mode (on paid plans) sizes your first message as small, normal or deep on your machine, picks the model to match, plans deep work on the strongest model and builds on the next one down, and says in the chat what it picked and why. Choosing a model yourself always wins.

Whichever way you do it, keep the escape hatch: a person should always be able to pin the big model for a task they know is hard.

FAQ

How much can model routing save?

It depends on your mix of tasks. Teams whose work is mostly small and normal tasks save the most. Measure it: compare spend on the same kind of work before and after.

Is the biggest model always the most accurate?

On hard problems, usually. On simple, well-specified tasks, smaller models often do equally well, faster.

Should I use a small model for code review?

For a first pass on small diffs, yes. For changes to money, auth or personal data, use a strong model and a person; see how to review AI-written code.

Can I switch models in the middle of a task?

Yes in most agents. Moving up a size when a task turns out harder than expected is often cheaper than letting a small model keep retrying.

About Zokie

Zokie is a lightweight agent IDE for macOS and Windows. Claude Code, Codex, Gemini CLI or Cursor CLI write the code, each task on its own git worktree, and you review the diff. It is invite only for now: enter an invite code or join the waitlist.

Read next

  • A desktop app for Claude Code: four ways to use it outside the terminal
  • Claude Code vs Codex vs Gemini CLI vs Cursor CLI: which coding agent in 2026?
  • How to review AI-written code: a 10-point checklist
Zokie·Blog·Changelog·Privacy·Terms