All newsArticle

Your AI Stack: How We Choose the Right Model for Each Job

Your AI Stack: How We Choose the Right Model for Each Job

The right AI model for writing a production script is almost never the right model for debugging a custom analytics dashboard. Picking the wrong one wastes time, produces mediocre output, and sometimes makes a mess a human has to clean up. At Mainstage, we treat model selection the way we treat any other production decision: match the tool to the job, test it against real work, and let results drive the choice.

Simon Willison, one of the most trusted practical voices in applied AI, frames this well in his opinionated guide to using large language models: there is no single best model, only the best model for a specific task at a specific moment. We agree. Here is how that plays out in our actual studio work.

TL;DR: We run a short internal evaluation for every new model that matters. We test it against real studio tasks, scripting, web dev assist, voiceover brief writing, and analytics work, before it touches a client project. Right now our stack is split across a few models, each earning its place on merit.

Why We Do Not Have One "Official" AI Tool

The landscape shifts fast. Claude Opus 5 landed with meaningfully stronger reasoning and long-document comprehension. Kimi K3 brought competitive coding performance at a price point that makes it worth testing for high-volume dev tasks. New capable models arrive regularly, and locking into one platform because it was best six months ago is a production mistake.

Our working principle: every model is on probation. It earns a role in the stack by performing well on tasks that actually matter to our clients, not by scoring well on benchmarks we will never run in practice.

The Four Task Categories We Evaluate Against

When a new model arrives, we run it through the same four categories. These map directly to where AI touches client work at Mainstage.

1. Script and Copy Writing

This is where tone, voice consistency, and structural instinct matter most. A brand film script needs a clear narrative arc, not just good sentences. A corporate training module needs a specific learner voice. A voiceover brief needs to communicate feel and pacing to a voice actor who has never met the client.

What we test: give the model a creative brief we have actually worked with, ask it to produce a first draft, then evaluate how much rewriting a senior producer needs to do before it is usable. We also test whether the model holds a brand voice across a long document or drifts toward generic language by page three.

Our current finding: models with strong reasoning and long-context fidelity, like Claude Opus 5, perform better here than models optimized primarily for speed or code. The difference shows up in the second half of a long script, where weaker models start losing the thread.

2. Web Development Assist

Our AI-accelerated web design and development work involves real code: custom builds, analytics dashboard integrations, and client-owned systems that need to be maintainable after we hand them off. AI assist is a force multiplier here, but only if the output is clean and the model can reason about architecture, not just autocomplete syntax.

What we test: a representative component build, a debugging task on a real codebase, and a refactor request where the model has to understand intent before touching anything. We look for how often the model introduces new bugs while fixing the one it was asked to fix.

Our current finding: Kimi K3 has earned a real look here. Its coding accuracy on mid-complexity tasks is strong, and the cost structure makes it practical for higher-volume dev work. We run it alongside a senior developer who reviews every output before it goes anywhere near a client repo.

3. Analytics and Data Interpretation

We build custom analytics dashboards that clients own outright. Part of that work involves helping clients interpret what the data is actually telling them. We use AI to assist with pattern recognition, anomaly flagging, and plain-language summaries of complex reporting.

What we test: give the model a synthetic dataset with a known pattern, ask it to describe what it sees, then check whether it flags the right signals or chases noise. We also test how confidently it states things it does not actually know, because overconfidence in analytics is expensive.

Our current finding: reasoning-heavy models outperform faster, smaller models here by a meaningful margin. Calibrated uncertainty matters. A model that says "this could indicate X or Y, and here is what to check" is more useful than one that confidently picks the wrong explanation.

4. Voiceover Brief Writing

This one surprises people. Writing a good voiceover brief is a specific craft. It has to communicate pacing, emotional register, emphasis, and character to a voice actor in plain language. Vague briefs produce takes that require multiple rounds of correction. Good briefs get it right fast.

What we test: give the model a finished script and a client brand profile, ask it to write the brief, then have a producer score it on specificity and usability. We also check whether the model understands the difference between "warm and conversational" as a style note and an actual direction a voice actor can act on.

Our current finding: this task rewards models that can shift register and understand craft context. Willison's point about models having different "characters" in their outputs is visible here. Some models write briefs that sound like they were written by a project manager. The better ones write like a producer who has been in a recording booth.

How We Make the Final Call

After testing, we assign each model a primary role and a set of tasks it does not touch. We are not running a research lab, we are running a production studio. The goal is a stable, reliable stack where every team member knows which tool to reach for and why.

We also track where AI saves real time versus where it creates review overhead that cancels out the gain. If a model generates output that requires more correction than writing from scratch would take, it is not in the stack, regardless of how it benchmarks.

Willison's broader argument, that the best way to use these tools is to stay curious, test constantly, and hold opinions loosely, matches how we actually work. Our stack today is not our stack from six months ago. That is not inconsistency. That is the job.

What This Means for Clients

You do not need to think about any of this. That is the point. When you bring a project to Mainstage, the model choices, the prompt engineering, the review loops, all of it is our responsibility. You get the output: a tighter script, a cleaner site, a dashboard that actually tells you something, a brief that gets the voice right on take two.

We use AI to move faster and to handle the parts of production work that benefit from scale. Humans at Mainstage craft every final decision. The two things are not in tension. Used right, AI makes the human judgment more valuable, not less.

If you want to see how a producer-led AI workflow plays out in a real project, explore our web and development work or see how we approach brand film and video production. Or just book a call and we can walk through it with your specific brief in hand.

Have a project worth telling?

Let's produce something worth watching.