Back to Journal

Stop Chasing Every New AI Model. Do This Instead

A few weeks ago, three or four AI labs discreetly released model upgrades that were noteworthy enough to make their own news cycle. You missed at least one of them…

What this looks like in practice

A few weeks ago, three or four AI labs discreetly released model upgrades that were noteworthy enough to make their own news cycle. You missed at least one of them if you blinked. That is no longer an exaggeration; tracking websites now estimate that the rate of significant model releases is around four times faster than it was in 2023, and it doesn’t seem to be slowing down.

Keeping up with each release seemed like due diligence for a time. It’s beginning to feel like a full-time job that you didn’t sign up for. The point that no one explicitly states in the model-comparison threads is that it shouldn’t be a full-time job for the majority of startups. It’s not the founders who have read every release note that are quietly making progress this year. They are the ones who developed a system for determining, once and rationally, what to employ where instead of chasing the leaderboard.

Why “which model is best” is the wrong question

If you ask ten people which model is the best right now, you’ll get ten different answers because the question is useless without context. Excellent at what? What is the best price? Ideal for an activity that must be completed once, as opposed to one that must be completed a thousand times a day?

Right now, the model race is actually three races going on at once: a distribution battle, a pricing war, and a speed race. Open-weight models that you can run on your own hardware, closed frontier models that strive for sheer reasoning excellence, and everything in between that compete on cost per token as much as functionality are all constantly being released. You wind up rebuilding your stack every six weeks for insignificant gains that no one using your product would notice if you treat this as a single scoreboard to climb.

It is not a better question to ask “which model is best.” It asks “which model is best for this specific task, at this price, with these privacy requirements” for each repetitive task that your team or product performs.

A framework that actually holds up

  1. Make a list of your top five repetitive jobs. Not speculative future features, but the things your team or product accomplishes on a daily basis. Triage for customer service. code examination. summary of the contract. Put them in writing, whatever they may be.
  2. Align every assignment with a model on four axs rather than just one. Along with speed, context length, and whether the output requires multimodal assistance (text, image, audio, and video), reasoning quality is important. A work that runs overnight in a batch job has different criteria than a task that must feel instantaneous to the user.
  3. Determine if it must be open-weight or closed. The majority of the field is eliminated before “quality” ever comes up if the task involves sensitive data that you are unable to provide to a third-party API. In 2026, privacy-conscious consumers and regulated sectors are asking this topic earlier rather than later in sales talks, making it a bigger deal than it was two years before.
  4. Set the price at your real volume rather than the volume of a demo. When you run a model thousands of times a day, it’s a bad trade if it performs slightly better on a benchmark but costs three times as much every call. Instead of using the volume you’re at now, do the arithmetic using the volume you’ll actually reach in six months.
  5. Establish a review schedule and cease checking in between. Every few weeks, a quick assessment is sufficient to identify any changes in pricing or new options that merit testing. Save the more in-depth discussion on “should we actually switch” for once every quarter. Anxiety disguised as diligence is something that occurs more frequently than that.

The mistake this framework fixes

The founders who spend the most time on this are not being lazy; rather, they are acting in the opposite way. When they notice a new release, they think it might be better, so they spend an afternoon comparing it to their present configuration “just in case.” You’ve lost weeks you’ll never get back when you multiply that by each significant release this year, usually for a change that wouldn’t have affected a single business statistic.

New releases are not being ignored by the fix. It’s determining ahead of time when you’ll assess them, so the launch of a new model doesn’t necessarily mean that you should abandon your work. Even if it’s trending, it can wait if it’s not your quarterly review or your two-week check-in.

What this looks like in practice

Let’s say your top five tasks include responding to customer service inquiries, searching internal documentation, reviewing code, creating a tool for summarizing research, and testing a voice feature. That’s probably five distinct models, not just one. Due to the enormous volume and low stakes each message, support replies may be quick and inexpensive. Given that it takes engineering time to identify a poor recommendation, code review may support a more robust reasoning model. While the other features don’t require multimodal support, the voice feature clearly must.

A new model release becomes a five-minute query, such as “does this beat what I’m using for task three, at the price I need?” after that mapping is created, rather than a whole day of unrestricted exploration.

This is the real unlock. failing to locate the ideal model. constructing a system that turns “best” from an existential question into a quick, tedious, task-specific one.


Send this framework to the team member that opens a new tab each time a model shipping if it saves you an afternoon. Next week: what your product truly changes when you switch out a model, as well as the three things you should test first.

Get the next issue

One email, every issue. No spam, unsubscribe anytime.