Skip to content

A new model is out. Do you really need to switch?

Every few weeks a new model comes out and someone on the team asks the same question: should we switch?

My honest answer is usually “probably not yet”.

When you have an agent running in production, changing the model is a lot more than updating one line of code. You have to test everything again, check that the prompts still behave the way you expect, and make sure nothing breaks in the outputs. All of that takes time and money.

What makes it even harder is that most of the use cases I work with don’t have an automated test set. Someone with domain knowledge has to sit down and review the answers one by one. In healthcare that person is usually a clinician or an expert whose time is expensive and limited. So every migration has a real cost that doesn’t show up in any benchmark.

That’s why I try not to decide based on the launch announcement. I look at three things. How much would the migration actually cost us, counting review hours and the risk of something breaking. How much better would the new model really be for our specific use case, and whether users would even notice. And whether something is forcing us to move, like the old model being deprecated, or the provider raising its price (which happens more often than people think: older models tend to get more expensive once the new one is out).

If the improvement is small, the cost is high and nothing is pushing us, we stay where we are. A lot of tasks, like classifying tickets or extracting data from documents, already work well with models that are a couple of generations old.

← All notes

Dealing with something similar? Email me