What Changed in Model Capability Over the Last Two Years, and What Did Not

What Changed in Model Capability Over the Last Two Years, and What Did Not - editorial illustration

It is tempting to describe every new model release as transformative. Looking back honestly, some changes mattered a great deal for business-development work, and others mattered less than the announcements suggested.

What genuinely improved

The ability to work reliably with long documents, to follow multi-step instructions without losing track partway through, and to produce writing that reads as considered rather than generic, all improved meaningfully. Tasks that were unreliable enough to require heavy babysitting became reliable enough to run with a lighter review pass.

What improved less than advertised

The core problem of a model stating something false with full confidence did not go away. It became somewhat less frequent on well-documented, factual questions, and it remains a live risk on anything where the model has thin or ambiguous source material. No release changed the basic rule that unverified output needs a human check.

The practical implication for a team

The right response to two years of steady improvement is not to relax review discipline; it is to keep the review discipline exactly as strict while letting the volume of work that discipline can cover grow. More gets automated at the front end; the back-end check does not get to shrink.

A short habit for staying current without overreacting to every release

Not every new model release changes how a team should work. A useful filter is asking whether the release changes the answer to a specific, previously tested question, does it change which task size warrants the smaller model, does it change whether a careful multi-step approach is still worth the extra cost, rather than reacting to every announcement as if it demands an immediate change.

A short closing note on expectations

None of this argues that model improvement is unimportant, only that its practical value to a BD team shows up as expanded automation coverage, not as reduced need for review. Teams that internalize this distinction get more benefit from each new release than teams that treat every announcement as a reason to loosen a discipline that was working.

A short closing note on humility about the pace of change

Predictions about how much further model capability will improve over the next two years are worth holding loosely. The more durable lesson from the last two is procedural: build workflows that can absorb a better or cheaper model without a redesign, and keep the review discipline steady regardless of how capable the tools become in the meantime.

A short note on avoiding vendor lock-in through model choice

Building a workflow around a single vendor's specific model, rather than a task-based routing layer that can call different models depending on need, creates a switching cost that grows every month the workflow runs. Teams that keep the routing logic in their own hands, even if it means slightly more setup work up front, retain the ability to move to a better or cheaper option later without redesigning how the whole team works.

This is worth raising directly with any vendor during evaluation: can the underlying model be swapped without disrupting the team's workflow, or is the routing decision locked inside the vendor's own product in a way that ties the firm to whatever choices that vendor happens to make going forward. The answer says as much about the vendor's own incentives as it does about the technology.

A final word on avoiding analysis paralysis

None of this framework is meant to turn every drafting task into a formal model-selection exercise. Most day-to-day work should default to whatever the standard workflow already routes to, and the more deliberate task-by-task thinking described here is reserved for genuinely new task types or periodic reviews, not for every single message a team sends. Overthinking a routine task defeats the purpose of having a sensible default in the first place.

What to watch for going forward

The most useful thing a team can do is not try to predict every future capability change, but keep a lightweight habit of periodically re-testing its own workflows against current tools. A workflow built two years ago on the assumption that heavy human review is needed at every step may be able to safely automate more of that review today, and a team that never re-checks will keep paying for caution it no longer strictly needs. The discipline that does not change is that every client-facing claim still needs a traceable source.

Key takeaways

  • Long-document handling and multi-step instruction-following improved substantially.
  • Confident, unverified claims remain a persistent risk regardless of model generation.
  • No amount of model improvement replaces a human verification step.
  • Improvement should expand what gets automated, not shrink the review discipline.

Questions, answered

What is the short answer on What Changed in Model Capability Over the Last Two Years, and What Did Not?

A grounded look back at which improvements actually mattered for BD work, without overselling the pace of change.

What are the key takeaways?

Long-document handling and multi-step instruction-following improved substantially. Confident, unverified claims remain a persistent risk regardless of model generation. No amount of model improvement replaces a human verification step. Improvement should expand what gets automated, not shrink the review discipline.

How does VIPMarketing approach model comparison?

VIPMarketing focuses on the work around the model: finding accounts that fit, matching buying signals to your past work, drafting in your voice and syncing results to your CRM.