One Provider, On Purpose

Justin Erswell
Photo by Zulfugar Karimov on Unsplash
Why I standardised an entire regulated-industry stack on a single AI vendor, against the 2026 consensus, with my bias declared up front
Let me get the conflict of interest out of the way in the first paragraph, because everything that follows is worthless if I bury it.
My company is part of the Anthropic Claude Partner Network. Anthropic's Claude is the sole, deliberate AI provider across select products we ship. We have a commercial relationship with them. So this is not a neutral buyer's guide, and you shouldn't read it as one. It's an argument, with my bias on the table and, towards the end, a precise list of the things that would make me tear the decision up. If the reasoning doesn't survive contact with that list, ignore me. That's why the list is there.
I'm doing it this way because I'm tired of "why we chose X" posts that are plainly the output of a partnership agreement dressed up as independent analysis. You can smell them. They never tell you what would change their mind, because the honest answer is "a better commission." So I'll show the working, including the parts that argue against me, and you can decide whether the conclusion holds for you. It almost certainly holds differently for you than it does for me, and that difference is the real subject of this post.
Everyone says run a portfolio. They're right.
If you've read anything about AI architecture in 2026, you already know the received wisdom: don't route everything through one model. Run a portfolio. Use the cheap, fast model for high-volume classification and summarisation, the strong reasoning model for the hard problems, the open-weight model where you need to self-host or squeeze the unit economics, and keep your provider abstraction layer thin enough that you can swap any of them out the week a competitor leapfrogs. Model selection becomes a configuration decision rather than an engineering project. Vendor lock-in becomes the enemy.
This is good advice. For most teams, in most domains, it's simply correct. The frontier has fragmented into a handful of genuinely excellent labs whose models trade the lead on a roughly monthly cadence. Whichever one topped your benchmark of choice this morning will have been overtaken by the time your procurement cycle closes. In that world, marrying a single provider is daft. You're locking yourself to a snapshot of something that, by design, won't be the same thing in sixty days.
The canonical example everyone reaches for is the engineering org that buys its AI coding tools one year at a time, refuses long commitments, and reserves the right to switch the moment something better ships. That's a rational policy. It works because two things are true in that context: the switching cost is low, and the downside of a bad output is low. If the model writes a duff function, your tests catch it, an engineer fixes it, and you've lost twenty minutes. Flexibility is cheap to hold and the failure mode is rework.
Hold onto those two assumptions, low switching cost and low downside of a wrong answer, because my whole argument is that in my domain both are false. And when they're false, the maths inverts.
What it costs me to be wrong
I build software for Medical Affairs, the part of a pharmaceutical company responsible for the scientific narrative around a medicine. Publication planning, evidence strategy, the governance of what gets claimed and where, the medical-legal-regulatory review that sits between a draft and anything a human outside the company ever sees. Our suite spans the lifecycle: plan, write, govern, deliver. Four products that hand work to one another in sequence.
In this world, a wrong output is not a dent in someone's user experience. A fabricated citation in a publication plan, a claim that drifts a half-step beyond what the evidence supports, a summary that quietly invents a result: none of these is an annoyance to be tidied up in code review. Depending on where they surface, they're compliance events. Regulatory review exists because the cost of getting scientific communication wrong is measured in patient harm and enforcement action, not in story points.
So when I choose an AI architecture, I'm not optimising for which model tops the reasoning benchmark this quarter. I'm optimising for something else entirely: behaviour I can predict, govern and audit across the whole stack, so I can stand in front of a client's compliance function and defend it. That's a different objective from the coding-throughput case, and it points the other way.
The carrying cost of flexibility nobody prices in
The multi-model orthodoxy underweights one thing in regulated domains: flexibility isn't free. It carries a cost, and in a validated environment that cost is brutal.
Every provider you add is another behavioural surface you have to characterise and defend: a different safety posture, different refusal behaviour around clinical and scientific content, a different prompt-injection profile, different data-handling terms, a different failure mode when it hallucinates, because each one hallucinates in its own characteristic way. On that refusal point, if you've ever watched a general-purpose model decline to engage with a perfectly legitimate oncology endpoint because a keyword tripped a filter, you'll know it isn't hypothetical.
In a domain with qualification and monitoring obligations, every one of those differences is something you must test, document, and keep testing as the models update underneath you. Run four providers and you're not four times more flexible. You're maintaining four divergent behavioural contracts, each drifting on its own release schedule, with your governance layer chasing a moving target in four dimensions at once.
This is the cost the "stay flexible, avoid lock-in" line quietly omits. It assumes switching is cheap, which in commodity coding it is. In my world switching includes re-validation, and the abstraction layer that lets you treat providers as interchangeable is the one thing you can't have when their behaviour isn't interchangeable and you're the one accountable for characterising it.
Consolidation here isn't laziness, and it isn't lock-in anxiety dressed up as principle. It's a deliberate cut to the number of behaviours I have to understand, predict and defend. One provider gives me one behavioural spine running the length of the suite. Plan, write, govern and deliver all reason in a consistent, characterised way, which lets the governance product audit against something stable instead of refereeing four models that each think differently. That coherence is worth more to me than the marginal capability I lose by not always running whichever model won this month.
I bet on the provider
This is the distinction I'd most want a fellow CTO to take away, because it's the one that resolves the apparent contradiction with the multi-model crowd.
They're right that you shouldn't marry a model. Models are snapshots, the lead changes monthly, and commitment to a specific version is a bet that decays. I agree completely. But that isn't the bet I made. I bet on a provider, which is a bet on institutional posture, trajectory and incentives rather than a leaderboard position.
What I was buying was a set of priorities. A lab whose stated reason for existing is to make these systems governable and safe; structured as a public benefit corporation, so the mandate sits in its constitution rather than its press releases; putting real money into interpretability and into the open standards that make agent behaviour auditable. In a domain where governability matters more than being a few points smarter, posture lasts and benchmark position doesn't. Models churn every few weeks. A lab's reason for getting up in the morning churns on a scale of years, if at all. That's the thing I was willing to anchor to.
That reframing mostly dissolves the obvious objection, the "but they'll fall behind on a benchmark" one. I don't need my provider to win every month. I need them to stay at the frontier and keep prioritising what my domain requires. The day those two come apart is the day the bet is wrong, which brings me to the part these posts always skip.
What would make me reverse this
If I can't tell you what would falsify the decision, then I haven't really made a decision. I've made a purchase and written a justification for it. So, concretely, here's what would make me unwind single-sourcing.
If the behaviour became erratic enough to disrupt legitimate work. If safety or refusal behaviour started reliably blocking valid clinical and scientific content, the over-cautious filtering problem at production scale, then the governability advantage would flip into an operational tax I couldn't justify, and the case for diversifying to route around it would become overwhelming.
If a competitor opened a durable, domain-specific gap. Not a benchmark win; those are noise. A real, sustained advantage on the exact thing I optimise for: faithful, auditable reasoning over regulated content that holds its citations. If another provider became clearly better at being governable in my domain, my own argument would force me to move. The logic was never loyalty to Anthropic. It's loyalty to governability.
If single-sourcing became a continuity risk. This one is live rather than theoretical. We're in a period where frontier models can be pulled from availability at short notice; export-control directives have already suspended specific top-tier models this year. Concentration risk is real, and the moment that supply threat outweighed the coherence benefit, prudence would force me to qualify at least a fallback provider, however much it complicated the governance story.
If the commercial relationship ever bent a product decision. This is why I put the partnership in the second paragraph and why I'm repeating it now. The day I catch the partnership influencing a call that should have been made purely on what's right for the client, the arrangement has become a liability to the thing it was meant to serve. Saying the bias out loud, in public, is part of how I stay honest about it.
If you're a CTO weighing your own version of this, that list matters more than the conclusion. The conclusion is mine, and it's contingent. The discipline of writing your falsification criteria down is the part that transfers.
What it actually comes down to
I'm not telling you to consolidate. I'd be a hypocrite to extract a universal rule from a decision I've spent two thousand words insisting is domain-specific.
The multi-model crowd and I don't actually disagree about AI. We optimise for different variables because we're exposed to different downsides. They want capability per pound in a world where switching is cheap and a wrong answer costs a retry. I want governable, consistent behaviour in a world where switching means re-validation and a wrong answer can become a regulatory finding. Hand them my exposure and they'd consolidate too. Hand me theirs and I'd run a portfolio quite happily.
So the real question was never "one provider or many." It's narrower and harder: which variable does my domain punish me for getting wrong? Answer that honestly and the architecture mostly falls out of it on its own. Answer it by copying whatever the loudest engineering blog did this quarter and you'll end up flexible where you needed coherence, or tied down where you needed to move.
Work out what you're actually optimising for. Then be honest, in public, about what would prove you wrong. The second part is harder than the first, and it's the part that earns anyone the right to an opinion.
Disclosure, once more, because it bears repeating: my company is a partner in the Anthropic Claude Partner Network, and Claude is the sole AI provider across some of our products. Read all of the above as the argument of an interested party who has tried to show his working, and judge it against the falsification criteria. That's why they're there.
Prefer to listen?
Share this post