Bourne Forge AI
← Notes
ArchitectureAI

Model Independence and the Portability Layer

How to avoid rebuilding an application every time the preferred model changes.

By Mark Bourne

Introduction

Securing Tool-Using AI Systems covered the risk of a model doing something it shouldn't. This article covers a quieter risk: an application that can no longer change its mind about which model to use.

Internet Architecture Lessons for AI Systems argued that loose coupling wins and that protocols matter more than products. This article is the practical follow-through: what a portability layer actually looks like, what it costs to build, and when skipping it is the right call.

The model you chose eighteen months ago is rarely the model you'd choose today. The question is whether changing your mind costs an afternoon or a quarter.

Model releases, price changes, and quiet capability shifts happen on a timeline measured in months, not years. An application wired directly to one provider's SDK, prompt format, and response shape pays for that convenience later, usually at the worst time to be paying for anything.

The Internet Lesson: Interfaces Outlive Implementations

Nobody rewrites their email client when they change Internet providers. SMTP doesn't care who carries the packets.

That durability wasn't an accident. It came from a deliberate choice to define stable interfaces—protocols—and let implementations underneath them change freely. The interface was the contract. The implementation was replaceable machinery.

AI application code rarely enjoys that separation today. It's common to find provider-specific request formats, response parsing, and even prompt phrasing scattered directly through business logic—the equivalent of writing an email client that only works with one ISP's mail servers.

A Realistic Failure Scenario

A team builds a document-processing product directly against one provider's SDK. Prompts are tuned to that model's specific phrasing habits. Response parsing assumes that provider's exact JSON shape. Tool definitions use that provider's schema format verbatim.

The provider deprecates the model eighteen months later, with six weeks' notice. The replacement model handles the same prompts differently enough that output quality drops noticeably. Fixing it means touching prompt logic, response parsing, and tool schemas throughout the codebase—not because a better option didn't exist, but because nothing in the architecture made switching cheap.

The six-week deadline was the provider's. The six-week scramble was a design choice made much earlier.

Provider Abstraction, Concretely

A portability layer doesn't mean avoiding provider SDKs. It means confining them to one place.

Application
Internal Model Interface
Provider A adapter
Provider B adapter
Provider C adapter
Future adapters

Business logic talks to the internal interface. Each adapter translates that interface into one provider's specific SDK calls, request shape, and response format. Switching providers means writing or swapping an adapter—not auditing every call site in the application.

Common Request and Response Formats

Different providers structure messages, tool calls, and streaming responses differently, even when the underlying concepts are the same.

A common internal format—one shape for a message, one shape for a tool call, one shape for a streamed chunk—means the rest of the application only ever deals with that shape. Translation to and from a specific provider's format happens once, in the adapter, instead of wherever a response happens to get parsed.

Capability Detection, Not Capability Assumption

Not every model supports every feature: vision input, native tool calling, structured output modes, extended context. Assuming a capability instead of checking for it is how a provider swap turns into a silent feature regression.

  • Maintain a capability registry per model or provider
  • Check capabilities before routing a request that depends on one
  • Fail explicitly, with a clear message, when a required capability isn't available

Model Routing

Once providers sit behind a common interface, routing becomes a configuration decision instead of a code change: send reasoning-heavy tasks to one model, high-volume simple tasks to a cheaper one, vision tasks to whichever model actually supports vision this month.

This is also where cost and latency trade-offs get made deliberately, rather than by whichever model happened to be wired in first—a theme worth its own article later in this series.

Prompt Portability Has Limits

Prompts are not as portable as request formats. Models respond differently to the same phrasing, formatting, and instruction style.

Treat prompts as provider-specific configuration behind the common interface, not as a single string reused everywhere. A prompt template per provider, tuned and tested independently, produces better results than one prompt forced to perform equally well across models it was never written for.

Testing Across Providers

A portability layer that has never been tested with a second provider isn't proven portable—it's untested in exactly the scenario it exists for.

  • Run the same golden test suite against every supported provider
  • Compare quality, latency and cost side by side, not just pass/fail
  • Add a new provider by running the suite before it ever reaches production traffic

Data Portability

Portability isn't only about swapping which model answers the next request. It includes whether conversation history, fine-tuned artifacts, and embeddings can move too.

Embeddings generated by one model are generally not compatible with another. If a provider switch requires re-embedding an entire knowledge base, that's a real cost worth knowing about in advance, not discovering during a migration.

Exit Plans

An abstraction layer that was never exercised end to end is a theory, not a safety net.

Write down, and periodically rehearse, what switching primary providers actually involves: which adapters exist, what needs re-testing, how long re-embedding would take, who signs off. The exit plan is worth writing before it's needed—that's the only time it's cheap to write.

The Portability Layer, Piece by Piece

ComponentWhat it decouples
Common request/response schemaApplication code from each provider's SDK shape
Capability registryFeature use from assumptions about what a given model supports
Model routerTask routing from a hard-coded model choice
Prompt templates per providerPrompt quality from a single provider's quirks
Golden test suiteConfidence in a swap from manual, one-off spot checks
Data export pathHistory and embeddings from one vendor's storage format

When Vendor Lock-In Is Acceptable

None of this is an argument for abstracting everything. Excessive abstraction has its own cost: slower adoption of genuinely new provider capabilities, more code to maintain, more surface area for bugs that have nothing to do with portability.

Lock-in is a reasonable trade when:

  • A provider-specific feature is core to the product and has no equivalent elsewhere
  • The application is small enough that a rewrite is genuinely cheap
  • Switching costs are low regardless, because usage or stakes are minimal
  • The team lacks the capacity to maintain an abstraction layer well

The goal isn't avoiding dependency. It's choosing dependency deliberately, instead of accumulating it by default.

Trade-offs and Anti-Patterns

  • Provider SDK calls scattered through business logic instead of confined to adapters
  • Assuming a capability instead of checking for it, then discovering the gap in production
  • One prompt forced across every provider instead of tuned per adapter
  • An abstraction layer that's never been tested against a second real provider
  • Building a full portability layer for a tool where lock-in was always the sensible choice

Final Thoughts

The Internet didn't become durable by avoiding dependency on any particular hardware or company. It became durable by deciding, deliberately, which interfaces had to remain stable while everything behind them was free to change. AI applications need the same deliberate line drawn—not because any provider is untrustworthy, but because the landscape underneath will keep moving regardless.

Lock-in by default is an accident.

Lock-in by choice is a decision you can defend.

A Practical Checklist

Before committing to a provider more deeply than intended, work through these questions:

  1. 1Is provider-specific SDK code confined to one thin layer, or scattered through the application?
  2. 2Do you detect model capabilities (vision, tool use, structured output) rather than assuming them?
  3. 3Can a new model be added by writing one adapter, not by editing business logic?
  4. 4Does a golden test suite run against every supported provider before a swap ships?
  5. 5Can conversation history and embeddings be exported in a provider-neutral format?
  6. 6Is there a written runbook for switching primary providers, tested at least once?
  7. 7Where you've deliberately used a vendor-specific feature, is that decision documented as a conscious trade-off?
  8. 8Have you weighed the ongoing cost of the abstraction layer against the risk it protects against?

About This Series

This article is part of the AI Infrastructure & Architecture series on Bourne Forge AI. It follows Securing Tool-Using AI Systems by turning loose coupling from a principle into a buildable layer, ahead of the next article on human-in-the-loop as an architectural component.

More from the Notes

Short technical notes and observations, written up as experiments produce something worth documenting.

Back to Notes

Was this useful?

…

Published