The first AI integration in a product often looks deceptively simple: send a prompt, receive text, display the answer. The architecture becomes harder when the product must support different hosted providers, local models, changing capabilities, enterprise constraints, and predictable behavior.
The real boundary is product behavior
A useful abstraction does more than rename one provider's API methods. It defines what the product needs: generate a response, stream tokens, call tools, produce structured data, or create embeddings. Provider SDKs belong behind that boundary.
This keeps the rest of the application focused on user workflows. A document assistant should ask for a grounded answer with citations; it should not need to know which provider names its token limit field differently.
Product workflow
→ AI capability interface
→ provider adapter
→ hosted API or local model
Shared layers
→ policy and validation
→ observability
→ evaluation
→ fallback and error handlingNormalize the contract, not every feature
Providers differ in tool calling, structured output, streaming, context limits, safety behavior, and error semantics. Hiding all differences behind one enormous interface usually produces a misleading lowest common denominator.
A better approach is capability-aware design. Keep the common contract small, advertise optional capabilities explicitly, and let the product choose a supported path. The system can remain portable while still using stronger provider features when they matter.
Configuration should select adapters
Provider choice should come from configuration, not conditionals distributed across controllers and interface components. A central factory or dependency-injection boundary can validate credentials, select an adapter, and expose a stable application-level interface.
This design also makes local or air-gapped deployments realistic. The application behavior stays recognizable even when the underlying model endpoint changes.
Reliability belongs outside the provider SDK
Retries, timeouts, rate limits, response validation, tracing, and cost reporting are product concerns. If each adapter implements them independently, behavior drifts. Shared middleware can apply consistent policies while adapters translate provider-specific errors into a small internal vocabulary.
Log enough to debug the system, but never treat prompts, documents, or model responses as harmless telemetry. Privacy and retention need deliberate policies.
Evaluation makes portability measurable
Swapping models safely requires a representative evaluation set. For a document workflow, that can include answer quality, citation correctness, structured-output validity, latency, and failure recovery. Use those results to check whether the replacement still does what your users need.
What I would build first
- Define two or three product-level capabilities instead of copying a provider API.
- Implement one adapter and a deterministic fake for tests.
- Add schema validation, timeouts, and structured error translation.
- Create a small evaluation set before adding a second provider.
- Add capability negotiation only when a real workflow needs it.
The result
Before switching providers, run the same evaluation cases against both adapters. Check the outputs, latency, and failure handling. Keep the adapter boundary small enough that you can explain what changed and why.
This article describes general engineering principles from my experience and independent study. It does not disclose proprietary architecture, source code, customer information, or confidential details from kW Engineering.