Model routing & provider independence

Use the right model for the work.Do not build the company around the model.

Models improve, prices move, regions differ, providers change policy and some workloads need controls that others do not. The durable architecture is the layer that decides which models are eligible, why they are eligible and what happens when the preferred route is unavailable.

The TEMRIK AI control plane is designed to keep company policy, workflow state, permissions and human authority outside the foundation model so model choice can remain a workload decision rather than an organisational dependency.

capability · privacy · region · latency · cost · tools · structure · availability

Reference routing path

TEMRIK task

Goal · tenant · workflow · required evidence

Policy engine

Tenant
Data class
Task
Region
Security
Capability
Budget
Latency

Model router

Provider A
Provider B
Private model
Customer-hosted / confidential endpoint

Normalised result

Provider-specific response translated into workflow state.

Policy / action

Allow · retry · fallback · escalate · require approval.

Reference architecture. Provider support and portability depend on the actual integrations, APIs and deployment configuration.

First principle

Model independence does not mean pretending every model is interchangeable.

Models differ materially in capability, context handling, tool interfaces, structured output behaviour, latency, price, geography and provider controls. A credible provider-independent architecture preserves the freedom to choose while still respecting those differences.

The model should be selected by the workload.The business should not be rewritten for the model.

Routing policy

Optimisation starts after eligibility.

Cost and speed matter, but an inexpensive fast model is not a valid route if the workload requires a region, security control, tool capability or data policy that endpoint cannot satisfy. TEMRIK's architecture direction is policy-first routing: eliminate ineligible routes, then optimise among the remaining options.

01

Capability

Can the model reliably perform the required reasoning, extraction, coding, vision or tool work?

02

Privacy

What provider controls apply to prompts, outputs, logs and abuse monitoring?

03

Retention

What is stored, for how long, and under which commercial configuration?

04

Region

Where can inference and supporting data processing occur for this workload?

05

Latency

How quickly must the result arrive for the business process to remain useful?

06

Cost

What budget or unit economics apply to this class of task?

07

Context

Does the model support the required input size without unnecessary context exposure?

08

Tools

Does it support the required tool-calling and agent interface for this workflow?

09

Structure

Can it reliably return the constrained or structured output the downstream system expects?

10

Confidential compute

Does the selected endpoint provide the required data-in-use protections where relevant?

11

Availability

Is the model and endpoint available in the required region and service tier?

Six different routing problems

“Model routing” is not one feature.

Fallback

Use a secondary eligible model when the primary route is unavailable or returns a defined failure.

Load balancing

Distribute traffic across equivalent or intentionally pooled endpoints for capacity or resilience.

Capability routing

Send work to a model class selected for the skill the task actually requires.

Cost routing

Prefer lower-cost eligible models when quality thresholds and policy allow it.

Security routing

Constrain model eligibility using workload security requirements before optimisation.

Data-class routing

Route public, internal, confidential or restricted work only to endpoints approved for that data class.

These controls can coexist.

A confidential contract-analysis task might first be restricted to approved regional endpoints, then routed by capability, then fall back only to another endpoint inside the same approved policy set. A public classification task might instead prioritise price and latency.

Policy before provider

Keep business rules outside provider-specific prompts.

When routing logic, approval rights, data classification and stop conditions live only inside one provider's prompt format, changing models becomes a business-process migration. The stronger pattern is to keep durable company policy in the orchestration layer and translate only the model-specific interface.

Architecture directionProvider dependent

Business policy

Persistent rules, authority, data classification and escalation.

Routing policy

Which endpoints are eligible for this specific workload?

Provider adapter

Translate messages, tools, structured-output requirements and response conventions.

Model endpoint

Perform the authorised inference task.

Normalisation

Return a stable result shape to the workflow where practical.

Provider choice

Provider choice is useful only when the business can explain the choice.

One model may be selected for deep reasoning, another for high-volume extraction, another because a customer requires a particular cloud or region, and a customer-hosted model for a workload that should not leave a defined boundary.

Frontier provider

Use when the workload benefits from current high-end capability and approved provider controls.

Efficient provider

Use where high-volume economics matter and quality thresholds remain satisfied.

Cloud-managed model

Use where cloud identity, networking, geography or procurement requirements influence the route.

Private / customer-hosted

Use where architecture, sovereignty, customisation or isolation requirements justify the operational burden.

TEMRIK does not claim universal live support for every model or provider shown conceptually on this page. Integrations must be verified for the specific deployment.

Fallback is not the same as portability

Fallback

A defined alternative route.

When a primary endpoint fails, rate-limits or becomes unavailable, a workflow may retry or move to another approved endpoint. The alternative still needs compatible capability, policy and output behaviour.

Portability

The workflow is not structurally owned by one provider.

Portability means company rules, evidence, identity and action rights can remain intact while the model implementation changes. It does not mean an instant hot swap will preserve identical behaviour.

Failover can be automatic. Trust should not be.

What must be revalidated when a model changes

A new endpoint is a deployment change, not just a configuration toggle.

Task quality
Structured output reliability
Tool-call compatibility
Prompt / instruction behaviour
Context limits
Latency
Unit cost
Region and residency
Retention controls
Safety behaviour
Failure modes
Observability

Microsoft's current model-router guidance explicitly recommends workload evaluation rather than assuming a routed model set will outperform a direct deployment. The same principle should apply whenever an enterprise changes provider, endpoint, model family or routing policy.

Model lifecycle

Models are versioned dependencies with lifecycles.

Providers introduce models, update versions, deprecate older endpoints and alter regional availability. An enterprise architecture should know which workflows depend on which model capabilities and have a controlled path for testing replacements.

01

Register model dependency

02

Record provider + version

03

Map workflows

04

Monitor lifecycle notice

05

Evaluate replacement

06

Approve migration

07

Observe production

08

Retire old route

Free field guide

27 Rules of Peace

Provider independence only matters if authority remains independent too.

Get TEMRIK's free field guide on decision rights, evidence, escalation, playbooks and keeping people in authority as AI becomes more capable.

Model-routing policy

Map which models are allowed to do which work—and why.

Start with one real workflow. Define its data class, security boundary, region, capability requirement, latency, cost ceiling, tool needs, fallback behaviour and human decision rights before selecting the model.

The intelligence can change. The company's authority model should remain.

TEMRIK does not claim universal seamless hot-swapping, identical behaviour across providers or support for every model shown conceptually. Production capability is provider- and implementation-dependent.