Topics
Recent articles

AI Agents

Testing LLM Fallback Tiers Before Outages

Learn how to keep an LLM fallback tier proven with async shadow traffic, response-shape contracts, and scheduled failover drills.

Table of Contents5 sections
A close-up shot of warmly glowing firewood burning in a cozy fireplace, creating a soothing atmosphere.
A close-up shot of warmly glowing firewood burning in a cozy fireplace, creating a soothing atmosphere.

When a primary language model provider rejects requests due to location or transient overloads, many teams rely on an in-platform fallback tier. Logging response bodies and separating deterministic blocks from transient failures solves the diagnosis layer, but it leaves a critical question unanswered. A fallback that never runs in production is merely a hope. If you do not test the backup path continuously, your first time seeing it fire will be during an outage, which is the worst possible time to discover a configuration bug.

Running a reliable secondary tier requires more than basic routing rules. You need a reliable warming pattern that does not inflate latency or cost, a strict response-shape contract so consumers never know which provider answered, and a scheduled drill cadence with a clear readiness metric. This guide covers how to implement these mechanisms in a serverless environment without breaking your budget.

Fractional Shadow Traffic With Diff Logging

To ensure a fallback model is ready, you need to send it realistic payloads. However, routing user traffic blindly to an unproven secondary provider risks increased latency and unexpected error rates. The solution is fractional shadow traffic.

In a serverless worker, you can mirror a small, sampled percentage of live requests to the fallback tier asynchronously. By using asynchronous execution primitives like waitUntil, the shadow request fires in the background without affecting the user-visible critical path. Both the primary response and the shadow response are logged, and their normalized shapes are compared.

export async function handleRequest(request: Request, env: Env, ctx: ExecutionContext): Promise<Response> {
  const primaryResponse = await callPrimaryProvider(request, env);
  
  // Sample 2 percent of traffic for shadow warming
  if (Math.random() < 0.02) {
    ctx.waitUntil(runShadowFallback(request, primaryResponse, env));
  }
  
  return primaryResponse;
}

async function runShadowFallback(request: Request, primary: Response, env: Env): Promise<void> {
  try {
    const fallbackRaw = await callFallbackProvider(request, env);
    const diff = compareResponseShapes(primary, fallbackRaw);
    if (diff.hasMismatch) {
      await logShapeMismatch(env, diff);
    }
  } catch (error) {
    await logShadowError(env, error);
  }
}

This approach gives you continuous coverage. If the fallback provider alters its error format or response structure, your diff logs catch the mismatch on a quiet Tuesday rather than during a traffic spike.

Response-Shape Adapter Contracts

Different language model providers return unique JSON envelopes, usage metadata, and finish reasons. If your application code directly handles provider-specific payloads, switching to a fallback tier requires branching your business logic everywhere.

To eliminate this complexity, introduce a response-shape adapter. Neither provider’s raw envelope should ever reach your core application. Instead, every provider output passes through a normalization layer that maps data to a single internal structure containing only the generated text, token counts, finish reason, and standardized error classes.

interface NormalizedResponse {
  text: string;
  inputTokens: number;
  outputTokens: number;
  finishReason: 'stop' | 'length' | 'error';
  errorClass?: string;
}

function adaptProviderResponse(raw: any, provider: string): NormalizedResponse {
  if (provider === 'primary') {
    return {
      text: raw.candidates[0].output,
      inputTokens: raw.usage.prompt_tokens,
      outputTokens: raw.usage.completion_tokens,
      finishReason: raw.candidates[0].finish_reason === 'MAX_TOKENS' ? 'length' : 'stop'
    };
  } else {
    return {
      text: raw.choices[0].message.content,
      inputTokens: raw.usage.input_tokens,
      outputTokens: raw.usage.output_tokens,
      finishReason: raw.choices[0].finish_reason
    };
  }
}

By forcing both primary and fallback responses through this adapter, failover becomes a simple routing decision. The rest of your application remains entirely agnostic of which provider fulfilled the prompt.

Scheduled Failover Drills and Readiness Metrics

Asynchronous shadow traffic proves that responses match, but it does not verify that your routing logic actually triggers when the primary provider fails. For that, you need scheduled failover drills.

A cron-triggered background worker can force a small batch of synthetic probe traffic through the fallback tier for a bounded window. The pass bar for these drills is binary. Every probe must succeed, and every response must conform to the normalized shape contract. If a single probe fails, the drill fails. For a related implementation, see Bounded Llm Fallback Chains.

To give the on-call engineer immediate confidence during an incident, combine your telemetry into a single readiness score. This score factors in three underlying data series:

  1. Shadow success rate over the last twenty-four hours.
  2. Shape-conformance rate from differential logging.
  3. Days elapsed since the last successful green drill.

If the drills are skipped or if the shadow diff rate degrades, the readiness score drops automatically. Your incident runbook should require the on-call engineer to verify that this readiness score meets a minimum threshold before manually forcing a failover switch, removing guesswork from high-stress moments. For a related implementation, see Audit Macos System Data Before Deleting.

Bounding Cost and Complexity

Warming a fallback tier and running regular synthetic drills consumes billable API requests. Without cost controls, your reliability engineering efforts can inadvertently trigger a billing incident.

Keep shadow traffic tightly sampled, typically between one and five percent of total production volume. Cap the number of synthetic probes executed during weekly drills to the minimum necessary for statistical confidence. By bounding these operations and scheduling them during off-peak windows, you maintain a continuously verified backup system while keeping expenses entirely predictable.

Continue Exploring

You Might Also Like

View all articles
Debugging Silent Skips in Poll-Based Reply Bots
5 min read

Debugging Silent Skips in Poll-Based Reply Bots

An analysis of why poll-based reply bots fail silently in production due to bare continue statements and fixed lookback windows, with architectural solutions for structured skip reporting.

Managing Agent Worktrees in Git
4 min read

Managing Agent Worktrees in Git

An architectural guide on how to safely manage cleanup and ephemeral lifecycles for parallel AI coding agents in Git worktrees without corrupting shared reflogs.

Bounded LLM Fallback Chains
5 min read

Bounded LLM Fallback Chains

Learn how to build bounded LLM fallback chains that prevent cost overruns, respect rate limits, and stop on billing errors.