Topics
Recent articles

DevOps & Cloud

Half-Open Circuit Breaker Probes in Serverless Workers

Learn how to prevent thundering-herd recovery in serverless workers by coordinating half-open circuit breaker probes through a SQLite-backed Durable Object instead of eventually consistent key-value stores.

Table of Contents5 sections
An illustrative circuit-breaker diagram showing Closed, Open, and Half-Open states, with five concurrent invocations sharing one SQLite-backed Durable Object so one probes and four
Cooldown expiry permits one controlled health probe while a shared lease keeps concurrent callers from overwhelming a recovering dependency.

Managing dependency recovery in serverless architectures requires coordinating multiple concurrent invocations to prevent a thundering-herd problem when a service heals. Quick solution: Coordinate the half-open lease through a single SQLite-backed Durable Object per dependency, keep the network probe outside the transaction, and defer competing callers.

async function acquireProbeLease(storage: DurableObjectStorage, now: number): Promise<boolean> {
  return await storage.transaction(async (txn) => {
    const state = await txn.get("breaker_state");
    if (state && state.status === "open" && state.cooldownUntil > now) {
      return false;
    }
    const lease = await txn.get("probe_lease");
    if (lease && lease.expiresAt > now) {
      return false;
    }
    await txn.put("probe_lease", { expiresAt: now + 5000 });
    return true;
  });
}

A circuit breaker protects downstream dependencies by failing fast while an upstream service is unhealthy. The recovery window introduces a new risk when independent request handlers attempt recovery simultaneously. This race condition happens whenever multiple concurrent invocations share a dependency and coordinate through shared state. A scheduled invocation runs at its configured time rather than automatically fanning out into multiple triggers, but a sudden influx of traffic can create the same competitive pressure.

Understanding the Circuit Breaker States

A reliable serverless circuit breaker relies on three distinct operational states:

A successful probe closes the circuit breaker. A failed or timed-out probe reopens it and schedules another recovery attempt. A cooldown period represents a probe opportunity rather than a guarantee of full recovery.

Choosing a Coordination Boundary

Choosing the right coordination backend is essential for safe probe admission. Eventually consistent storage layers cannot provide the atomic guarantees required for single-probe coordination.

Option When to Use It Trade-off or Failure Mode Recommendation
Workers KV Global read-heavy configurations Eventually consistent, lacks atomic read-modify-write Do not use for transactional lease locks.
SQLite-backed Durable Object Single-region state consistency and transactions Bound to a single location, requires routing Recommended for single coordination points.
External Relational Database Existing enterprise database clusters Adds network latency to serverless invocations Avoid if serverless cold-start and latency matter.

A SQLite-backed Durable Object provides strong consistency and transactional guarantees for a single coordination point. Use a stable Durable Object identity for each protected dependency to manage lease admission in a short transaction.

Managing Recovery and Failure Handling

Leases must include an expiration timestamp so a crashed or timed-out claimant does not permanently block recovery. Callers that do not receive the lease should return a retryable or deferred outcome rather than spinning in a tight retry loop.

On a successful probe, close the breaker and clear the active lease. On a failure or timeout, reopen the breaker with an appropriate delay and jitter. For further reading on operational safety mechanisms, see Self-Expiring Kill Switches and Human-Readable Alerts.

Verification Checklist

Continue Exploring

You Might Also Like

View all articles
Fix pnpm Migration Phantom Dependency Errors
3 min read

Fix pnpm Migration Phantom Dependency Errors

Learn how to diagnose and resolve ERR_MODULE_NOT_FOUND phantom dependency failures when transitioning from flat npm hoisting to strict pnpm in CI/CD pipelines.

Preventing ReDoS in Frontmatter Parsers
2 min read

Preventing ReDoS in Frontmatter Parsers

Learn how to avoid regular expression denial of service vulnerabilities in lightweight Markdown frontmatter parsers by replacing complex regex with bounded line scans.

Nightly Analyst Digest from Slot Activity Logs
5 min read

Nightly Analyst Digest from Slot Activity Logs

Learn how to build a reliable scheduled digest email for high-frequency serverless cron systems, featuring local-timezone logging, fallback raw-stats reporting, and advisory-only recommendations.