Skip to main content

Canary Deployments

Aegis supports canary traffic splitting — a deployment strategy where a percentage of live traffic is routed to a new upstream version (the “canary”) while the rest continues to the current primary. Aegis monitors error rates and latency for both groups independently and can automatically roll back to the primary if the canary starts failing.

How It Works

  1. You deploy a new version of your backend alongside the current version
  2. Mark the new upstream as canary in Aegis
  3. Set a traffic percentage (e.g., 10% to canary)
  4. Aegis routes traffic according to the split and tracks metrics for both groups
  5. If the canary is healthy, gradually increase the percentage
  6. If the canary fails, Aegis automatically rolls back to 100% primary

Upstream Roles

Each upstream for a proxy host has a role: Roles are assigned per upstream in the host editor. A host can have multiple primary upstreams (load balanced normally) and one or more canary upstreams (also load balanced among themselves).

Configuration

Where to Configure

  • Admin UI → Hosts → edit a proxy host → Upstream section → set upstream roles → Canary Deployment card

Sticky Canary Routing

Without sticky routing, each request independently rolls the dice — a user might hit primary on one request and canary on the next. This can cause inconsistent behavior for stateful applications. With sticky routing enabled, routing is deterministic per client IP:
The same client IP always gets the same routing decision for a given traffic percentage. No cookies, no session state — just a hash of the IP.

Metrics and Monitoring

Aegis tracks metrics independently for primary and canary groups: These metrics are computed over the configured evaluation window (default 5 minutes) and are available in real-time through the admin UI and API.

Live Dashboard

The canary dashboard shows a side-by-side comparison:

Auto-Rollback

When auto-rollback is enabled, Aegis continuously evaluates canary metrics:
  1. Wait until the canary has received at least min_sample_size requests
  2. Compute the canary error rate over the evaluation window
  3. If the error rate exceeds the threshold → rollback
  4. Compute the canary P95 latency (if latency threshold is set)
  5. If P95 exceeds the threshold → rollback
Rollback sets the traffic percentage to 0% immediately. All subsequent requests go to primary upstreams. The rollback is persisted to the database so it survives restarts.

Rollback Alert

When auto-rollback triggers:
  • A warning banner appears in the admin UI
  • The event is logged at WARN level
  • If notifications are configured, an alert is sent

Canary Lifecycle

Typical Workflow

  1. Deploy the new version to a separate server
  2. Add it as an upstream with role canary
  3. Enable canary routing at 5%
  4. Monitor for 10-30 minutes
  5. Increase to 25%, then 50%, then 100%
  6. Promote the canary to primary (swap roles)
  7. Remove the old primary upstream

Manual Actions


Graceful Degradation

If all canary upstreams become unhealthy (health checks fail), Aegis automatically routes 100% of traffic to primary upstreams. This is transparent — no rollback is triggered, and when canary upstreams recover, traffic splitting resumes. If all primary upstreams become unhealthy, Aegis routes to canary upstreams as a fallback (same behavior as standard load balancing failover).

Difference from Load Balancing

Both systems coexist. Within the primary group, load balancing distributes traffic normally. Within the canary group, the same load balancing applies. The canary split happens first, then load balancing routes within the selected group.

API Reference

Canary Status Response

Upstream Role

Upstream role is set in the proxy host configuration: