Q: Your company wants automated canary deployments and blue-green testing in Kubernetes, but leadership forbids deploying a heavyweight service mesh (Istio/Linkerd) due to CPU overhead. How do you design and execute a progressive canary rollout using native Ingress-Nginx annotations (canary-weight, canary-by-header, canary-by-cookie)?
Engineering a zero-downtime canary deployment pipeline using native Ingress-Nginx canary annotations, weight-based traffic splitting, and header/cookie matching with zero service mesh dependencies.
Want to master this scenario in a live sandbox? KodeKloud's Enterprise GitOps with ArgoCD & Kubernetes Rollouts covers this exact problem with hands-on terminal drills.
🛠️ Production Runbook & Step-by-Step Resolution
Deploy Dual Stable (Blue) and Canary (Green) Workloads
Maintain two parallel production workloads in the same namespace:
- Primary Deployment: Deployed
payment-stable(v1.0.0) with Servicepayment-stable-svc. - Canary Deployment: Deployed
payment-canary(v1.1.0) with Servicepayment-canary-svc. - Primary Ingress: Configured standard Ingress pointing to
payment-stable-svcfor hostapi.example.com.
Enable Header-Based Canary Testing for Internal QA
Route internal testing requests to the canary without exposing public traffic:
- Canary Ingress Manifest: Created secondary Ingress with identical host
api.example.compointing topayment-canary-svc. - Header Annotations: Added annotations:
nginx.ingress.kubernetes.io/canary: 'true'andnginx.ingress.kubernetes.io/canary-by-header: 'X-Canary'withcanary-by-header-value: 'always'. - Verification: Requests with header
X-Canary: alwaysroute to Green; all other public traffic routes strictly to Blue.
Execute Progressive Weight-Based Traffic Shifting (10% -> 50% -> 100%)
Gradually shift public customer traffic while monitoring error rates:
- Weight Annotation: Updated canary Ingress annotation:
nginx.ingress.kubernetes.io/canary-weight: '10'(routes 10% of randomized requests to canary). - Progressive Ramp: CI/CD pipeline steps through 10% -> 25% -> 50% over 15 minutes while polling Prometheus for HTTP 5xx error spikes.
Promote Canary to Stable & Clean Up Secondary Resources
Complete the release lifecycle and reclaim compute capacity:
- Promotion: Updated
payment-stableimage to v1.1.0. - Delete Canary Ingress: Deleted the secondary canary Ingress manifest, returning 100% of traffic to the stable service.
- Teardown: Scaled
payment-canarydeployment to 0 replicas. - Overhead Benchmark: Delivered full progressive delivery with 0% sidecar CPU overhead.
- Deploy dual stable and canary Deployments with independent Kubernetes Services.
- Use nginx.ingress.kubernetes.io/canary: 'true' with canary-by-header for internal QA.
- Shift public traffic progressively using canary-weight (10% -> 50%) while monitoring Prometheus.
- Promote stable deployment to the new image and delete canary ingress with zero sidecar overhead.