Project Name

Scaled to 200+ Microservices Using Consul Service Mesh

Scaled to 200+ Microservices Using Consul Service Mesh
Industry
E-Commerce
Technology
HashiCorp Consul, Envoy Proxy, Grafana, Prometheus, mTLS (Mutual TLS), Canary Deployment Pipelines, and Kubernetes

Loading

Scaled to 200+ Microservices Using Consul Service Mesh
Overview

Our client is a global eCommerce retailer operating across multiple continents with approximately 5,000 employees. The platform processes millions of transactions daily, with traffic surging up to tenfold during seasonal sales events. Over three years, the organisation had aggressively decomposed its legacy monolith into over 200 independently deployed microservices running on Kubernetes.

 

As the architecture expanded, the lack of a unified service mesh to govern secure communication, observability, and traffic management became a critical operational risk, especially during the high-stakes peak shopping seasons that drive the majority of annual revenue.

 

A single slow downstream service could cascade failures across unrelated domains, and diagnosing any incident required manually correlating logs across five separate tools.

Key Challenges

Two hundred microservices communicating without centralised encryption, no cross-service visibility during peak loads, and deployments that hit the entire user base in a single blast radius.

  • No Centralized Service-to-Service Encryption: With hundreds of microservices communicating across the network, there was no consistent mTLS enforcement. Sensitive customer and payment data moved unencrypted between services, increasing security and compliance risks.
  • Limited Visibility Into Inter-Service Traffic: During flash sales and peak traffic, teams lacked end-to-end request tracing across microservices. Identifying the root cause of latency or failures required manually correlating logs across multiple services, delaying incident resolution.
  • Risky All-at-Once Deployments: Every release was deployed to the entire production environment at once. Without canary deployments or traffic splitting, a faulty release could impact all users, leading to costly rollbacks and service disruptions.
  • Inconsistent Service Discovery Across Environments: Development, staging, and production relied on different service discovery methods. Hardcoded endpoints and environment-specific configurations caused deployment failures and made production issues difficult to reproduce.
  • No Built-In Resilience Mechanisms: Services lacked circuit breaking and other resilience patterns. Failures in downstream services often triggered cascading timeouts, resulting in widespread performance degradation during high-demand periods.
  • Fragmented Observability: Metrics, logs, and traces were spread across multiple monitoring tools with no unified dashboard. SRE teams had to switch between platforms during incidents, slowing troubleshooting and increasing customer impact.
Our Solution

Ksolves, an AI-first DevOps consulting services company, deployed a HashiCorp Consul service mesh powered by Envoy sidecar proxies, creating a zero-trust communication layer across the retailer's 200+ microservices.

  • Consul Connect With Automatic mTLS: Deployed Consul service mesh with Envoy sidecars across all Kubernetes clusters. Every inter-service call is now automatically authenticated and TLS-encrypted, securing east-west traffic without modifying application code.
  • Distributed Tracing With Envoy Telemetry: Integrated Envoy tracing with Consul observability to centralize request traces, latency, and error metrics. SRE teams can now identify performance bottlenecks within seconds instead of manually correlating logs.
  • Canary Deployments With Traffic Splitting: Configured L7 traffic splitting to gradually route production traffic to new releases. This enables safe canary deployments, continuous monitoring, and instant rollbacks without disrupting all users.
  • Unified Service Discovery: Standardized service registration and discovery using Consul's catalog and DNS. Dynamic service discovery and health-based routing eliminated hardcoded endpoints while ensuring consistency across all environments.
  • Circuit Breaking and Outlier Detection: Implemented Envoy circuit breaker policies through Consul to limit unhealthy connections and automatically reroute traffic. This prevents cascading failures and improves application resilience during service disruptions.
  • Single-Pane Observability: Built unified Grafana dashboards combining Consul service health, Envoy metrics, and distributed traces. Teams now monitor latency, traffic, errors, and saturation from a single interface, significantly accelerating incident response.

Technology Stack

Category Technology
Service Mesh HashiCorp Consul
Sidecar Proxy Envoy Proxy
Observability Grafana + Prometheus
Security mTLS (Mutual TLS)
Deployment Canary Deployment Pipelines
Infrastructure Kubernetes
Impact

From cascading failures, fragmented observability, and risky deployments to a resilient, secure, and fully observable service mesh delivering reliable performance during peak demand.

  • Zero Mesh-Related Incidents During Peak Sales: The next major seasonal sales event was completed with zero service mesh-related incidents across 200+ microservices, replacing recurring outages caused by poor visibility and unsecured service communication.
  • 50% Faster Root-Cause Identification: Unified dashboards and distributed tracing reduced mean time to diagnose incidents from around 90 minutes to under 45 minutes, enabling faster resolution and minimizing business impact.
  • 100% Encrypted Inter-Service Traffic: All east-west traffic is now protected with mutual TLS, ensuring authenticated and encrypted communication while closing critical security and compliance gaps.
  • Zero-Downtime Deployments: Canary deployments with progressive traffic splitting enabled safe production releases, automated monitoring, and instant rollbacks, eliminating the risks of all-at-once deployments.
  • Eliminated Cascading Failures: Circuit breaking and outlier detection automatically isolate unhealthy services and reroute traffic, preventing cascading failures and significantly improving platform resilience during peak workloads.
Solution Architecture
stream-dfd
Conclusion

Managing 200+ microservices without centralized security, observability, and traffic control created significant operational and business risk for the retailer. Ksolves addressed these challenges by implementing a HashiCorp Consul service mesh with automated mTLS, intelligent traffic management, circuit breaking, and unified observability. The result was a more secure, resilient, and scalable microservices platform, enabling zero mesh-related incidents during peak sales, faster incident resolution, and consistent governance across every new service deployed.

Ready to Secure Your Microservices with a Zero-Trust Service Mesh?

Copyright 2026© Ksolves.com | All Rights Reserved
Ksolves USP