Project Name

70% Reduction in Unplanned Downtime for a Global Logistics Platform with HashiCorp Consul Health Checks

70% Reduction in Unplanned Downtime for a Global Logistics Platform with HashiCorp Consul Health Checks
Industry
Logistics and Supply Chain
Technology
DevOps

Loading

70% Reduction in Unplanned Downtime for a Global Logistics Platform with HashiCorp Consul Health Checks
Overview

A mid-market logistics platform coordinating shipments across 40+ countries for shippers, carriers, and third-party providers was detecting service failures the wrong way: customers were calling before engineers knew anything was wrong. Service instances degraded or failed without triggering automated alerts. On-call engineers followed manual runbooks, averaging 25 minutes from detection to escalation. Load balancers lacked real-time health signals so traffic continued routing to unhealthy instances until someone manually drained them. Each service team defined health differently, and monitoring was split across five disparate tools with no single source of truth.

 

Ksolves deployed a layered health-checking framework anchored on HashiCorp Consul, replacing every manual detection step with automated, policy-driven health signals and cutting unplanned downtime by 70% in the first quarter post-rollout.

Challenge
  • Silent Instance Failures Caught by Customers, Not Engineers: Service instances would degrade or fail without triggering any automated alert. The first indication was typically a customer opening a support ticket, delaying incident response by up to 25 minutes while the on-call team scrambled to confirm the issue.
  • Manual Runbook-Driven Detection Too Slow: On-call engineers followed static, checklist-based runbooks to identify unhealthy nodes, averaging 25 minutes from detection to escalation. By that time the blast radius had already expanded to downstream services and affected multiple shippers.
  • No Automated Failover Routing: Even when an unhealthy instance was identified, traffic continued routing to it because load balancers lacked real-time health signals from the service layer, forcing engineers to manually drain and restart instances during every incident.
  • Inconsistent Health Definitions Across Services: Each service team defined health differently, with some checking HTTP status only, others using ICMP ping, and several having no health endpoint at all, making it impossible to build a unified observability layer across the platform.
  • Fragmented Monitoring with No Single Source of Truth: A patchwork of cron-based scripts, Nagios checks, and CloudWatch alarms left the team piecing together signals from five different tools during an incident, delaying root-cause identification and exhausting the on-call rotation.
  • Stale Health Data Driving Service Mesh Routing: As the platform adopted service mesh patterns, the absence of native health-check integration with Consul meant sidecar proxy routing decisions were based on stale information, compounding the blast radius of each partial failure across the service graph.
Solution

Ksolves deployed a layered health-checking framework anchored on HashiCorp Consul, replacing every manual detection step with automated, policy-driven health signals. The governing principle was that no traffic should ever reach an instance that cannot serve it correctly.

  • Layered Health Checks (HTTP, TCP, and Script-Based): Triple-mode health probes were implemented on every registered Consul service, combining shallow HTTP endpoint checks, deep TCP port verification, and custom script-based probes validating application logic, ensuring no failure mode from network partition to application deadlock went undetected.
  • Automated Failover with Consul Service Discovery: Unhealthy instances are automatically deregistered from Consul's service catalog within seconds of a failed health check, with upstream routing tables updating in real time so traffic is seamlessly redirected to healthy nodes without human intervention.
  • Consul Agent Deployment via Terraform: Consul agents were shipped via Terraform to every service host with standardised health-check definitions baked into the provisioning pipeline, eliminating the per-service inconsistency that had previously blocked unified observability.
  • Centralised Health Dashboard and Alerting Pipeline: All health-check telemetry is streamed into a single Prometheus console with alerting rules routed through PagerDuty, giving the on-call team one pane of glass instead of five with precise instance-level context on every alert.
  • Script-Based Health Checks for Business-Logic Validation: Custom Consul health scripts validate actual business flows including round-trip database queries and end-to-end order-status lookups, catching logic-layer degradations that TCP and HTTP checks alone would miss.
  • Consul Connect Service Mesh Integration: Consul Connect was configured as the service mesh data-plane, tying health-check status directly to sidecar proxy routing decisions so traffic never reaches an instance already flagged as unhealthy, containing blast radius at the network layer.

Technology Stack

Category Technology
Service Discovery HashiCorp Consul
Monitoring and Alerting Prometheus, PagerDuty
Agent Deployment Consul Agent
Failover Automation Automated Failover Pipeline
Infrastructure as Code Terraform
Service Mesh Consul Connect
Results: 70% Less Downtime, Incident Response from 25 Minutes to Under 5, Zero Customer-Reported Outages
  • 70% Reduction in Unplanned Downtime: Unscheduled outages dropped from 8 incidents per quarter averaging 35 to 90 minutes each to 2 to 3 incidents per quarter with mean duration under 15 minutes, a 70% reduction in total downtime hours in the first quarter post-rollout.
  • Incident Response from 25 Minutes to Under 5: Automated health-check failover initiates remediation in under 5 minutes with PagerDuty alerts providing precise instance-level context, replacing a 25-minute manual runbook process that allowed blast radius to expand unchecked.
  • Zero Customer-Reported Outages Post-Rollout: 100% of incidents are now detected and escalated internally before any customer notices a service degradation, replacing a model where over 60% of degradations were first reported through customer support tickets.
  • MTTD Improved by Over 80%: Unified Consul health telemetry in Prometheus reduced mean time to detect from over 20 minutes across five disparate tools to under 4 minutes, an improvement of over 80%.
  • Standardised Health Checks Across 30+ Services: All 30+ services adopted the three-tier HTTP, TCP, and script-based health standard, providing uniform health telemetry and enabling platform-wide monitoring for the first time across the logistics network.
Data Flow Diagram
stream-dfd
Client Testimonial

“The moment our on-call team stopped firefighting and started trusting automated health signals, everything changed. Incidents now resolve before most users even notice.”

– VP of Engineering, Global Logistics Platform

Conclusion

The client chose Ksolves DevOps Consulting Services to replace a manual, runbook-driven incident response model that was leaving customers as the first line of detection. The result was a HashiCorp Consul health-check framework that automated failover, standardised observability across 30+ services, and cut unplanned downtime by 70% in the first quarter.

 

Before this engagement, the on-call team was the last line of defence between a failing instance and thousands of stalled shipments. After deploying the Consul framework, incidents are detected and contained before downstream impact materialises; response time fell from 25 minutes to under 5, and no customer has reported a service degradation since go-live.

 

The standardised health-check framework is now the platform’s operational governance baseline, with health telemetry serving as a single source of truth for service availability across audits and SLA reporting.

Is Your On-Call Team Still Detecting Failures from Customer Support Tickets?

Copyright 2026© Ksolves.com | All Rights Reserved
Ksolves USP