Observability at Internet Scale: How Zscaler Unified 30+ Tools with Kloudfuse
Executive Summary
Zscaler, the leader in cloud security, processes over 600 billion transactions daily. Behind that scale sat a fragmented observability stack: more than 30 tools across the organization, with some critical product lines running on signal-based monitoring that collected no underlying data at all.
By standardizing on Kloudfuse, Zscaler onboarded [XX] teams and services, retired New Relic and Sumo Logic with Datadog scheduled to follow, and gained near-real-time telemetry across the board, all while keeping 100% of its data in-house through Kloudfuse's Self-SaaS model. The platform is now the foundation for proactive, AI-driven observability using Kloudfuse Enterprise MCP Server and cross-signal correlation.
By switching to Kloudfuse, Zscaler achieved:
Observability tools consolidated
30+
Unified daily data ingestion
300+TB
Telemetry landing in under
40s
About Zscaler
Zscaler is the global leader in cloud security, operating the world's largest inline security cloud through its Zero Trust Exchange platform. Distributed across more than 160 data centers globally, the platform processes over 600 billion transactions daily, more than half a trillion and serves over 9,400 customers across more than 185 countries, protecting nearly 45% of the Fortune 500. By the third quarter of fiscal 2026, Zscaler had surpassed $3.5 billion in Annual Recurring Revenue, with a stated long-term ambition of reaching $10 billion. For a platform this mission-critical, where every packet from every customer flows through Zscaler, reliability and rapid incident response are non-negotiable, which is what put observability at the center of the company's transformation.
The Challenge: Observability at Internet Scale
As Zscaler embarked on a transformation to support its rapid growth, the leadership team set clear reliability objectives:
Mean Time to Detect (MTTD): Less than 5 minutes
Mean Time to Mitigate (MTTM): Less than 30 minutes
However, the existing observability landscape presented significant obstacles.
Fragmented Tooling Across 30+ Platforms
Zscaler's observability stack had grown organically, resulting in over 30 different tools across the organization — Datadog, New Relic, Sumo Logic, Grafana, OpenSearch, ELK Stack, LGTM Stack, Nagios, and various open-source solutions. Each team used different tools for different data points, creating silos where "when a problem comes, you never know where the data is."
Signal-Based Monitoring with No Data Collection
For one of Zscaler's largest product lines, the team relied on Nagios — a signal-based tool that creates alerts but doesn't collect or analyze underlying data. Effective for basic alerting, but not built to collect high-volume telemetry for deeper analytics, correlation, or AI-driven trend detection. This limited Zscaler's ability to move beyond "known knowns" and proactively detect anomalies and unknown failure patterns.
Manual, Time-Consuming Troubleshooting
When issues arose, engineers had to SSH into individual servers, manually check logs, and attempt to correlate data across disconnected systems. Teams spent up to 30 minutes just proving "it's not my problem" before actual troubleshooting could begin.
Unknown Scale and Cardinality
Processing 600+ billion daily transactions generates massive telemetry signals with high cardinality and dimensionality, far beyond what traditional monitoring assumptions were built for. The observability team faced "unknown unknowns" around:
Telemetry volume growth
Cardinality and dimensionality explosions
Concurrency and user experience under incident load
Long-term platform performance at scale
Zscaler needed a platform that could handle unknown, and potentially enormous, data volumes as they instrumented more systems.
Data Sovereignty Requirements
As a cybersecurity company, Zscaler couldn't send sensitive telemetry data to external SaaS vendors. They required complete control over their observability data and a platform that ensured:
Sensitive telemetry stays under strict governance controls
Data can be reused beyond alerting, for analytics, ML models, and product improvements
Compliance readiness for federal environments (FIPS 140-3, FedRAMP trajectory)

Kishore Thakur
Senior Director, Cloud Platform Engineering, Zscaler
"At Zscaler's scale, the cost of reactive observability isn't just engineering hours, it's customer trust. We needed to move from a world where teams were manually hunting through individual systems and hoping they were looking in the right place, to a world where we can anticipate issues before they reach our customers. That required a platform that could unify our telemetry, keep it inside our perimeter, and give us the AI-driven correlation to shift from firefighting to prevention. That's the transformation we're building with Kloudfuse."
Why Kloudfuse
After evaluating five to six vendors, including enterprise incumbents and emerging startups, Zscaler selected Kloudfuse. The decision came down to several key factors:
Data Control & Security
Zscaler needed telemetry to stay within its security perimeter and governance controls, especially as a cybersecurity company handling sensitive signals. Kloudfuse's Self-SaaS deployment model ensures all observability data remains in-house.An AI-Ready Platform, Not Just Monitoring
Zscaler's goal is proactive reliability and unknown detection, anticipating issues earlier, identifying trends, and reducing customer impact. Kloudfuse aligned with requirements around analytics, correlation, and AI-driven observability.OpenTelemetry-First Telemetry Pipeline
Zscaler standardized on OpenTelemetry as its telemetry foundation, a CNCF Graduated project and the second-most active in the cloud-native ecosystem after Kubernetes, with over 12,000 contributors from 2,800+ companies. This ensures their telemetry strategy remains durable and vendor-flexible over time. Kloudfuse became the unified platform layer to store, query, and operationalize that data at speed.Proven, Scalable Backend Architecture
The underlying architecture, including scalable columnar storage patterns like Apache Pinot, supported confidence for long-term scale and performance, even as data volumes and cardinality grow unpredictably.Familiar Experience for Broad Adoption
Kloudfuse's Grafana-compatible experience reduced adoption friction. Teams didn't need to relearn observability basics to get started. This familiarity accelerated onboarding across 78 teams and supported migration from legacy tools.Roadmap Alignment & Partnership Commitment
Kloudfuse's roadmap, including MCP Server, AI-driven correlation, and enhanced analytics, aligned with Zscaler's vision. The "extended dev team" model — shared roadmap, responsiveness, and iterative delivery — stood out during evaluation.
Quantifying Success: Results & Impact
Tool Displacement
Zscaler is consolidating 30+ tools into a single unified platform. New Relic, Sumo Logic, and Datadog have been retired, and open-source tools including Elasticsearch and Grafana have been decommissioned, replaced by one platform that streamlines operations and improves the engineering experience.Near-Real-Time Telemetry
With telemetry landing in under 40 seconds, Zscaler has moved from an inconsistent, partly signal-only picture to near-real-time visibility across teams. Alerts can now be triggered almost immediately upon data ingestion, a step change from the previous environment where some critical product lines had no underlying data collection at all.Centralized Troubleshooting
Teams no longer SSH into individual servers or manually correlate logs. All data is queryable from a central platform with pattern recognition and time-based correlation.Self-Service Dependency Analysis
Teams that previously spent 30+ minutes proving dependencies can now instantly check upstream and downstream dependencies themselves, dramatically accelerating the troubleshooting process.Full Coverage Across the Organization
Zscaler now collects telemetry from systems that previously had no observability at all. The result is significantly more data flowing through the platform than the legacy stack ever handled, a productivity and reliability gain, not a cost-reduction exercise.Before and After
What changed when Zscaler moved from a fragmented stack to a single platform:
Observability tools: From 30+ tools spread across teams to consolidating on one unified platform.
Data coverage: From gaps where some product lines had no observability at all to full telemetry collected and queryable.
Telemetry availability: From inconsistent and signal-based for some critical lines to near-real-time, landing in under 40 seconds.
Total data ingestion: From siloed, unknown volumes to 300+ TB/day unified.
Data ownership: From mixed, with some data sitting with external vendors, to 100% in-house with Self-SaaS.
Compliance readiness: From mixed tooling and inconsistent posture to FIPS 140-3 certified (#5186 / #5209) and deploying into federal environments.
The Kloudfuse Difference
With Kloudfuse, Zscaler is building a modern observability foundation designed for reliability outcomes, not just dashboards.
Unified Observability for Faster Troubleshooting
By centralizing telemetry in a single platform, teams can correlate logs, metrics, and traces in one place, reducing manual "server hopping," scattered searches, and cross-team back-and-forth.
Teams reported meaningful workflow improvements, including:
Faster pattern discovery via centralized querying
Easier upstream/downstream dependency validation
Less back and forth across teams, since dependencies become visible and verifiable
More Coverage, Not Just Consolidation
Systems that previously had no observability at all are now instrumented and queryable. Zscaler is collecting more telemetry than ever before — trading a fragmented, partial picture for full, unified visibility across the organization.AI-Ready Observability for Proactive Detection
Zscaler's observability strategy prioritizes unknown detection and proactive reliability: anticipating issues earlier, identifying trends, and reducing customer impact before it occurs.
Kloudfuse supports this shift by pairing a unified telemetry store with capabilities aligned to:
Cross-signal correlation across logs, metrics, and traces to surface patterns invisible in siloed tools
AI-driven analysis for proactive anomaly detection and trend identification
Structured investigation workflows via MCP Server, enabling automated RCA alongside Zscaler's existing agentic AI platform
Compliance-Ready for Federal Environments
With FIPS 140-3-certified images (NIST Certificates #5186 and #5209), Zscaler is actively deploying Kloudfuse into federal environments, enabling a single, consistent observability platform across the entire organization, including regulated workloads.
"When you're running observability for [XX] teams and services, the hardest part isn't the tooling, it's the toil. Every team had their own dashboards, their own alerting logic, their own way of proving a problem wasn't theirs. Now we have one telemetry backbone with consistent correlation across signals, which means my team spends less time stitching data together during incidents and more time building the automation that prevents them."

Naresh Gambirapuram
Senior Manager, Zscaler
Looking Ahead
With 78 teams/services onboarded and the core platform consolidation underway, Zscaler's observability journey is shifting from migration to leverage, extracting maximum value from the telemetry it now collects at scale.
The immediate priorities are clear. FIPS 140-3-certified Kloudfuse is deploying into Zscaler's federal environments, bringing government and commercial workloads onto a single, consistent observability platform for the first time. With MCP Server, Zscaler's engineering teams can bring AI-driven investigation directly into their existing workflows, structured root cause analysis, proactive anomaly detection, and cross-signal correlation that would have been impossible when data lived in 30+ disconnected tools.
The longer-term vision goes further: using the unified telemetry store not just for alerting and troubleshooting, but for trend analysis, AI-driven insights, and reliability improvements that reduce customer impact before it occurs. It's the shift from reactive to proactive that Zscaler's leadership set out to achieve, and the foundation is now in place.
"The thing people underestimate about recording rules is the damage they do unsupervised: expensive functions running in serial, high-cardinality labels nobody thought to drop, all compounding as steady-state load. Kloudfuse handles multi-resolution rollups natively, so we're not managing a sprawl of brittle rules just to keep dashboards responsive."

Benjamin Goodrich
Staff Software Engineer, Zscaler