Zscaler, the leader in cloud security, processes over 600 billion transactions daily. Behind that scale sat a fragmented observability stack: more than 30 tools across the organization, with some critical product lines running on signal-based monitoring that collected no underlying data at all.
Zscaler’s observability footprint today
Five minutes with the team that consolidated more than thirty observability tools onto one platform.
By standardizing on Kloudfuse, Zscaler onboarded 78 teams and services, retired commercial tools including New Relic, Sumo Logic, and Datadog, and gained near-real-time telemetry across the board, all while keeping 100% of its data in-house through Kloudfuse’s Self-SaaS model. The platform is now the foundation for proactive, AI-driven observability using Kloudfuse Enterprise MCP Server and cross-signal correlation.
Zscaler is the global leader in cloud security, operating the world’s largest inline security cloud through its Zero Trust Exchange platform. Distributed across more than 150 data centers globally, the platform processes over 500 billion transactions daily and serves nearly 8,000 customers worldwide, protecting more than 45% of the Fortune 500. By the third quarter of fiscal 2026, Zscaler had surpassed $3.5 billion in Annual Recurring Revenue, with a stated long-term ambition of reaching $10 billion.
For a platform this mission-critical, where every packet from every customer flows through Zscaler, reliability and rapid incident response are non-negotiable. That is what put observability at the center of the company’s transformation.
As Zscaler embarked on a transformation to support its rapid growth, the leadership team set clear reliability objectives:
However, the existing observability landscape presented significant obstacles.
Zscaler’s observability stack had grown organically, resulting in over 30 different tools across the organization: Datadog, New Relic, Sumo Logic, Grafana, OpenSearch, ELK Stack, LGTM Stack, Nagios, and various open-source solutions. Each team used different tools for different data points, creating silos where, as the team put it, when a problem comes you never know where the data is.
For one of Zscaler’s largest product lines, the team relied on a signal-based tool that creates alerts but doesn’t collect or analyze underlying data. Effective for basic alerting, but not built to collect high-volume telemetry for deeper analytics, correlation, or AI-driven trend detection. This limited Zscaler’s ability to move beyond “known knowns” and proactively detect anomalies and unknown failure patterns.
When issues arose, engineers had to SSH into individual servers, manually check logs, and attempt to correlate data across disconnected systems. Teams spent up to 30 minutes just proving “it’s not my problem” before actual troubleshooting could begin.
Processing 500+ billion daily transactions generates massive telemetry signals with high cardinality and dimensionality, far beyond what traditional monitoring assumptions were built for. The observability team faced “unknown unknowns” around:
Zscaler needed a platform that could handle unknown, and potentially enormous, data volumes as they instrumented more systems.
As a cybersecurity company, Zscaler couldn’t send sensitive telemetry data to external SaaS vendors. They required complete control over their observability data and a platform that ensured:
At our scale, reliability depends on how quickly teams can identify issues, understand service dependencies, and take action with confidence. Kloudfuse has helped simplify that by giving our teams a more unified view of production behavior across the platform. With Kloudfuse 4.0’s workload isolation, we can also scale observability infrastructure more deliberately as demand grows, without creating new operational bottlenecks. That combination strengthens both reliability execution and long-term resilience.
After evaluating five to six vendors, including enterprise incumbents and emerging startups, Zscaler selected Kloudfuse. The decision came down to several key factors:
Zscaler needed telemetry to stay within its security perimeter and governance controls, especially as a cybersecurity company handling sensitive signals. Kloudfuse’s Self-SaaS deployment model ensures all observability data remains in-house.
Zscaler’s goal is proactive reliability and unknown detection, anticipating issues earlier, identifying trends, and reducing customer impact. Kloudfuse aligned with requirements around analytics, correlation, and AI-driven observability.
Zscaler standardized on OpenTelemetry as its telemetry foundation, a CNCF Graduated project and the second-most active in the cloud-native ecosystem after Kubernetes, with over 12,000 contributors from 2,800+ companies. This ensures their telemetry strategy remains durable and vendor-flexible over time. Kloudfuse became the unified platform layer to store, query, and operationalize that data at speed.
The underlying architecture, including scalable columnar storage patterns like Apache Pinot, supported confidence for long-term scale and performance, even as data volumes and cardinality grow unpredictably.
Kloudfuse’s Grafana-compatible experience reduced adoption friction. Teams didn’t need to relearn observability basics to get started. This familiarity accelerated onboarding across 78 teams and supported migration from legacy tools.
Kloudfuse’s roadmap, including MCP Server, AI-driven correlation, and enhanced analytics, aligned with Zscaler’s vision. The “extended dev team” model — shared roadmap, responsiveness, and iterative delivery — stood out during evaluation.
At Zscaler’s scale, the cost of reactive observability isn’t just engineering hours, it’s customer trust. We needed to move from a world where teams were manually hunting through individual systems and hoping they were looking in the right place, to a world where we can anticipate issues before they reach our customers. That required a platform that could unify our telemetry, keep it inside our perimeter, and give us the AI-driven correlation to shift from firefighting to prevention. That’s the transformation we’re building with Kloudfuse.
Zscaler is consolidating 30+ tools into a single unified platform. New Relic, Sumo Logic, and Datadog have been retired, and open-source tools including Elasticsearch and Grafana have been decommissioned, replaced by one platform that streamlines operations and improves the engineering experience.
With telemetry landing in under 40 seconds, Zscaler has moved from an inconsistent, partly signal-only picture to near-real-time visibility across teams. Alerts can now be triggered almost immediately upon data ingestion, a step change from the previous environment where some critical product lines had no underlying data collection at all.
Teams no longer SSH into individual servers or manually correlate logs. All data is queryable from a central platform with pattern recognition and time-based correlation.
Teams that previously spent 30+ minutes proving dependencies can now instantly check upstream and downstream dependencies themselves, dramatically accelerating the troubleshooting process.
Zscaler now collects telemetry from systems that previously had no observability at all. The result is significantly more data flowing through the platform than the legacy stack ever handled, a productivity and reliability gain, not a cost-reduction exercise.
What changed when Zscaler moved from a fragmented stack to a single platform:
With Kloudfuse, Zscaler is building a modern observability foundation designed for reliability outcomes, not just dashboards.
By centralizing telemetry in a single platform, teams can correlate logs, metrics, and traces in one place, reducing manual “server hopping,” scattered searches, and cross-team back-and-forth.
Teams reported meaningful workflow improvements, including:
Systems that previously had no observability at all are now instrumented and queryable. Zscaler is collecting more telemetry than ever before — trading a fragmented, partial picture for full, unified visibility across the organization.
Zscaler’s observability strategy prioritizes unknown detection and proactive reliability: anticipating issues earlier, identifying trends, and reducing customer impact before it occurs.
Kloudfuse supports this shift by pairing a unified telemetry store with capabilities aligned to:
With FIPS 140-3-certified images (NIST Certificates #5186 and #5209), Zscaler is actively deploying Kloudfuse into federal environments, enabling a single, consistent observability platform across the entire organization, including regulated workloads.
When you’re running observability for 78 teams and services, the hardest part isn’t the tooling, it’s the toil. Every team had their own dashboards, their own alerting logic, their own way of proving a problem wasn’t theirs. Now we have one telemetry backbone with consistent correlation across signals, which means my team spends less time stitching data together during incidents and more time building the automation that prevents them.
With 78 teams/services onboarded and the core platform consolidation underway, Zscaler’s observability journey is shifting from migration to leverage, extracting maximum value from the telemetry it now collects at scale.
The immediate priorities are clear. FIPS 140-3-certified Kloudfuse is deploying into Zscaler’s federal environments, bringing government and commercial workloads onto a single, consistent observability platform for the first time. With MCP Server, Zscaler’s engineering teams can bring AI-driven investigation directly into their existing workflows, structured root cause analysis, proactive anomaly detection, and cross-signal correlation that would have been impossible when data lived in 30+ disconnected tools.
The longer-term vision goes further: using the unified telemetry store not just for alerting and troubleshooting, but for trend analysis, AI-driven insights, and reliability improvements that reduce customer impact before it occurs. It’s the shift from reactive to proactive that Zscaler’s leadership set out to achieve, and the foundation is now in place.
The thing people underestimate about recording rules is the damage they do unsupervised: expensive functions running in serial, high-cardinality labels nobody thought to drop, all compounding as steady-state load. Kloudfuse handles multi-resolution rollups natively, so we’re not managing a sprawl of brittle rules just to keep dashboards responsive.
Run Kloudfuse against your own telemetry, in your own cloud, and see what a single correlated view does to your mean time to resolution.