Customer stories · Cloud security

Observability at Internet Scale: How Zscaler Unified 30+ Tools with Kloudfuse

Zscaler, the leader in cloud security, processes over 600 billion transactions daily. Behind that scale sat a fragmented observability stack: more than 30 tools across the organization, with some critical product lines running on signal-based monitoring that collected no underlying data at all.

Zscaler’s observability footprint today

30+Observability tools consolidated
300+ TBUnified daily data ingestion
Telemetry landing in under40s

Hear the story from Zscaler.

Five minutes with the team that consolidated more than thirty observability tools onto one platform.

Kishore ThakurLeads SRE, AIOps and cloud platform engineering, Zscaler

Executive Summary

By standardizing on Kloudfuse, Zscaler onboarded 78 teams and services, retired commercial tools including New Relic, Sumo Logic, and Datadog, and gained near-real-time telemetry across the board, all while keeping 100% of its data in-house through Kloudfuse’s Self-SaaS model. The platform is now the foundation for proactive, AI-driven observability using Kloudfuse Enterprise MCP Server and cross-signal correlation.

About Zscaler

Zscaler is the global leader in cloud security, operating the world’s largest inline security cloud through its Zero Trust Exchange platform. Distributed across more than 150 data centers globally, the platform processes over 500 billion transactions daily and serves nearly 8,000 customers worldwide, protecting more than 45% of the Fortune 500. By the third quarter of fiscal 2026, Zscaler had surpassed $3.5 billion in Annual Recurring Revenue, with a stated long-term ambition of reaching $10 billion.

For a platform this mission-critical, where every packet from every customer flows through Zscaler, reliability and rapid incident response are non-negotiable. That is what put observability at the center of the company’s transformation.

The Challenge: Observability at Internet Scale

As Zscaler embarked on a transformation to support its rapid growth, the leadership team set clear reliability objectives:

  • Mean Time to Detect (MTTD): Less than 5 minutes
  • Mean Time to Mitigate (MTTM): Less than 30 minutes

However, the existing observability landscape presented significant obstacles.

  1. Fragmented Tooling Across 30+ Platforms

    Zscaler’s observability stack had grown organically, resulting in over 30 different tools across the organization: Datadog, New Relic, Sumo Logic, Grafana, OpenSearch, ELK Stack, LGTM Stack, Nagios, and various open-source solutions. Each team used different tools for different data points, creating silos where, as the team put it, when a problem comes you never know where the data is.

  2. Signal-Based Monitoring with No Data Collection

    For one of Zscaler’s largest product lines, the team relied on a signal-based tool that creates alerts but doesn’t collect or analyze underlying data. Effective for basic alerting, but not built to collect high-volume telemetry for deeper analytics, correlation, or AI-driven trend detection. This limited Zscaler’s ability to move beyond “known knowns” and proactively detect anomalies and unknown failure patterns.

  3. Manual, Time-Consuming Troubleshooting

    When issues arose, engineers had to SSH into individual servers, manually check logs, and attempt to correlate data across disconnected systems. Teams spent up to 30 minutes just proving “it’s not my problem” before actual troubleshooting could begin.

  4. Unknown Scale and Cardinality

    Processing 500+ billion daily transactions generates massive telemetry signals with high cardinality and dimensionality, far beyond what traditional monitoring assumptions were built for. The observability team faced “unknown unknowns” around:

    • Telemetry volume growth
    • Cardinality and dimensionality explosions
    • Concurrency and user experience under incident load
    • Long-term platform performance at scale

    Zscaler needed a platform that could handle unknown, and potentially enormous, data volumes as they instrumented more systems.

  5. Data Sovereignty Requirements

    As a cybersecurity company, Zscaler couldn’t send sensitive telemetry data to external SaaS vendors. They required complete control over their observability data and a platform that ensured:

    • Sensitive telemetry stays under strict governance controls
    • Data can be reused beyond alerting, for analytics, ML models, and product improvements
    • Compliance readiness for federal environments (FIPS 140-3, FedRAMP)
At our scale, reliability depends on how quickly teams can identify issues, understand service dependencies, and take action with confidence. Kloudfuse has helped simplify that by giving our teams a more unified view of production behavior across the platform. With Kloudfuse 4.0’s workload isolation, we can also scale observability infrastructure more deliberately as demand grows, without creating new operational bottlenecks. That combination strengthens both reliability execution and long-term resilience.

Why Kloudfuse

After evaluating five to six vendors, including enterprise incumbents and emerging startups, Zscaler selected Kloudfuse. The decision came down to several key factors:

  1. Data Control & Security

    Zscaler needed telemetry to stay within its security perimeter and governance controls, especially as a cybersecurity company handling sensitive signals. Kloudfuse’s Self-SaaS deployment model ensures all observability data remains in-house.

  2. An AI-Ready Platform, Not Just Monitoring

    Zscaler’s goal is proactive reliability and unknown detection, anticipating issues earlier, identifying trends, and reducing customer impact. Kloudfuse aligned with requirements around analytics, correlation, and AI-driven observability.

  3. OpenTelemetry-First Telemetry Pipeline

    Zscaler standardized on OpenTelemetry as its telemetry foundation, a CNCF Graduated project and the second-most active in the cloud-native ecosystem after Kubernetes, with over 12,000 contributors from 2,800+ companies. This ensures their telemetry strategy remains durable and vendor-flexible over time. Kloudfuse became the unified platform layer to store, query, and operationalize that data at speed.

  4. Proven, Scalable Backend Architecture

    The underlying architecture, including scalable columnar storage patterns like Apache Pinot, supported confidence for long-term scale and performance, even as data volumes and cardinality grow unpredictably.

  5. Familiar Experience for Broad Adoption

    Kloudfuse’s Grafana-compatible experience reduced adoption friction. Teams didn’t need to relearn observability basics to get started. This familiarity accelerated onboarding across 78 teams and supported migration from legacy tools.

  6. Roadmap Alignment & Partnership Commitment

    Kloudfuse’s roadmap, including MCP Server, AI-driven correlation, and enhanced analytics, aligned with Zscaler’s vision. The “extended dev team” model — shared roadmap, responsiveness, and iterative delivery — stood out during evaluation.

At Zscaler’s scale, the cost of reactive observability isn’t just engineering hours, it’s customer trust. We needed to move from a world where teams were manually hunting through individual systems and hoping they were looking in the right place, to a world where we can anticipate issues before they reach our customers. That required a platform that could unify our telemetry, keep it inside our perimeter, and give us the AI-driven correlation to shift from firefighting to prevention. That’s the transformation we’re building with Kloudfuse.

Quantifying Success: Results & Impact

  1. Tool Displacement

    Zscaler is consolidating 30+ tools into a single unified platform. New Relic, Sumo Logic, and Datadog have been retired, and open-source tools including Elasticsearch and Grafana have been decommissioned, replaced by one platform that streamlines operations and improves the engineering experience.

  2. Near-Real-Time Telemetry

    With telemetry landing in under 40 seconds, Zscaler has moved from an inconsistent, partly signal-only picture to near-real-time visibility across teams. Alerts can now be triggered almost immediately upon data ingestion, a step change from the previous environment where some critical product lines had no underlying data collection at all.

  3. Centralized Troubleshooting

    Teams no longer SSH into individual servers or manually correlate logs. All data is queryable from a central platform with pattern recognition and time-based correlation.

  4. Self-Service Dependency Analysis

    Teams that previously spent 30+ minutes proving dependencies can now instantly check upstream and downstream dependencies themselves, dramatically accelerating the troubleshooting process.

  5. Full Coverage Across the Organization

    Zscaler now collects telemetry from systems that previously had no observability at all. The result is significantly more data flowing through the platform than the legacy stack ever handled, a productivity and reliability gain, not a cost-reduction exercise.

  6. Before and After

    What changed when Zscaler moved from a fragmented stack to a single platform:

    • Observability tools: From 30+ tools spread across teams to consolidating on one unified platform.
    • Data coverage: From gaps where some product lines had no observability at all to full telemetry collected and queryable.
    • Telemetry availability: From inconsistent and signal-based for some critical lines to near-real-time, landing in under 40 seconds.
    • Total data ingestion: From siloed, unknown volumes to 300+ TB/day unified.
    • Data ownership: From mixed, with some data sitting with external vendors, to 100% in-house with Self-SaaS.
    • Compliance readiness: From mixed tooling and inconsistent posture to FIPS 140-3 certified (#5186 / #5209) and deploying into federal environments.

The Kloudfuse Difference

With Kloudfuse, Zscaler is building a modern observability foundation designed for reliability outcomes, not just dashboards.

  1. Unified Observability for Faster Troubleshooting

    By centralizing telemetry in a single platform, teams can correlate logs, metrics, and traces in one place, reducing manual “server hopping,” scattered searches, and cross-team back-and-forth.

    • Faster pattern discovery via centralized querying
    • Easier upstream/downstream dependency validation
    • Less back and forth across teams, since dependencies become visible and verifiable

    Teams reported meaningful workflow improvements, including:

  2. More Coverage, Not Just Consolidation

    Systems that previously had no observability at all are now instrumented and queryable. Zscaler is collecting more telemetry than ever before — trading a fragmented, partial picture for full, unified visibility across the organization.

  3. AI-Ready Observability for Proactive Detection

    Zscaler’s observability strategy prioritizes unknown detection and proactive reliability: anticipating issues earlier, identifying trends, and reducing customer impact before it occurs.

    • Cross-signal correlation across logs, metrics, and traces to surface patterns invisible in siloed tools
    • AI-driven analysis for proactive anomaly detection and trend identification
    • Structured investigation workflows via MCP Server, enabling automated RCA alongside Zscaler’s existing agentic AI platform

    Kloudfuse supports this shift by pairing a unified telemetry store with capabilities aligned to:

  4. Compliance-Ready for Federal Environments

    With FIPS 140-3-certified images (NIST Certificates #5186 and #5209), Zscaler is actively deploying Kloudfuse into federal environments, enabling a single, consistent observability platform across the entire organization, including regulated workloads.

When you’re running observability for 78 teams and services, the hardest part isn’t the tooling, it’s the toil. Every team had their own dashboards, their own alerting logic, their own way of proving a problem wasn’t theirs. Now we have one telemetry backbone with consistent correlation across signals, which means my team spends less time stitching data together during incidents and more time building the automation that prevents them.

Looking Ahead

With 78 teams/services onboarded and the core platform consolidation underway, Zscaler’s observability journey is shifting from migration to leverage, extracting maximum value from the telemetry it now collects at scale.

The immediate priorities are clear. FIPS 140-3-certified Kloudfuse is deploying into Zscaler’s federal environments, bringing government and commercial workloads onto a single, consistent observability platform for the first time. With MCP Server, Zscaler’s engineering teams can bring AI-driven investigation directly into their existing workflows, structured root cause analysis, proactive anomaly detection, and cross-signal correlation that would have been impossible when data lived in 30+ disconnected tools.

The longer-term vision goes further: using the unified telemetry store not just for alerting and troubleshooting, but for trend analysis, AI-driven insights, and reliability improvements that reduce customer impact before it occurs. It’s the shift from reactive to proactive that Zscaler’s leadership set out to achieve, and the foundation is now in place.

The thing people underestimate about recording rules is the damage they do unsupervised: expensive functions running in serial, high-cardinality labels nobody thought to drop, all compounding as steady-state load. Kloudfuse handles multi-resolution rollups natively, so we’re not managing a sprawl of brittle rules just to keep dashboards responsive.

More customer stories.

Bring an incident. We'll bring the platform.

Run Kloudfuse against your own telemetry, in your own cloud, and see what a single correlated view does to your mean time to resolution.