Case studies

Observability for Next-Gen AI Applications: How Automation Anywhere Scaled with Kloudfuse

Observability at Internet Scale: How Zscaler Unified 30+ Tools with Kloudfuse

cover image for zscaler case study
icon for a document

Executive Summary

Zscaler, the leader in cloud security, processes over 600 billion transactions daily. Behind that scale sat a fragmented observability stack: more than 30 tools across the organization, with some critical product lines running on signal-based monitoring that collected no underlying data at all.

By standardizing on Kloudfuse, Zscaler onboarded [XX] teams and services, retired New Relic and Sumo Logic with Datadog scheduled to follow, and gained near-real-time telemetry across the board, all while keeping 100% of its data in-house through Kloudfuse's Self-SaaS model. The platform is now the foundation for proactive, AI-driven observability using Kloudfuse Enterprise MCP Server and cross-signal correlation.

By switching to Kloudfuse, Zscaler achieved:

Observability tools consolidated

30+

Unified daily data ingestion

300+TB

Telemetry landing in under

40s

icon for information

About Zscaler

Zscaler is the global leader in cloud security, operating the world's largest inline security cloud through its Zero Trust Exchange platform. Distributed across more than 160 data centers globally, the platform processes over 600 billion transactions daily, more than half a trillion and serves over 9,400 customers across more than 185 countries, protecting nearly 45% of the Fortune 500. By the third quarter of fiscal 2026, Zscaler had surpassed $3.5 billion in Annual Recurring Revenue, with a stated long-term ambition of reaching $10 billion. For a platform this mission-critical, where every packet from every customer flows through Zscaler, reliability and rapid incident response are non-negotiable, which is what put observability at the center of the company's transformation.

hazard icon for challenges

The Challenge: Observability at Internet Scale

As Zscaler embarked on a transformation to support its rapid growth, the leadership team set clear reliability objectives:

  • Mean Time to Detect (MTTD): Less than 5 minutes

  • Mean Time to Mitigate (MTTM): Less than 30 minutes

However, the existing observability landscape presented significant obstacles.

  1. Fragmented Tooling Across 30+ Platforms

Zscaler's observability stack had grown organically, resulting in over 30 different tools across the organization — Datadog, New Relic, Sumo Logic, Grafana, OpenSearch, ELK Stack, LGTM Stack, Nagios, and various open-source solutions. Each team used different tools for different data points, creating silos where "when a problem comes, you never know where the data is."

  1. Signal-Based Monitoring with No Data Collection

For one of Zscaler's largest product lines, the team relied on Nagios — a signal-based tool that creates alerts but doesn't collect or analyze underlying data. Effective for basic alerting, but not built to collect high-volume telemetry for deeper analytics, correlation, or AI-driven trend detection. This limited Zscaler's ability to move beyond "known knowns" and proactively detect anomalies and unknown failure patterns.

  1. Manual, Time-Consuming Troubleshooting

When issues arose, engineers had to SSH into individual servers, manually check logs, and attempt to correlate data across disconnected systems. Teams spent up to 30 minutes just proving "it's not my problem" before actual troubleshooting could begin.

  1. Unknown Scale and Cardinality

Processing 600+ billion daily transactions generates massive telemetry signals with high cardinality and dimensionality, far beyond what traditional monitoring assumptions were built for. The observability team faced "unknown unknowns" around:

  • Telemetry volume growth

  • Cardinality and dimensionality explosions

  • Concurrency and user experience under incident load

  • Long-term platform performance at scale

Zscaler needed a platform that could handle unknown, and potentially enormous, data volumes as they instrumented more systems.

  1. Data Sovereignty Requirements

As a cybersecurity company, Zscaler couldn't send sensitive telemetry data to external SaaS vendors. They required complete control over their observability data and a platform that ensured:

  • Sensitive telemetry stays under strict governance controls

  • Data can be reused beyond alerting, for analytics, ML models, and product improvements

  • Compliance readiness for federal environments (FIPS 140-3, FedRAMP trajectory)

image for bill burton

Kishore Thakur

Senior Director, Cloud Platform Engineering, Zscaler

"At Zscaler's scale, the cost of reactive observability isn't just engineering hours, it's customer trust. We needed to move from a world where teams were manually hunting through individual systems and hoping they were looking in the right place, to a world where we can anticipate issues before they reach our customers. That required a platform that could unify our telemetry, keep it inside our perimeter, and give us the AI-driven correlation to shift from firefighting to prevention. That's the transformation we're building with Kloudfuse."

icon for a clock

Why Kloudfuse

After evaluating five to six vendors, including enterprise incumbents and emerging startups, Zscaler selected Kloudfuse. The decision came down to several key factors:

  1. Data Control & Security
    Zscaler needed telemetry to stay within its security perimeter and governance controls, especially as a cybersecurity company handling sensitive signals. Kloudfuse's Self-SaaS deployment model ensures all observability data remains in-house.

  2. An AI-Ready Platform, Not Just Monitoring
    Zscaler's goal is proactive reliability and unknown detection, anticipating issues earlier, identifying trends, and reducing customer impact. Kloudfuse aligned with requirements around analytics, correlation, and AI-driven observability.

  3. OpenTelemetry-First Telemetry Pipeline
    Zscaler standardized on OpenTelemetry as its telemetry foundation, a CNCF Graduated project and the second-most active in the cloud-native ecosystem after Kubernetes, with over 12,000 contributors from 2,800+ companies. This ensures their telemetry strategy remains durable and vendor-flexible over time. Kloudfuse became the unified platform layer to store, query, and operationalize that data at speed.

  4. Proven, Scalable Backend Architecture
    The underlying architecture, including scalable columnar storage patterns like Apache Pinot, supported confidence for long-term scale and performance, even as data volumes and cardinality grow unpredictably.

  5. Familiar Experience for Broad Adoption
    Kloudfuse's Grafana-compatible experience reduced adoption friction. Teams didn't need to relearn observability basics to get started. This familiarity accelerated onboarding across 78 teams and supported migration from legacy tools.

  6. Roadmap Alignment & Partnership Commitment
    Kloudfuse's roadmap, including MCP Server, AI-driven correlation, and enhanced analytics, aligned with Zscaler's vision. The "extended dev team" model — shared roadmap, responsiveness, and iterative delivery — stood out during evaluation.

graph-like arrow for showing success

Quantifying Success: Results & Impact

  1. Tool Displacement
    Zscaler is consolidating 30+ tools into a single unified platform. New Relic, Sumo Logic, and Datadog have been retired, and open-source tools including Elasticsearch and Grafana have been decommissioned, replaced by one platform that streamlines operations and improves the engineering experience.

  2. Near-Real-Time Telemetry
    With telemetry landing in under 40 seconds, Zscaler has moved from an inconsistent, partly signal-only picture to near-real-time visibility across teams. Alerts can now be triggered almost immediately upon data ingestion, a step change from the previous environment where some critical product lines had no underlying data collection at all.

  3. Centralized Troubleshooting
    Teams no longer SSH into individual servers or manually correlate logs. All data is queryable from a central platform with pattern recognition and time-based correlation.

  4. Self-Service Dependency Analysis
    Teams that previously spent 30+ minutes proving dependencies can now instantly check upstream and downstream dependencies themselves, dramatically accelerating the troubleshooting process.

  5. Full Coverage Across the Organization
    Zscaler now collects telemetry from systems that previously had no observability at all. The result is significantly more data flowing through the platform than the legacy stack ever handled, a productivity and reliability gain, not a cost-reduction exercise.

  6. Before and After

    What changed when Zscaler moved from a fragmented stack to a single platform:

    • Observability tools: From 30+ tools spread across teams to consolidating on one unified platform.

    • Data coverage: From gaps where some product lines had no observability at all to full telemetry collected and queryable.

    • Telemetry availability: From inconsistent and signal-based for some critical lines to near-real-time, landing in under 40 seconds.

    • Total data ingestion: From siloed, unknown volumes to 300+ TB/day unified.

    • Data ownership: From mixed, with some data sitting with external vendors, to 100% in-house with Self-SaaS.

    • Compliance readiness: From mixed tooling and inconsistent posture to FIPS 140-3 certified (#5186 / #5209) and deploying into federal environments.

icon for a bulb showing a unified vision

The Kloudfuse Difference

With Kloudfuse, Zscaler is building a modern observability foundation designed for reliability outcomes, not just dashboards.

  1. Unified Observability for Faster Troubleshooting
    By centralizing telemetry in a single platform, teams can correlate logs, metrics, and traces in one place, reducing manual "server hopping," scattered searches, and cross-team back-and-forth.

Teams reported meaningful workflow improvements, including:

  • Faster pattern discovery via centralized querying

  • Easier upstream/downstream dependency validation

  • Less back and forth across teams, since dependencies become visible and verifiable


  1. More Coverage, Not Just Consolidation
    Systems that previously had no observability at all are now instrumented and queryable. Zscaler is collecting more telemetry than ever before — trading a fragmented, partial picture for full, unified visibility across the organization.


  2. AI-Ready Observability for Proactive Detection
    Zscaler's observability strategy prioritizes unknown detection and proactive reliability: anticipating issues earlier, identifying trends, and reducing customer impact before it occurs.

Kloudfuse supports this shift by pairing a unified telemetry store with capabilities aligned to:

  • Cross-signal correlation across logs, metrics, and traces to surface patterns invisible in siloed tools

  • AI-driven analysis for proactive anomaly detection and trend identification

  • Structured investigation workflows via MCP Server, enabling automated RCA alongside Zscaler's existing agentic AI platform

  1. Compliance-Ready for Federal Environments
    With FIPS 140-3-certified images (NIST Certificates #5186 and #5209), Zscaler is actively deploying Kloudfuse into federal environments, enabling a single, consistent observability platform across the entire organization, including regulated workloads.


"When you're running observability for [XX] teams and services, the hardest part isn't the tooling, it's the toil. Every team had their own dashboards, their own alerting logic, their own way of proving a problem wasn't theirs. Now we have one telemetry backbone with consistent correlation across signals, which means my team spends less time stitching data together during incidents and more time building the automation that prevents them."

image for karthic seetharaman shankaran

Naresh Gambirapuram

Senior Manager, Zscaler

compass logo for adoption of observability

Looking Ahead

With 78 teams/services onboarded and the core platform consolidation underway, Zscaler's observability journey is shifting from migration to leverage, extracting maximum value from the telemetry it now collects at scale.

The immediate priorities are clear. FIPS 140-3-certified Kloudfuse is deploying into Zscaler's federal environments, bringing government and commercial workloads onto a single, consistent observability platform for the first time. With MCP Server, Zscaler's engineering teams can bring AI-driven investigation directly into their existing workflows, structured root cause analysis, proactive anomaly detection, and cross-signal correlation that would have been impossible when data lived in 30+ disconnected tools.

The longer-term vision goes further: using the unified telemetry store not just for alerting and troubleshooting, but for trend analysis, AI-driven insights, and reliability improvements that reduce customer impact before it occurs. It's the shift from reactive to proactive that Zscaler's leadership set out to achieve, and the foundation is now in place.

"The thing people underestimate about recording rules is the damage they do unsupervised: expensive functions running in serial, high-cardinality labels nobody thought to drop, all compounding as steady-state load. Kloudfuse handles multi-resolution rollups natively, so we're not managing a sprawl of brittle rules just to keep dashboards responsive."

image of sudhanshu sharma

Benjamin Goodrich

Staff Software Engineer, Zscaler