Build the agent fleet that investigates live production incidents, and the guardrails that make it safe to run.
What you’ll work on
Build the agent fleet that investigates live production incidents: gathering context, forming hypotheses, testing them against real telemetry, and synthesising a root cause.
Design multi-agent orchestration: coordinator loops, specialised sub-agents, and the handoffs between them.
Wire agents to real observability data through MCP tools: metrics, logs, traces, Kubernetes state, profiles.
Build the evaluation harnesses that tell us whether an investigation was actually correct, not merely plausible.
Ship guardrails: least-privilege tool surfaces, read-only boundaries, and permission models that hold under autonomous operation.
What we look for
Strong backend engineering first. This is distributed systems work that happens to involve LLMs, not prompt-tuning.
Experience building agentic systems: tool use, function calling, multi-step planning, long-running loops.
Judgment about autonomy and safety: what an agent should do unattended, and what it must never do.
Ability to reason about production incidents yourself. You can’t build an SRE agent without SRE instincts.
Bonus
SRE, on-call or incident-response background. You’ve been paged and know what a real investigation feels like.
Observability depth: correlating metrics, logs and traces to isolate a failure.
Experience with agent SDKs, MCP, or building tools for LLMs to consume.
Eval tooling and LLM-as-judge methods, and a healthy scepticism about both.
Prompts-as-code discipline: versioned, reviewed and tested like any other source.
How we work
Small team, short feedback loops, real ownership from week one.
You’ll talk to customers, engineers debugging real incidents, and that shapes what you build.
We ship, then iterate; bias toward hands-on building over process.
Modern tooling is encouraged, including AI-assisted development.