Featured image of post Meet Cody: coding agents that live in our cluster

Meet Cody: coding agents that live in our cluster

By spring most of our engineers had an agent in their terminal, and writing code was no longer the slow part. A lot of time still went into work that nobody sits down to do in a terminal. Someone posts “the notification service is crash looping in dev” in Slack, and somebody else drops what they’re doing, pulls logs, checks the last deploy and finds that a secret got renamed. A security finding stays open for weeks because the fix touches three repos. This kind of work is small and important, and it keeps interrupting people.

So in May we started Cody: our internal platform for running coding agents in the background, inside our own Kubernetes cluster, with the same tools and the same guard rails as the rest of our platform.

The ground rules

We wrote down four ground rules early on. They’re copied straight from the first spec:

Kubernetes is the execution environment. GitOps is the change-control mechanism. Pull requests are the mutation path. Cluster access is read-only by design.

Cody works on non-prod. It can read pods, events, logs, HelmReleases and the rest of what you’d look at when debugging. It cannot kubectl apply, cannot exec into a pod and cannot change anything in the cluster. If it wants to change something, it opens a pull request against the GitOps repo or the service repo, and a human reviews it like any other PR.

That rule made a lot of later decisions easy. We never had to discuss what Cody is allowed to do to the cluster, because the answer is nothing.

How it works

Everything in Cody is a Kubernetes resource. A TaskSpawner says when to start an agent (a Slack mention, a GitHub webhook, a cron schedule), an AgentConfig says who the agent is and how it works, and every run is a pod with a Codex agent inside.

This is roughly the first spawner we deployed. It turns @cody debug notification-service in dev into an agent run:

apiVersion: cody.alpheya.com/v1alpha1
kind: TaskSpawner
metadata:
  name: cody-debug-slack
spec:
  when:
    slack:
      triggers:
        - pattern: '(?i)^debug\s+\S+(?:\s+in\s+\S+)?'
  taskTemplate:
    type: codex
    agentConfigRef:
      name: cody-debugger
    promptTemplate: |
      A user has asked you to debug a service. Read your AGENTS.md and
      the `cody-debug-service` skill before doing anything else, then
      follow the workflow there.

      Slack message:
        {{.Body}}

      Slack thread: {{.URL}}      
    ttlSecondsAfterFinished: 3600
    podOverrides:
      serviceAccountName: cody-debugger

The debug workflow itself lives in a skill, not in the prompt. Cody collects cheap cluster evidence first (pod status, events, recent deploys) before cloning anything, then decides what kind of problem it is. An operational problem gets the exact next step in the Slack thread. A config problem gets a PR to the GitOps repo, and a code problem a PR to the service repo. If it’s unclear, Cody asks one clarifying question instead of guessing.

The agent image

Most of my own time in May went into the image the agent runs in. An agent debugging a platform needs the same tools a platform engineer has, so we baked them in: kubectl, flux, kustomize, helm, psql, redis-cli, grpcurl, gh, plus our own engineering skills. It also needs to call our internal services, which expect a signed service JWT.

My first version was a small sign-jwt helper the agent was supposed to call before talking to a service. That repeats a failure mode we already knew: agents silently skip auth steps they weren’t reminded about, the same way they skip anything else that isn’t in front of them. So I replaced it with a curl wrapper that signs every request transparently. With the wrapper, signing is a property of the tool, and the agent’s prompt doesn’t need to mention it.

Two days later I found the hole in that. The wrapper was installed at /usr/local/bin/curl, shadowing the real one on PATH. Any agent or script that typed /usr/bin/curl explicitly went straight past it. The fix was to take the real binary away. Simplified, the Dockerfile now does this:

COPY bin/cody-curl /usr/local/bin/curl

# Move the real curl out of PATH. Runtime calls always resolve the wrapper,
# and the wrapper finds the real binary through CODY_CURL_REAL.
RUN mkdir -p /usr/local/libexec \
    && mv /usr/bin/curl /usr/local/libexec/curl-real
ENV CODY_CURL_REAL=/usr/local/libexec/curl-real

Now /usr/bin/curl returns “No such file or directory” instead of silently bypassing auth. The general lesson: if a safety property depends on the agent remembering to do something, it isn’t a safety property. Put it in the environment.

Capabilities, not secrets

The same thinking applied to credentials. The first version of Cody’s pods held our GitHub App’s private key so they could mint installation tokens. That works, but it puts a long-lived key inside a shell that the agent can inspect.

We moved that into a small tools service next to Cody. The pod asks it for a short-lived token, and the private key never leaves the server side. Jira, Confluence and our security scanner got the same treatment. One line from the spec became our rule for every integration since:

Cody should receive a capability, not a raw upstream secret.

Context isn’t continuity

At first, every @cody mention created a brand new run. For a follow-up in the same Slack thread we passed the whole thread in as text. The spec described the problem better than I can:

This gives Cody context, but not continuity: each follow-up is a separate Kubernetes Job and a separate Cody execution.

So we added sessions. One Slack thread maps to one session, and turns run in order inside it. Humans can keep talking in the thread without waking Cody up, and when they mention @cody again it gets only what happened since its last answer, plus its own earlier state.

Babysitters

The part I’m most excited about is Cody working without anyone asking.

In June we added a cron-driven infra health check for our non-prod clusters, and a security remediation flow. The remediation flow follows a shape we now use for anything long-running:

  • a deterministic controller discovers the work (it queries the scanner itself, not the agent);
  • one stable session owns one durable problem;
  • scheduled heartbeats advance the session over time;
  • Slack gets one root message per session with the latest status, and the details go in the thread.

A dependency bump in a shared package can take days: open a PR, wait for CI, wait for the package to publish, bump the consumers, wait again, then confirm with a fresh scan that the finding is gone. A single agent run can’t wait that long, but a session with a heartbeat can.

This split matters: discovery, deduplication, scheduling and credentials are plain code, and the agent only gets the part that needs judgment.

Two months in

Between May 1 and today, Cody opened a bit over 300 pull requests: GitOps fixes, service fixes, dependency bumps, CI changes. Close to two thirds were merged. The rest were closed, which I’m fine with. A closed PR costs a reviewer a minute; a PR that never got opened would have cost someone an afternoon.

This month we also started pointing Cody at pull requests itself, as a reviewer. It’s early. More on that when we have data.

The catch is that nearly every one of those 300 PRs went through a human reviewer. We built an agent that does useful work in the background, and its output lands in exactly the queue that was already our bottleneck. Making Cody write more is easy. The hard part is making it easier to trust what Cody (and everyone else) writes, and that’s where we’re heading next.

As always, if you have any questions or remarks, feel free to ping me on twitter @bobby_donchev.