Architecting and Herding Coding Agents

6 min read
I trusted coding agents to work autonomously on my computer until on 14 July 2026, I discovered Claude Code exploring my computer looking for context1. When an agent is missing something, it gets resourceful. A clearer prompt helps. But it isn't a boundary.
Two days later, Hugging Face disclosed a security incident. OpenAI's agents had escaped an evaluation sandbox and compromised parts of its infrastructure. Different scale, same pattern: when something is missing, the agent gets resourceful.
The next day, I started designing an autonomous coding agent system for my personal software development workflow.
Herding Coding Agents
Herding Coding Agents

Agentic Coding Today

Agentic coding today mostly means a coding agent running on a developer's laptop. On that machine, it technically has access to everything the developer has: SSH keys, cloud credentials, API tokens, and every repository checked out locally.

There are generally two ways to run a session:

  1. Interactively. You watch the agent work, approve its changes, and step in when it misbehaves. In other words, you babysit. That scales with your attention, not with the number of agents.
  2. Autonomously. You hand the agent a task and walk away, which more and more developers are doing. The agent keeps working with your credentials and your access, and nobody is watching. That's how mine ended up searching my file system.

Tools for managing many sessions at once are popping up, like Herdr and Agent Manager, many of which are designed around interactive sessions. They make babysitting more efficient, but it still requires my attention to orchestrate. And that doesn't scale.

My Setup: Two Roles, One Workflow

I wanted a third option that was both secure and scalable: define the work once, let agents run in parallel without my credentials, and review what comes back. With boundaries that hold whether I'm watching or not.

And this is my design, which splits the work into two roles.

Two roles, one workflow On the left, me as the human architect, working interactively on my machine in a dev container, pairing with Claude in Zed on design. On the right, the geese, autonomous agents running headless in Kubernetes as Jobs, one goose per task, for example a developer agent and a review agent, each with many instances in parallel. I define the work as GitHub issues. The geese git push branches back to GitHub. mehuman architectinteractive, in the loopgeeseautonomous agentsheadless, nobody watchingHost (my machine)Dev Container+ZedClaudepair on designKubernetesJobdeveloperagentJobreviewagentone goose per task, many in paralleldefines the workgit pushGitHubissues in, branches out

The architect is me. I work interactively in a dev container in Zed, pairing with Claude on design. I break the work down into GitHub issues and review the pull requests that come back.

The geese are the agents. The harness is goose, an open-source agent that is part of the Agentic AI Foundation, so a group of them is a flock. Each goose picks up one issue, runs headless in its own pod, and pushes a branch. Nobody is watching.

My job shifts from operating agents to designing the workflow they run in, and the boundaries around it.

AI Native Builds on Cloud Native

Organizations already have rules for who can do what: identities and roles, read and write access, branch protection, code review, scoped secrets, approval steps.

A headless agent is a new actor, and it needs to be mapped into that system.

We've built this foundation before. CI/CD pipelines were the first headless actors in the software lifecycle. They run under their own service identities, with scoped secrets and environment approvals, and a human triggers them with a git push or a button. I authored Microsoft's official best-practice guidance on governance for CI/CD pipelines. Agents build on that same foundation.

Agents add one new requirement. A pipeline runs code that someone reviewed. An agent decides what to do at runtime. The foundation holds, but the scoping gets tighter.

The failure mode is an agent running in my session, with my credentials. It gets around every control the organization has. It acts with my access, and Git records that I did it.

In my design, each goose gets:

  • its own GitHub account, julieio-goose
  • its own credentials to services in the SDLC
  • write access only to branches prefixed agent/*
  • signed commits, so authorship can be verified (designed, not yet implemented)

Then the existing controls apply to it like any other actor. Branch protection still protects. Code review still reviews. And the record is accurate.

Solid Boundaries vs Soft Guardrails

If identity decides what an agent may do, isolation decides what it can reach.

Coding agents ship with permission settings, like allow and deny rules in a settings.json. They help, but they're a deterrent, not a boundary. The agent runs in the same environment the rules are trying to protect. The container is the boundary.

Soft guardrails vs hard boundaries Two views of the same laptop. On the left, a coding agent sits inside a dashed boundary of settings.json allow and deny rules. Its arrows cross the dashed line and reach ~/code and ~/.ssh on the laptop. On the right, the coding agent sits inside a solid container or VM boundary. Its arrow stops at the wall, and ~/code and ~/.ssh are out of reach. soft guardraila request the agent can work aroundlaptopsettings.jsoncoding agent~/code~/.sshhard boundarya wall the agent can't crosslaptopcontainer or VMcoding agent~/code~/.ssh

So: one task, one goose, one pod. Each runs as a Kubernetes Job that starts, does its work, and exits.

Right-Size the Model per Task

No nested agents either. A goose doesn't spawn its own subagents. Specialization comes from what I dispatch (e.g. a frontend-dev goose or a qa-checker goose), not from agents delegating to each other.

That matters for cost. When I'm driving Claude Code, for example, with Opus and hand off a task to run in the background, that task inherits Opus unless I configure it otherwise. My most expensive model ends up doing work a cheaper one would handle just fine.

Inherited vs right-sized models On the left, an Opus session hands off three background tasks, implement, write tests and update docs, and all three inherit Opus, each costing three dollar signs. On the right, I dispatch the same three tasks, each with a model that fits: implement on Sonnet for two dollar signs, write tests and update docs on Haiku for one dollar sign each. inherited modelevery task runs on the session's modelOpus sessionimplementOpus$$$write testsOpus$$$update docsOpus$$$right-sized modeleach task gets a model that fitshuman dispatchedimplementSonnet$$write testsHaiku$update docsHaiku$

This is a familiar cloud-native lesson. We don't run every workload on the biggest VM. We size each one for what it needs. Models are no different. Small, decomposed tasks run fine on cheaper models, so each goose gets the model that fits its task, chosen when I dispatch it.

important

Headless means API pricing. Once an agent runs headless in the cloud, most providers' terms require API keys instead of a subscription plan. Every token is billed, so a task running on a larger model than it needs is avoidable cost.

What's Next in AI Transformation

My biggest takeaway this year is that none of the cloud-native work was wasted. Securing agents comes down to identity, least privilege and isolation. Scaling them comes down to short-lived jobs that run in parallel, each sized for its work. Governing them builds on how we already govern headless actors like CI/CD pipelines. These are cloud-native practices we've refined for years.

Agentic coding started on the laptop, in the developer's terminal and IDE, and that was the right place to start. It's where we learned what agents can do. The next step is moving them onto infrastructure built for headless actors, where they scale beyond one developer's attention and run inside boundaries that hold whether anyone is watching or not.

That step doesn't require new practices. Agents are a new kind of actor, but they fit into systems we already know how to run.

That's the foundation. Follow along as I build it now from scratch. My working notes, spikes and decision log are in julie-ng/autonomous-agents-setup.

Footnotes

  1. I asked Claude Code to reference a project on my public GitHub to see how I structure my Makefiles. My context window was filling up, so I asked it to compact, and didn't realize the URL to the example was lost in the summary. Instead of asking, it went looking for the repository on my local file system. I describe it in this interview: watch from 21:10. ↩