← Back to news
Coding agents need human policy, not just a prompt

Photo: Pexels via Wikimedia Commons

02/09/2026

Coding agents need human policy, not just a prompt

The prompt is not the policy

AI-assisted development is no longer a toy autocomplete box. Coding agents can inspect repositories, edit files, run tests, call CLIs, open pull requests, and sometimes chain those actions into a single session. That is a real productivity gain, but it changes the unit of control. Once an agent can do more than suggest a completion, the important question is no longer whether the model can produce good text. It is what the model is allowed to touch, when it must stop, and how a human can still understand the chain of decisions after the session is over. OpenAI’s recent guidance on running Codex safely makes the shift plain: keep the agent inside a bounded environment, make low-risk work smooth, and force higher-risk actions into the open. That is the real control problem in 2026. Human in command is not a slogan; it is the operating model.

The temptation is to treat safety as a prompt-writing problem. It is not. The prompt can tell the model what the team wants. It cannot define the power the model has to act. Once the agent can write files, call shell commands, query internal systems, or talk to an MCP server, the boundary has to live in configuration, permissions, and policy. A good prompt is still useful, but it is only one layer. If the system lets the agent reach secrets, production, or arbitrary network destinations, then the team has already lost the important argument. The model may still be helpful, but it will be helpful inside a weak container. The work of keeping a human in charge begins before the first token is generated. It begins with deciding which actions are safe enough to be frictionless and which actions are not.

The model can carry out work, but the team must own the rules of the road.

Boundaries need to be technical, not ceremonial

The practical controls are unglamorous, and that is exactly why they matter. Sandboxes define where the agent can write. Approval policies define when it must ask. Network rules define which domains it can reach. Credential storage defines how the agent authenticates. Workspace pinning keeps the activity tied to the right organizational boundary. None of that is flashy, but each part prevents an invisible shortcut from becoming a security incident. OpenAI says directly that low-risk everyday actions should be frictionless while higher-risk actions stop for review. That is the right instinct. It keeps the agent moving on routine tasks without giving it blanket permission to improvise where the downside is irreversible.

That distinction should show up in code, not just in policy slides. If a team cannot encode the boundary, it probably does not have one. The simplest version looks boring on purpose:

if action in {"secrets", "production", "new network domain", "permission change"}:
    require_human_approval()
else:
    allow_and_log()

This is not bureaucracy for its own sake. It is a way of saying that the agent may accelerate work, but it may not silently move the team into a new risk class. The human does not need to approve every keystroke. The human does need to approve any step that can change who can access what, what can be deployed, or what might be lost if the action is wrong. The boundary is what makes delegation legitimate. Without it, the session is not collaboration. It is just a faster way to create ambiguity about responsibility.

What the agent should do, and what the human should keep

  • The agent should read, draft, run local checks, prepare diffs, and summarize what it is uncertain about.
  • The human should decide scope, approve permission changes, review secrets, production changes, schema changes, and final merges.
  • The agent should be able to stop and ask when context is missing instead of guessing across a risky boundary.
  • The human should be able to resume the session without having to reconstruct every step from scratch.

MCP makes the boundary more important, not less

Once agents start using MCP servers, they are no longer just completing a prompt; they are speaking to external services. That is where the permission problem becomes concrete. A server that can reach a database, a ticketing system, a build system, or an internal API is not a harmless extension of the prompt. It is a new path out of the sandbox. That is why credentials, scopes, allowlists, and audit logs matter so much. OpenAI’s safe-deployment guidance points to keyring-stored OAuth credentials, workspace pinning, network policies that allow expected destinations, and approvals for unfamiliar domains. In other words, the interesting question is not whether the prompt said “be careful.” The interesting question is which identity the agent is using, what it can actually do with that identity, and how quickly a human can see when the boundary is being crossed.

This matters because modern agent systems often mix local and remote actions in the same flow. The agent may read a file, request a tool, call a server, and return a draft all inside one interaction. If the workflow does not separate those steps, a developer can end up approving a harmless-looking action that quietly opens a door to something more powerful. Good agent design treats network access as a policy object, not as a convenience. Unfamiliar domains should stop and ask. Credentials should live in hardened system storage, not be pasted into prompts or scattered through config files. And if the agent can infer too much about the environment from accidental context, that is a clue that the boundary is too porous.

Telemetry turns trust into evidence

Safety without visibility is theatre. OpenAI is clear that traditional logs tell you what happened, but not always why. Agent-native telemetry closes that gap by logging user prompts, approval decisions, tool results, MCP usage, and network allow-or-deny events. That matters because humans can then review intent after the fact, especially when something looks suspicious. OpenAI also describes internal monitoring for coding agents that examines conversations and tool traces, surfaces anomalies to humans, and treats the monitor as one layer in a defense-in-depth stack rather than as a replacement for approval. That is the right mental model. Monitoring is not a magic shield. It is a way to shorten the distance between an agent making a questionable move and a human noticing it.

For engineering teams, the implication is simple: if you cannot reconstruct the session, you probably should not have let it operate unsupervised. Logs are not only for compliance. They are how the next reviewer understands what the agent was trying to do, what it asked for, what it was allowed to do, and where the human intervened. Without that evidence, every postmortem becomes guesswork. With it, the team can distinguish a reasonable action taken in the wrong context from a policy failure that should never have been possible. That distinction is important because it tells you whether to tighten the prompt, tighten the permissions, or tighten both.

System-level review still needs human judgment

Datadog’s story shows why this is not just a security conversation. Datadog integrated Codex into a large repository and had every pull request automatically reviewed. Then it built an incident replay harness: it reconstructed historical pull requests that had contributed to incidents, ran Codex against them as if it were part of the original review, and asked the engineers who owned those incidents whether the feedback would have changed the outcome. The result was not that the agent replaced human reviewers. The result was that it surfaced risks humans had not caught at the time. Datadog reported that Codex found more than 10 cases, roughly 22% of the incidents it examined, where engineers confirmed the feedback would have made a difference. It flagged cross-module interactions, missing test coverage, and API contract changes that carried downstream risk.

That is the best-case shape of AI review. The agent is not the authority. It is the lens. It can hold more of the codebase in context than a reviewer can during a single diff, and that is useful because many real bugs do not live in the touched lines. They live in the relationship between systems, in the expectations of downstream services, or in the places where a change is technically valid but operationally dangerous. The human still has to decide whether to ship, but the human no longer has to pretend that a limited attention span is enough to catch every systemic risk. That is a better division of labor, not a replacement for labor.

The strongest AI reviewer is the one that makes human judgment more informed, not less necessary.

A practical policy template

If a team wants to keep the human in command, the policy has to be concrete enough to survive a stressful week. A useful starting point is to classify actions by impact, not by convenience. Read-only inspection, local tests, and draft diffs should stay fast. Anything that touches secrets, production, permissions, data retention, or network access outside the allowlist should pause for review. Any change that alters an API contract, a schema, a migration, a deployment pipeline, or a dependency boundary should trigger human sign-off even if the tests look fine. The reason is simple: those changes can pass a machine check and still be wrong in the world the team actually ships into.

That same policy should require the agent to summarize uncertainty. What changed? What did not change? What remains unverified? Which assumptions were inferred rather than observed? Which tools were used? Which approvals were requested? Those questions sound basic, but they are what make a long session legible. They also make it easier for a human to resume work without rereading every prompt and every output. If the session summary is missing, the system is asking the human to trust memory instead of evidence. That is backward. Human command means the human gets evidence, not just a chat transcript.

1. Keep low-risk actions fast.
2. Pause for explicit approval before secrets, production, or new external domains.
3. Log every tool call, approval, and network decision.
4. Require a human for API, schema, permission, and migration changes.
5. Make the agent summarize uncertainty before handing off.

Human in command is what keeps autonomy honest

The point of all this is not to slow agents down. It is to make their speed usable. Agents are excellent at accelerating the work between decisions. They are not yet trustworthy as the source of the decisions themselves. The teams that benefit most from coding agents will not be the ones that say yes to everything. They will be the ones that build a system where the model can move quickly, the human can intervene immediately, and every important action remains legible, reversible, and accountable. That is the difference between using an agent and surrendering to one.

OpenAI’s safe-deployment guidance, its monitoring work, and Datadog’s system-level review example all point in the same direction. Keep the boundary technical. Keep the evidence rich. Keep the approval path clear. Keep the human in command. That is not anti-automation. It is the only way to scale automation without pretending that speed has replaced judgment.

Sources