← Back to news
MCP Needs a Human at the Controls

PXHERE via Wikimedia Commons: https://commons.wikimedia.org/wiki/File:Computer_developer.jpg

27/08/2026

MCP Needs a Human at the Controls

Every new MCP server is a new permission boundary

The appeal of MCP is easy to understand: a model that can reach real systems stops being a demo and starts being useful. With one protocol, a coding agent can inspect a repository, query a knowledge base, open an issue, fetch logs, or prepare a pull request inside an actual workflow. That is the point. MCP reduces integration friction and makes tool use feel uniform. But the same simplicity also turns a single chat surface into a gateway to many tools, many identities, and many failure modes. The moment a model can do more than describe a plan, the question changes from “Is the agent smart enough?” to “Who controls what the agent is allowed to touch?”

That shift matters because tool access is not just another feature. It is a trust decision. A coding assistant connected to a Git server is already operating in a privileged environment. Connect it to a ticketing system, a secret store, a CI provider, and a production log viewer, and the blast radius grows with every integration. The same automation that saves twenty minutes can also leak a token, follow a malicious prompt buried in a document, or turn a small mistake into a large, hard-to-reverse one. The risk is not theoretical; it is the normal price of giving software the ability to act on your behalf.

One useful way to think about MCP is that it is not just a transport layer. It is a delegation layer. The protocol makes tool discovery and context retrieval easier, but it also makes privilege aggregation easier. If a team treats “we use MCP” as the end of the design problem, it will discover the missing parts the first time an agent asks for a write scope it should never have had in the first place.

MCP is not just a transport layer. It is a delegation layer.

Why the protocol itself is not the safety model

MCP is useful precisely because it standardizes how agents discover tools and context. Anthropic’s engineering post on code execution with MCP explains the upside of that standard: when there are many tools, loading all definitions into the prompt wastes context, and a code-based harness can fetch only what is needed. That is good engineering. But a standard is not a safety policy. A protocol can make a risky workflow elegant without making it safe.

The security issues are familiar to any team that has managed production credentials. Data exfiltration, impersonation, prompt injection, and unclear attribution do not disappear when the interface becomes more polished. In fact, they often become harder to notice because the agent is operating inside a workflow that looks productive. A tool call that returns a neat summary can hide the fact that the agent saw more data than it needed. A well-formed action can still be the wrong action. A correct-looking change can still be unauthorized.

Prompt injection becomes more dangerous when agents can reach multiple systems. A poisoned issue comment, a malicious document, or a crafted response from a connected service can steer a model toward actions a human never intended. If the agent can also write back into repositories, deployment systems, or external APIs, then the attack surface grows from “bad text” to “bad text plus privileged action.” That is why the safety model has to sit above the protocol, not inside it.

GitHub’s recent agentic security principles say the quiet part out loud: the more agentic a product becomes, the more it needs a human-in-the-loop element. OpenAI’s Codex guidance says the same thing in operational language: keep the agent inside clear technical boundaries, let low-risk work flow, and make higher-risk actions explicit. Those are not conservative footnotes. They are the operating system of responsible delegation.

What good boundaries look like

Start with read-only by default

If an MCP server is mainly for context retrieval, make it read-only. If a workflow needs write access, do not grant it because it is convenient; grant it because a human has reviewed the risk and the fallback. A read-only server can summarize logs, inspect tickets, and retrieve documentation. A write-capable server can change state. Those two jobs belong in different trust tiers.

Scope identity to the task, not the organization

An agent should not inherit a giant human-shaped set of permissions just because a person has them. The smallest useful identity wins. If the task is to update a README, the agent does not need a secret store. If the task is to propose a fix in a branch, the agent does not need branch-merge rights. GitHub’s guidance on attribution and authorized context is important here: the initiating user and the agent should both be visible in the chain of responsibility, and the agent should only gather context from users who are allowed to provide it.

Put approvals where the blast radius is real

Approval should not be ceremonial. It should happen where an action becomes expensive to undo: before network egress to an unfamiliar host, before reading or writing secrets, before altering deployment settings, before publishing a pull request that can trigger a pipeline, before calling a production API, before turning a suggestion into a state change. If every action asks for approval, the system becomes unusable. If no action asks for approval, the system becomes a liability. The art is to align approval with irreversible or high-impact steps.

Log enough to reconstruct intent

Traditional logs tell you that a process changed a file. Agent-native logs should tell you why the agent chose that file, what tool input it sent, which approval gates were crossed, and which result it returned. OpenAI’s Codex article highlights telemetry and audit trails for exactly this reason. When things go wrong, security teams need more than a diff. They need a story they can replay. Good logs also help teams tune the boundaries: if the agent keeps requesting a network path it never uses, or a server that adds no value, the telemetry tells you where to simplify or tighten policy.

Human-in-the-loop is not a speed bump

Some teams still describe approval as friction, as if the ideal product were one that never makes a person think. That framing is backwards. The point of human review is not to punish the user with delay. The point is to keep the machine from making decisions that only a human can responsibly make. A model can compare tool outputs faster than any person. It cannot decide whether the output belongs in a release candidate, whether a data path crosses a policy boundary, or whether a compromise is acceptable for the business.

That is why the best agent workflows make judgment visible. A good flow says: here is the plan, here are the tools, here are the scopes, here is the risk, here are the tests, here is what changed, and here is the human who approves the last step. A bad flow says: trust me, the agent handled it. Software teams already know how that story ends. Every incident review that follows is a lesson in why invisible delegation is still delegation.

In practice, that means splitting work into three layers. First comes exploration: the agent gathers context, narrows the search space, and drafts options. Second comes verification: the agent runs tests, checks assumptions, and produces evidence. Third comes authorization: a human decides whether the proposed action belongs in the real system. The first two layers can be fast. The third layer must stay deliberate.

MCP makes tool composition easier, so review discipline has to get stricter

Anthropic’s discussion of code execution with MCP is useful because it shows the upside of composability. When tools become plentiful, direct tool calls can waste context. Code execution can load only what is needed and process the rest outside the model. That can reduce token use and make workflows more efficient. But the same pattern also reinforces the central lesson for software teams: the more powerful the tool graph, the more important the harness around it.

A coding agent that can chain servers together is effectively writing operational code on your behalf. That code may be transient, but the consequences are not. If the agent can read a transcript from one system and write it into another, then your approval policy must cover both the source and destination. If the agent can search across repositories and then open a pull request, then the reviewer needs to know whether the search results were sufficient or whether the agent simply took the most convenient path. If the agent can connect a doc tool to a deployment tool, then the human has to ask whether a clean-looking handoff also crosses a policy line.

This is why “agentic” should never mean “unreviewed.” It should mean that the software can operate in a limited envelope while the human remains the ultimate decision-maker. There is nothing anti-AI about that. It is how you make the AI useful enough to trust. The team that can say “the agent can propose, but only a human can commit” is much better positioned than the team that lets the machine decide when a change is done.

The policy that teams can actually ship

  • Define trust tiers for every server. Read-only, write-only, secret-access, and production-touching systems should not share the same defaults.
  • Require explicit owner approval for new servers. If a team wants to connect a new MCP server, someone must sign up for the risk and the rollback plan.
  • Make write actions visible before they happen. A human should see the intended mutation, the target, and the consequence before execution.
  • Prefer short-lived credentials. Revoke agent tokens when the task ends, and scope credentials to the smallest environment possible.
  • Record enough telemetry to audit decisions. Tool calls, approvals, denials, and outputs should be reviewable after the fact.
  • Use tests as gates, not decorations. If an agent suggests code, the default question is not “does it look clever?” but “what evidence says it is safe?”

Those rules are not special to MCP. They are the same rules teams should have used for privileged scripts, deployment bots, and service accounts years ago. MCP simply makes the gap harder to ignore because it packages more power into a cleaner interface. Clean interfaces are good. Clean interfaces also make it easier to forget how much authority you just handed over.

A practical rollout should start small. Put one server behind a narrow scope. Test a read-only workflow first. Add a write path only when the logging, review, and rollback story is already working. If the team cannot explain what happens when the agent is wrong, the team is not ready to let the agent act. That rule sounds simple because it is simple.

The real promise of MCP is controlled delegation

The strongest version of the case for MCP is not that agents can replace developers. It is that they can become better assistants when the surrounding system is designed for control. A coding agent that can gather context, draft a fix, and propose tests can save real time. A tool-enabled agent that can reach a dozen systems without guardrails can save a few minutes and create a months-long cleanup project. The difference is not intelligence. It is governance.

That is why the human in the loop needs to remain visible in the workflow, not hidden in a policy document no one reads. Humans decide the goal, set the boundaries, review the evidence, and own the merge. Agents can accelerate the work. They should not inherit the authority. The future of coding agents is not full autonomy. It is disciplined delegation with human command at the center.

MCP makes that future practical. It does not make it optional.

Sources