What Can Your Agent Do When No One Is Watching?
AGENTS·OPINION·October 5, 2026·4 min read

What Can Your Agent Do When No One Is Watching?

The practical risk of AI agents is not in the model — it's in the permissions you gave it. How to apply the minimum access principle before your agent ecosystem grows without governance.

Every week a new agent appears. OpenAI has several under different names, xAI has Grok operating as an agent running complete processes, Anthropic has expanded Claude's capabilities to execute autonomous tasks, Meta is adding its own, DeepSeek is heading in that direction, and in the local ecosystem there are already people using tools like Ollama to run their own agents without leaving the network. The list is growing so fast that today it is hard to know how many agents a single person has active in parallel, let alone how many a company has.

That growth brings a question that is not always asked before activating something: what can this agent do when no one is watching?

It is not a rhetorical question. It is a technical and operational question with very concrete consequences. An agent that can read, write, send, approve, and connect to external systems without defined access limits is an agent that can cause a significant problem without anyone having instructed it to do so. Not because of the model's malice, but because of an absence of design.

The risk is not in the model. It is in the permissions you gave it.

Before getting into that, it is worth distinguishing the types of agents circulating today, because not all are equal and not all require the same considerations. There are agents that function as personal assistants: they help you write, search, organize, summarize. Their radius of action is essentially that of a conversation. There are more operational agents that connect to systems, execute process steps, send notifications, create records, or read databases. And there are agents that run in the background, with no visible interface, monitoring conditions, executing periodic actions, or responding to events. These last ones are the ones that require the most attention, because they are the ones that are least visible.

The question of whether a work agent and a personal agent require the same controls does not have a single answer, but it does have a clear principle: control must be proportional to access. A personal agent that helps you organize your private email operates in a different context than an agent that reads your company's CRM, processes customer requests, or has permissions to execute actions in an ERP. The context determines how much governance that agent needs, not the name of the model powering it.

What does apply in all cases, especially in the work context, is the logic of minimum access. An agent should be able to do exactly what its task requires, and nothing more. If its function is to check the status of an order, it does not need to be able to modify records. If its function is to draft response templates, it does not need access to the customer's financial information. This logic is not new: it is the same logic applied to onboarding any new employee who arrives without knowing the processes, without history inside the organization, and without verified credentials of what they can or cannot do well. You give them access to what they need to perform their function, with supervision, and you expand it as they demonstrate judgment. With agents, the logic should be the same.

The second control that makes a difference is human approval for high-impact actions. Not every action an agent takes needs confirmation, but those that are hard to reverse do. Sending an email to a customer, modifying a contract, making a publication, deleting a record, transferring a file to an external system: these are the kinds of actions that should go through human validation before being executed. Not because the agent cannot perform them, but because the margin of error in those actions has consequences that are not easily corrected.

The third element is the audit trail. An agent that leaves no record of what it did is an agent that, if something goes wrong, turns the investigation into a manual reconstruction exercise. Knowing what it did, when, under what instruction, and on what data is not just an auditing practice — it is the minimum foundation for being able to supervise and improve.

These three criteria — minimum access, approval for sensitive actions, and a log of what was executed — do not require a dedicated security team or complex infrastructure. They require that whoever deploys the agent asks that question before activating it: what can this agent do when no one is watching? If the answer is "I'm not sure," that is the starting point.

The agent ecosystem is going to keep growing. More models, more platforms, more use cases, more access to real systems. The governance of those agents cannot trail behind adoption. It has to lead it.