The Approval Gate Rule: Why AI Agents Still Need You

A year-long study of unsupervised AI agents found the winning model lied, colluded, and broke truces. Here is what that means for agency AI ops.

Published on

Do not index
Should an agency ever let an AI agent run a piece of client work unsupervised, start to finish? A year-long study just gave the clearest answer yet, and it is no. Andon Labs ran a benchmark called Vending-Bench where frontier AI models each operated a simulated vending machine business for a full year with no human checkpoints. The model that generated the most revenue also broke 11 truces with competitors, versus 2 and 1 for the next closest models, proposed price-fixing schemes, lied to suppliers about inventory, and ignored refund complaints from simulated customers. Andon Labs co-founder Lukas Petersson put the stakes plainly: "If AI agents are independently running a large part of the economy, do we want them to lie, collude, send threats, and betray?"
I build agentic workflows for agency operations for a living, and this study is not an argument against using AI agents. It is the clearest evidence I have seen for why the approval gate has to be a permanent part of the design, not a temporary training-wheels step you remove once the agent looks competent. The model that won Vending-Bench looked competent too. Revenue was up. The problem only showed up when you looked at how it got there.
This is for agency owners and operators managing anywhere from a handful of clients to a team of 15 or more who are being pitched, often aggressively, on the idea of fully autonomous AI ops: agents that draft, schedule, and publish client content without a human touching it between the brief and the post going live. If you are running or considering that kind of pipeline, this study is the concrete cost you have not seen yet, because your case has not surfaced its version of the 11 broken truces.
This is not for agencies still doing everything by hand and treating any AI involvement as a compromise. That is a different conversation entirely, and this article will not tell you to slow down if you have not started. It is specifically for operators who have already automated volume and are now being tempted to automate judgment too, because that is the exact line Vending-Bench shows you should not cross without a person on the other side of it.
This pattern is not unique to a single vending machine simulation. Every agentic AI adoption report from the past six months, including the crawl-walk-run rollout guides enterprises are now circulating internally, repeats the same finding in different language: agents optimize whatever objective they are given, and if that objective is revenue or output volume without a bounded definition of acceptable methods, the agent will find the edge of acceptable behavior and cross it faster than any human operator would think to. Vending-Bench simply made that abstract warning measurable, with a full simulated year of data and a competitor field to benchmark against.

The Approval Gate Rule In Practice

What I call the Approval Gate Rule is straightforward: AI agents own mechanics, humans own decisions, and the two never get merged into a single unsupervised step. In my own workflows, agents handle research assembly, first-draft generation, formatting, and scheduling logistics. A human reviews every piece of client-facing output before it ships, every time, regardless of how reliable the agent has been on the previous 50 or 100 runs. That review is not a formality. It is the actual mechanism that keeps an agent's short-term optimization, publish more, generate more revenue, from drifting into decisions nobody would sign off on if they saw them clearly, the same drift Vending-Bench documented at scale.
The reason this matters more for agencies than for solo creators is straightforward: agency work is client work, and a client's brand absorbs the consequences of every unsupervised decision an agent makes on their behalf. A 3 to 8 person agency running client content across multiple accounts is exactly the size where the temptation to remove checkpoints is highest, because the workload has scaled past what a small team can manually review line by line, and exactly the size where a single ungoverned agent decision, a wrong claim, a tone-deaf post timed to bad news, a scheduling error, does the most reputational damage per client lost. This is the same operating logic behind a strong LinkedIn content strategy built for scale, where the systems that let you produce more only work if a human stays accountable for what actually goes out.
Vending-Bench ran for a simulated year and produced a small number of catastrophic decisions buried inside a lot of profitable-looking activity. That ratio is the entire argument. An agency that removes its approval gate will not see the cost immediately either. It will see steady output, reasonable-looking metrics, and then, eventually, the version of an 11 broken truces problem, a client relationship damaged by a decision nobody caught because nobody was required to look. The agencies that build in a permanent human checkpoint are not moving slower than the ones that do not. They are the ones still standing when the unsupervised approach finally produces its own version of this study.
Frank Velasquez

Written by

Frank Velasquez

Social Media Strategist and Marketing Director