# Agent Review Lab

Canonical: https://www.danielnetzer.com/lab/agent-review
Author: Daniel Netzer
Model version: 1
Source frame checked: 2026-10-05

A synthetic simulation of task arrivals, human review and completion. Starting values are hypothetical; no vendor data calibrates the inputs.

## Question
How does human review capacity affect completion and waiting as agent output grows? More completed work and a growing decision queue can coexist.

## Why this lab exists
Give an agent a job. Give its decisions an owner. Then look at what happens when the work accelerates.

In my note on building a personal agent team, I started with the seat contract and the approval boundary. Before adding another bot, define the job it owns, the evidence it brings back, and the point where someone has to decide. This lab follows that thread into a capacity question.

A boundary can be clear on paper and still become a queue. If more agent work arrives at that boundary, what changes on the other side? Do we have more time to inspect it? Is the decision easier to make? Or have we simply made the waiting less visible?

The visualization makes that tension inspectable. Mint tasks can finish while amber tasks accumulate. The total looks productive; the decision queue tells another part of the story. Both can be true at once.

I’m using a deliberately small model so its assumptions can be challenged. The purpose is to explore where the next unit of effort belongs: more agent output, more review capacity, or better preparation for the decision. It is not a safe staffing ratio, a forecast, or an argument that people must review everything forever.

## What the recorded runs show
### More completed work. A longer decision queue.
The baseline produces 640 tasks across eight hours. Under its hypothetical review rules, 639 finish and one remains in review at 17:00. Double the agents at the same review capacity: 1,231 finish, 48 wait, and one remains in review.

Completion almost doubles. The review queue also grows. A total-output chart alone would hide the second result. These are outputs of this model, not measurements of a real team.

### Preparation changes a different part of the system.
Keep 40 agents and the same review requirement. Either increase available review effort to 120 minutes per hour or assume each review takes three minutes instead of six. Both configurations finish 1,279 tasks, with none waiting and one still in review at the cutoff.

The outcomes match because the modeled service rate matches. The real costs may differ. An evidence packet, tests or a clearer scope could help a reviewer, but this experiment has not measured that saving or whether decision quality holds.

### A quiet queue at 17:00 can hide a difficult morning.
The burst scenario sends 40% of the day’s work into its first hour. It ends with the same completion and waiting counts as the baseline, but mean waiting among reviews that started is about 30 minutes instead of zero.

Look at time as well as totals. The mean excludes tasks still waiting; their count and oldest age are shown separately. Neither an empty queue nor a short wait establishes that the right decisions were made.

## Who has to approve the work?
An approval policy creates review demand. The team still needs enough capacity to carry it out.

SOC 2 Type 2 is sometimes described as a ban on fully agentic merges. The reviewed evidence does not support that blanket claim. The change-management criterion, CC8.1, addresses authorization, testing, approval and controlled implementation. It does not prescribe a person clicking every merge button.

The controls an organization implements can be more specific. If its process requires an independent person’s approval, an agent bypassing that step breaks the stated control. A bot merging after the required approval is a different workflow. Whether an entirely automated approval path is appropriate depends on its risks, controls and evidence, assessed with the responsible control owner and auditor.

Type 2 examines how controls operated over a stated period. Generating code, reviewing it, authorizing a change, merging a branch and deploying production are separate actions to trace. A passing test suite or a second agent identity does not, by itself, establish effective approval or independence.

The workbench’s approval-policy presets change only the fraction requiring human review. Every task sets it to 100%; selected tasks sets it to an illustrative 10%. Your other inputs stay fixed and the previous scenario is pinned for comparison. Neither percentage is a SOC 2 requirement. The model assigns tasks mechanically; it does not classify risk, inspect approval evidence or measure missed defects.

Use the comparison to ask whether the workload fits a chosen policy. A shorter queue cannot tell you whether that policy is sound. Before changing a real approval boundary, examine the applicable controls and the quality of the decisions alongside waiting time.

- [CC8.1 and an organization’s implemented controls](https://wac-cdn.atlassian.com/misc-assets/pdfs/Atlassian_New_Products_SOC_2_Type_2_Final_Report.pdf#page=70): Atlassian / Coalfire, report period 30 June–30 September 2024, page 70. An issued report illustrates the distinction; it does not establish a current AI policy or a universal rule.
- [AI within the existing SOC criteria](https://schneiderdowns.com/our-thoughts-on/ai-in-soc-examinations-what-the-aicpas-new-tqa-section-9561-means-for-service-organizations/): Bill Deller, Schneider Downs, 30 September 2026. Professional interpretation of AICPA TQA 9561, not approval of a particular autonomous pipeline.

Checked 5 October 2026. The full AICPA criteria and new AI TQA downloads require account access and were not read in this review. The interpretation draws on reproduced CC8.1 wording in an issued report and the attributed practitioner analysis. This lab does not assess SOC 2 compliance.

## Connected notes
- [Personal agent teams need a contract](https://www.danielnetzer.com/build-log/how-im-setting-up-a-personal-agent-team-on-grok-bot): Roles, evidence, approval boundaries. The starting point for this experiment.
- [Evaluate the work inside your harness](https://www.danielnetzer.com/build-log/same-day-sol-and-opus-cost-curves): A model ranking cannot tell you the economics or quality of your own workflow.
- [The work continues after the code ships](https://www.danielnetzer.com/build-log/cursor-after-spacex): Operate and learn belong in the system, with accountable owners.

## Bring the question to your workflow
- **Name the decision:** Which output needs a person, what are they deciding, and who owns the consequence? Separate routine checks from consequential approval.
- **Observe the queue:** Record arrival, review start and completion times. Include work that is still waiting, and look at bursts rather than only the daily average.
- **Test one intervention:** Try a smaller batch, clearer evidence or more review time. Measure decision quality and rework alongside delay before changing the approval boundary.

## Inputs
Agents (0–80), tasks/agent/hour (0–12), required review (0–100%), pooled reviewer-minutes/hour (0–240), minutes/review (1–30), steady or burst arrivals. Defaults: 20, 4, 10%, 60, 6, steady.

## Method and limits
- Capsules are a bounded illustration of recorded events. Their travel, lift and landing time are visual pacing, not modeled task latency. The counts always include every task.
- One task is agent output arriving at a decision boundary. It is not a token, API call, transcript or pull request benchmark.
- The workday lasts 480 minutes. Arrivals are evenly spaced at interval midpoints; the burst option puts 40% in the first hour and 60% in the next seven.
- The review fraction is assigned cumulatively across the full stream, rounding the final review count down to a whole task. Required tasks complete only when their human review finishes. Other tasks complete on arrival under this model's assumption.
- Review uses first-in, first-out ordering and fixed effort. Available reviewer-minutes are one equivalent pooled service rate, not separate people's calendars. Completion at exactly 17:00 counts; unfinished service stays visible.
- Mean waiting time includes reviews that started, including work still in review. Still-waiting tasks are excluded from that mean and shown separately with their count and oldest age.
- Prepared review changes the assumed minutes per review. It does not prove that better context or automated checks achieve that saving.
- Approval-policy presets use hypothetical review fractions of 100% and 10%. They preserve other inputs and compare against the prior scenario; neither represents a SOC 2 requirement or an assessment of control effectiveness.
- Task quality, review accuracy, routing errors, fatigue, dependencies, rework and business value are outside the model. Lowering review requirements does not establish safe automation.

## Recorded synthetic scenarios
| Scenario | Arrivals | Completed | Waiting | In service | Mean wait among started reviews (minutes) |
| --- | ---: | ---: | ---: | ---: | ---: |
| Baseline | 640 | 639 | 0 | 1 | 0.00 |
| Double the agents | 1280 | 1231 | 48 | 1 | 88.88 |
| More review capacity | 1280 | 1279 | 0 | 1 | 0.00 |
| Prepared review (assumed) | 1280 | 1279 | 0 | 1 | 0.00 |
| Work arrives in bursts | 640 | 639 | 0 | 1 | 30.20 |

These are deterministic model outputs, not real-team observations. Compare configurations and rerun the model before generalizing.

## In conversation with
- [OpenAI: DevDay 2026 recap](https://openai.com/index/devday-2026-recap/) (2026-09-29; Vendor announcement). Ongoing agents make work that continues to arrive a timely operational question. An announcement does not establish review capacity or delivered business value.
- [Anthropic: Measurements for understanding the pace of AI development](https://www.anthropic.com/institute/measuring-pace-of-ai-development) (publication date unverified; Internal measurements · August 2026). Monitoring coverage, review latency and escalation describe different parts of oversight. Internal reporting does not calibrate this model. Agent actions, transcripts and operational tasks are different units.
- [Addy Osmani: Brownfield Agentic Engineering](https://addyosmani.com/blog/brownfield-agentic-engineering/) (2026-09-14; Practitioner argument). Parallel agent work still needs verification, review effort and accountable ownership. The page identifies an Anthropic affiliation; this is practitioner guidance, not independent validation of our simulation.
- [Claire Vo: LinkedIn triage, CXO email, and getting through the list](https://clairevo.com/notes/linkedin-triage-cxo-email-and-the-list) (2026-09-21; Workflow account). Checking source context before presenting a focused human decision suggests review preparation as another lever. A direct account without a measured review-time saving.
- [Gergely Orosz: What is happening with code reviews?](https://newsletter.pragmaticengineer.com/p/what-is-happening-with-code-reviews) (2026-09-08; Reported engineering practice · accessible preview). Teams are exploring different review boundaries: risk, plans, tests, smaller changes and human review of agent feedback. Reported approaches are context, not calibration. The paid remainder was not reviewed. This lab does not evaluate which boundary is safe.
- [Lenny Rachitsky / How I AI: How Warp ships with AI factories](https://www.lennysnewsletter.com/p/how-i-ai-metas-muse-review-how-warp) (2026-09-21; Interview summary · Zach Lloyd’s account). The interview places review delay and human interactions alongside agent output. That gives this capacity question a concrete operational setting. The reported timings are not independently verified, and are not used as this lab’s input values. The newsletter byline is Lenny; the workflow account is Zach Lloyd’s.
- [METR: We are Changing our Developer Productivity Experiment Design](https://metr.org/blog/2026-02-24-uplift-update/) (2026-02-24; Independent research update). Selection and concurrent-agent time measurement complicate estimates of productivity. Task flow alone cannot establish productivity, quality or business value; the update does not establish a precise current effect.

[Scenario data](https://www.danielnetzer.com/lab/agent-review/scenarios.json)
[Follow Daniel's writing](https://www.linkedin.com/in/daniel-netzer)
