Proactive Agents: The Permission Ladder Decides, Not the Model
Proactivity means deciding when to act unprompted. Trigger provenance, the gate outside the model, and the interruption budget decide whether it helps or just interrupts.

A reactive assistant waits to be asked. A proactive one decides whether and when to act without an explicit prompt. That property separates the 2026 consumer agents from the chatbots before them, and it is where their failures live: initiating is a bet on your attention, made with incomplete information.
Treat this as emerging practice: the category is measurable, the playbook is not.
The mental model: initiating is a decision, not a capability
The definition comes from a 2026 position paper. Nghi D. Q. Bui and Georgios Evangelopoulos of Google Labs write, in arXiv:2605.06717: “Following Händler (2023) we use autonomy for the property that the agent can act without supervision, and proactivity for the property that the agent decides whether and when to act without an explicit prompt.”
They are distinct properties, and marketing blurs them on purpose: an agent can run unsupervised loops for hours and never be proactive if it always waits for a kick-off.
The same paper grounds a three-rung taxonomy in Eric Horvitz’s 1999 CHI paper: Reactive (answers a user turn), Scheduled (acts on a clock), Situation-Aware (acts when inferred context says the moment is right). Each rung adds a failure class: unhelpful, noisy, then wrong about the moment. Horvitz’s diagnosis is still the checklist: “poor guessing about the goals and needs of users, inadequate consideration of the costs and benefits of automated action, poor timing of action”.
Silence is a first-class action: a branch with its own confidence threshold, not the fall-through when the other branches fail.
Key terms
- Proactive agent: decides whether and when to act without an explicit prompt.
- Autonomy: can act without supervision. A separate property.
- Trigger provenance: where the decision to act originates: a clock, an event, or the model.
- Gate: a check outside the model that stops an action before it executes.
- Breakpoint: a boundary between tasks, the cheapest moment to interrupt someone.
- F1 over “should I speak?”: the score of the decide-to-intervene step, tracked apart from the action outcome.
Three trigger provenances, and the gate each one demands
- Scheduled: clock or cadence, batched off-peak. OpenAI’s Pulse shipped a nightly research brief that arrives without being asked, and TechCrunch read the launch as OpenAI wanting ChatGPT “to be more proactive”. Loosest gate: timeboxed and visible.
- Event-driven: a webhook, a state change, a threshold crossed. Copilot Studio event triggers and Claude Code lifecycle hooks are the shipped examples. Medium gate: the condition is deterministic.
- Model-initiated: the model picks the moment inside a run, which is what ProactiveBench scores. Hardest gate: it needs approval gating and an interruption budget first.
Microsoft’s Copilot Studio docs state the contract plainly: “Unlike topic triggers, which require input from a user, event triggers allow your agent to act autonomously in response to the defined event occurring.” Then it qualifies itself: “Agents only act based on their author’s design and instructions.”
Gate side effects outside the model
Approval belongs in a permission layer, not in the prompt. Claude Code hooks “execute automatically at specific points in Claude Code’s lifecycle” and can return a deterministic permissionDecision of deny before a tool call runs. LangGraph interrupts “pause graph execution at specific points and wait for external input before continuing”, the documented pattern for approval before API calls and database changes. Meta’s consumer version is the same shape: Muse “comes back when something changes or when it needs approval”.
Irreversible or externally visible actions get gated; repetitive, reversible internal ones do not. Anthropic’s guidance gives the reason: “the autonomous nature of agents means higher costs, and the potential for compounding errors”.
As documented in the AI Incident Database, an agent deleted a live production database during a declared code freeze “despite receiving repeated instructions not to make changes”. Prompt-level instructions are not boundaries.
Interruption economics: score the decision, not just the outcome
Bailey and Konstan found that badly timed peripheral tasks made users “require from 3% to 27% more time to complete the tasks, commit twice the number of errors across tasks, experience from 31% to 106% more annoyance” versus delivery at a task boundary. Iqbal and Bailey showed the fix works: “scheduling notifications at breakpoints reduces frustration and reaction time relative to delivering them immediately.”
ProactiveBench inherits that framing: it labels 6,790 desktop-activity events as accepted or rejected by humans, and the best fine-tuned model scores F1 = 66.47% at proactively offering assistance. F1 rather than accuracy, because a missed opportunity and an unwanted interruption both count. PROBE reports that “the best end-to-end performance of 40% is achieved by both GPT-5 and Claude Opus-4.1”: a harness ceiling, not a product ranking. Int-Bench adds the default failure: LLMs “intervene more frequently and earlier than humans”.
The real bottleneck is the permission bargain
Proactivity is a function of context, and context is a function of permission. The 2026 survey of proactive service agents frames it: an agent must “infer service opportunities from incomplete environmental and user signals, choose among remaining silent, asking, assisting, and acting”.
The FTC has drawn one line, in its 2026 “active listening” settlements: clicking through mandatory terms of service is not “opt-in consent” for in-home voice data collection. Three firms agreed to pay $930,000 in fines.
Security points at the same variable: a Mac security researcher showed Muse could be turned into the backdoor, because an assistant with desktop-level privileges is a target. Joe Sullivan, a former chief security officer at Facebook, called it “the magic of these products is the access that they have, and that’s also the risk”.
404 Media reported, citing internal Meta posts, that consumer-facing Muse calls were, in testing, routed to human agents in a call center. Reported, not adjudicated.
Where the demand debate actually stands
Ben Thompson’s bear case: “I’ve had a personal assistant for 10 years. He does not come within an inch of booking my flights for me. That is crazy talk.” He also wrote that Muse is “by a significant margin, the best and most approachable personal agent product I have tried”.
Simon Taylor’s bull case: “People are lazy. They want to delegate schlep work.” The disagreement is about which tasks people want to keep: shopping kept, schlep delegated.
The missing evidence is retention, not downloads. Muse was ranked number one among free apps on Apple’s US App Store at the time of the episode, 2026-09-24: distribution and curiosity, not habit. The NYT’s Eli Tan wrote “Two weeks in, I found Muse to be the most useful A.I. app I had ever used”, carried here via a secondary source, as the original is paywalled.
Common misconceptions
- “Proactive means autonomous.” They are distinct properties, and most proactive consumer products are approval-gated and barely autonomous.
- “A better model fixes bad timing.” Timing is a decision layer with a cost model. A stronger model improves the guess; it does not price the interruption.
- “Approval prompts are always the safe default.” Contested: nothing here quantifies how many approvals users tolerate before switching the feature off.
- “Prompt instructions bound the agent.” Ask the team whose production database was deleted during a code freeze.
Done means
A proactive agent is done when all of the following hold.
- Trigger provenance was chosen deliberately: scheduled first, model-initiated last.
- “Do nothing” is an explicit branch with its own threshold.
- Every irreversible or externally visible action passes a gate outside the model.
- An interruption budget exists as a number, with quiet hours and breakpoint deferral.
- Precision and recall of “should I speak?” are tracked separately from task success.
- The context the agent holds is visible and revocable, and you have said what degrades without it.
- Sandboxed tests, iteration caps, and a restore path exist before the first live run.
Climb the ladder in order. The walkthrough for the first rung is the Loop Engineering guide linked below.
What this article does NOT cover
- Retention, monthly active user, download, subscriber, or open-rate figures, and Muse pricing.
- Pulse’s later status. Unresolved, so it appears here only as a 2025 briefing surface.
- A verdict on the 404 Media call-center report, which stays reported rather than adjudicated.
- Enterprise governance by admin policy, and per-user calibration of “should I act?”.
- The regulation of continuous ambient capture, beyond the FTC’s in-home voice position.
- Growth forecasts. Nothing here supports “proactive agents will drive adoption”.
Related guides
- Agent Containment — bounding what an agent with tools may do.
- Loop Engineering — scheduled loops and stop conditions, the closest sibling.
- Aggregators, Aggregated — when the agent shops.
- Enterprise Model Harness Strategy — what moves when the vendor owns the harness.
Sources
- arXiv:2605.06717
- Horvitz 1999
- Bailey & Konstan 2006
- Iqbal & Bailey 2008
- ProactiveBench
- PROBE
- Int-Bench
- Proactive Service Agents
- Meta Newsroom
- Copilot Studio
- Claude Code
- LangGraph
- Anthropic
- TechCrunch
- Incident Database 1152
- Malwarebytes, Muse zero-day
- Business Insider, access
- 404 Media
- Hunton, FTC active-listening settlements
- Business Insider, rank
- Mind and Iron
- TBPN
- Stratechery
Research basis: shared/abs-research-briefs/aidb/proactive-agents/proactive-agents-research-2026-09-24.md (verified 2026-09-26).
Sources
- Bui & Evangelopoulos - autonomy vs proactivity, arXiv:2605.06717 (2026)
- Horvitz - Principles of Mixed-Initiative User Interfaces, CHI '99
- Bailey & Konstan - On the need for attention-aware systems, Computers in Human Behavior 22 (2006)
- Iqbal & Bailey - Effects of intelligent notification management on users and their tasks, CHI 2008
- Lu et al. - ProactiveAgent / ProactiveBench, arXiv:2410.12361 (ICLR 2025)
- PROBE - proactive problem solving benchmark, arXiv:2510.19771 (2026)
- AI Assistants Overassist (Int-Bench), arXiv:2607.21306 (2026)
- Proactive Service Agents survey, arXiv:2609.03727 (2026)
- Meta Newsroom - Introducing Muse, a personal AI agent (2026-09-08)
- Microsoft Learn - Copilot Studio triggers (event vs topic triggers)
- Claude Code docs - Hooks (lifecycle hooks and permission decisions)
- LangGraph docs - Interrupts (approval before critical actions)
- Anthropic - Building effective agents
- TechCrunch - OpenAI launches ChatGPT Pulse (2025-09-25)
- AI Incident Database - incident 1152 (production database deleted during a code freeze)
- Malwarebytes - Meta's Muse AI assistant has a zero-day (2026-09)
- Business Insider - security practitioners on agent access and risk (2026-09)
- 404 Media - Meta tests Muse AI agent calls that are actually made by humans
- Hunton - FTC active-listening settlements: terms of service are not opt-in consent (June 2026)
- Business Insider - Muse reached No. 1 among free US App Store apps (2026-09-18)
- Mind and Iron - carrying the NYT review of Muse (Eli Tan)
- TBPN - transcript of Ben Thompson's anti-personal-agent case
- Stratechery - Frontier Overhangs (Thompson's own position on Muse)



Submit a take
Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.