Survey finding: As of October 3, 2026, AI agents have moved from answering questions to reading data, using tools, changing files, and checking results. Software development is among the best-documented uses. But can run for a long time, can finish a task, and should be trusted with authority are different claims. This article synthesizes public primary sources. It is not an independent vendor test, and neither a demo nor a user count proves reliability.
Day 4 covered September's frontier incidents and security failures. Day 8 asks a more ordinary question: What work can an agent complete today, how do we move from one agent to a cluster, how do we optimize it, and when is training an agent model justified? General readers can start with the definition, maturity model, and five developments. Builders can continue to the project map, six implementation gates, model development, and evaluation. Here, “can do” means a product or study has demonstrated a capability; “reliable” requires evidence on the particular task.
What an AI agent actually is
A chat model responds to a prompt. A fixed automation follows predefined steps, such as sending a weekly report. An AI agent has a goal, tools, and feedback, and chooses some next steps during execution: after reading a file it may search for missing evidence, run code, and check the result. This describes behavior; it does not mean independent understanding or freedom from human accountability. Anthropic's study of deployed autonomy notes that there is no single accepted definition and uses “AI systems equipped with tools that take actions” as a measurable one. OpenAI's explanation of the agent loop describes the recurring interaction among user, model, and tools.
Suppose the job is to analyze ten public documents. A chat model answers from the text it was given. A retrieval workflow follows preset search steps. An agent might decide which source to inspect first and revisit the original page when dates conflict. If the job is always “download A, extract B, save C,” conventional code is usually easier to test and avoids paying for step-by-step model reasoning. Anthropic's architecture guide distinguishes predetermined workflows from systems in which the model dynamically directs tool use, and argues for the simplest structure that works.
| Form | Who decides the next step? | Appropriate task | Main limit |
|---|---|---|---|
| Chat response | A person asks again | Explanation, draft, one-off analysis | No independent check of external outcomes |
| Fixed workflow | The programmer | Stable, repetitive steps and exceptions | New cases require new rules |
| Bounded agent | The model, within allowed tools | Narrow tasks requiring search, judgment, retries | Wrong tools or treating data as instructions |
| Long-running agent | Model plus runtime, checkpoints and checks | Multi-step software or research work | Context, cost, error accumulation and handoff |
A maturity map for 2026
This is this article's classification, not a universal score for products. Moving up a level requires more than a better model: the runtime, acceptance test, permissions, and accountable owner also change.
Model writes; a person performs the action
Code sets the steps; model handles a local judgment
Model reads, searches, and retries
Persists artifacts across time and contexts
Changes external state inside approved boundaries
The most defensible snapshot is that tool use is widespread, longer tasks are improving rapidly, and production still demands task-specific permissions and measured success. A browser-capable agent should not automatically be trusted with payments, medical work, deletion, or publication.
Five visible developments
1 Coding agents now work through whole loops
Coding agents can search a repository, edit several files, run tests, read errors, and return a reviewable change. OpenAI's Codex app overview shows parallel work and isolated worktrees. Anthropic's long-running harness research reports that even a frontier model with context compaction may attempt too much, hand off poorly, or declare completion too early across sessions; durable progress artifacts and tests help. These are product and engineering advances, not evidence that arbitrary changes are safe to merge without review.
Usage is growing too. OpenAI's own usage analysis documents longer, harder Codex delegation. Anthropic's Claude Code and API study found that the longest few Claude Code turns grew from under 25 minutes to over 45 minutes in three months, while the median turn was about 45 seconds. These are observations from specific products and windows. Longer runtime can also mean harder work, so it is not itself a success rate.
2 Knowledge work is moving from answers to inspectable artifacts
Research, data preparation, reports, slides, and cross-tool tasks now appear in agent products and enterprise pilots. OpenAI reported more than five million weekly Codex users by mid-2026, about one fifth in knowledge-work roles. These are vendor-reported usage figures, not independent industry penetration or proof that every artifact is correct. More useful questions are whether quotations lead back to the source, table figures can be recalculated, and a coworker can inspect the files.
Long-running infrastructure is also becoming a product. OpenAI's September 2026 Agents API public beta offers a hosted harness and a choice of execution environments; Anthropic's Managed Agents engineering account addresses persistent tasks and evolving harnesses. A beta, platform capability, or customer example proves availability or adoption, not fitness for your particular workflow.
3 Browser and computer use broaden reach and failure modes
A screen-using agent can theoretically cross systems without APIs. It must still handle changed pages, logins, dialogs, stale data, and actions with external effects. The Odysseys paper, built from realistic long web tasks, reports 44.5% success for its strongest tested model. WeaveBench, requiring GUI, command-line and coding actions together, reports 41.2% for its best pairing. These are different tasks, settings, and scoring rules—not a shared leaderboard. Together they show that tested long-horizon, cross-interface tasks remain far from consistently correct.
4 Multi-agent systems and protocols mature, but coordination is costly
The MCP 2026-07-28 release documents evolving tool, resource, authorization, and service interfaces. Google Cloud's agent guide positions MCP as a way to connect tools and data, and A2A as a way for different agents to communicate. These are interoperability interfaces. They do not establish whether source data is trustworthy, privileges are narrow, or two agents are independently checking each other's errors.
Several agents help when work can be separated—for example, one gathers sources, another checks arithmetic, and a third compares claims with the originals. If all three rely on the same wrong source or exploit the same grading loophole, extra agents do not create independent evidence. Measure coordination, duplicated calls, and human arbitration. Anthropic's evaluation guide emphasizes tool trajectories and environment state, rather than a final “done” message alone.
5 Organizations are governing agent identity and lifecycle
An agent that reads company documents, calls APIs and edits records is another actor requiring identity and audit. Microsoft Agent 365's documentation groups observation, governance and protection in a control plane. Its access guide distinguishes acting on behalf of a user, acting as an application, and an agent with persistent identity. This indicates a design direction; actual safety depends on deployed privileges, tool behavior, and recovery.
In Anthropic's API tool-call sample, software engineering represented nearly half of observed calls, and most actions appeared low-risk and reversible. The researchers warn that this is one provider's tool-call sample, not proof that half of all industry agents write code. Higher-risk uses appear, but providers cannot always determine from individual calls whether an action happened in production or an evaluation.
From one agent to a cluster: coordination is the new problem
A single agent here means one model-led tool loop. An agent cluster means multiple agent instances with explicit roles working on one task. Several instances may use the same model; a single agent may call a model served by a distributed GPU cluster. Agent count, model count, and GPU count are separate dimensions. The new engineering questions are who assigns work, which data each worker may see, how work is handed over, who accepts the final result, and who recovers failures. The OpenAI Agents SDK orchestration guide distinguishes a manager calling specialists from a handoff that transfers control. Google Cloud's architecture guide covers sequential, parallel, and loop patterns.
| Shape | Control flow | Example: surveying public product updates | Reason to move up |
|---|---|---|---|
| Fixed workflow | Code searches, extracts, deduplicates, and writes in a set order | Same sources and fields each week | Stay here when sources and rules are stable |
| One agent | The model chooses which source to inspect next and when to stop | Read-only search, date comparison, explicit unknowns | Dynamic investigation with checkable results |
| Manager and specialists | A manager delegates bounded tasks and accepts structured outputs | Separate source search and date checking | Measured single-agent failures and separable subtasks |
| Parallel cluster | Workers investigate independently; a system schedules and merges | Assign vendors to workers; verify without copying their conclusions | Broad tasks, time pressure, and measured quality or cost gains |
| Cross-service cluster | Services have identities, queues, state, and permissions | Search, analysis, and audit services across teams | Multiple data domains and clear operating owners |
The upgrade test is not “can we attach more agents?” It is whether accepted completion, human rework, end-to-end latency, and total cost improve on the same task set. Anthropic's account of its multi-agent research system finds parallel exploration especially useful for broad, independent investigations, while warning about much higher token use and tasks that require tightly shared context. Its internal eval does not establish a general return on investment. Coding workers editing the same files need isolated workspaces, an integration order, and a conflict owner; otherwise parallelism creates merge conflicts.
A handoff can fit on one card: task_id, objective, allowed tools, forbidden actions, input provenance, output schema, deadline and budget, acceptance criteria, evidence links, unknowns. A manager must inspect evidence before accepting a specialist's “done.” Magentic-One's source illustrates an orchestrator, web and file workers, coding and terminal roles, and a progress ledger. It is a research reference, not a safe production default.
A project and reading map, beginner to frontier
These projects are grouped by the part of the system they teach, not GitHub popularity. This is an October 3, 2026 entry map; check current versions, licenses, model requirements, sandboxes, and costs before following a tutorial. Pick one project per layer and produce a repeatable result before adding another framework.
| Layer and primary project | What to inspect | Deliverable |
|---|---|---|
| First concepts: Hugging Face Agents Course, smolagents | The think/act/observe loop and tool contracts; compare structured calls with generated code, which needs isolation | Answer a question with two read-only tools and retain every call and error |
| Single-agent application: OpenAI Agents SDK or Google ADK | Tools, structured output, state, human intervention, and traces | One bounded agent with sources and a stop condition, compared with a script |
| Explicit flow and handoff: LangGraph, Microsoft Agent Framework | State graphs, checkpoints, branches, workflows, and handoffs | Search → calculate → review, with pause and resume |
| Specialist collaboration: CrewAI Crews/Flows, Magentic-One | Role-based collaboration and manager planning; compare on identical tasks rather than assuming more roles help | One manager, two specialists, one independent checker, and measured overhead |
| Real coding work: OpenHands, SWE-agent | Isolated workspaces, file and terminal tools, executable tests, reviewable diffs | Repair a known bug in a disposable repository and run held-out tests |
| Evaluation environments: BrowserGym/AgentLab, τ²-bench, OSWorld 2.1 | Reproducible browser, support-tool, and desktop tasks; scores across different suites are not interchangeable | Pin the benchmark version and split; report pass rate, traces, and failure types |
| Model research: TRL, verl, Agent Lightning | SFT, preference or outcome feedback, multi-turn tool rollouts; GPUs, data, and verifiable environments are separate prerequisites | Demonstrate a training gain on a small simulated task against an untrained baseline |
The AutoGen repository now says it is in maintenance mode and directs new users to Microsoft Agent Framework. AutoGen tutorials and Magentic-One still matter for research, but version checks are essential. OpenHands and Magentic-One can execute code or change their environments: study them in disposable sandboxes with dummy data, not with production permissions copied from an example.
Six gates from a first loop to a measurable cluster
These are this article's proposed experiments, not official project curricula. Use the same set of public product updates, human-created answers, and acceptance rules. Change one architectural feature per gate. Make tasks resettable, keep a held-out subset untouched by prompt tuning, and record model, prompt, tool, data, and code versions.
- L0, no-agent baseline: A script or human records publication dates, original URLs, and changes. Measure minutes, citation errors, and omissions. If the fixed method is good enough, stop here.
- L1, one tool loop: Use smolagents or Agents SDK with only search and read tools, a step cap, timeout, and an explicit “cannot verify.” Measure valid calls, factual and citation accuracy, fabricated links, and repeated failures.
- L2, resumable state: Add a durable task ID, checkpoint, retries for read-only tools, and idempotent behavior. Interrupt a run. The resumed run must retain sources, avoid duplicate submission, and not mistake stale state for completion. LangGraph's checkpoint interfaces are one reference.
- L3, two agents: A researcher finds sources, a separate checker verifies dates and citations, and a manager accepts only supported items. Compare with L2 using the same cases, model, and budget. Revert if tokens rise without fewer errors.
- L4, parallel cluster: Divide independent sources among workers; each returns evidence, unknowns, and cost. A merger handles duplicates and contradictions rather than masking a shared error by majority vote. Measure end-to-end latency, cost per accepted result, conflicts, and human arbitration; bound concurrency with a queue.
- L5, permissioned service: Test writes, retry, rollback, audit, and human approval in dummy-data environments, then observe real tasks in shadow mode. Verify no unauthorized actions in the test set, recoverability, rollback, and a named owner before enabling external writes.
Anthropic's agent eval guide recommends inspecting the final output, tool trajectory, and environment state. BrowserGym, τ²-bench, and OSWorld offer useful testbeds, but their tasks and scoring differ. Public scores supplement a held-out set from your own workflow.
Optimize the workflow before adding agents or training a model
Optimize for accepted tasks / total cost and actual human minutes saved, not agent or token counts. Change one variable, rerun the same held-out cases, report results by scenario, and inspect failed traces.
| Order | Concrete change | Regression to watch |
|---|---|---|
| 1 Instrument | Record each input, tool arguments, result provenance, model, latency, tokens, error, and human intervention; redact sensitive data | Pass rate up but review time worse, or traces exposing private data |
| 2 Tool contracts | Fewer tools, typed arguments and errors, separate read and write rights, programmatic output validation | Wrong tool choice, invalid parameters, endless error retries |
| 3 Context | Remove duplicate retrieval, keep source URL and time, summarize old progress while retaining original evidence | Summaries losing constraints, sources, or cross-turn references |
| 4 Flow | Put predictable steps in code, parallelize only independent work, cap retries and spending | Shared-state conflicts, cost spikes, correlated errors |
| 5 Model selection | Compare small-model routing, stronger models for hard cases, and reasoning budgets on the same suite; tune only when needed | Lower price but more human repair or long-tail failures |
| 6 Verification | Prefer executable tests, state checks, and source matching; use another model plus human sampling for subjective judgments | Generator and judge sharing an error, or learning to game the judge |
Anthropic's architecture guide describes chaining, routing, parallelism, orchestrator-workers, and evaluator-optimizer patterns. They are alternatives selected by task, not mandatory maturity levels. MCP can standardize tool access and A2A can connect services; neither supplies acceptance criteria or authorization. More agents do not automatically provide more independent evidence.
The agent model and the external runtime have different jobs
The model interprets the goal, chooses actions under uncertainty, emits valid tool arguments, and updates its plan after results. The harness executes those actions, enforces permissions, preserves state, times out, retries, audits, and rolls back. A model saying “approved” does not grant permission; JSON output alone does not satisfy business rules. TRL's SFT documentation specifies tool-call messages, tool responses, and available tool schemas in training examples. Transformers' chat-template guide explains why inference formatting must match what a model learned.
| Capability | What the model should learn and be tested on | What the runtime must enforce |
|---|---|---|
| Goal interpretation | Constraints, stopping, when to ask a person | Task contract, timeout, step budget |
| Tool use | Correct choice and arguments, error handling, knowing when to stop | Schema validation, least privilege, sandbox, timeout |
| Multi-turn work | Revise after observations, never invent completion | Durable state, versions, provenance, tool logs |
| Evidence and verification | Trace claims to sources, express uncertainty, distinguish passed tests from belief | External tests, database-state checks, human acceptance |
| Domain work | Code and tests, GUI perception, or business policies depending on task | Actual tool and data rights, expert labels |
| Refusal and authorization | Ignore malicious commands in webpages or files, seek approval for risky actions | Identity, policy engine, runtime blocking, audit |
| Efficiency | Solve within token, time, and tool budgets | Routing, caching, concurrency, cost caps |
Start with an existing model that can already call the needed tools, then identify whether the gap lies in tool design, data, domain skill, or model weights. Qwen3's original project documents open-weight tool-use options; compare it and closed APIs under the same tools and cases, not by stitching vendor leaderboards together. Screen use additionally needs visual input, localization, and action alignment. Training on text-only tool calls cannot establish reliable GUI operation.
The model development and training chain, first fine-tune to frontier research
“Complete” means the whole decision and verification chain, not that one person must pretrain a foundation model. An individual can start with an existing model and a small tool environment, then improve it against a held-out suite. Training a base model from scratch is another scale of project: OLMo 3's released training records show the data, compute, checkpoints, and configurations involved. They are a research reference, not a solo-project budget.
The vocabulary matters: a tokenizer divides text into model input units; a chat template formats roles, tool calls, and replies; SFT learns from demonstrated correct behavior; DPO and related methods use paired preferences; RL adjusts behavior from rewards after repeated attempts in an environment; a rollout is one recorded sequence of actions, tool observations, and outcomes. More execution logs alone do not mean the model has learned.
- Define task and success: Pick one job—read-only verification, fixing a bug, or policy-bound support—and specify tools, permitted side effects, terminal state, and failure cost. Record a no-agent and off-the-shelf-model baseline. τ²-bench illustrates how policy, tools, tasks, and environment jointly define an eval.
- Build and govern data: Gather licensed human or real-task trajectories; deduplicate, anonymize, and verify usage rights. For each, store
goal → available tools → model action → tool observation → terminal environment state → human/programmatic grade. Include errors, refusals, timeouts, and cases where the agent should ask. Split train, development, and locked test sets by source or time; prevent near-duplicate and answer leakage. - Choose base model and format; pretrain from scratch only as a separate research program: An individual should compare valid tool calls, long-turn reliability, languages, latency, and cost per accepted result under one tool schema. Pin tokenizer, chat template, serialization, and server parser. If these disagree, the model may misread tool output or emit unparseable actions. Transformers' tool format guide provides concrete checks. Building a base model from scratch additionally requires licensed corpora, deduplication and filtering, tokenizer training, architecture and data-mixture choices, next-token pretraining, domain and long-context midtraining, loss and downstream monitoring, and recoverable checkpoints. OLMo 3's public configurations expose these research dependencies; earlier steps do not somehow yield a trained base model.
- SFT on demonstrations: Teach when to call, wait, recover from invalid arguments, stop, and say “unknown” with high-quality multi-turn demonstrations. TRL SFTTrainer supports
tool_calls,toolresponses, andtoolsschemas; PEFT/LoRA can reduce trainable parameters and memory, without guaranteeing full-fine-tuning quality. Compare on a small task before ingesting large, unchecked synthetic traces. - Evaluate trajectories and preferences: Check tool choice, arguments, citations, steps, stopping, and authorization as well as the final answer. Where two plausible outputs differ in quality, collect preference labels and investigate TRL's DPO and related trainers. Preference means an annotator chose an output, not that the environment goal was achieved.
- Multi-turn RL only with a trustworthy grader: Let the model act in resettable sandboxes, reward executable tests or terminal environment state, and cap steps, cost, and risky actions. Research entry points are TRL agent GRPO, verl agentic RL, and Agent Lightning. Watch for reward hacking, hidden-answer leakage, test edits, and safety sacrificed for score. Re-evaluate on a locked set after training. This stage needs compute, environment engineering, and research experience; it is not an entry requirement.
- Add advanced modalities only as required: Multi-file and long tasks need long-context and handoff data; browser/desktop work needs paired screens, actions, and outcomes; voice needs latency and interruption tests; multi-agent work needs clean input/output contracts. Evaluate each separately. A longer context does not itself teach memory or authorization. BrowserGym and OSWorld 2.1 are browser and computer-use research environments; pin versions and tasks.
- Serve, monitor, retrain: Begin with a permissioned shadow deployment. Track cost per accepted result, human takeover, latency, unauthorized actions, and privacy events, with version rollback. With permission and cleaning, turn real failures into the next training set; rerun blind and regression tests. Agent Lightning's implementation explores collecting rollouts from a real harness for training; the developer still owns data rights, reward quality, and release decisions.
Train when the diagnosed gap is in model behavior. Bad permissions, poor source data, missing verifiers, and tightly coupled workflows are system problems that weight updates rarely solve. A repeatable, labelable tool-choice or format deficit may justify SFT. Reliable, resettable outcome rewards and an SFT plateau may justify RL research. The frontier is a reproducible loop linking model, tools, environment, feedback, evaluation, and operating cost, not merely a larger collection of agents.
Two numbers that people often confuse
Capability evaluations measure a controlled task. METR's software-task time-horizon work uses estimated human time to finish a task as the difficulty scale and examines the span that an agent solves at 50% reliability in a specified environment. A “five-hour task” means an estimated five hours of human work, not five hours of agent runtime or a production success rate. Change the task mix, tools, budget, or pass threshold and the number may not transfer.
Product outcomes require a separate test. When a report says “80% of tasks solved,” ask where the tasks came from, whether training data leaked, how hidden tests work, whether humans helped, how many runs were tried, and whether mistakes were recoverable. OpenAI's 2026 review of SWE-bench Verified described contamination and flawed tests; its later SWE-bench Pro audit estimated roughly 30% problematic tasks. This does not make all agent benchmarks useless. It means a leaderboard cannot replace an acceptance set from your workflow.
Evidence can also point the other way: a METR randomized study of early-2025 tools assigned 246 tasks to 16 experienced developers in mature open-source projects and found 19% longer average completion time with the then-current AI tools. That result is bounded to its tools, people, and tasks; it cannot establish that all 2026 agents slow developers down. It does show why feeling faster, scoring higher, and saving human time are different outcomes.
| Public number | What it measures | What it cannot prove |
|---|---|---|
| Users or tool calls | Adoption and use | Artifact correctness or time saved |
| Benchmark pass rate | One dataset under one scoring rule | Success in every company workflow |
| 50% human-time task horizon | Capability on tasks of an estimated difficulty | Agent runtime or production availability |
| A demo or one success | Possibility under demonstration conditions | Repeatability, exception handling, safety |
| Accepted completion rate in a pilot | User acceptance on specified tasks | Unchanged performance on new data or sites |
The cost denominator that matters
The model bill is one line. For each deliverable accepted by the user, count model and tool calls, sandboxes and storage, human review, failure repair, permissions, and trace retention. If an agent attempts 100 tasks but only 60 results are accepted, 100 is the wrong denominator for value. A simple measure is:
cost per accepted output = (inference + tools/runtime + review + rework + operations) / accepted outputs
Suppose 100 narrow tasks a month each take 12 human minutes: 20 hours total. A pure illustration: 80 agent results each require two review minutes, 20 still take 12 human minutes each, and setup and maintenance take four hours. Human time is 10 hours 40 minutes, saving 9 hours 20 minutes before model and infrastructure costs. If errors are hard to spot or expensive, expected incident cost belongs in the numerator too. This example does not claim any product achieves an 80% acceptance rate.
Security boundaries belong in the system
Webpages, email, files, and tool results are data, and may contain instructions planted by someone else. Prompt injection tries to make the agent treat such third-party material as the user's order. Anthropic's browser injection research says browser agents are not immune. OpenAI's account of running Codex safely describes access boundaries, approvals, and external telemetry. Writing “do not leak secrets” in a prompt is not sufficient protection.
A workable progression is read only → writes in a test environment → narrow data and tool scopes → human authorization for irreversible actions → records the agent cannot rewrite, with recovery plans. Tools need explicit input schemas, timeouts, provenance, and error states. Separate “may view,” “may edit,” and “may send” permissions rather than handing over an entire account. The MCP maintainers' tool annotations note says labels such as read-only or destructive are risk hints, not replacements for runtime access control.
How to evaluate a first agent in one team
The first pilot need not send mail or modify a database. Choose summarizing public product updates with citations, a task a person can check. Weekly output might include the change, publication date, source link, scope, and an explicit “cannot verify” marker. If a fixed script plus human check already works, an agent may add no value.
- Collect 30–50 real past cases, including stale announcements, conflicting pages, missing links, duplicates, and pages containing malicious instructions. Label acceptable answers or scoring rules. This is a proposed pilot size, not a statistical guarantee.
- Measure a human or fixed-workflow baseline: minutes per case, omissions, bad citations, and rework. Hold some cases back for acceptance testing; do not tune prompts on them.
- Give the agent a minimum set of read tools and require provenance and uncertainty in each output. Begin in shadow mode: it drafts a result but publishes nothing.
- Record accepted completion, citation accuracy, human review minutes, failure types, latency, and total cost. Test whether malicious pages cause unauthorized actions. A final-answer-only score misses hazardous intermediate tool calls.
- Expand only if held-out cases and fresh weekly data repeatedly beat the baseline. Sending, publishing, or editing later requires separate privileges, approval, and recovery tests.
These steps align with Anthropic's agent evaluation guide and Google Cloud's production guide. They are engineering methods, not a requirement to buy either platform.
Signals to watch over the next twelve months
The following is inference, not an event that has already happened. Longer-running services, tool standards, and identity governance could move more agents from chat windows into background work. But a longer runtime without higher accepted-completion rates or lower human rework does not establish value. Useful public signals will be cross-provider, reproducible long-task results; costs reported alongside errors; incidents and recovery for external actions; and comparisons against human or fixed automation in the same workflow.
The current assessment is therefore: narrow, reversible, measurable agents already have practical uses; multi-agent clusters must show a gain over one agent; high-stakes, cross-system, long-term autonomy still needs case-by-case evidence. Ask “what output can we accept and verify?” before “how many agents and do we need new model weights?” This survey has no same-condition independent test across products, and it did not run the listed projects or train a model. Project status, features, protocols, and figures were checked October 3, 2026.