In one line: AI can already draft a fix for most vulnerabilities (Google says that is already the case in Chrome). The hard part is proving the fix is actually right. So build a pipeline where every step is checked by actually running programs, and at the end a human reads it, signs off and submits.
Yesterday I kept turning one question over: if I wanted AI agents to fix vulnerabilities in code the size of Chromium, how would I do it? Chromium is the open-source project behind the Chrome browser, and the git history of its GitHub mirror alone is about 64 GB. Can AI handle a project that big? And once it has made a fix, how do you know the fix is right?
Today I went through a round of research, and the conclusion is clear: writing the fix is no longer the hardest part. The Chrome Security Team says large language models now produce candidate fixes for most vulnerabilities, though it does not say how many were merged as the AI wrote them. The hard part is verification. I learned this the hard way in Day 3: the AI said "done", and an independent review found 37 problems. Fixing vulnerabilities is the same, except a wrong fix costs more.
This post is for people who use AI to write code but have never done security work. By the end you will know:
- Where the industry is: what Google, a DARPA competition and a few AI companies have each built
- Why "the crash is gone" does not mean "fixed", and how big the gap is in the studies
- Which practices have numbers showing they actually help
- What real gates appear once the target is Chromium
- How one person can start, step by step
- The realities and red lines to know before you begin
This is research as of 2026-10-09, and every number links to its source. Numbers that a vendor published about itself and that nobody has independently verified are marked as such. The whole post is about defence only: reproduce, patch, verify, disclose responsibly. It does not cover how to exploit a bug.
Eight Terms First
These eight terms come up again and again, so here they are in plain words:
- Memory-safety vulnerability: the program reads or writes memory it should not touch, for example writing past the end of a block of space, or using memory it has already given back. Like stuffing a letter into the neighbour's mailbox.
- Fuzzing: hitting a program with huge numbers of random or slightly mangled inputs to see whether it crashes. Like pressing every button on a vending machine tens of thousands of times.
- Sanitizer / ASan: a checker compiled into the program that stops the moment memory goes wrong and reports which line it was and how the program got there. Like a sensor in a warehouse that beeps as soon as something lands on the wrong shelf.
- Crash: the program dies. With ASan, it comes with the call path at the moment of failure (the stack). Like a car stalling mid-trip with a fault code left on the dashboard.
- PoC: the input file that reliably makes the problem happen again, for example a specially crafted image. Like the stretch of road where the mechanic hears the strange noise every single time.
- Root cause: where the problem really starts, which is often not the line that crashed. Like a dripping ceiling: where the water falls is not necessarily where the roof is broken.
- Patch / CL: the fix, meaning the diff of which lines changed; Chromium calls one change sent for review a CL. Like a red-pen edit handed to a reviewer.
- Regression test: a set of tests that reruns automatically every time the code changes, to make sure things that used to work still work; after fixing a vulnerability you add its PoC to the set, so the problem cannot quietly come back. Like checking every corner of the house each time it rains after fixing a leak.
One Pipeline That Explains It
Each station in the picture does one job and hands the result to the next. For AI fixing vulnerabilities, it comes down to roughly these six steps:
Use the PoC to make the crash happen again in a clean environment
Trace back from the crash to the assumption that was wrong
Have different models each write a fix
Compile, rerun the PoC, fuzz, run the tests
Another agent hunts for problems with a checklist
Understand every line before sending it for review
Why split it this finely? Because every handoff is a chance to let evidence talk.
Reproduce first, then find the root cause. Without a PoC that reproduces the problem, there is no way to tell whether a fix worked. The root cause matters because where a program crashes is often not where the mistake starts: block it right at the crashing line and the original PoC passes, but a slightly different input breaks it again. The later section "A Gone Crash Is Not a Fix" walks through a small example.
Write several, then filter by running them. Write only one and you cannot tell when it is wrong. Chrome already has several agents propose candidates and pick each other's work apart; the next section has the details. The filtering has to come from actually running things: it compiles, the PoC no longer triggers, the tests pass. An agent saying "verified" counts for nothing until a program has run.
Review and the human gate come last. Independent review is a second pair of eyes, but it cannot replace running the code. The last step is always a human, who has to understand every line before pressing submit. Chrome's process also ends with a developer reviewing.
I did not invent this shape. Google DeepMind's CodeMender also checks with tools first and only passes patches that fix the root cause and break nothing else on to human review. DARPA's AIxCC is a competition where AI systems find and fix vulnerabilities on their own, and inaccurate patches cost points; according to a paper that reviews the whole competition, the cautious top teams did not submit a patch unless they had an input that proved the vulnerability exists. The three started from different places and ended up with the same line: AI writes, programs verify, a human signs.
Where the Industry Is
First, a look at how the big players actually do it. Many of the numbers below are published by the companies themselves, and I will say so; where I could not find independent verification, I will say that too.
The Chrome In-House Fix Pipeline
In a post on 2026-07-30, the Chrome Security Team described how they fix vulnerabilities today:
- A "fixing agent" first proposes several candidate patches.
- A "critic agent" picks each one apart, and the two go back and forth, like a human code review.
- "Test-writing agents" write and check tests on every platform Chrome supports.
- At the end, a developer reviews the result.
Google says LLMs now produce candidate fixes for most vulnerabilities. The post does not say how many fixes were finally merged exactly as the AI wrote them.
The finding side is wired in too. Every 24 hours, Big Sleep and CodeMender (introduced below) scan all CLs in Chrome's CI (the system that runs checks automatically whenever someone submits code). In May, that setup blocked more than 20 vulnerabilities, one of them rated critical.
The volume is large as well: Chrome 149 and 150 together fixed 1,072 security bugs, and by March 2026 Chrome had already received more bug reports than in all of 2025.
The safety measures are concrete: scans run on locked-down machines with restricted network access, and each component's SECURITY.md tells the critic agent where the trust boundaries are, that is, which data comes from outside and cannot be taken at face value.
Google CodeMender and OSS-Fuzz
CodeMender is a fixing agent Google DeepMind announced on 2025-10-06. It uses Gemini Deep Think models together with a full set of tools:
- static analysis (reading the code without running it to find problems) and dynamic analysis (running it and watching)
- differential testing (feeding the same input to the program before and after the change and comparing the results)
- fuzzing
- SMT solvers (tools that work out by logic whether a condition can ever hold) and a debugger
On top of that, an LLM critique tool compares the original and the modified code. Only patches that fix the root cause, are functionally correct and cause no regressions (things that used to work breaking) go to a human. DeepMind says it upstreamed 72 security fixes to open-source projects in about six months, and a human reviewed every one before it was submitted.
Then it moved beyond Google's walls:
- On 2026-07-22, the Google Cloud version opened in preview, for a limited set of customers only, with human approval of every patch.
- On 2026-07-29, it was connected to OSS-Fuzz. OSS-Fuzz is Google's service that keeps running fuzzing for open-source projects.
Eligible reports, that is, memory-safety bugs in C/C++ projects, get a CodeMender patch attached. Each patch is first tested on its own: it compiles, the original crash is gone, and the functionality tests pass. During the beta, Google engineers review every patch, and projects can opt out; no acceptance rate by the projects has been published.
Earlier experience is worth a look too. In a 2024 study, Google had Gemini fix sanitizer bugs caught by internal unit tests. Gemini fixed about 15%; of the patches sent to code owners after human filtering, about 95% were accepted. The study also recorded bad fixes: simply deleting the failing test, or making the code run one step at a time to dodge a data race (two threads changing the same piece of data at once).
Big Sleep, the Finder
Big Sleep is an agent built by Google Project Zero and DeepMind. It only finds bugs; it does not fix them. It starts from recent code commits and looks for vulnerabilities that look similar. Chrome's release notes (2025-08-26) credit it with CVE-2025-9478 (a CVE is a public vulnerability ID): a critical use-after-free (memory that has already been given back but the program keeps using) in the ANGLE component, fixed in Chrome 139.
The DARPA AIxCC Competition
AIxCC is an AI security competition run by DARPA. The final was held at DEF CON 33 on 2025-08-08, and the competing systems had to find and fix vulnerabilities on their own. First place went to Team Atlanta's Atlantis ($4M), second to Trail of Bits' Buttercup ($3M), and third to Theori's RoboDuck ($1.5M).
The competition planted 63 synthetic vulnerabilities in the projects. The finalist systems together found 54 of them (86%) and patched 43 (68%); they also found 18 real vulnerabilities and supplied 11 patches. A patch took about 45 minutes on average, and a task cost about $152.
Two things to keep in mind. First, the targets were 24 OSS-Fuzz-style projects, such as curl, libxml2, dav1d, OpenSSL and Wireshark; Chromium, V8 and the Linux kernel were not among them. Second, the scoring punished inaccurate patches, so the cautious top teams did not submit a patch without a PoV (proof of vulnerability, the competition's name for a PoC). A paper reviewing the competition also notes that for a Claude Code baseline used for comparison, 37.7% of the patches that passed the automatic checks were rejected in human review.
The competition also left open-source tools you can use. Buttercup has a standalone version that needs at least 8 cores, 16 GB of RAM and 100 GB of disk. OSS-CRS is a project of OpenSSF (the Open Source Security Foundation) that runs AIxCC systems on your own machine and uses a LiteLLM proxy (a relay that sits between the systems and the model APIs) to set a dollar budget per system; according to the OSS-CRS paper, its port of Atlantis found 10 new bugs across 8 OSS-Fuzz projects, 3 of them high severity. The original competition versions are tied to the competition cloud, which has been shut down, so they are hard to reuse directly.
Other Companies
These companies have published less, and their numbers are mostly self-reported, so read them as direction only.
- OpenAI: Aardvark (announced 2025-10-30, renamed Codex Security on 2026-03-06) first builds a threat model (a list of who might attack and where they could get in), then scans code commits, tries to trigger the problem in a sandbox (an isolated environment where breaking things does no harm), has Codex draft a patch, and ends with human review. OpenAI says it found 92% of the known and synthetic bugs in its test repos; I could not find independent verification.
- Anthropic: Claude Code Security opened as a limited research preview on 2026-02-20. It scans code, runs multi-stage verification and suggests patches, and every patch needs human approval; open-source maintainers can apply for free access.
- Mozilla: a post on 2026-05-07 says they built their AI bug-finding process on top of their existing fuzzing infrastructure, require the agent to produce a test file that reproduces the problem, and run every job in throwaway virtual machines. Firefox 150 fixed 271 bugs found with Claude Mythos Preview. That figure comes from Mozilla and its partner, and some people have questioned how those bugs map to CVEs.
| System | Find bugs | Write patch | Verify by running | Human gate |
|---|---|---|---|---|
| Chrome pipeline | Yes | Yes | Yes | Yes |
| CodeMender | Yes | Yes | Yes | Yes |
| Big Sleep | Yes | No | Undisclosed | Undisclosed |
| AIxCC teams | Yes | Yes | Yes | Undisclosed |
| Codex Security | Yes | Yes | Partial | Yes |
| Claude Code Security | Yes | Yes | Undisclosed | Yes |
One thing they share: every product and company pipeline that writes patches keeps a human gate at the end (AIxCC was a competition; the material does not say whether it had one, only that its scoring punished inaccurate patches). The difference is how strict the verification is: some require the patch to compile, the crash to disappear and the functionality tests to pass, all actually run before it counts; others only say "multi-stage verification" without saying how.
A Gone Crash Is Not a Fix
The previous section showed that the products and pipelines that write patches keep a human gate; the difference is how strict the verification is. This section explains why it has to be that strict. The most common problem with an AI-written patch is not "it didn't fix anything" but "it looks fixed." The original crash is gone, yet the problem is still there.
A Small Example
Imagine a library that reads images. It first reads the header (the short block at the start of the file that records width and height), allocates a chunk of memory of "width × height × 4" (4 bytes per pixel), then copies the pixels in, one row at a time. It trusts the numbers in the header completely.
Fuzzing turns up a file whose width and height are both very large, so large that their product does not fit in a 32-bit integer and wraps around to a tiny number. This is called integer overflow, like an old car's odometer rolling back to zero.
So the program allocates only a small chunk of memory but copies using the real width, and after a few rows it writes past the end of that chunk. ASan halts on the spot and reports a heap-buffer-overflow (a write outside the memory you were given); the top line of its crash stack points at the pixel-copying line. The file that triggered it is this bug's PoC.
The easiest fix is to add an if right before the line that crashed:
// Fix A: block it right before the crashing line
uint32_t row = w * 4; // bytes per row
for (uint32_t y = 0; y < h; y++) {
if ((y + 1) * row > buf_size) // stop if this row won't fit
return 0; // 0 means "success"
memcpy(buf + y * row, src + y * row, row);
}
The original PoC really does stop crashing: the if ends the function before the row that does not fit. But the real problem, "trusting the header too much," has not changed at all:
- It only checks the side being written to. Take a different file whose width and height do not overflow but which is shorter than its header claims, and the program still reads past the end of the input data. ASan reports again, just from a different spot.
- The
(y + 1) * rowinside theif, and the multiplication that computesrowabove it, can overflow themselves when the numbers are large. Pick a larger set of numbers and the check stops working. - Even when nothing crashes, it returns 0, meaning "success." The calling code gets an image that was only half copied and thinks it is fine. This kind of silent wrong output is harder to notice than a crash.
The real fix goes back to the root cause: the mistake is not in the pixel-copying line but in trusting the header the moment it was read.
// Fix B: check the header when it is read
int parse_header(const uint8_t *p, size_t len, struct hdr *hd) {
if (len < HDR_SIZE) return ERR_TRUNCATED;
hd->w = read_u32(p); hd->h = read_u32(p + 4);
if (hd->w == 0 || hd->h == 0 || hd->w > MAX_DIM || hd->h > MAX_DIM)
return ERR_BAD_SIZE; // bad width/height
if (ckd_mul(&hd->size, (size_t)hd->w * 4, hd->h)) return ERR_BAD_SIZE;
if (len - HDR_SIZE < hd->size) return ERR_TRUNCATED; // file too short
return OK;
}
It does four things: it rejects an incomplete header; it rejects a width or height of 0, or one above the limit MAX_DIM that the project sets for itself; it computes the memory size with ckd_mul, the overflow-checking multiplication built into the C standard, and reports when the result does not fit; and finally it confirms the file really contains that much pixel data. If any check fails, it returns a clear error code, so the layer above can show "corrupted file" instead of handing back half an image. Every later piece of code that uses the width and height gets numbers that have been checked, so the protection covers more than the pixel-copying line.
For an AI, Fix A is tempting: the ASan report points at that line, adding an if there is the least work, and rerunning the original PoC passes. But the place that really needs changing is not on the crash stack.
What the Research Says
This is not an imagined worry. The University of Maryland's PatchBench (2026-09-03) deliberately picked 213 C/C++ tasks where the correct fix is not on the crash stack. Across 11 agents, on average 83.1% of patches made the original PoC stop crashing, but only 45.3% actually solved the problem, an overestimate of about 1.8×; the best combination reached only about 59%. Remove the "compare outputs and state with the original program" check, and the solved rate inflates by about 8 more percentage points. My reading is that this check is what catches Fix A-style patches that "don't crash, but give the wrong result."
Meta's AutoPatchBench (2025-04-29) ran a case study on a 113-bug subset with 2025-era models: about 60% of patches compiled and stopped the original PoC from crashing; after adding 10 minutes of fuzzing and white-box differential testing (feeding the same input and comparing the internal state of the patched version with the correct fix), only about 5–11% were judged correct. One model fell from 61.1% to 5.3%. The authors themselves say the case study is not statistically rigorous.
Two more studies say the same. CodeRover-S had people check 88 PoC-passing patches one by one, and only 40 were correct (45.5%). Team Atlanta used 10 coding-agent configurations to fix 63 crashes from the AIxCC competition and reviewed all 630 patches by hand: even patches that passed both the PoV and the project tests were semantically wrong (the code runs, but does the wrong thing) about 20–40% of the time, and about 20% even for the best setups.
Common Wrong Fixes
Several papers list much the same wrong fixes. In plain words:
- Add an
ifwhere it crashed: blocks the original PoC but not files that look almost the same. Fix A is this one. - Patch only the symptom: make the crash or the ASan report go away without dealing with what causes it. Fix A counts here too.
- Delete the feature or the operation: that code no longer runs, so of course it no longer crashes.
- Check too strictly: reject normal files along with the bad ones.
- Hard-code a limit when it runs too long: when the program runs so long it is flagged as a timeout, simply cap how many times the loop may run.
- Fix a different bug nearby: the neighbouring problem gets fixed, the original one is still there.
- Get the root cause wrong: the real fix needs changes in three files, and it changes one.
- Introduce new bugs: for example a memory leak (memory borrowed and never returned, so the program takes up more and more space as it runs).
- Claim success without verifying: the agent reports "fixed" without having run any check.
Team Atlanta sorted the 145 semantically wrong patches by type. The biggest group is not "symptom only" but patches that broke the original functionality:
Why Tests and Agent Claims Fall Short
Can a project's existing tests act as the gate? A study that came out this October (2026-10-07) looked at 112 historical bugs: one frontier model's patches passed the PoC 90.2% of the time, but passed the developers' own tests only 60.7% of the time. Meanwhile the project's existing regression tests passed about 95–97% of the time whether the patch was good or bad, so they barely tell the two apart. The paper is a preprint posted to arXiv this October, so treat it as a pointer. My guess at why: those tests were written before anyone knew about this bug, so naturally they do not exercise that path.
The same goes for an agent saying "verified." Wrong fix number 9 above is exactly this: it says the bug is fixed, but nothing was actually run. When fixing vulnerabilities, "I verified it" is just a sentence. To count, a program has to actually run it: compile, replay the PoC, fuzz the patched build, run the project tests, compare outputs on normal inputs, with every step's result kept in a log that people can see. The next section covers which practices really make patches more reliable.
What Actually Helps
The previous section showed that a gone crash is not a fix. The good news is that researchers have measured a few things that really do make AI-written patches right more often. Each of the seven practices below has three parts: what to do, why (with a number), and how to do it in practice.
One thing first: these numbers come from different studies with different bugs, models and scoring. They tell you whether something helps. They cannot be compared with each other.
Give It Enough Material
What: Before the agent starts, hand it three things: the PoC, a cleaned-up sanitizer report, and build and test commands that run as written.
Why: A study published this October found that with build and test instructions added, the share of C/C++ patches that passed the original PoC rose from 72% to 92%. It is a fresh arXiv preprint, and it counts "passes the PoC", which is not the same as "really fixed", so treat it as a pointer. A more direct example: PatchAgent (USENIX Security 2025) removed its "clean up the report" step, and its fix rate dropped from 77.3% to 64.0%.
In practice: Run the commands yourself first and make sure they work when copied and pasted, then put them in the task description. Keep the parts of the report that relate to this crash; do not dump the whole log in.
Find the Root Cause Before Touching Code
What: Ask the agent to find the root cause first, and only let it change code once it has explained that.
Why: VulDebugger has the agent trace the bug step by step with the LLDB debugger (a tool that pauses a program at a line so you can look at the values of its variables). On 50 real C bugs it fixed 60%; AutoCodeRover, another bug-fixing system used for comparison, fixed 28%. When both the root cause and the fix location were right, 75.8% were fixed.
In practice: Split the task in two. In the first half the agent may only read code and run the debugger, and it hands back one paragraph: which value starts going wrong, where, and why. Only when you understand it and it makes sense does it move on to writing the patch.
Generate Several Candidates at Once
What: For the same bug, have several different models each write a patch, then choose among them.
Why: In the PatchAgent experiments the best single model fixed 84.8%; five models pooled (counted as fixed if any one of them fixed it) reached 92.1%. Team Atlanta won AIxCC, and their system, Atlantis, ran 8 patch agents in parallel; their later study also found that the choice of model mattered more than the choice of agent framework, with three used together being the best value. The Chrome Security Team pipeline works the same way: a fixing agent proposes several candidates and a critic agent evaluates them.
In practice: Start with three different models. Every candidate goes through all the verification gates below, and only the ones that pass go on to review and to a human.
The numbers from the first three practices side by side:
Retry in a Fresh Conversation
What: When a patch fails, do not tell the agent to "try again" in the same conversation. Open a new conversation with the original material and a short summary of the last attempt.
Why: CyberGym-E2E (ICML 2026) measured that retrying in a fresh context (everything the agent remembers in this conversation) improved results by 4.8 to 7.1 percentage points. The other way round, Meta's AutoPatchBench observed that models were more likely to "cheat" when retrying in the same conversation: they made the crash disappear without fixing its cause.
In practice: Keep the summary to a few lines: what was tried, which gate it failed, and what the error message said. Do not paste the whole previous conversation.
Verify by Running Things
What: Whether a patch is right is decided by what programs report when they run, not by what the agent says.
Why: PatchBench, mentioned in the previous section, found that dropping the check that compares outputs and state with the original inflates the solved rate by about 8 percentage points. Before Google attaches a CodeMender patch to an OSS-Fuzz report, it also tests the fix on its own: it compiles, the crash is gone, and the functionality tests pass.
In practice: Turn the five gates below into one script and keep it where the agent cannot change it; the agent only sees "passed" or "failed" and the error messages. The second gate checks UBSan as well as ASan (UBSan is another checker that catches C/C++ code whose result cannot be predicted); the third gate runs fuzzing starting from the PoC.
Independent Review, but Not Instead of Execution
What: Set up a separate reviewer agent that looks at the patch with a checklist. Its only job is to find problems; whether the patch passes is still decided by running things.
Why: Automated checks miss things: as mentioned earlier, in the AIxCC review paper 37.7% of a Claude Code baseline's auto-check-passing patches were still rejected by human review.
But an AI is not a reliable judge. CodeRover-S used an LLM to decide whether patches were correct, and its precision (of the patches it called correct, the share that really were) was only 0.47 to 0.57, about half. Another 2026 study (secondary source; I did not check the original) reports that several AI judges agreed closely with each other (κ=0.75; κ measures how much two sides agree, and the closer to 1, the more they agree) but poorly with the results of running the code (κ≤0.26).
In practice: Build the checklist straight from the common wrong fixes in the previous section: Does it just add an if where the crash happened? Does it delete a feature or a test? Is the check so broad that it rejects normal input? The reviewer agent hands back a list of doubts, not the word "passed", and a human judges the doubts.
Set Budgets and Hard Rules
What: Give every bug a money cap and a time cap, and use tool rules to lock away what the agent must not touch.
Why: PatchBench raised the budget per bug from 5 to 25 US dollars and it helped only a little; taken together, PatchBench and CyberGym-E2E show the gains flattening at around 10 to 25 dollars and about 90 minutes per bug. Rules matter too: as mentioned earlier, Google's 2024 study saw the AI simply delete a failing test to get through. Atlantis had a rule: a patch may only touch source code and must leave the fuzz harnesses (the glue code that feeds random input into the program) alone.
In practice: Make the tests, the fuzz harnesses, and the sanitizer and CI configuration read-only. If a patch touches any of those files, or deletes a large chunk of code at once, flag it for a human automatically. Stop when the cap runs out, and let the agent answer "cannot fix". In my opinion, only when "cannot fix" is an acceptable answer will the agent stop forcing out a patch just to have something to hand in.
Of these seven, I think "verify by running things" is the one you cannot skip: the other six make the candidates better, and this one decides whether you can trust them.
The Extra Gates in Chromium
The pipeline above runs fine on a small library. Move to Chromium and the steps stay the same, but every step gets an extra gate: the build is heavy, not everyone can see the bugs, the rules are detailed, and Google's own AI is already running down the same road.
Building It Is a Gate by Itself
For an agent to reproduce a crash, the first step is building Chromium. The official Linux build instructions set the bar at: an x86-64 machine, at least 8 GB of RAM (16 GB or more strongly recommended), and at least 100 GB of free disk.
First you install depot_tools (Chromium's official toolkit for downloading and building), then fetch the source with fetch --nohooks chromium; adding --no-history skips the full history and saves time. Then you build with GN (the tool that generates build settings) and Siso, using autoninja -C out/Default chrome.
Fixing vulnerabilities needs the ASan build. You set is_asan=true is_debug=false in the GN args; details are in the ASan docs. My own estimate is that building the ASan version comfortably takes about 32–64 GB of RAM and an SSD with 200 GB or more. That is an estimate, not an official number.
Remote building does not help much. Google's remote build service, RBE, is invitation-only for external contributors and offered on a best-effort basis; the Mac docs say outright that remote execution is not supported for external contributors. So most of the time you rely on your own machine.
Two things save effort. First, prebuilt ASan builds of Chrome can be downloaded (tools/get_asan_chrome), though not for every revision. Second, tools/bisect-builds.py can bisect with ready-made builds (cut the range in half each time to find the version where things started breaking); the set the public can use is Chrome for Testing. When a ready-made build exists, you do not have to compile every step yourself.
The Bugs You Can Get Are Limited
Chromium security bugs are restricted by default; only a small group of relevant people can see them. The reason is simple: publishing a problem before it is fixed is a heads-up to anyone who wants to abuse it.
ClusterFuzz is Google's platform for running fuzzing automatically. For a restricted bug it found, your account has to be CC'd on that bug before you can download the testcase, which is the PoC. The full local reproduction script is available only to Google employees. Please play by these rules and do not try to get around them.
Disclosure timing has rules too. Security bugs usually become public 30 days after the fix, and AI-found bugs are no different; bugs marked WontFix (will not be handled) or Invalid become public after 14 weeks; if disclosure needs to wait, a SecurityEmbargo hotlist can hold it.
For outsiders, this means what you can practise on is mostly old bugs that are already fixed and already public. That is also why the "How One Person Can Start" section later suggests replaying old bugs first.
Chromium Already Has Tools for Agents
The good news: the Chromium repo has an //agents/ folder dedicated to things for AI agents: shared prompts, approved MCP servers (the interface that lets an agent call outside tools), and 56 skills (written-down procedures an agent loads when it needs them). These are the ones related to fixing vulnerabilities:
- fuzzing: find problems with FuzzTest (a fuzz-testing framework) plus ASan.
- bisect: find the commit where things started breaking.
- apply-fix-from-crbug-with-diff: apply a fix from the diff attached to a bug page.
- multi-agent-code-review: several agents review code together.
- multi-agent-engineering-workflow: a full development process split among several agents.
That last workflow has two design choices worth learning from. First, security-related work goes down a stricter "rigor path" with extra reviewers dedicated to security and auditing. Second, it does not let the agent retry forever: the number of rounds is capped (fewer than 3 in total), and when the cap is hit or it notices the agent flip-flopping between fixes, it hands the decision back to a human.
One caution: the workflow's release manager step can upload a CL by itself. I would leave that step for a human to press. The skills README says agents like Claude can use them, but I have not tried that myself yet.
There is also a document called security-for-agents.md, published around the first quarter of 2026, written for agents reviewing Chromium code. It is not the reward rules; it spells out which cases are not security bugs and what a report must include.
A built-in check fails and the program halts on purpose
Two kinds the document names as not counting
Only stalls or slows the program
Without these it is hard to reproduce
A few terms in plain words: a CHECK is a check written into the code that stops the program when its condition fails; a DCHECK is a check that only runs in debug builds. A null pointer points at "nothing". A UAF (use-after-free) is memory used again after it was freed, and MiraclePtr is a protection Chromium adds to its pointers. DoS means making a service slow down or stop. A symbolized ASan stack is an error report that shows function names and line numbers.
The Rules Are Spelled Out
Chromium has a formal AI coding policy. The key points:
- Authors must review and understand every change themselves and be able to answer reviewers' questions.
- Parts written with AI help that you are unsure about must be flagged.
- Every change needs two human committers (people allowed to merge code into the main branch); an author who is a committer counts as one of them.
- When people comment on a bug or CL an agent filed, the human operating the agent replies personally.
- Submitting code you do not understand can cost you committer status; repeating it after a warning can get you blocked.
For handling AI-found security bugs there is a separate FAQ (April 2026). S0 and S1 are severity levels: S0 is to be handled within a week, S1 within four weeks. You may land a mitigation first (a temporary measure that blocks the harm) and track the root cause in a follow-up. This is the same point the earlier section "A Gone Crash Is Not a Fix" made: blocking first is allowed, but someone has to keep track of the root cause.
Another document, written for triagers, Shepherding AI Reports, says shepherds, the people who triage security reports, may close AI reports as WontFix when there is no working PoC and no plausible ASan stack. Red flags include citing CVEs, APIs or class names that do not exist.
How an outsider submits a fix is laid out concretely in the contributing docs, and the chart below lists the steps in order. A few terms first: the CLA is the contributor license agreement; Gerrit is Chromium's code review site; OWNERS files list who is responsible for each folder; authors who are not committers need a Code-Review +1 vote from two committers. The guide for security fixes also asks you to state precisely which object lifetime or assumption was wrong, and to check for the same kind of problem elsewhere.
Sign the contributor license agreement with a Google account; add yourself to AUTHORS in your first patch
Reproduce with the ASan build and fix the root cause, not just the line that crashed
Add a test so the problem cannot come back; look for the same kind of problem elsewhere
Send it to Gerrit; say which assumption was wrong and flag AI parts you are unsure about
Find reviewers with git cl owners and answer comments yourself
No details go public before then; disclosure can be delayed if needed
Colliding With Google Internal AI
As the "Where the Industry Is" section showed, Chrome already runs a large internal AI pipeline every day: models draft candidate fixes for most vulnerabilities, and Big Sleep and CodeMender scan every CL every 24 hours. An outsider working on a known Chrome bug can easily end up racing it.
Outside rewards are changing too. According to secondary sources, Chrome's Vulnerability Reward Program (VRP) updated its reward structure in April 2026 to account for what internal AI tools already find, and reports may be judged duplicates of internal-tool findings. Separately, BleepingComputer reports that from 2026-10-01 Google paused new product-vulnerability reports to its Open Source VRP (OSS VRP), after a surge of automated and AI-generated reports that were mostly invalid. Patch Rewards, which rewards submitted fixes, stays open, and a new plan is expected in the first quarter of 2027.
This is my view: a known Chrome bug is quite likely already being handled inside Google. An outsider spending time on it is mostly duplicate work. The real value is in what the internal tools miss. Before getting there, practise the whole pipeline on smaller projects first, which is where the next section starts.
How One Person Can Start
My plan is five steps. The first four steps practise on old, risk-free bugs, and I only move on once the numbers are measured.
Step 0 Build Your Own Exam
Leave Chromium alone at first. ARVO is a public dataset of more than 5,000 reproducible C/C++ memory bugs, all from OSS-Fuzz; the newer version has 6,138 bugs in 311 projects. Every bug comes with two Docker images (packaged, ready-to-run environments), one vulnerable and one fixed, plus the input that makes the program crash and the developer's original fix.
That makes a ready-made exam. Pick 20–50 bugs, starting with dav1d, libxml2, freetype, libwebp and sqlite: they all sit in Chromium's third_party (the folder for outside libraries) and are all OSS-Fuzz projects; dav1d and libxml2 were also targets in AIxCC.
Record four numbers per bug:
- Does it compile after the patch?
- Does the original PoC still trigger ASan?
- After fuzzing for a while with the PoC as the starting point, does a new, mutated crash turn up?
- Do the developers' tests pass, and does normal input produce the same output as the original?
The most important rule: the developer's real fix, the fixed image and the git history after the fix all go somewhere the agent cannot reach. Otherwise it may simply copy the answer, and the score means nothing.
Step 1 Build the Pipeline
No need to start from scratch. Buttercup and OSS-CRS, both mentioned earlier, can run on your own machine. Buttercup is AGPL-3.0 licensed, needs at least 8 cores, 16 GB RAM and 100 GB of disk, needs your own model API keys, and lets you set a cost limit. OSS-CRS is MIT licensed, works with OSS-Fuzz-format projects and ships a patcher based on Claude Code.
With the skeleton in place, add the five verification gates from earlier and three hard rules: the agent may not edit tests, fuzz harnesses, sanitizer or CI settings; large deletions get flagged; and "I can't fix this" is an allowed answer. Each bug also gets a hard cap: PatchBench and other studies saw gains level off at roughly $10–25 and about 90 minutes per bug.
Then run the step 0 exam once. Those four numbers become the baseline for every later change.
Step 2 Build PDFium, ANGLE and V8 on Their Own
Once the exam scores are stable, switch to parts of Chromium. PDFium (the component that displays PDFs), ANGLE (the graphics layer under 3D drawing on web pages) and V8 (the engine that runs JavaScript) can each be built on their own. Their git histories are about 0.15 GB, 0.2 GB and 1.2 GB, far smaller than all of Chromium.
The approach is the same as in step 1: same pipeline, same four numbers, with only the build and test commands swapped for each project's own.
Step 3 Full Chromium, Replaying Public Old Bugs First
Only at this step does hardware become the gate. The section on Chromium's extra gates listed it earlier: officially at least 8 GB RAM and 100 GB of disk; my estimate for an ASan build is 32–64 GB RAM and 200 GB+ of SSD. Where an official prebuilt ASan build exists (tools/get_asan_chrome), download it instead of compiling every time.
Replay only old bugs that are already public, and leave restricted bugs alone as the rules require. I plan to try plugging the ready-made fuzzing and bisect skills in //agents/skills/ into the step 1 pipeline.
Step 4 Real Contributions
Only after all of that is stable do I send anything to Chromium. I would start with low-risk hardening (changes that fix no specific bug but make the code harder to break), such as adding a FuzzTest (the framework Chromium uses to write fuzz tests) to a component, rather than fixing security bugs straight away.
Follow the steps in the earlier chart, "An outsider submitting a security fix"; the AI coding policy points are in the same section: understand every line yourself, be able to answer reviewers, and flag the AI-assisted parts you are unsure about.
So my rule is: the agent never uploads on its own. The //agents/ workflow has a step that can upload a CL directly, and that step stays with a human. I answer reviewers' comments myself, too.
Wiring It Into My AI Cluster
This pipeline plugs straight into the AI cluster this series is building. Split the roles, from reproducing and finding the root cause to patching and validation, so that each agent does exactly one job:
Re-runs the crash in a disposable container; if it cannot, it stops there
Traces back with a debugger to the first thing that went wrong and states which assumption was wrong
Three different models each write one, without seeing the others
Writes and edits no code and only runs things: compile, PoC, fuzz, tests, output comparison
Uses a checklist to look for deleted features, edited tests, overly broad checks and similar mistakes; its opinion is advice and never replaces execution
Submit only when I understand every line, and answer reviewers myself
The patchers use three different models because a study by Team Atlanta, the AIxCC winner, found that the choice of model mattered more than the choice of agent framework, and that a group of three was the best value.
Rules for the sandbox:
- Each bug gets a fresh disposable container (an isolated environment thrown away after use), and the whole thing is deleted when done.
- No keys inside the container. All model calls go through a LiteLLM proxy outside it; the keys stay with the proxy, which also sets a dollar cap per agent. OSS-CRS also uses LiteLLM to manage budgets.
- Network on an allowlist only: it can reach the proxy and the places that serve source code, and everything else is blocked.
- Testcases and bug comments are always treated as untrusted data. They may hide instructions written for the AI that try to trick it into doing something else; this is called prompt injection. An agent that reads such text may only report it, never follow it.
The big players do the same. The Chrome Security Team runs its scans on locked-down machines with restricted network access, and Mozilla runs its jobs in throwaway virtual machines.
There are two human checkpoints: I choose which bugs to work on, and any action that leaves the sandbox, such as uploading, sending a report or sending email, stops and waits for me.
This is the concrete version of the security-first plan from Day 1: limit what agents can touch first, then let them do more. Keeping the validator and the reviewer separate from the patchers is the lesson from Day 3: the AI said it was done, and an independent review still found 37 problems. Whoever does the work cannot be the one who declares it done.
Realities and Red Lines
- You will probably collide with Google. As covered earlier, Chrome already scans and patches with AI every day, and its rewards reportedly account for what internal AI already finds. An outsider's value lies in what internal tools miss; otherwise, treat it as practice.
- Unverified AI reports hurt your reputation. Under Chromium's Shepherding AI Reports guide, an AI report with no working PoC or plausible ASan stack may be closed as WontFix; and since 2026-10-01 Google has paused new product-vulnerability reports to its OSS VRP after a surge of mostly invalid automated and AI reports (BleepingComputer, 2026-10-05).
- No details of unfixed bugs. Write only after the fix has become public, usually about 30 days later; do not try to see restricted bugs.
- Respect small projects. libxml2 stopped offering security embargoes (an agreement to keep a bug private until it is fixed) in mid-2025, and its long-time maintainer stepped down in September 2025. Maintainers reacted badly to AI patches thrown straight at them as PRs; the better way is to open an issue or report privately first, then offer the patch (OpenSSF podcast, February 2026). Practising on old bugs locally is fine, but do not dump piles of AI patches on under-staffed projects.
- Do not paste restricted bug details into third-party AI services unless the project explicitly allows it. I found no written rule, so ask first.
- This post is defensive research. It covers only reproducing, finding the root cause, patching, verifying and disclosing responsibly. A PoC here is just "the file that makes the crash happen again"; how to exploit a bug is not discussed.
Closing
Drafting a patch with AI is no longer hard; proving the patch is right is. What one person can do is let programs verify every step and keep the last gate for themselves. I will start at step 0: pick 20 old ARVO bugs, measure the four numbers, and write up the results in a log once I have them.
References
Industry systems
- Chrome Security Team: Chrome stronger with every update, 2026-07-30
- Google DeepMind: Introducing CodeMender, 2025-10-06
- Google Cloud: CodeMender in preview, 2026-07-22
- Google: CodeMender patches attached to OSS-Fuzz reports, 2026-07-29
- Google Research: AI-powered patching, 2024-01
- Chrome Releases: CVE-2025-9478, found by Big Sleep, 2025-08-26
- DARPA: AIxCC final results, 2025-08-08
- SoK: a paper surveying the AIxCC team systems, 2026-02
- Trail of Bits: Buttercup, accessed 2026-10-09
- OpenSSF: OSS-CRS, accessed 2026-10-09; OSS-CRS paper, 2026-03
- OpenAI: Aardvark, now Codex Security, 2025-10-30
- Anthropic: Claude Code Security, 2026-02-20
- Mozilla Hacks: behind the scenes of hardening Firefox, 2026-05-07
Research on whether patches are right
- PatchBench, 2026-09-03
- CodeRover-S, ICSE-SEIP 2026, 2026
- Meta: AutoPatchBench, 2025-04-29
- Team Atlanta: patch ensemble study, 2026-03-11
- Northwestern patch study, 2026-10-07
- PatchAgent, USENIX Security 2025, 2025
- VulDebugger, 2025-04
- CyberGym-E2E, ICML 2026, 2026-06
- ARVO dataset, IEEE EuroS&P 2026, 2026
Official Chromium docs
- Linux build instructions, accessed 2026-10-09
- ASan docs, accessed 2026-10-09
- agents/skills directory, accessed 2026-10-09
- AI coding policy, accessed 2026-10-09
- Security for agents, accessed 2026-10-09
- AI-generated security bugs FAQ, 2026-04
- Shepherding AI Reports, accessed 2026-10-09
- Contributing guide, accessed 2026-10-09
News and community