My verdict, up front: cybersecurity can be drawn as a map of 7 broad categories (I call them "families") and 22 domains. Over the past two years, two things have run through almost every domain: speed, and people passing themselves off as someone else. Vulnerabilities are often exploited before a fix is out, and a single phone call to an IT help desk, or one stolen login token (the temporary pass a system gives you once you have logged in), can be enough to let someone walk straight in. AI is making both attack and defence faster, and AI systems themselves have become a new target, which the industry calls an "attack surface". In the short term, the organisations and individuals who have prepared least stand to lose the most.
This is a survey of the whole field of cybersecurity as it stands in early October 2026, written so that, as far as possible, someone without a computing background can follow it. AI is changing attack and defence at the same time, but the more I looked, the more convinced I became that protecting AI systems themselves (often called AI security) is just one corner of cybersecurity, and only makes sense once you put it back in the context of the field as a whole.
The figures in this post come from public annual reports, official data and media coverage. For a few of the reports I could not get hold of the full text and had to rely on how the media reported them; some figures I recalculated myself from publicly available raw data. The details are in "About the Sources".
How to read this post: this post is long. It works more like a map you can look things up in, so feel free to skip to the parts you need.
- For the big picture first: read "Key Takeaways" and "The Whole Field on One Map".
- To see what attacks look like today: see the section "Who Attacks and How They Get In".
- To understand where AI fits in: see "The Four Roles of AI in Security" and "Who Has the Upper Hand, Attackers or Defenders".
- To do something straight away: jump to "What Individuals and Teams Can Do".
Key Takeaways
- Vulnerabilities are the most common way in: a vulnerability is a flaw in software that someone can take advantage of. According to M-Trends 2026, the annual report from the security firm Mandiant, exploiting a vulnerability came top among all the known ways attackers got in, at 32%. Mandiant also calculates that, on average, attackers start exploiting a vulnerability 7 days before the vendor releases a patch (an update that fixes it). Google's threat intelligence team, for its part, found that of the 90 vulnerabilities exploited in 2025 before a patch was available, 21 targeted security and networking equipment such as firewalls and VPNs, the very devices meant to guard a company's gates.
- Attackers log in rather than break in: according to several media outlets, 82% of the attacks detected by the security firm CrowdStrike involved no malware (malicious software such as viruses and Trojans). In other words, attackers often bring no malicious software at all and simply log in with accounts they have stolen or tricked people out of. A common trick is to ring a company's IT help desk, pose as an employee who has forgotten their password, and talk the staff into resetting it, or into moving the phone verification step used at login over to the attacker's phone. Reportedly, this is how the UK retailer Marks & Spencer was broken into, and its online sales were halted for about 46 days.
- Extortion — fewer victims pay, so the tactics get harsher: ransomware locks your files by encrypting them and demands payment to unlock them. According to several media outlets, the blockchain analytics firm Chainalysis estimates that ransomware payments worldwide came to about US$820 million in 2025, down for the second year running. Yet in the same year, the number of attacks claimed by the ransomware gangs themselves rose by about 50%. With fewer victims paying, the gangs' answer is to destroy backups first, so that victims cannot recover on their own; some skip encryption altogether, simply steal the data and then threaten to publish it.
- Poison one spot and everyone downstream is hit: a supply chain attack does not go after you directly; instead it contaminates software you are going to install. A package is a ready-made module of code, written by someone else, that you can install and use straight away. In March 2026, LiteLLM, an AI-related package, was tampered with, and its malicious versions were installed about 119,000 times within 2 hours 32 minutes; each of those installations could have carried malicious code onto a computer.
- AI is already helping attackers, but not yet on autopilot: in November 2025, Anthropic, the company that makes the AI Claude, disclosed that a hacking group it believed with high confidence to be backed by the Chinese state had used Claude Code, a coding tool, to have the AI carry out 80% to 90% of the hands-on steps of its attacks. As of September 2026, however, Google's threat intelligence team still said it had not seen a fully automated attack workflow used against real targets.
- AI helping defenders — finding vulnerabilities is no longer the bottleneck, fixing them is: Anthropic runs a vulnerability-hunting programme called Glasswing that is open only to its partners. By Anthropic's own account, of the 530 vulnerabilities it had reported to software maintainers by 22 May 2026, only 75 had been fixed. As I see it, the three companies building the most advanced AI models (the industry calls them frontier labs) are all handing their most cyber-capable models to vetted defenders first, so that defenders can use them a step ahead of attackers; but that head start will probably last months, not years.
- AI itself has become an attack surface: large language models (the kind of AI behind ChatGPT) cannot tell "instructions" apart from "data". Say someone sends you an email with the sentence "send this person's files out" hidden inside it: when your AI assistant sorts through your inbox for you, it may simply do as it is told. This is called indirect prompt injection. The EchoLeak vulnerability in Copilot, Microsoft's AI assistant for its office software, was of this type. There is a connection standard that lets AI assistants plug into outside tools, and the publicly disclosed vulnerabilities related to it numbered about 72 for the whole of 2025 and had already reached about 514 in 2026 up to 3 October (approximate figures from a keyword search).
- Taiwan — relentless attempts, old holes: according to several media outlets, Taiwan's National Security Bureau counted an average of 2.63 million intrusion attempts a day against the country's critical infrastructure (systems such as electricity and water supply) in 2025. That counts people knocking at the door, not actual break-ins. Microsoft's 2026 Digital Defense Report says that, in the Asia-Pacific region, Taiwan is where it saw the most security incidents. The US government's "Known Exploited Vulnerabilities" catalogue only lists vulnerabilities that hackers are confirmed to have used, and at least 78 of its entries are products from Taiwanese brands, mostly routers and network storage devices. The revised Cyber Security Management Act took effect on 1 December 2025, and its accompanying regulations require the organisations it covers to notify the government within 1 hour of becoming aware of an incident.
What Security Actually Protects
Before drawing the map, let me set out a few ideas that will keep coming up later. If you already know them, skip straight to "How We Got Here Over Forty Years".
Three Basic Goals
The most common definition says that security protects three things. Put their initials together and you get CIA, which has nothing to do with the American spy agency.
- Confidentiality: only the people who should see something can see it. If medical records, passwords or customer lists leak out, confidentiality has been broken.
- Integrity: data and systems have not been secretly changed, for example by swapping the account number on a bank transfer. Once that happens, you no longer know which records you can still trust.
- Availability: things work when you need them. Ransomware locks up your files and demands a ransom to unlock them, and what it attacks is availability. According to reports, a ransomware attack on the University of Mississippi Medical Center in the US in February 2026 closed all 35 of its clinics for 9 days.
Another goal that often comes up alongside these is authenticity: the other party really is who they say they are. According to media reports, in early 2024 an employee at the Hong Kong office of the engineering consultancy Arup joined a video call in which the "chief financial officer" and the other colleagues on screen were all deepfakes (fake video and voices made with AI). After the call he sent out HK$200 million in 15 separate transfers. No system was broken into; what was broken was authenticity.
Factories and power grids have different priorities. An ordinary company fears a data leak most, so confidentiality often comes first. But in places that control physical equipment, such as factories, power grids and reservoirs, the top priorities are usually human safety and availability. Stuxnet, discovered in 2010, damaged centrifuges at Iran's Natanz enrichment plant and proved that code can break physical equipment. In January 2024, according to reports, a piece of malicious software called FrostyGoop cut off the heating to more than 600 apartment buildings in Lviv, Ukraine, for about two days in sub-zero temperatures. As I see it, that attack hit availability and human safety, not data.
Security, safety and privacy are three different things. In Chinese, a single word covers both "security" and "safety", so in this post I use them as follows:
- Security: defending against a deliberate adversary.
- Safety: preventing accidents. In July 2024, the security company CrowdStrike pushed out a faulty update that crashed huge numbers of Windows computers around the world. It was not an attack, but the consequences looked very much like one.
- Privacy: people's rights over their own data. Data can be kept perfectly safe from leaks and still be used for things the person never agreed to.
How Risk Is Calculated
Security is not about locking everything up; it is about deciding where to spend limited time and money. Roughly speaking, risk is about likelihood times impact: how easily something could happen, multiplied by how much you would lose if it really did.
Take a home network-attached storage device (NAS) as an example. It is basically a private hard drive plugged into your home network. If its admin page is open to the internet, and its system has not been updated and still has known vulnerabilities (flaws in a program that someone can exploit), the likelihood of it being found is high. If it holds the only copy of ten years of family photos, the impact is large. When both are high, that is the risk to deal with first.
What you want to protect
Who might come for it
Where the gaps are
How much you lose if it happens
Lower the likelihood or the loss
Likelihood is the hardest part to estimate. Many thousands of vulnerabilities are made public every year, but the Google Threat Intelligence Group found that only about 0.23% of those disclosed in 2026 were actually used in attacks. Since nobody can fix them all straight away, the industry pays more and more attention to the Known Exploited Vulnerabilities catalogue (KEV), kept by the US Cybersecurity and Infrastructure Security Agency. It lists only vulnerabilities that are confirmed to have been used in attacks, so it works as a "fix these first" list.
Once you know the risk, there are four ways to deal with it:
- Reduce: add controls, for example by keeping the NAS off the internet and making offline backups.
- Transfer: buy cyber insurance, or hand the job to a managed service provider. What you outsource is the work; the responsibility is still yours.
- Avoid: simply don't do it. If you don't need to reach the NAS from outside, don't switch on remote access.
- Accept: if the risk is small, or dealing with it would cost more than the loss, write that down explicitly so you know what you have accepted.
Principles That Run Through This Post
The seven principles below come up again and again in every domain later on.
- Defence in depth: stack several independent layers of protection, so that if one fails the next is still there. The counter-example is the mass takeover of customer accounts on the Snowflake cloud data platform in 2024. After investigating, Mandiant, the security firm owned by Google, listed three causes: multi-factor authentication (MFA) was not switched on, passwords were never changed, and nothing restricted where logins could come from. Multi-factor authentication means that when you log in, besides your password you also confirm it is you in a second way, for example with your phone.
- Least privilege: every person, every program and every AI agent (an AI program that can carry out a task step by step on its own) gets only the permissions it needs to do its job. Many companies keep their customer data in Salesforce, a cloud service for managing customer relationships, and then let a third-party tool, Salesloft Drift, connect to it and access that data. In August 2025, attackers stole this tool's authorisation tokens (passes that let one service get into another on your behalf) and exported large amounts of data; according to several media outlets, more than 700 organisations were affected. A single integration tool could reach an entire customer database, and that is what giving too much permission looks like.
- Separation of duties: no important action should be carried out from start to finish by the same person. If those Arup transfers had needed a second person's approval, followed by a call back to a pre-registered phone number to confirm, fooling one person would not have been enough. Against deepfakes, I think what works is process, not trying to tell real video from fake by eye.
- Secure by default and secure by design: users are safe even if they do nothing. For example, since April 2024 the UK has banned connected devices from using default factory passwords.
- Reduce the attack surface: the fewer services you expose to the outside, the fewer ways in there are. "Edge devices" that guard the edge of a company network, such as virtual private networks (VPNs, the tunnels staff use to connect back to the office from outside) and firewalls, were meant to keep people out, yet they are now one of the main ways in. In the KEV list, new entries for vulnerabilities from edge device vendors numbered 33 for the whole of 2024, and had already reached 49 in 2026 by early October.
- Zero trust and assume breach: don't trust a connection just because it comes from inside the company; check identity on every access, and design as if the attacker is already inside. This is not pessimism: in M-Trends 2026, Mandiant's report on the intrusions it investigated, 1 in every 10 intrusions started from a system that had already been broken into earlier.
- Prevention will sometimes fail: so you need to spot trouble early (detection), and to stop the damage and recover once something has happened (response). The good news is that in M-Trends 2026, the share of intrusions discovered by the victim organisation itself rose from 43% to 52%. The bad news is that ransomware gangs have started deliberately destroying backups so that you cannot restore your data. Backups therefore need to be kept offline or somewhere they cannot be altered, and restoring from them needs to have actually been rehearsed.
Why Security Is So Hard
You often hear that defenders have to guard every door, while attackers only need to find one way in. That is only half right. For getting in, one way is broadly enough. But once inside, attackers still have to gain higher privileges, hop from one machine to another (called lateral movement), find something valuable and get the data out, and every one of those steps can be spotted. The "Cyber Kill Chain", put forward in 2011 by the major US defence contractor Lockheed Martin, breaks an intrusion into seven stages. The attacker has to succeed at every stage; the defender only has to catch them at one.
In a blog post in September 2026, Google Cloud's chief information security officer went further, arguing that defenders actually have the advantage, because they know better than anyone what their own house looks like. My view is that this only holds if you can actually see your own house: for an organisation that doesn't know what devices it has, and whose systems keep no record of what has been done on them, this home advantage is zero. More on this in "Who Has the Upper Hand, Attackers or Defenders".
Too many systems, many of them old. Every organisation has been built up layer by layer over decades: new cloud services are wired to ten-year-old servers, and any piece of it could have vulnerabilities that need patching. New vulnerabilities are also turning up faster and faster: the Google Threat Intelligence Group found that the number disclosed in August 2026 was more than double the number in January.
The people who decide are not the people who pay. When a software vendor's product has a vulnerability, its customers bear the cost; when one company is hit, its suppliers and the businesses that depend on it grind to a halt too. When the UK carmaker Jaguar Land Rover was attacked in 2025, the media quoted an estimate from the UK's Cyber Monitoring Centre: the incident affected more than 5,000 UK organisations and cost the UK economy about £1.9 billion.
When defence works, the result is that nothing happens, and that is hard to justify in a budget meeting.
People get fooled. In M-Trends 2026, voice phishing, meaning tricking people over the phone, has become a more common starting point for intrusions than phishing emails (fake messages designed to scam you). More training sessions also have only a limited effect; see "22 Human Factors and Social Engineering".
Speed. CrowdStrike says criminal groups take on average just 29 minutes from breaking in to starting lateral movement. Patching is much slower: the 2026 edition of Verizon's Data Breach Investigations Report found that the median time to patch a vulnerability was 43 days (reports do not make clear whether this covers only edge devices).
How We Got Here Over Forty Years
Here are the landmark incidents of nearly forty years in date order, with the oldest at the top.
Reading down this stack, I see one trend: the door defenders have to guard, known as the "perimeter", keeps moving inwards. At first it was enough to guard the front gate of the company network (the firewall). Later, attackers were already inside the gate, so the focus shifted to guarding every single computer. After that, attackers simply logged in with employees' accounts, or tampered with software you were going to install. Now AI agents (AI programs that can carry out a task step by step on their own) also hold accounts and do work with them, so they have become another door to guard.
The network. The Morris worm of 1988 (a worm is malicious software that copies and spreads itself) brought a sizeable share of the computers then on the internet to a halt. The answer at the time was intuitive: put a firewall between inside and outside. After Snowden revealed state-level surveillance in 2013, encrypted connections gradually became the default, and the network itself was no longer treated as a place that could be trusted.
The computers themselves. The ILOVEYOU worm of 2000 hid in an email attachment and relied on recipients opening it themselves. In 2017, once WannaCry got inside a company network, it used a Windows vulnerability to sweep through the whole organisation. A firewall cannot stop something that is already inside, so the focus moved to antivirus and detection tools on each computer.
Identity. The 2013 break-in at the US retailer Target started from an account belonging to an air-conditioning contractor. At the US fuel pipeline company Colonial Pipeline in 2021, the starting point was an old VPN password without multi-factor authentication. The ransomware only reached the company's IT systems, yet the company shut down the pipeline as a precaution. In 2025, attackers posing as employees phoned the outsourced IT help desk of the UK retailer Marks & Spencer and tricked the staff into resetting a password; reportedly, its online sales stopped for about 46 days as a result. Attackers less and less often "break in" and more and more often "log in": of the attacks CrowdStrike detected, 82% involved no malicious software at all.
The software supply chain. Most software today is assembled from parts that other people have already written (called packages), and many of those parts are maintained by just a handful of people. A supply chain attack doesn't go after you directly; instead, it contaminates software or parts that you are going to install.
In 2020, attackers planted a backdoor (a secret entrance left in on purpose) in an official update of the SolarWinds network management software. With the xz backdoor in 2024, someone spent several years winning the trust of the maintainers of this open-source package before slipping the backdoor in. In March 2026, a North Korean team tricked the maintainer of the popular package axios out of their account and published a malicious version. The package is downloaded about 70 million to 100 million times a week.
AI agents. In November 2025, Anthropic, the company that makes the AI assistant Claude, said it was highly confident that a Chinese state-sponsored hacking group had used its coding tool, Claude Code, to carry out 80% to 90% of the hands-on work in an attack campaign. In May 2026, the Google Threat Intelligence Group said it had found the first attack program that it believes was developed by AI to exploit a zero-day vulnerability (one the vendor does not yet know about and has not patched). As of September, though, it still said it had not seen a fully autonomous attack used against real targets.
On the defending side, Anthropic has opened an unreleased AI model only to partners, for finding vulnerabilities; this is called Project Glasswing. By Anthropic's own published figures, of the 530 high- or critical-severity vulnerabilities it had reported as of 22 May 2026, only 75 had been fixed. My reading is that fixing cannot keep up with the pace at which vulnerabilities are being found. The details of these events are in "The Four Roles of AI in Security".
Old perimeters don't disappear; new ones are simply added on top. In 2023, the Cl0p gang used a zero-day vulnerability in the file-transfer software MOVEit to steal data on a massive scale, yet it did not lock up files the way typical ransomware does (see "Extortion Shifts from Encryption to Data Theft"). In 2025, the same playbook turned up again under the Cl0p name, this time against Oracle's business management software. Even today, people are still trying to prise open the outermost gate. Worms haven't gone away either: in 2025, the self-replicating Shai-Hulud worm appeared on npm, a package library that developers use all the time.
The Whole Field on One Map
Security covers so much ground that it is hard to take in at a glance. At the top, a board decides how much risk the company is willing to carry; at the bottom sits a single line of code in a medical device. In between are cloud settings and a phone call to the help desk. Everyone draws the map differently, and the industry has no classification that everyone accepts.
I have divided the field into 7 families and 22 domains. Each domain is a speciality the industry can name, and domains in the same family answer the same kind of question. The order of the families roughly follows the Cybersecurity Framework from the US National Institute of Standards and Technology (NIST). A framework here is simply a list of the things an organisation should get done.
Who is responsible, how much risk is acceptable, who must be told when things go wrong
Who you are, and why anyone should believe you
Where data and code actually run
The whole line from writing code to releasing and patching it
Assuming you have already been broken into
Authorised testing from the attacker's side
Where standard IT practice does not fit
The families are bound to overlap. AI system security sits in the fourth family because an application built on a large language model (LLM, the kind of model behind ChatGPT) is, first of all, a piece of software, and its vulnerabilities still have to be found and patched. That domain covers only how to protect AI itself; how AI is changing all 22 domains comes at the end of this section.
How the Seven Families Are Divided
The NIST Cybersecurity Framework 2.0 (CSF 2.0 from here on) was released in February 2024. It splits an organisation's security work into six functions: Govern, Identify, Protect, Detect, Respond and Recover. The CSF only describes the outcomes an organisation should achieve and does not say which products to buy, so I think it makes a good skeleton for grouping the families.
Govern maps roughly onto the first family, Identify and Protect onto the second to fourth, and Detect, Respond and Recover onto the fifth. Offensive security, and the physical world and people, do not fit into any single function, so each gets a family of its own.
Family 1 Governance and Compliance. Who answers for it when something goes wrong, and whether the regulator has to be told, are not things engineers can decide on their own. The three domains here, governance, regulation and privacy, do not touch code directly; questions like these are their job. Regulation turns part of governance into obligations with penalties and reporting deadlines, while privacy puts the focus on the person behind the data.
Family 2 Identity and Cryptography. Identity and access management is like a company handing out door passes: it decides which card opens which door. Zero trust architecture checks again every single time, and does not wave you through just because you are on the office network. Cryptography here means encryption, which scrambles data so that only someone holding the key can read it; it has nothing to do with the password you log in with. Cryptography is the mathematical foundation for all of this, and post-quantum migration means replacing that foundation before quantum computers mature.
Family 3 Infrastructure. This family is about places: edge devices such as firewalls, which guard the border between a company's internal network and the outside world; endpoints such as computers and phones; the cloud and software as a service (SaaS, software like Gmail that you use straight over the internet without installing anything); and data and backups. The homework is much the same for all of them: know what you have, set it up correctly, patch promptly, and be able to restore things when something goes wrong.
Edge devices have been a favourite target in recent years. A zero-day vulnerability is one the vendor does not yet know about, so it has not been patched. Google's threat intelligence team counted 33 zero-days aimed at enterprise products in 2024, and 20 of them were in security or networking products. The very devices meant to keep people out have become the way in.
Family 4 Software and AI Systems. Application security looks after the code a company writes itself. The software supply chain looks after the ready-made building blocks of code (packages) that other people wrote and you simply use, and it guards against someone poisoning software you are going to install. Vulnerability management decides which of the known flaws, or vulnerabilities, gets fixed first. AI system security looks after models, prompts (the instructions you type into ChatGPT) and AI agents, meaning AI assistants that can operate tools themselves rather than just answering questions.
All four domains rely on CVE numbers, the standard identifier each publicly known vulnerability is given, rather like an ID number.
Family 5 Detection and Response. This family starts from the assumption that hackers are already inside. The security operations centre is like a building's CCTV control room: it watches alerts and looks for anything unusual. Incident response and forensics deals with what has already happened; forensics is like police gathering evidence, working out how the hackers got in and what they took. Threat intelligence tells the other two who the adversary is and what to look for.
Family 6 Offensive Security. Penetration testing means being hired to attack a client's systems, find the holes and write them up in a report. Red teaming is closer to a live exercise: it plays the part of a real attacker to see how long the company takes to notice. Vulnerability research means hunting for holes in software that nobody knows about yet; a bug bounty is a company publicly offering rewards to outsiders who report vulnerabilities. The techniques are the same ones attackers use, and the difference is written authorisation with a clear scope. Without authorisation, the same actions are a crime.
Family 7 The Physical World and People. Operational technology (OT) means the systems that control physical equipment such as factory machines and power grids. What it fears most is people getting hurt and things grinding to a halt. You cannot just restart it the way you would a laptop, so patches have to wait for maintenance shutdowns. Equipment stays in service for 15 to 30 years, and many of the rules machines use to pass commands to each other (communication protocols) never check who is on the other end.
The Internet of Things (IoT) means connected household appliances and cameras; medical devices and vehicles each have their own regulations. Last come people, and human carelessness cannot be "patched". Social engineering means leaving the computer alone and fooling a person instead, for example by ringing the company's IT help desk and pretending to be an employee who has forgotten their password.
Next, "Who Attacks and How They Get In" looks at the attackers first. After that, each family gets a section of its own, covering the 22 domains one by one.
Which Part Each Common Framework Covers
Security frameworks have many similar-sounding names, which often leads people to think they must complete every one of them before they count as compliant. In fact each one answers a different question. If you are just starting out, recognising these few is enough:
- NIST CSF 2.0: which security outcomes an organisation should achieve, namely the six functions described above.
- ISO 27001: a certificate awarded after a third-party audit, proving to customers that the company has a proper system for managing information security.
- MITRE ATT&CK: an encyclopaedia of attack techniques compiled by MITRE, a US non-profit, recording how attackers actually operate.
- OWASP Top 10: the ten biggest risks for web applications, listed by the Open Worldwide Application Security Project (OWASP), with separate editions for LLMs and AI agents.
- Common Vulnerability Scoring System (CVSS): gives the severity of a vulnerability itself a score from 0 to 10.
- Exploit Prediction Scoring System (EPSS): estimates the probability that a vulnerability will be exploited in the next 30 days.
- Known Exploited Vulnerabilities catalogue (KEV): maintained by the US Cybersecurity and Infrastructure Security Agency, it lists only vulnerabilities that have actually been exploited.
CVSS looks only at how severe a vulnerability is in itself, not at whether anyone will actually use it. The software company Datadog says in its own 2026 report that of the vulnerabilities in packages rated "critical" by CVSS, only 18% still count as critical once you take into account the environment the software really runs in, along with EPSS. So a high score does not mean most urgent.
As I see it, these frameworks work as a relay rather than replacing one another. The CSF sets the goals, ISO 27001 proves to others that you have done the work, ATT&CK lets you check whether your monitoring would catch these attacks, and KEV and EPSS decide which one to fix first today. If I were one person looking after the devices at home, I would not bother with an ISO certificate, but I would borrow the CSF's six functions as a checklist and use KEV to decide the order of patching.
Which Domains AI Touches
How much has AI changed each of these 22 domains? Below, every domain gets three scores, each from 0 to 3:
- Attack boost: how much stronger AI makes attackers, for example by finding vulnerabilities faster or producing scams more cheaply.
- Defence boost: how much labour AI saves defenders, and how much it lets them see that they could not see before.
- New risk: how many new openings for attack appear once AI is brought in, such as the permissions and data that agents are given.
| Attack | Defence | New risk | |
|---|---|---|---|
| 1 Governance | Low | Medium | Medium |
| 2 Compliance | None | Medium | Medium |
| 3 Privacy | Low | Low | High |
| 4 Identity | Medium | Medium | High |
| 5 Zero Trust | Low | Low | Medium |
| 6 Crypto | None | Low | Low |
| 7 Network | High | Medium | Low |
| 8 Endpoints | Medium | Medium | Low |
| 9 Cloud/SaaS | Medium | Medium | Medium |
| 10 Data | Low | Medium | High |
| 11 AppSec | High | High | High |
| 12 Supply | Medium | Medium | High |
| 13 Vuln mgmt | High | High | Medium |
| 14 AI systems | High | Medium | High |
| 15 SOC | Low | High | Medium |
| 16 IR | Low | Medium | Medium |
| 17 Intel | Low | Medium | Low |
| 18 Red team | High | High | Low |
| 19 Research | High | High | Low |
| 20 OT/infra | Low | Low | Low |
| 21 IoT | Low | Low | Low |
| 22 People | High | Low | Medium |
Later on, the last part of each domain, "What AI changes", refers back to this table.
The hottest area is code and vulnerabilities. Application security, vulnerability management, penetration testing and red teaming, and vulnerability research (11, 13, 18 and 19) score 3 in both the attack and defence columns. The likely reason is that the daily work in these domains is reading code and finding mistakes, which happens to be exactly what today's models do best.
According to several media outlets, in the final of the AI Cyber Challenge run by the US Defense Advanced Research Projects Agency in 2025, the competing systems found 18 real vulnerabilities that had been unknown until then. When a vendor ships a patch, it is in effect telling everyone where the hole used to be. Google's threat intelligence team found that attackers get LLMs to compare the code before and after a patch so that they can build attack tools faster; one vulnerability was exploited within 4 days of being made public. So in my view, finding vulnerabilities is no longer the bottleneck; patching them is.
The other 3 in the defence column goes to the security operations centre (15). Having AI agents sift through the alerts first, before a person looks at them, appears to be the direction vendors are pushing hardest in 2026.
Identity and access management (4) scores 3 for new risk. The likely reason is that every AI agent is like a new employee: it needs its own door pass (an account, keys, and the temporary pass a system hands out after you log in), and it is often given a pass that opens too many doors.
A login entrance that is not properly guarded will also draw attackers who use AI. Single sign-on is a main entrance where one username and password get you into many systems; multi-factor authentication means confirming on your phone as well as typing your password. Taiwan's Ministry of Digital Affairs confirmed that in July 2026, someone used a tool that can direct several AI agents at once to break into Taiwanese government systems and take over 85 accounts. Of those, 84 came through a single sign-on that did not have multi-factor authentication switched on. The details are in "Where Taiwan Stands".
Human factors and social engineering (22) scores 3 in the attack column. According to several media outlets, in February 2026 the then chair of Italy's Fideuram received WhatsApp messages impersonating the chief executive of its parent group, backed up by an AI-cloned voice of a senior lawyer, and ended up handling transfers of about €95 million. Yet this row scores only 1 for defence. What really stops scams like this is mostly process, such as hanging up and calling back yourself to check, or requiring two people to approve large transfers, and AI cannot help much there.
AI system security (14) is a domain that AI has created by itself, and it scores 3 for both attack and new risk. The Model Context Protocol is a common connector that lets AI plug into outside tools, like a USB socket for AI. CVEs related to it numbered about 72 for the whole of 2025 and had already reached about 514 in 2026 by 3 October; these are rough figures based on keyword matching.
Privacy (3) and data security (10) also score 3 for new risk. Once company data is pasted into a tool like ChatGPT, or used to train a model, there is one more way for it to leak. Someone could also deliberately slip fake data in to teach the model bad habits.
By contrast, cryptography (6), OT (20), and IoT, medical devices and vehicles (21) are almost all 0 or 1. The hard part in these domains is probably not finding the problems but replacing old equipment and old algorithms that have been in use for many years, and that is a step where AI cannot help much.
Who Attacks and How They Get In
The map in the previous section was drawn along the lines of how defenders divide up the work. This section moves over to the attackers' side and looks at who is attacking, where they get in, how fast they move, what they are after, and how much damage they end up doing.
The data comes mainly from a handful of reports published every year: those of two security companies, Mandiant and CrowdStrike; the data breach reports of the telecoms company Verizon and the technology company IBM; statistics from Microsoft and from the Google Threat Intelligence Group (GTIG); and the EU Agency for Cybersecurity. Each uses its own samples and categories, so their answers often don't match; where that happens, I set out the differences as well.
Who the Attackers Are
Sorted by motive, attackers fall roughly into four groups: criminal gangs out for money, state-backed teams working for a government, hacktivists acting for a cause, and people who are already inside the organisation. The lines between them seem to be getting blurrier: state teams steal money too, and criminal gangs' tools end up in state hands.
Why the names are so odd. You will meet a lot of strange names below. They are code names that security companies give to attacking groups. Each company has its own naming scheme; Microsoft, for example, uses the weather, with Typhoon for China and Sleet for North Korea. The result is that the same team doesn't match up from one report to the next. The North Korean team that attacked axios, which comes up later, is Sapphire Sleet to Microsoft and MIDNIGHT NEPTUNE to GTIG. The more practical approach is not to memorise names but to look at the methods a group uses.
Criminal gangs have turned extortion into a franchise business. Ransomware is malicious software that locks your files by encrypting them and demands payment for the key. Today most of it is run as "ransomware-as-a-service": one group writes the software, handles the negotiations and runs a leak site, where it publishes the data of victims who won't pay in order to pile on the pressure; another group, the "affiliates", does the breaking in, and the ransom is split by an agreed share. Some groups take a different route. Cl0p, for example, uses zero-day vulnerabilities in business software, meaning flaws that are attacked before the vendor has had a chance to patch them, to steal data on a large scale and then demand money.
The ones who make phone calls. Scattered Spider and ShinyHunters don't rely on clever code. Their speciality is the phone call, and their members are mostly young English speakers. They pose as employees and ring the company's IT help desk (the people staff call when something goes wrong with their account), then talk whoever answers into resetting a password or re-linking multi-factor authentication (where, as well as typing your password, you confirm the login again on your phone). Once inside, they work their way through to the virtualisation hosts and the backup systems. A virtualisation host is a single physical machine running many "virtual computers" at once, so taking it is like taking many machines in one go. The UK's National Crime Agency arrested four people in July 2025, the youngest only 17. How to defend the IT help desk is left for "22 Human Factors and Social Engineering".
State-backed teams. The four countries named most often each have different aims.
- China wants to lie low for the long term. In 2023 Microsoft described Volt Typhoon as doing its work almost entirely with the management tools already built into computers, which makes it very hard to catch as malicious software. Microsoft judged, with moderate confidence, that its aim is to be able to disrupt critical communications infrastructure between the United States and Asia during a future crisis.
- North Korea steals money, and tricks its way into jobs. In February 2025 the cryptocurrency exchange Bybit was robbed; CrowdStrike puts the haul at about US$1.46 billion and calls it the largest single financial theft on record. axios is a ready-made building block that many engineers drop straight into their code when building websites; publicly shared building blocks like this are called open-source packages. On 31 March 2026, a North Korean team tricked its way into the account of one of the people who maintain axios and published a malicious version that let hackers control computers remotely. It stayed up for about 3 hours. North Korea also applies for remote IT jobs under fake identities, which is covered below. GTIG did not attribute a single zero-day to North Korea in 2025; it looks as though its teams rely mainly on deceiving people rather than on digging up vulnerabilities.
- Russia is after sabotage. Sandworm has been active in and around Ukraine for years, and in October 2022 it used the power system's own control software to trip circuit breakers and cut the power.
- Iran is good at intelligence and deception. GTIG has observed one of Iran's state-backed teams using large language models (AI such as ChatGPT) to research targets and write messages designed to trick people into clicking. Iranian and North Korean teams also use AI to create convincing fake personas and fake CVs.
GTIG's September 2026 report says all four countries are now using large language models at every stage of an attack. But in real attacks, nobody has yet seen AI exploit a zero-day entirely on its own, without people involved. The details are in "The Four Roles of AI in Security".
There is one more kind: companies that sell hacking tools to governments, known as commercial spyware vendors. In GTIG's review of the zero-days exploited in 2025, more were attributed to these vendors than to state espionage teams, for the first time. Zero-days are now goods with a price tag.
Hacktivism. These are groups that act out of political or ideological conviction. The EU Agency for Cybersecurity found that more than half of the incidents in the EU in 2025 were denial-of-service attacks (flooding a website with junk traffic until it falls over), mostly carried out by groups of this kind. A few go as far as tampering with physical facilities. In April 2025, the web control panel of a dam in Norway was broken into because its password was too weak, and a discharge valve was opened for about 4 hours; fortunately no damage was done. Norway's Police Security Service believes a pro-Russian group seeking to "create fear" was behind it.
Insiders and impostor employees. The traditional insider threat is an employee who has been bribed. The new version is the impostor employee: North Korea uses fake identities to apply for remote IT jobs and, once it has legitimate accounts, draws a salary while stealing data. Mandiant found that in cases such as espionage and North Korean IT workers, the median time from intrusion to discovery was 122 days. Another situation isn't malicious but has similar consequences: employees pasting company data into AI tools that haven't been approved, known as "shadow AI".
Where They Get In
"Initial intrusion" means the step by which attackers first set foot inside an organisation. Companies often call in Mandiant to investigate break-ins, and its figures are the most direct, because they come from cases it has actually handled.
Exploiting a flaw means using a defect in software as a key to open the door. More surprisingly, the phishing email everyone knows best is down to just 6%, while tricking people over the phone has climbed to second place. "Prior compromise" is a way in opened by one set of attackers and then reused by the next; in ransomware cases, it comes first.
Verizon's 2026 Data Breach Investigations Report points the same way. Exploited vulnerabilities account for 31%, overtaking credential abuse (logging in with stolen usernames and passwords), which Verizon says is a first in the report's 19 years. Another striking figure is that 48% of breaches involved a third party such as a supplier or contractor, up from just 30% in the previous edition.
Why the reports contradict each other. The security company Sophos says 79% of ransomware attacks started with compromised identities (that is, usernames, passwords and login permissions), the opposite direction to the two reports above. The reason is usually not that someone got the sums wrong, but that they are measuring different things. Sophos asks managers at victim organisations, so its figures are self-assessments; Mandiant looks at cases it investigated itself, which lean towards serious incidents that needed outside experts. The categories matter too: a phishing email that tricks someone out of a password, which is then used to log in, can be recorded as "phishing" or as "credentials". So when you see a news story quote a percentage, check first which report it came from.
Taken together, the two doors opened most often are probably these: unpatched systems exposed to the internet, and people and their identities.
Edge devices are the beachhead. Edge devices sit where a company's network meets the internet: firewalls, for example, and the virtual private network gateways that let staff connect back to the office from home. They are there to keep attacks out, yet by their nature they are exposed to the internet, hold powerful privileges and usually can't run monitoring software. GTIG found that of the 90 zero-days exploited in 2025, 21 targeted security and networking equipment.
Identity: logging in, not breaking in. In the attacks CrowdStrike detected, 82% involved no malicious software at all; the attackers simply worked with legitimate accounts and tokens. Think of a token as the temporary pass a website gives you after you log in: with it, you don't have to type your password again. Microsoft's 2025 report says most identity attacks involve trying common passwords everywhere or guessing with huge numbers of passwords, and that multi-factor authentication blocks more than 99% of them. So attackers are switching to routes that multi-factor authentication can't block:
- Stealing tokens. Many companies use Salesforce, a cloud system for managing customers, and authorise an outside tool, Salesloft Drift, to connect to it and read the data, so Drift holds tokens for getting into Salesforce. In August 2025, attackers stole these tokens and exported large amounts of victim companies' Salesforce data. More than 700 organisations were affected.
- Phoning people. A passkey is a newer way of logging in that replaces the password with a key stored on your phone or computer, and it resists phishing by design. Yet a ransomware group that Google tracks under the code name UNC6671 rang employees' personal mobiles, told them they had to complete an "urgent, mandatory passkey enrolment", and led them to a fake web page that stole their usernames and passwords. To me this is the most ironic part: the very tool defenders are pushing has become the bait.
- Getting users to do it themselves. ClickFix is a fake error message that asks users to copy a command, paste it into the small Windows "Run" box and press Enter, claiming this will "fix the problem". In reality, they are running the hacker's program with their own hands. Microsoft counted more than 1.1 million devices that did exactly that between February and May 2026 alone.
Speed Becomes the Main Measure
It used to be that after a vulnerability was made public, it usually took a month or two before anyone exploited it, which gave defenders time to schedule a patch. Mandiant measures how long, on average, it takes for a vulnerability to be used in attacks, counting from the day the vendor releases a patch: in 2018–19 it was 63 days, and in the 2026 report it is −7 days. In other words, typically a vulnerability is already being exploited before the patch has even come out.
Mandiant: on average, exploited 7 days before the patch is released
Microsoft: the median time from a vulnerability being found in the wild (in real attacks on the internet) to being turned into an attack tool is well under 24 hours
Microsoft: cloud systems exposed to the internet are attacked within 5.3 hours on average
Mandiant: the median time for an intrusion team to hand its access to the next team is 22 seconds; in 2022 it took more than 8 hours
CrowdStrike: criminal groups spread beyond the first computer in 29 minutes on average, and in 27 seconds at the fastest
Verizon: the median time to patch is 43 days; the year before it was 32 days, but the two years may have been counted differently
Two of the steps on this clock are worth unpacking. The first is a vulnerability becoming a weapon. Publishing a patch effectively tells attackers where the vulnerability is, and anyone who hasn't updated becomes a target. GTIG believes large language models are speeding up this step, which is called "patch diffing": comparing the code before and after a patch to work backwards to where the vulnerability lies.
Another example involving AI is a vulnerability in BeyondTrust's remote support software (a tool that lets IT staff operate other people's computers from a distance), which was found by an AI research program. Attackers were exploiting it 4 days after it was made public, and within 7 days the number had grown to 6 groups of attackers.
The second is access changing hands. The median time for the team that breaks in to pass its access on to the next one, such as a ransomware affiliate, is just 22 seconds. With so little time, the handover has probably been largely automated.
Defenders are much slower. Verizon's median time to patch this year is 43 days. The previous year's figure was 32 days, but that year counted only edge devices, and it isn't clear whether this year's figure was worked out the same way, so the two can't be compared directly. The good news is that detection is improving: in Mandiant's cases, the share of organisations that discovered the intrusion themselves rose from 43% to 52%.
The US Cybersecurity and Infrastructure Security Agency (CISA) maintains a list called the Known Exploited Vulnerabilities catalogue (KEV), which includes only vulnerabilities confirmed to have been exploited in real attacks; many organisations use it to decide what to patch first. In the whole of 2025, 245 entries were added to the list; in 2026, up to 2 October, 249 have already been added. GTIG's own statistics agree: more vulnerabilities are being exploited, mainly because more n-days are being exploited. An n-day is a vulnerability that is already public, and may even have a patch, but that people haven't yet got round to installing. On the other hand, only about 0.23% of the vulnerabilities published in 2026 have actually been exploited.
What these numbers mean is that there are far more vulnerabilities than anyone can patch and very few are actually exploited, but that small fraction arrives very fast. So the order of patching should depend on information like KEV, which tells you whether something is being exploited, not just on how severe a vulnerability has been rated. My view is that for edge devices exposed to the internet, patching should be measured in hours. More fundamentally, any service that doesn't have to be exposed to the internet shouldn't be: if attackers can't reach it from outside, you don't have to race them.
Extortion Shifts from Encryption to Data Theft
The EU Agency for Cybersecurity considers ransomware still the threat with the greatest impact, but fewer victims are paying. Verizon found that 69% of victims didn't pay the ransom. The blockchain analytics firm Chainalysis estimates that global ransomware payments in 2025 came to about US$820 million, the second annual fall in a row (the figure may later be revised up to around US$900 million). Over the same period, though, the number of attacks the ransomware gangs themselves claimed rose by about 50%.
Making sure you can't recover on your own. Mandiant says ransomware gangs are shifting towards "recovery denial": first tracking down the backup systems, the virtualisation hosts and the systems that manage accounts and permissions across the whole company, and destroying whatever would let you recover by yourself. The reason isn't hard to imagine. Encrypt one virtualisation host and every virtual computer on it stops at once; delete the backups first, and it becomes very hard for the victim not to pay.
Skipping encryption and just stealing data. Microsoft found that 63% of intrusions involved data theft. UNC6671, mentioned earlier, works exactly this way: after stealing data it asks for US$1 million to US$3 million, and the deals end up settling at about US$750,000 on average. Cl0p took the same "steal the data, then send the extortion letter" route against customers of Oracle's business software.
Why the shift? There are probably several reasons. More organisations now have good backups, so the threat of encrypting files carries less weight. Stealing data is also quieter than encrypting it, and less likely to get caught. On top of that, many countries require data breaches to be reported, so once data gets out, a company faces legal liability and damage to its reputation, which is exactly the leverage extortion needs.
So "if you have backups, you needn't fear ransomware" no longer seems to hold. Backups can bring back data that has gone missing, but they can't stop data from being published. The only ways to guard against that are to keep less unnecessary data in the first place, be clear about which data is most sensitive, and encrypt sensitive data in advance, so that even if it is stolen it is harder to read.
Major Incidents from 2025 to 2026
Apart from the one item flagged in the caption, every incident on the timeline below is backed by an official primary source or by coverage from several media outlets. Hover over a label to see the date and details.
These twenty months of incidents point to several things.
- One intrusion can drag down hundreds or thousands of others downstream. This is called a supply chain attack: instead of attacking you directly, the attacker poisons software you are going to install. The malicious version of axios was up for only about 3 hours, yet GTIG saw victims in 13 countries. Verizon's finding, mentioned earlier, that nearly half of breaches involve a third party very likely has incidents like this behind it.
- One phone call or one intrusion can halt operations for more than a month. The IT help desk that the UK retailer Marks & Spencer (M&S) had outsourced was tricked by someone posing as an employee into resetting a password, and online sales were halted for about 46 days. The UK carmaker Jaguar Land Rover (JLR) stopped production at the start of September 2025 and only began a phased restart on 8 October. The UK's Cyber Monitoring Centre estimated the overall economic loss at about £1.9 billion and called it the costliest cyber incident in the country.
- States and criminals come in through the same doors. A vulnerability in Microsoft SharePoint (a system companies use to share documents internally) was exploited at the same time by two Chinese state-backed teams and by a group that deploys ransomware. For defenders, not being able to tell whether the other side is a state or a criminal gang usually doesn't matter much, because the doors to guard are the same either way.
- AI has appeared in incidents, but is not yet the main cause. Mandiant's conclusion is that 2025's breaches were not "caused" by AI, even though malicious software that calls large language models has already appeared. The root causes of most incidents still seem to be the old problems, such as unpatched systems and IT help desk procedures that are easy to fool.
Losses and Costs
What one breach costs. IBM's 2026 Cost of a Data Breach Report puts the global average cost of a single breach at US$4.99 million, the highest in the report's history. On average it took 247 days from a breach happening to it being brought under control, longer than in the previous edition, which ends a five-year run of that figure getting shorter.
Far more money is lost to scams than to ransoms. The Internet Crime Complaint Center of the US Federal Bureau of Investigation received reports of more than US$20 billion in losses for 2025. The biggest losses came from investment fraud, followed by business email compromise, meaning emails that pose as the boss or a supplier and ask for money to be transferred. These are mostly scams rather than system intrusions, so they can't be added to IBM's breach costs. But set beside Chainalysis's estimate of US$820 million in ransomware payments, they are more than twenty times larger; for victims, far more money seems to be lost to scams than is ever paid to ransomware gangs.
Who gets hit most. Microsoft's report says government is the most attacked sector. The same report also notes that Taiwan was the country with the most observed incidents in the Asia-Pacific region; the details are in "Where Taiwan Stands".
Overall, ransoms are falling but total costs are rising. My view is that the real bill is downtime and recovery, as JLR's £1.9 billion shows. So rather than agonising over whether to pay a ransom, the money is better spent on being able to patch quickly and, if you are hit, being able to recover on your own.
Family 1 Governance and Compliance
This family doesn't touch network equipment, and it doesn't involve writing code. It deals with questions that are more about people: who is responsible, how much risk is acceptable, and whom you have to tell when something goes wrong. As I see it, governance, regulation and privacy together decide how deep the work in every other domain needs to go.
1 Governance and Risk Management
What it is: governance has to answer a few practical questions: who is responsible for security, how much risk the company is willing to take on, where the money goes, and how you know whether any of it is working. Risk management turns these questions into a table that can be tracked, called a risk register.
Each row in the table is one risk, such as "customer data gets stolen". Next to it sit how likely it is, how big the impact would be, who is responsible for it, and the decision on how to handle it: find ways to reduce it, buy insurance to pass it on to someone else, stop doing the risky thing altogether, or write down in black and white "we accept this". The table rests on an asset inventory, a list of the computers, systems, data and accounts the company actually has. If you don't even know what you have, you can't talk about risk.
This work is usually led by the chief information security officer, who reports to the board. The board doesn't need to understand the technology, but it does have to decide how much risk the company can tolerate and, when something goes wrong, judge whether "this counts as material", in other words serious enough to matter.
Companies often borrow a ready-made framework. A framework is a checklist that tells you what security work needs doing and who is responsible for it. Version 2.0 of the Cybersecurity Framework from the US National Institute of Standards and Technology (NIST) adds a new "Govern" section, which I think makes it well suited to conversations with the board. ISO/IEC 27001 is an internationally recognised security certification, while SOC 2 is a report issued after an audit by accountants. Both are used to reassure customers that "this company takes proper care of security". How the frameworks divide up the work is covered in "Which Part Each Common Framework Covers".
Where things stand: one obvious change is that third parties have moved to centre stage. Verizon's 2026 Data Breach Investigations Report found that 48% of breaches involved a third party such as a supplier or contractor, up from 30% the year before. As I see it, vetting suppliers and writing security clauses into contracts have therefore become everyday governance work.
The other pressure is time. Since December 2023, the US Securities and Exchange Commission has required listed companies to disclose a security incident publicly within 4 business days of deciding that it is "material". Four business days is not long, so it is best to write down in advance what counts as material and who makes the call.
What AI changes: the first problem is new assets that never make it into the inventory. "Shadow AI" means AI tools that employees find and use on their own without the company's approval, such as pasting meeting notes into a free chatbot. The same Verizon report says the share of employees using shadow AI has roughly tripled, to 45%. In my view, the AI tools and browser extensions you install on a whim count as shadow AI too, and they belong in the inventory as well.
I think the way we look at AI agents needs to change too. An agent is an AI program that can plan its own steps and call tools to get a task done; it holds permissions and makes its own decisions. The risk register should give agents rows of their own rather than treating them as ordinary software.
2 Regulation and Compliance
What it is: the security rules a company has to follow come from three sources: laws (break them and you are fined), standards (mostly voluntary, but often named in laws or contracts) and contracts (extra demands from customers). Compliance means being able to produce evidence that you have met all of them.
I think there have been two obvious changes over the past two years. The first is reporting, meaning telling the authorities within a deadline after something goes wrong: the deadlines keep getting shorter, and every country sets different ones. The second is that the EU has started getting serious about enforcing the rules it wrote.
Where things stand: start with the EU. The NIS2 directive requires essential and important organisations to have proper security in place. An EU "directive" only sets the direction; each member state has to pass its own law before it counts, and the deadline for doing so was 17 October 2024. This is what getting serious looks like: on 8 July 2026, the European Commission took Ireland, Spain, France and the Netherlands, which still had not finished, to the Court of Justice of the EU.
For the financial sector, the Digital Operational Resilience Act has applied since 17 January 2025; resilience here means being able to keep running after something goes wrong and to recover quickly. Europe's financial regulators have also designated 19 critical technology service providers, including AWS, Google Cloud and Microsoft, to be supervised directly, which in effect means regulating the financial institutions' suppliers themselves.
Another EU law, the Cyber Resilience Act (CRA), covers "products with digital elements", which simply means anything with software inside. Its reporting obligations came into force on 11 September 2026: once a vulnerability in a product (a flaw in the software that attackers can exploit) is being exploited, or a serious incident happens, an early warning is due within 24 hours and a formal notification within 72 hours.
In the US, the law on incident reporting for critical facilities such as water, power and hospitals has already been passed, but the detailed rules have still not appeared. According to reports, the final version was sent to the White House for review around 1 October 2026: significant incidents would have to be reported within 72 hours and ransom payments within 24 hours. The rules have not been published yet, so the details may still change.
China's amended Cybersecurity Law took effect on 1 January 2026, and the accompanying rules require "especially serious" incidents to be reported within 1 hour. Taiwan's comprehensively overhauled Cyber Security Management Act took effect on 1 December 2025, and requires incidents to be reported within 1 hour of being discovered. Penalties and other details are in "Where Taiwan Stands".
A company that falls under several sets of rules at once may have to file reports on the same incident at 1 hour, 24 hours and 72 hours, each to a different recipient and in a different format.
What AI changes: AI itself has become something that is regulated. The EU AI Act requires high-risk AI, meaning AI that could significantly affect people's safety or rights, to withstand adversarial manipulation and data poisoning. The first means deliberately crafting inputs to fool a model; the second means slipping bad data into its training data. However, the EU has already decided to postpone these obligations in two batches, to December 2027 and August 2028.
What worries me more is reporting deadlines meeting the speed of AI. Microsoft's 2026 Digital Defense Report says that the median time from a vulnerability being discovered to attackers turning it into a working attack tool is "well under 24 hours". Attacks now move on a timescale of hours. The 24- and 72-hour deadlines originally assumed that people would investigate at a steady pace; from now on, even deciding whether to report at all will depend on systems that pull the timeline and the evidence together automatically.
Small teams that aren't covered by these laws can still copy the approach: make an initial assessment and preserve the records within 1 hour, and hand in a full report within 72 hours.
3 Privacy and Data Protection
What it is: security asks "has the data been touched by someone who shouldn't touch it?"; privacy asks one more question: even with permission, is this use right? The EU's General Data Protection Regulation (GDPR) and Taiwan's Personal Data Protection Act both require a legal basis for collecting data, limits on what it is used for, the right of individuals to see and delete it, and notification when it leaks.
On the engineering side, the approach is to collect as little as possible and to de-identify what you do collect, for example by replacing names with codes. A step further are privacy-enhancing technologies, which let a system compute results from data while seeing as little of the data itself as possible. Homomorphic encryption, for instance, can compute directly on encrypted data. Apple uses it for caller ID lookups on the iPhone: the server looks up a number for you without knowing which number you asked about.
Where things stand: GDPR fines in Europe are still heavy. According to compilations by the International Association of Privacy Professionals and others, the largest in 2025 was the €530 million fine the Irish Data Protection Commission imposed on TikTok for transferring European users' data to China.
Amendments to Taiwan's Personal Data Protection Act were formally published in November 2025. They set up an independent Personal Data Protection Commission and make it mandatory to notify affected people and report to the authorities when data leaks. But publication is not the same as taking effect: as of mid-September 2026, the amendments had still not come into force. The office preparing the commission has drafted detailed accompanying rules, proposing that notification and reporting be completed within 72 hours.
What AI changes: AI conversations have become a new channel for leaks. Microsoft's 2026 Digital Defense Report mentions a malicious browser extension, installed more than 600,000 times, that was built specifically to harvest users' conversations on ChatGPT and DeepSeek. Every piece of text you paste into an AI leaves one more copy stored somewhere else.
Models themselves can leak too: attackers can try to get them to spit out data they saw during training. The Open Worldwide Application Security Project, an international community that catalogues software security risks, publishes a list of the top ten risks for AI models like ChatGPT, and in its 2026 edition "sensitive information disclosure" ranks second.
My approach is to sort data into levels and give the most sensitive data only to models running on my own computer, never sending it to the cloud. An agent's memory and conversation logs are themselves a store of personal data: they need encryption and a retention limit, and they must be deletable in bulk.
Family 2 Identity and Cryptography
This family answers who you are and what you are allowed to do, and how mathematics can guarantee that those answers haven't been overheard or tampered with. Cryptography here means the science of encryption, which is a much bigger subject than the password you type when you log in.
Companies used to rely on a wall between the internal network and the outside world to keep people out. Now employees use cloud services from home and from cafés, so there is no wall left to defend, and the focus of gatekeeping has moved to the checkpoint that "confirms who you are". Cryptography is the foundation of all this, and that foundation is being rebuilt because of quantum computers.
4 Identity and Access Management
What it is: identity and access management deals with two things. One is authentication, confirming that you are who you say you are; the other is authorisation, deciding what you are allowed to do. Besides people, it also has to manage servers, programs, and AI agents that call tools on their own to get things done.
Multi-factor authentication adds another checkpoint on top of the password, and some forms are stronger than others. Codes sent by text message are the weakest; authenticator apps (phone apps that generate changing codes themselves) and phone push notifications are in the middle; the strongest are passkeys and physical security keys that plug into the computer like a USB stick. A passkey is a digital key stored on your phone or computer that you unlock with your fingerprint or face to sign in.
A passkey is tied to the real website's domain, the part of the web address that names the site, and the browser does the checking. One attack is called adversary-in-the-middle phishing: a fake site sits between you and the real site and relays the password and code you type, in real time. Faced with a passkey, the fake site's domain doesn't match, so it gets nothing.
Another key term is token, the temporary pass a website gives you after you sign in. When you click "Allow this app to read my Google Calendar", the app also receives a token; this authorisation standard is called OAuth.
A one-time code on top of the password
A phone app generates changing codes itself
Your phone shows an alert and you type in the number on the sign-in screen
The key is tied to the real site, and the browser checks it
The pass the site issues works only on this computer
Where things stand: the focus of attacks appears to be shifting from stealing passwords to stealing the tokens and permissions that come after sign-in. Defenders are following suit: Chrome has started tying the sign-in state to the computer, so a stolen token is useless on another machine; for now this works only on Windows.
Password attacks are still the bulk of it. Microsoft's 2025 Digital Defense Report says more than 97% of identity attacks are password guessing, for example trying a few common passwords against large numbers of accounts. The same report says multi-factor authentication cuts the risk of an account being compromised by more than 99%.
But getting around multi-factor authentication has become a business. Tycoon2FA, an off-the-shelf adversary-in-the-middle phishing kit, can bypass text message codes, authenticator app codes and phone push notifications, and renting it for 10 days costs just US$120.
Even the push to adopt passkeys is being used as bait. In August 2026, Google Threat Intelligence Group revealed that a hacking group had phoned employees and told them to "urgently register a passkey" on a fake site, using the chance to steal their sign-in sessions and then carry off data from the company's cloud services. In my view, passkeys can stop phishing at sign-in, but they cannot stop people being tricked at the "registration" and "reset" steps; details are in "22 Human Factors and Social Engineering".
OAuth tokens have become the new master keys. If a company had clicked "allow the Salesloft Drift integration service to read data" in Salesforce, its customer data system, then Drift held a token. In August 2025, attackers stole a batch of these tokens and, without needing a single password, carried off data from the Salesforce accounts of more than 700 organisations; the full story is in "9 Cloud and SaaS".
Non-human identities far outnumber people: service accounts used by programs, for example, API keys (strings of characters that programs use to prove who they are to one another) and agents. The security vendor CyberArk's own count puts them at 82 times the number of humans; every vendor counts differently, but even on a conservative view there are more than ten times as many. Each one is a key that could be stolen, and someone has to keep track of them all.
What AI changes: agents are a new kind of identity and need tokens of their own. The Model Context Protocol (MCP), a shared specification that lets AI call external tools, went through three versions of its authorisation rules within a single year.
The safer design is for each layer to be smaller than the one above it: you give an agent a job, the agent asks another tool for help, and with each hand-off down the chain, what can be done should shrink a little further. It is like a manager delegating to an assistant, who then passes the task to a student temp; the temp should never have more access than the assistant. The records should also be able to answer "who actually told the agent to do this".
Attackers are using agents too. In July 2026, an attack in which several AI agents divided up the work broke into Taiwanese government systems. Of the 85 accounts compromised, 84 came through the same entrance: a single sign-on (one username and password that opens several systems) with no multi-factor authentication. The full account is in "Where Taiwan Stands".
I would give each agent its own identity rather than letting it borrow a person's account. Its tokens would be short-lived, good only for the task at hand, and would expire as soon as that task ends.
5 Zero Trust Architecture
What it is: zero trust is a premise: you are no longer trusted just because you are "on the company's internal network". Picture an office building. The old approach let you into every office once you were past the security gate on the ground floor; with zero trust you have to swipe your card again at every door, and each door also checks who you are and which computer you are using.
A 2020 document from the US National Institute of Standards and Technology (NIST) set out the architecture clearly: a central "security control room" decides whether you should be let in this time, and the card reader at each office door opens according to its decision.
The traditional way to connect back to the company from outside is a virtual private network (VPN). A VPN is like a pass to the whole building; zero trust network access gives you the key to just the one room you need.
Where things stand: perhaps the strongest argument for zero trust is that the edge devices meant to keep people out have themselves become the way in. Edge devices are machines such as firewalls and VPN servers that sit at the point where the network meets the outside world. Verizon's 2025 report found that 22% of the vulnerabilities attackers exploited were in devices like these, yet the median time to patch them was 32 days. Each vendor's record is in "7 Network and Edge Devices".
What AI changes: there is less time left for patching. Google Threat Intelligence Group says that in the first eight months of 2026, clearly more vulnerabilities were exploited each month than the year before, mainly because vulnerabilities that were already public are being turned into attack tools faster, very likely with help from AI.
In my view, if an edge device takes 32 days to patch, the door is effectively left open. Rather than racing to patch faster, it is better not to put devices' management pages directly on the internet at all, and to have employees connect through the zero trust network access described above.
Zero trust has to apply to agents as well. Every time an agent calls a tool, that is a request, and each one should be judged on its own rather than handing over all the permissions at the start of the task. The US National Security Agency's security guidance on MCP follows the same thinking: anything a tool sends back should always be treated as untrustworthy, because it may contain hidden instructions meant to fool the AI (see "Role 3 AI Itself Becomes the Attack Surface"); and every tool call should be sent to the monitoring platform.
For ordinary people, the Wi-Fi router at home and the network-attached hard drive (NAS) you keep there are your edge devices. Turn on automatic updates, and switch off the remote management feature that lets you "change settings from outside the home".
6 Cryptography and Post-Quantum Migration
What it is: a key is the string of numbers used to encrypt and decrypt. In symmetric encryption both sides use the same key, like two people each holding an identical front-door key, and it is fast.
Asymmetric encryption is like a postbox: anyone can use the slot (the public key), but only you have the key that opens the box (the private key). Common examples are RSA and elliptic curves. This is how two computers on the internet that have never met can safely agree on a shared key. It can also produce digital signatures, which prove that data really came from the person it claims to come from and hasn't been altered.
Behind the padlock in your browser's address bar is a set of encryption rules: first agree on a key just for this connection, then use a certificate to prove that the other side really is that website. A certificate is like a website's identity card, issued by a certificate authority.
Where things stand: two timelines are running at once.
The first is that certificates are being given ever shorter lives. In April 2025, certificate authorities and browser makers voted to cut the maximum validity, which used to be 398 days, to 200 days from 15 March 2026, 100 days from 15 March 2027, and just 47 days from 15 March 2029. With 47-day validity, certificates have to be replaced several times a year. Relying on people to remember makes it easy to miss one, and a single miss means the website shows a "not secure" warning. So my conclusion is that renewal has to be fully automated.
The second is post-quantum migration. Quantum computers work on completely different principles and can solve certain maths problems far faster than today's computers, and encryption such as RSA and elliptic curves happens to rest on exactly those problems. New encryption methods designed so that even quantum computers cannot break them are collectively called post-quantum cryptography; replacing the old methods with them is called post-quantum migration.
Symmetric encryption is affected far less: it is only weakened, and it stays safe with a long enough key. The real urgency is "harvest now, decrypt later": adversaries record encrypted traffic today and decrypt it once quantum computers mature. That is why the way keys are agreed has to change now. Signatures are different: forging a signature is only useful if it fools someone today, and once quantum computers mature there is no point going back to forge today's signatures, so they can be replaced a little later.
In 2024 NIST finalised its first batch of post-quantum standards, including one new method for exchanging keys and one for signatures. Google has brought forward its target for completing its own migration to 2029, on the grounds that quantum computing is progressing faster than expected; the US government also requires high-value, high-impact systems to switch their key exchange by the end of 2030.
In fact, many people are already using it. Since November 2024, Chrome has used hybrid post-quantum key exchange by default, meaning it uses a traditional method and a post-quantum one together, so things only go wrong if both are broken.
The next battle is over signatures. Post-quantum signatures are much larger: about 2,420 bytes (a unit of data size) each, against just 64 bytes for the ones in common use today. Every connection to a website involves sending several signatures, and if they are too big, connections slow down. Chrome's preferred solution is to have a whole batch of certificates share a single signature.
What AI changes: the mathematics itself hasn't changed because of AI. I think two other things have.
First, there are more things that need protecting with signatures. The AI data security guidance published in May 2025 by the US National Security Agency and partner agencies in several countries recommends signing datasets and recording their hashes (a hash is like a fingerprint for data: change a single character and the fingerprint is completely different), and using post-quantum signatures to do so.
Second, keys are starting to travel with agents. Agents need tokens and keys to get anything done, and where those keys are kept and how often they are changed are more likely to go wrong than the choice of encryption method.
I would start by building a cryptographic inventory, listing which keys and certificates I have, which encryption method each one uses, where it is kept and when it expires. Then, when the time comes to switch to post-quantum methods, I will know where to begin.
Family 3 Infrastructure
Family 3 is what most people picture first when they think of security: networks, computers and phones, the cloud, and the data itself. What needs guarding stopped being just the company's front door a long time ago; cloud accounts and backups count too.
According to several media outlets, Verizon's 2026 Data Breach Investigations Report says that for the first time in 19 years, "exploiting software vulnerabilities" (31%) has overtaken "logging in with stolen usernames and passwords" as attackers' main way in. A vulnerability is a mistake in a program's code that someone can take advantage of. The four domains below work from the outside in, ending with data and backups.
7 Network and Edge Devices
What it is: network security is about which traffic is allowed to go where, whether an attack is hiding inside that traffic, and whether services can keep running. Edge devices are the machines placed where the internet meets a company's internal network, like the main entrance and security desk of an office building. Common tools include:
- Firewalls: decide whether to let traffic through based on where it comes from and where it is going, like a security guard checking visitors against the room number they are heading for.
- VPNs and zero trust network access (ZTNA): a VPN is an encrypted tunnel that lets employees connect back to the company from outside, and once connected they have effectively been let into the whole internal network. ZTNA opens up only one application at a time. And rather than the company's systems keeping a door open and waiting for someone to knock, a small program inside the company reaches outwards to make the connection. Outsiders cannot even find where the door is, so there is no door for them to force. It is like being given the key to one meeting room instead of a pass for the whole building; the idea is explained in "5 Zero Trust Architecture".
- Distributed denial of service (DDoS) protection: a DDoS attack takes control of a large number of compromised devices and has them all flood one website with traffic at the same time. It is like a huge crowd ringing a customer service line at once, so real customers cannot get through. According to reports, the largest attack recorded by the internet services company Cloudflare in the fourth quarter of 2025 reached about 31.4 Tbps, roughly the same as more than 300,000 homes running 100 Mbps broadband flat out at the same time. Most of the devices launching it were infected Android TVs, the kind found in ordinary living rooms.
Where things stand: the most obvious change of the past two years is that edge devices, which were meant to block attacks, have themselves become one of the main ways in. In Verizon's 2025 report, the share of exploited vulnerabilities found in edge devices and VPNs rose from 3% to 22%; the same report found that these vulnerabilities took a median of 32 days to fix.
The US Cybersecurity and Infrastructure Security Agency keeps a "Known Exploited Vulnerabilities" catalogue, which only includes vulnerabilities that attackers are confirmed to have used. Using the version of 2 October 2026, I counted ten vendors whose products are mainly edge devices, such as Cisco and Fortinet. With 2026 only into early October, more of their vulnerabilities had already been newly added to the list (49) than in the whole of 2025 (39).
In September 2025, the agency issued an emergency directive over zero-day vulnerabilities in Cisco firewalls, ordering US federal agencies to deal with them by a deadline. A zero-day is a vulnerability that is already being used in attacks before the vendor knows about it or has fixed it. According to reports, the attackers tampered with the device's lowest-level start-up code, its firmware (the software built into a device that controls its hardware), so neither rebooting nor upgrading could get rid of them. Home-grade equipment is a target too: Volt Typhoon, a hacking group Microsoft described in 2023, used hijacked home and small-office routers as stepping stones.
My view is that edge devices are especially hard to defend. They face the internet directly and usually cannot run detection software; malicious code hidden in their firmware survives a reboot; and if one goes down, the whole company goes offline, so you cannot simply update it whenever you like.
What AI changes: a report from the Google Threat Intelligence Group (GTIG) on 30 September says that from January to August 2026, an average of 18 vulnerabilities a month were exploited in attacks, against 10.5 a month in 2025. Most of the extra ones were not zero-days. GTIG believes the increase comes mainly from n-days, vulnerabilities that have already been made public and usually already have a fix, being turned quickly into attack tools, very likely with help from AI.
One way AI can help is by using a large language model (the kind of model behind ChatGPT) to compare a program's code before and after a fix. It is like reading the old and new versions of a contract side by side, clause by clause: once you see which clause changed, you know where the original loophole was. Another example is a BeyondTrust product that lets IT staff operate computers remotely. One of its vulnerabilities was found by an AI research agent; an agent is an AI program that can plan its own steps and use tools. Just 4 days after the vulnerability was made public, an attack group was using it, and within 7 days there were 6 groups.
The 2026 edition of Microsoft's Digital Defense Report says the median time from a vulnerability first being spotted out in the world to it being turned into an attack tool is "well under 24 hours". As I see it, the two sides run on very different clocks: attackers count in hours, while the median time to patch an edge device is 32 days, and a downtime window still has to be scheduled.
From what I could find, there are few public measurements so far of how much AI helps with network defence. I think the surer fix actually does not rely on AI: if you cannot patch in time, keep the device from being visible on the internet at all, which is exactly what ZTNA's "outbound only" design does.
8 Endpoints and Mobile Devices
What it is: endpoints are the machines that actually run programs, such as laptops, servers and phones. Endpoint security does three things: it sets machines up so they are hard to attack, spots malicious behaviour, and deals with it on the spot.
The core tool is endpoint detection and response (EDR). It keeps a continuous record of which files each program opens and where it connects to, judges from that behaviour whether something is wrong, and if necessary cuts the whole computer off from the network. Think of it as a dashcam installed in the computer, plus a security guard who can slam on the brakes.
Where things stand: a single crash got Microsoft to start reworking how Windows is built. To see everything happening on a computer, the protection software from the security company CrowdStrike runs in the Windows kernel, the lowest level of the operating system and the one with the most privileges. In July 2024, CrowdStrike pushed out a faulty update, and huge numbers of Windows computers could not start up as a result. It was not an attack, but it showed everyone the price of security software running in the kernel.
According to reports, Microsoft announced the Windows Resiliency Initiative in November 2024. One part of it lets antivirus and EDR software run outside the kernel, and from July 2025 some partner vendors have been able to try this out first. Public information does not yet show whether it has been fully launched. Security vendors, for their part, worry that once their software moves out of the kernel, it will see less and be easier for malware to switch off.
On phones, the adversaries are often well funded. One example is commercial spyware: software that companies develop and sell to clients to monitor particular people's phones. According to reports, Apple's Memory Integrity Enforcement feature, announced in September 2025, is always switched on in the iPhone 17 and iPhone Air. It uses hardware to check whether memory is being written where it should not be, and it is aimed squarely at this kind of spyware's exploit chains: attack code that links several vulnerabilities together to take over a phone step by step. GTIG says that among the zero-days of 2025 whose users could be identified, commercial spyware vendors were behind more of them than state-level spy groups for the first time.
According to press coverage of CrowdStrike's 2026 report, 82% of the attacks it detected used no malware at all; attackers used legitimate tools already on the system and stolen accounts instead. A key player here is the infostealer, a program built to steal the passwords and login sessions saved in a browser in one go. In the 2024 incident at the cloud data platform Snowflake (covered in more detail in the next section), the usernames and passwords used to log in came from several different infostealers.
Microsoft's 2026 Digital Defense Report found that in 30% of cases, the way attackers first got in was something users ran themselves. Take ClickFix: a web page pops up some normal-looking instructions telling you to copy a command and paste it into your computer to run, and in doing so you are really running the attacker's command for them. From February to May 2026, this trick was run on more than 1.1 million devices.
Once you log in, a website leaves a login pass (a cookie) in your browser so you do not have to type your password again. One new defence is to make that pass unable to leave the original computer. According to reports, Chrome 146 has officially launched this feature on Windows, tying the pass to a key held in the computer's security chip, so a stolen pass will not work on a different machine. The chart below puts the whole path together.
The user runs a disguised program, or pastes in a command following a web page's instructions
Passwords and login passes (cookies) saved in the browser are sent off in one bundle
The attacker logs in to cloud services from their own computer, with no malware needed at any point
Data is pulled out in bulk and then used for extortion
What AI changes: a new kind of malware waits until it is running to consult a large language model. GTIG has been publishing findings on programs like this since November 2025. One example is PROMPTFLUX, which asks Google's model to rewrite its own code. GTIG describes most of them as experimental, though, and the security company Mandiant also stresses that the intrusions of 2025 were not caused by AI; see "Role 1 AI Helps Attackers" for details.
I think these programs actually leave behind a useful clue. They all have to connect to some model's API, the interface that programs use to call on one another. A program that has no business using AI calling an AI service is suspicious in itself.
On the defence side, AI is starting to work as a reverse engineer, taking malware apart to understand what it does. According to Microsoft's own tests, Project Ire, which Microsoft unveiled in August 2025, had a recall (the share of real malware it catches) of 0.83 on a public dataset. On hard-to-judge real-world files, that fell to just 0.26, catching about one in four; see "Role 2 AI Helps Defenders".
AI agents that write code for people run on developers' computers. They can read files and run commands, and they install packages on their own; packages are ready-made building blocks of code that developers share. On 31 March 2026, the axios package on npm, a public package registry, was hijacked and replaced with a malicious version that planted a remote-control program on any computer that installed it. That version was up for only about 3 hours, and both Microsoft and GTIG believe North Korea was behind it. This technique, where attackers do not hit you directly but poison the parts you are going to install, is called a supply chain attack. In principle, any agent that installed packages automatically during that window would have been caught out just like a human.
9 Cloud and SaaS
What it is: cloud security rests on "shared responsibility". The cloud provider looks after the data centres, the hardware and the underlying systems; the customer looks after accounts, settings, data and code. It is a bit like renting a flat: the landlord takes care of the building and the main entrance, and you take care of whether your own door is locked and who has the spare keys. Most incidents happen on the customer's side, for example settings configured wrongly, permissions that are too broad, or leaked keys (a key here is a string of characters, like a password, that lets a program log in to a cloud service).
SaaS is software you rent directly over the internet, such as Salesforce or Microsoft 365. What is left in the customer's hands is mainly "who can log in" and "which third-party apps have been given access". A safer approach is not to write keys directly into code, but to hand them to a dedicated safekeeping service that lends out only a temporary key, one that expires quickly, each time.
Where things stand: in 2024, attackers took leaked usernames and passwords, logged straight in to companies' accounts on Snowflake and pulled out data in bulk. After Mandiant's investigation, about 165 organisations that may have been affected were notified; for at least 79.7% of the accounts used to log in, the credentials had leaked long before. The causes were covered in "Principles That Run Through This Post": multi-factor authentication (a second check on top of the password, such as confirming on your phone) was not switched on, passwords were not changed regularly, and there were no limits on where logins could come from.
Another route is the access that SaaS apps grant one another. When you click "allow access to my account" in an app, the other side gets a token, a temporary pass that lets it in on your behalf. In August 2025, attackers used tokens stolen from Salesloft Drift, an integration service, to export customers' Salesforce data in bulk, then searched that data for keys and passwords to other services. According to several media outlets, more than 700 organisations were affected.
The 2026 edition of Microsoft's Digital Defense Report says that cloud workloads (programs and servers running in the cloud) exposed to the internet are attacked within 5.3 hours on average.
What AI changes: getting from a break-in to large-scale harvesting now takes anything from a few hours to a day or two. GTIG's report of 8 September describes an attacker in the second quarter of 2026 who used a ready-assembled set of AI agent tools, starting from one cloud resource that had already been compromised, to "plan, build and execute" an entire large-scale operation to collect login details, all in under 6 hours.
Anthropic, the company that makes Claude, also recorded in its September report on hackers misusing AI that someone linked to the hacking group ShinyHunters had an AI agent "do almost all the work", collecting login tokens for more than 2,100 Microsoft business accounts in about 34 hours.
What worries me more is that one breach now snowballs into the next more easily. The Drift attackers searched the stolen data for the next key; with agents, that drudgery of digging through data costs almost nothing.
AI services also bring more keys of their own that need safekeeping. According to reports, the security company GitGuardian's 2026 report found that among the keys and passwords leaked publicly on the code-sharing site GitHub in 2025, those belonging to AI services rose by 81%.
Attackers are already fast enough to be measured in hours. Rather than trying to outrun them, I think it is more effective to make whatever they steal useless: give tokens short lifetimes, tie logins to devices, and allow connections only from approved networks.
10 Data Security and Backup Resilience
What it is: data security has to answer three questions: where the sensitive data is, who can get at it, and whether anyone has seen it who should not, or tampered with it, at any point from its creation to its deletion.
- Data loss prevention: first label data by how sensitive it is, then inspect content on computers, in email and in the cloud, blocking it or raising an alert when sensitive data is about to be sent out. It is now also starting to inspect prompts sent to generative AI, meaning the text you type to the AI.
- Encryption: stored data is often protected like this: each piece of data is locked with its own key, all those keys are then locked in one master safe, and the safe is handed to a dedicated key management service to look after.
- Backups: the usual rule is 3-2-1, meaning three copies, on two kinds of storage, with one kept off-site or offline, plus immutable backups that nobody can change for a set period after they are written. A backup you have never actually restored does not count.
Where things stand: ransomware originally worked by encrypting a company's files to lock them up, and handing over the key to unlock them only after payment. The focus is now shifting to "stealing the data" and "making sure you cannot recover". According to reports, in the security company Sophos's 2026 survey, only 56% of extortion attacks actually encrypted data, so nearly half did not lock anything at all. UNC6671, a criminal group tracked by GTIG, steals without locking; how that played out is in "Who Attacks and How They Get In".
Another trend is "recovery denial". To me the logic is simple: if you can restore your data, you do not need to pay, so attackers first make sure you cannot. Mandiant's M-Trends 2026 says ransomware gangs now go first for backups, virtualisation hosts (the layer that runs many virtual computers at once on one physical server) and the identity systems that manage accounts across the whole company. Once attackers take over the internal certificate authority, the part of the identity system that issues "electronic ID badges", they can log in as anyone.
Ransoms are shrinking, but recovery is getting more expensive. According to several media outlets, the blockchain analysis company Chainalysis estimates ransomware payments in 2025 at about US$820 million, down about 8%; ransoms are mostly paid in cryptocurrency, so they can be estimated from transactions on the blockchain. Also according to media reports, Sophos's survey found that recovery alone, not counting the ransom, cost an average of US$1.7 million, against US$1.53 million the year before. The chart below shows the usual steps in this kind of attack.
Using stolen accounts, exploited edge devices, or a call to the company help desk pretending to be an employee
Control the account system and the internal certificate authority, and log in as anyone
Delete or encrypt backups, and lock every virtual computer at once from the virtualisation host
Threaten to publish the stolen data, sometimes without encrypting anything at all
What AI changes: when employees paste company documents into AI tools they have found for themselves and the company has not approved (known as shadow AI), the data walks straight out of the company. According to several media outlets, Verizon's 2026 report says the share of employees using shadow AI has roughly tripled, to 45%.
AI chat histories are themselves worth stealing. "3 Privacy and Data Protection" mentioned a malicious browser extension (a small add-on for browsers such as Chrome), installed more than 600,000 times, that was built to harvest ChatGPT and DeepSeek chat histories. AI agents can open websites and upload files on their own, so they too may send data where it should not go.
Anthropic's September report also records one group that targeted more than 20 organisations and stole over 300,000 national identity records. My view is that when AI can comb through stolen data in a few hours and pick out the parts usable for extortion, "steal without locking" becomes even cheaper, and backups cannot save you from that kind of extortion. So I recommend keeping less data: regularly delete chat histories and temporary files you do not need, and encrypt everything you do keep.
Defenders have new approaches too. Apple has published the code of the server system it uses to run AI in the cloud (Private Cloud Compute), so outside researchers can check whether it really keeps its privacy promises.
Family 4 Software and AI Systems
This family is about software itself: the code you write yourself, the packages you bring in from outside (ready-made pieces of code written by other people that you can use straight away), the vulnerabilities that surface every day after release, and the newest additions, AI models and agents (AI programs that break a task into steps on their own and call tools to carry it out). The four domains can be seen as one production line: first you write the code, then you assemble parts made by other people, and after release you keep patching holes. On this line, AI is both a new tool and a new part.
11 Application Security
What it is: application security deals with whether "the software itself has holes". It isn't something you scan once before launch and then forget; it's a habit that has to be there at every stage of writing code. The US National Institute of Standards and Technology (NIST) has a Secure Software Development Framework that is often cited as the baseline.
A few things you run into day to day:
- Threat modelling: before you start writing, draw the system out and ask, one part at a time, "how could this be attacked?" It's a bit like walking round a house before you move in, checking which windows a burglar might climb through.
- Automated checks: some tools read the code directly, looking for dangerous ways of writing it; others wait until the program is running and then simulate attacks from outside. There are also tools that check specifically whether packages written by other people have known vulnerabilities, or whether someone has accidentally written a password into the code.
- Memory safety: languages such as C and C++ make programmers manage memory themselves, and one slip can write data where it shouldn't go. Mistakes like this are a way in for attackers. "Memory-safe languages" such as Rust block most of these mistakes through the design of the language itself.
- API security: most apps today are pieced together from many APIs (the windows through which programs call on one another). The classic vulnerability is "change the order number in the web address and you can see someone else's order". OWASP (an open community for application security) keeps a list of the top ten API risks; it is still the 2023 edition, and this kind of problem is number one.
The list the industry refers to most widely, though, is OWASP's Top 10, whose 2025 edition is the 8th. Number one is still "Broken Access Control", meaning the system doesn't properly control who can see or change what. Number three is the new "Software Supply Chain Failures" (covered in the next section). It turns up least often in the real-world data, but on average it scores highest on both "how easy it is to exploit" and "how much damage it does"; in the community survey, too, it was what most people worried about.
Number ten, "Mishandling of Exceptional Conditions", is also new. It asks whether a program that hits a timeout or badly formatted data stops safely or simply lets things through. Think of a door entry system that crashes: the door should stay locked, not swing open on its own.
Where things stand: most organisations are carrying holes that can be exploited. The security vendor Datadog, using data from its own customers, found that 87% of organisations have at least one known exploitable vulnerability. On top of that, the outside packages that programs use lag behind their latest major version by a median of 278 days. Falling behind on versions can mean missing holes that a newer version has already fixed.
Memory-safe languages are also making their way into the deepest layer of the operating system. The official documentation for the Linux kernel (the foundation underneath many servers and Android phones) has dropped its section on the "Rust experiment", which amounts to no longer treating Rust as an experiment.
Rust isn't a cure-all, though. CVE-2025-68260, disclosed on 16 December 2025 (a CVE is the standard ID number given to a publicly known vulnerability, rather like a person's ID number), was found in an unsafe block of the kernel's Rust code. That is the part of Rust where some of the language's checks are allowed to be bypassed for a while, and mistakes happen there just the same.
What AI changes: defenders are now using AI to review code at scale. Anthropic itself reports that its AI model Opus 4.6 found more than 500 vulnerabilities in open-source projects in real use (software whose code is public, so anyone can read and change it).
But careless scanning just produces junk. Mantis, a bug-hunting tool Google released as open source in September, specifically warns that sloppy AI scanning produces "hallucinated bugs": of the problems such scans report, fewer than 7% are real vulnerabilities. By contrast, Anthropic reports that in Project Glasswing, its programme for hunting bugs with AI together with partners, 90.6% of 1,752 findings that were each checked by a person turned out to be real.
The holes AI finds also tend to be serious. By the count of the Google Threat Intelligence Group (GTIG), 50% of the vulnerabilities AI discovers are "remote code execution", meaning an attacker can run programs on your machine from across the network; across all CVEs the average is only 26%.
Code written by AI should be treated as if an outsider wrote it. Datadog's report says outright that the output of vibe coding (letting AI write code on gut feel without looking at it closely yourself) "should not be implicitly trusted".
AI development tools themselves get attacked too. Take CVE-2026-47751: an attacker submits a request to change the code (a pull request) and slips a malicious MCP configuration into it. MCP is a common standard for connecting AI to outside tools, a bit like a universal plug adapter. When Anthropic's AI tool claude-code-action reads this configuration, it runs the attacker's code and leaks secrets as well.
12 Software Supply Chain
What it is: a supply chain attack doesn't hit you directly; it poisons software you are going to install. It's like leaving the restaurant's door alone and tampering with the ingredients at the supplier instead. On its way from the developer's keyboard to going live, software passes through several stops: the websites where code is stored; public package download sites such as npm and PyPI; the process that automatically packages and publishes the program every time the code is changed; developers' own computers; and now AI agents as well. Any of these stops can be tampered with.
This isn't a new problem. In 2020, the system the software company SolarWinds used to package its software was tampered with, and the official updates customers received had malicious code hidden inside, complete with a genuine digital signature (like a manufacturer's seal of approval). Log4Shell (CVE-2021-44228), a vulnerability in a package disclosed in December 2021, was added the same day to the Known Exploited Vulnerabilities catalogue (KEV) of the US Cybersecurity and Infrastructure Security Agency (CISA), the list of vulnerabilities "confirmed to be already used in attacks". It became the textbook example of why you need to know which packages you use.
The main defensive tools:
- Software bill of materials (SBOM): a software "ingredients list" showing which packages and which versions are used. When the next Log4Shell appears, having one lets you find out quickly whether "I'm using it".
- No long-lived tokens: a token can be thought of as the pass a website gives you after you log in. If a token that stays valid for a long time is stolen, an attacker can publish new versions in your name whenever they like. With "trusted publishing", a one-time permission is issued only at the moment of each release, so no long-lived pass is kept around.
- Pinning versions: GitHub is a website many people use to store code and work on it together, and its automated workflows often name a tool's version with a tag such as "version 4". But a tag can be re-pointed to a different piece of code, like a label on a bookshelf being moved onto another book. Pinning the full hash, which amounts to pinning the code's fingerprint, is what really fixes the version.
- Cooldown period: wait a few days after a new version is released before using it, giving malicious versions time to be spotted and taken down.
Where things stand: my view is that attackers now spend less time looking for a package that happens to have a hole, and instead go straight for the accounts and automated pipelines of maintainers (the people who manage and update a particular package). GTIG's analysis from July 2026 also points out that most high-impact incidents from 2025 to the first half of 2026 involved open-source packages and development tools being compromised. According to reports, Verizon's 2026 Data Breach Investigations Report found that breaches involving a third party (such as an outsourcing supplier or business partner) made up 48%, up from 30% the year before.
The most telling part of the chart is the chain of events in March 2026. The first to be hit was the security scanner Trivy: using stolen credentials (things that prove who you are, such as passwords and tokens), attackers published a malicious v0.69.4, and quietly re-pointed 76 of the 77 version tags in its GitHub automation component to a program that steals data. Trivy's v0.69.3 was left untouched because it used "immutable releases" (which cannot be changed once published).
Next, the attackers used tokens leaked in the Trivy incident to publish a malicious version of LiteLLM (a package that forwards requests to many AI models through a single entry point), which was installed about 119,000 times within 2 hours 32 minutes. PyPI's incident report says that about 40% to 50% of LiteLLM installs download the latest version directly. If you haven't pinned the version, you automatically accept whatever version the attacker publishes.
The number of malicious packages has also exploded. Counting the Open Source Security Foundation's malicious packages database by year, there were 11,811 in 2024 and 192,896 in 2025. npm alone accounted for 191,304 of those, mostly junk packages uploaded in bulk and floods uploaded to rack up token rewards (token farming).
The progress this year has come from package registries changing their defaults directly. npm 12.0.0, released in July 2026, blocks by default the "install scripts" of dependencies (the other packages a project uses), that is, code that runs automatically as soon as a package is installed. Since July 2026, PyPI no longer allows new files to be added to releases more than 14 days old, so attackers can't slip anything into old versions; the change was prompted precisely by LiteLLM and another poisoned package, telnyx. In Taiwan, the Regulations for Reviewing Products That Endanger National Cyber Security, in force since 1 December 2025, also use bills of materials or SBOMs to review products from brands other than mainland Chinese ones.
What AI changes: AI imagines package names that don't exist. According to reports, a study at USENIX Security 2025, an academic security conference, tested a set of code-writing AI models and found that 19.7% of the packages they recommended didn't exist at all. Worse, these fake names are consistent: 43% of them came up every time the same question was asked again. All an attacker has to do is register those names first, put malicious code in them, and wait for the AI to tell people to install them.
AI agents can also add malicious packages themselves. GTIG recorded a North Korea-linked malicious package that was added to a real cryptocurrency project as part of a code change; the co-author named on that change was an AI agent that writes code for people.
AI-related software has itself become a target. GTIG counts 2,076 CVEs related to AI software since January 2025, more than 1,500 of them in 2026. I think "gateways" like LiteLLM are especially dangerous: they forward requests to many models and also hold the keys used to call those models (passwords that let a program use a paid service). Once one is compromised, every service that goes through it, and every key it holds, is affected.
Models themselves are a link in the supply chain too. Downloaded model files, MCP servers that connect AI to outside tools, and "skill packs" that give agents new abilities are all things brought in from outside that end up being run, and all of them already have real cases of abuse.
13 Vulnerability Management
What it is: vulnerability management deals with "holes that are already public". It's a bit like dealing with car recall notices: new notices arrive every day, and you first need to know which car you drive, then decide which notices concern you and which need dealing with straight away. It runs on a few public systems:
- CVE: the vulnerability ID mentioned earlier, like the reference number on each recall notice. It is run by MITRE (a non-profit organisation that does research for the US government) under contract to CISA.
- Common Vulnerability Scoring System (CVSS): a severity rating from 0 to 10.
- Exploit Prediction Scoring System (EPSS): like a chance-of-rain forecast, it is recalculated every day and estimates how likely a hole is to be used in attacks in the near future.
- KEV: the list of vulnerabilities "confirmed to be already used in attacks" mentioned earlier, with 1,733 entries as of 2 October 2026.
Where things stand: the number of vulnerabilities is rising fast. According to reports, there were 48,185 CVEs in 2025, 20.6% more than the year before. 2026 is faster still:
The rise isn't all down to AI; it also depends on how the IDs are handed out. The Linux kernel team can assign CVE IDs itself, and between January and August 2026 alone it assigned about 5,000, not one of which was a "zero-day", a vulnerability that is used in attacks before it has been patched.
Very few vulnerabilities are actually used in attacks. By GTIG's count, only about 1 in 431 vulnerabilities disclosed in 2026 has been exploited. But the number being exploited is growing, from an average of 10.5 a month in 2025 to 18 a month from January to August 2026. Attacks are also coming sooner: Mandiant, the security company owned by Google, worked out in M-Trends 2026 that on average attacks begin about 7 days before the fix is even released.
Patching doesn't seem to be keeping up either. According to reports, Verizon's 2026 report says the median time to patch went from 32 days to 43 days. However, the previous year's 32 days counted only devices sitting at the front door of a company's network and facing the internet directly, such as firewalls and VPNs, so the two years may not be measuring the same thing.
In practice, you can stack four layers to decide what to fix first:
- KEV: anything already confirmed as exploited comes first.
- EPSS: the probability of exploitation in the near future. Datadog's approach is to downgrade "critical" when EPSS is below 1%.
- Reachability: whether you actually use the vulnerable piece of code at all. By Datadog's count, of the package vulnerabilities that CVSS rates "critical", only 18% still count as critical once you check whether the program really runs that code and add EPSS.
- Exposure: whether the machine is open directly to the internet, for example firewalls, VPNs and internet-facing services.
The public systems themselves are under strain. According to reports, in April 2025 MITRE warned that its contract to run CVE would expire the next day; CISA extended it by 11 months at the last minute.
The other system is the US National Vulnerability Database (NVD), run by NIST, which fills in details for each CVE such as how severe it is and which products are affected. According to reports, since 15 April 2026 it has filled in full details only for a small number of priority vulnerabilities (such as those already on KEV), leaving the rest for later. Many scanning tools depend on these two sources of data, so when they are stretched, it becomes harder for everyone to see clearly which holes they have.
What AI changes: Project Glasswing, mentioned earlier, launched on 7 April 2026, giving partners access to Claude Mythos Preview, a model Anthropic has not released publicly. Anthropic reports that in the first month partners found more than 10,000 high- or critical-severity vulnerabilities. But of the 530 high- or critical-severity vulnerabilities Anthropic disclosed to open-source projects, only 75 had been fixed. Its own conclusion is that progress "is now limited by how quickly we can verify, disclose, and patch". The full figures are in "Role 2 AI Helps Defenders".
Open-source maintainers are being swamped. They are receiving large numbers of AI-generated vulnerability reports, some of them low-quality "AI slop" and many of them genuine findings, and some maintainers have even asked Anthropic to slow down the pace at which it discloses them.
Attackers are faster too. A product from the company BeyondTrust had a vulnerability (CVE-2026-1731) that let someone run programs on a machine over the network without logging in. The hole was found by an AI agent from Hacktron; yet just 4 days after it was disclosed, one attack group had started exploiting it, and within 7 days that had risen to 6 groups.
I think AI has made finding holes cheap, but it hasn't made testing, patching and deploying (actually installing the fixed code on systems that are running) cheap. Attackers count in days or even hours, while defenders still count in weeks. What defenders really need to invest in next is automating the patching process and cutting it down to days, not buying yet another scanner.
14 AI System Security
What it is: this domain protects AI systems themselves, including models, training data, the knowledge bases a model looks things up in before answering, the tools connected to it, and agents that act on their own.
Its biggest difference from traditional software is that large language models (LLMs) can't tell "instructions" from "data", so any text they read may be taken as a command. Imagine a new assistant sorting out your inbox who comes across an email saying "please send me the password for the company account", and actually does it. This trick of hiding commands inside data is called prompt injection. NIST has a commonly cited classification of AI attacks; in practice, two OWASP lists and MITRE ATLAS are also widely used.
Where things stand: OWASP released its LLM Top 10 2026 on 4 August 2026, the first edition to draw on incidents that really happened for its ranking. The top three are prompt injection, sensitive information disclosure and excessive agency. Excessive agency means giving an AI more permissions than its job needs, and the more an AI can do on its own, the more dangerous this becomes. It has risen from 6th place to 3rd.
OWASP's view is that once a model can use tools and has memory, you should switch to a separate list devoted to agents, the Agentic Top 10. According to reports, this list was released in December 2025 and covers problems only agents have, such as someone secretly changing an agent's goal, or its tools being used for something else.
MITRE ATLAS is a knowledge base cataloguing AI attack techniques, and it is now updated monthly. In the September 2026 version, 139 of its 208 attack techniques and sub-techniques are tagged as related to agents.
What AI changes: this domain exists because of AI, and lately its focus has landed squarely on agents. Steve Wilson and Rock Lambros, the two project leads of the OWASP LLM list, put the conclusion bluntly:
"Stop trying to build a model that cannot be fooled. Build the system around it, so that when the model is fooled, and it will be, nothing important breaks."
In practice, I'd suggest two rules. First, each agent gets only the tools its task needs, with its permissions written into the code, rather than written into the prompt (the text instructions given to the AI) in the hope that it will behave itself. Second, imagine an AI assistant that reads an email from a stranger (untrusted input), can also see the company's account details (sensitive data), and can send emails out on its own (an external action); when all three come together, it should stop and let a person confirm. The details of attack techniques and defensive architecture are in "Role 3 AI Itself Becomes the Attack Surface".
Family 5 Detection and Response
The earlier families were mostly about keeping attackers out. This one is about what to do when that fails: spotting attackers while they are still inside your systems, driving them out, and then piecing together from the records what actually happened.
The thinking behind it is called "assume breach": you take it for granted that sooner or later someone will get through, and then aim to notice them and clean up faster than they can act.
15 Security Operations Centre
What it is: a security operations centre (SOC) is the team and platform within an organisation whose job is to keep watch. Computers, servers and cloud services leave behind huge numbers of records every day, a bit like the log of key-card swipes at a building's doors; software then applies rules to pick out anything suspicious in those records and raises an alert. The SOC's job is to sift through the alerts, pick out the few that are real incidents and deal with them.
Traditionally a SOC has three tiers: tier 1 does the first sift, tier 2 investigates in depth, and tier 3 goes hunting, which means not waiting for an alarm to go off but searching through the records for intruders who may have slipped past.
One of the SOC's standard tools is security information and event management, which gathers records from everywhere into one place and cross-checks them. In recent years it has become popular to manage detection rules the way programmers manage code: every change is saved as a new version and tested automatically, so if a change breaks something, it is spotted straight away and the old version can be put back.
Where things stand: the most common complaint about SOCs is alert fatigue: too many alerts, and too few of them about anything real. Surveys give widely differing figures, but the rough consensus is that about half of all alerts are false alarms, and many are never looked at at all.
Defenders are not losing on every front, though. M-Trends 2026, from the security company Mandiant, says that in 2025, 52% of intrusions were first discovered by the organisations themselves, rather than learned about from an outsider's tip-off or a hacker's ransom note; the year before, the figure was 43%.
What AI changes: an AI agent is an AI program that can look things up by itself and work through a task step by step. The big selling point for SOCs in 2025 and 2026 has been the "agentic SOC", which means handing the tier 1 sifting to AI agents. In 2025 Microsoft and then Google announced agents of this kind, which sift through phishing emails and investigate alerts one by one. Nearly all the figures on how well they work come from the vendors' own studies; Microsoft, for example, says the average time to resolve an incident fell by 30%. Details are in "Role 2 AI Helps Defenders".
AI brings new risks too. The Google Threat Intelligence Group (GTIG) has recorded malware built specifically to take on AI: hidden inside the program is a passage written for an AI to read, meant to fool defensive tools that use large language models (the kind of model behind ChatGPT) to scan files. This trick is called prompt injection: slipping instructions into data to trick an AI into following them.
Asking AI to check code for vulnerabilities is not very reliable either. In September 2026 Google released an open-source tool called Mantis (open source means the code is public and anyone can use it), and its documentation warns plainly that if you carelessly ask an AI to scan code, it will "hallucinate" vulnerabilities that do not exist, making them up with a perfectly straight face. By its account, fewer than 7% of the problems such scans report are real vulnerabilities.
16 Incident Response and Forensics
What it is: incident response is the standard procedure for after something has gone wrong, like a building's fire plan. Digital forensics is about working out exactly what happened, and preserving the records in a form that can serve as evidence.
In 2025 the US National Institute of Standards and Technology revised its incident response guidance for the first time in more than a decade. The message of the new version is that responding to incidents should not be a standalone manual locked away in a drawer, but should be tied into the way the organisation manages risk day to day.
The basics of forensics are gathering evidence, building a timeline and keeping the "chain of custody": a record of every handover that proves nobody has tampered with the evidence from start to finish.
In quieter times, organisations rely on tabletop exercises: managers, lawyers, communications staff and IT people sit around a table and rehearse in advance decisions such as whether to pay a ransom and whether to go public.
Where things stand: most victims no longer pay. The 2026 edition of the Data Breach Investigations Report from the US telecoms company Verizon says that 69% of ransomware victims did not pay the ransom.
But the time it takes to get from being broken into to finding and controlling the problem has grown. IBM's 2026 Cost of a Data Breach Report puts the average for "identifying plus containing" a breach at 247 days, slower than the year before, which ended five consecutive years of improvement. Containing means cutting off the affected computers so the damage does not spread.
Mandiant has also observed that ransomware gangs, as well as encrypting files, have started going straight for the very things you would need in order to recover: backups, virtualisation platforms (the layer that lets one physical machine run many virtual computers), and the identity systems that manage account logins.
Deadlines for reporting incidents are now written into law as well. Taiwan's Cyber Security Incident Reporting, Response and Drill Regulations, amended in January 2026, require the organisations they cover to report an incident within one hour of becoming aware of it. From September 2026, the EU's Cyber Resilience Act requires a brief early warning within 24 hours, followed by a formal notification within 72 hours.
What AI changes: AI is starting to reverse-engineer malware on its own, which means taking a program apart to see what it really does. Project Ire, which Microsoft unveiled in August 2025, can decide automatically whether a file is malicious. In Microsoft's own tests on a public set of test files, 98% of the files it flagged as malicious really were malicious, and it caught 83% of the genuinely malicious ones. On hard real-world files, though, it recognised only 26% of the genuinely malicious files and missed most of them.
AI may also make forensics harder. Suppose a company lets an AI agent handle its accounts automatically, and someone tricks it into sending money to the wrong place. In my view, the record of what the agent did is then the evidence, but agents can delete their own tracks and even fake conversation records, so the records must be written somewhere the agent cannot touch. It is like not letting a suspect look after the CCTV footage.
17 Threat Intelligence
What it is: threat intelligence means working out who your adversaries are, what methods they use and whom they target, and then turning that into information people can act on as soon as they receive it. Managers get trends; front-line staff get concrete signs of intrusion, such as a malicious web address or a suspicious file.
The reference book everyone shares is ATT&CK, a knowledge base of attack methods compiled by MITRE, a US non-profit research organisation. It calls what an attacker wants to achieve a "tactic" and the way they achieve it a "technique", and gives each one a number. It is a bit like hospitals using a standard set of disease codes: when different teams talk about the same kind of attack, they are not talking at cross purposes.
Where things stand: I think the easiest trap to fall into when reading threat intelligence is the muddle over names, and as AI churns out more and more reports, matching up the same adversary across them will only get harder. The same group of hackers often goes by different names in different companies' reports. In March 2026 a North Korean team attacked axios, an open-source package. A package is a ready-made building block of code that someone else has written and many programs simply reuse; attacking one usually means slipping malicious code into it so that everyone who installs it is infected too. Microsoft calls this team Sapphire Sleet; GTIG calls it MIDNIGHT NEPTUNE.
Another practical source of intelligence is the Known Exploited Vulnerabilities (KEV) catalogue kept by the US Cybersecurity and Infrastructure Security Agency. It looks at whether a vulnerability has actually been used in attacks, not just at how severe it is, and many organisations use it to decide what to patch first. By early October 2026 it had already added 249 entries for the year, more than in the whole of 2025.
What AI changes: threat intelligence has gained a new category: who is using AI to do harm. GTIG and the AI company Anthropic both publish reports tracking how attackers use large language models. In August 2025, one person used an AI coding tool to extort at least 17 organisations. That November, Anthropic revealed that a group it was highly confident was backed by the Chinese state was also using AI to launch attacks. How things escalated step by step after that is covered in "Role 1 AI Helps Attackers".
Family 6 Offensive Security
Offensive security means attacking yourself with the bad guys' methods before the real bad guys get the chance, so you can find the holes and fix them. The only thing that separates it from crime is whether you have written permission.
In my view, the biggest change in 2025 and 2026 is that AI has dramatically lowered the bar for finding vulnerabilities and writing attacks.
18 Penetration Testing and Red Teaming
What it is: attack testing comes in several kinds, from shallow to deep. A vulnerability assessment is a broad scan that lists every known weakness. A penetration test tries, within a set time and scope, to find as many holes as possible that can really be broken through, and the defenders know it is happening. A red team works towards a particular goal and operates in secret; what it is really testing is "can your detection and response catch me?"
Going further, a purple team has attackers and defenders sitting down together to fine-tune detection rules, while adversary emulation acts out, from start to finish, the methods of a particular known hacking group. A commonly used open-source tool is MITRE's CALDERA.
A broad scan that lists every known weakness
Looks for holes that can really be broken through, within a set time and scope, with the defenders aware
Goal-driven; tests whether detection and response can catch the attack
Attackers and defenders tune detection rules together
Replays the full set of methods of a known hacking group
Before starting, you also have to set out the scope and the rules of engagement in writing, which means agreeing in advance which systems may be attacked, at what times, and in what situations you must stop immediately. On the legal side, the US has the Computer Fraud and Abuse Act, and Taiwan has Chapter 36 of its Criminal Code, on offences against computer use. In 2022 the US Department of Justice stated that good-faith security research should not be prosecuted under the Computer Fraud and Abuse Act.
Where things stand: there are few public annual statistics in this line of work. Ironically, attack tools themselves can have serious vulnerabilities. In February 2025 CALDERA fixed a vulnerability rated "critical", the highest level on the severity scale.
What AI changes: Cybench is a set of test questions that measures how well AI can attack on its own, with challenges taken from professional capture-the-flag competitions. Capture the flag is a kind of security contest: the organisers hide a piece of text in a system as the "flag", and contestants have to break in and retrieve it. Stanford University's AI Index 2026 records that the share of challenges AI could solve without any hints rose from 15% in 2024 to 93% in 2026.
A more concrete example is XBOW, an AI tool that carries out penetration tests on its own. HackerOne is a bug bounty platform where companies announce "find a vulnerability in our systems and we will pay you", so that outside researchers come looking. According to several media outlets, in June 2025 XBOW reached number one on HackerOne's US leaderboard, the first non-human to do so, and later took first place worldwide as well. Under HackerOne's rules, though, a human reviewed every report before it was submitted.
Incalmo, published in June 2025 by the AI company Anthropic and Carnegie Mellon University, is a toolkit that helps AI carry out multi-step attacks, meaning first getting in, then finding a way to gain more access, then getting hold of the data. Without it, the AI models tested could not break into a single one of ten practice environments; with it, they broke into nine out of ten. These practice environments had no active defenders fighting back, though, so the results only show that the AI could get in, not whether it would have been caught.
19 Vulnerability Research and Bug Bounties
What it is: vulnerability research means looking for brand-new weaknesses that nobody knows about yet, known as zero-day vulnerabilities: the vendor does not know about them, so there is no fix yet. By contrast, there are also old vulnerabilities that the vendor fixed long ago but that are still being attacked, because many people have not installed the update.
The industry has a set of norms for how to go public once you have found one, called coordinated disclosure: you first tell the vendor privately, and usually give them about 90 days before going public. Bug bounties turn this into a formal system, with vendors paying researchers who report vulnerabilities by these rules.
Where things stand: every year the Google Threat Intelligence Group (GTIG) counts the zero-days exploited "in the wild", meaning really used in attacks rather than demonstrated in a lab. In 2025 there were 90, and 48% of them targeted enterprise products (the equipment and software that companies use internally), which GTIG says is the highest share on record. GTIG also points out that, for the first time, more zero-days were traced to commercial spyware vendors (companies that sell surveillance software to governments) than to state espionage groups.
Vulnerabilities also have a market. In a price list published in April 2024 by Crowdfense, a company that specialises in buying vulnerabilities, a full "zero-click" attack chain for the iPhone was priced at US$5 million to US$7 million. Zero-click means the victim is infected without having to tap or click anything; a full attack chain links several vulnerabilities together to go all the way from getting into the phone to taking complete control of it.
Public bug bounties are still growing, too. HackerOne paid out US$81 million in the year to June 2025. Among the reports, those about prompt injection (the trick mentioned earlier for fooling an AI into following instructions) shot up by 540%.
What AI changes: in July 2025 Google said that Big Sleep, its vulnerability-hunting AI agent, had found a vulnerability in SQLite (a small and very widely used database), and claimed this was the first time an AI agent had directly stopped a real attack that was about to happen. Anthropic, for its part, says its model Mythos Preview dug up a vulnerability in OpenBSD (an operating system known for its security) that had been hidden for 27 years.
This capability is spreading, too. GLM-5.3, from the Chinese company Zhipu, is an open-weight model, meaning the model itself is public and anyone can download it and run it themselves. Citing an evaluation by the US Center for AI Standards and Innovation, Anthropic says GLM-5.3's security capabilities trail the best US models by only about four months, and that the built-in safety layer that makes the model "refuse to help with harmful things" can be stripped out for about US$4,400 in computing costs.
But finding a hole is not the same as fixing it. Anthropic has reported that of the 530 high-risk or critical vulnerabilities disclosed by May 2026 through its Glasswing programme (which puts Mythos in defenders' hands first), only 75 had been fixed. Anthropic also mentions that maintainers (the people who look after and fix this software) are being swamped by AI-generated vulnerability reports. Some are poor-quality "AI slop reports", and some are genuine but too numerous to deal with; some maintainers have even asked Anthropic to slow down its disclosures.
My judgement is that the bottleneck has moved from finding holes to fixing them. So anyone who uses AI to find a vulnerability should at least first follow the steps themselves and confirm that the problem really happens, before going to the maintainers.
Family 7 The Physical World and People
The first six families are mostly about data and accounts, and when something goes wrong it can usually be rebuilt from backups. The seventh family is different. Here, computers directly control valves, heating and production lines, so a failure can mean a blackout, homes without heating, or even a threat to people's safety, and this equipment stays in service for 15 to 30 years. People, too, are a door that is often tricked open. This family has three domains: factories and power grids, the connected devices that are everywhere, and people.
20 OT and Critical Infrastructure
What it is: operational technology (OT) means systems that directly control the physical world, such as power grids, water supply, oil and gas pipelines, factory production lines and building air conditioning. You can think of it as "computers that make things physically move". On site, rows of small controllers read sensors and open and close valves and motors. There is also a separate safety instrumented system that automatically stops the process as soon as pressure or temperature goes over its limit; it is the last line of protection.
The biggest difference between IT and OT is what comes first. Ordinary IT fears data leaks most. OT cares most about people's safety and keeping things running, and someone seeing the data comes last, because a control system stopping for one minute can mean a blackout or damaged equipment.
OT also faces some practical constraints. Equipment often runs old operating systems that nobody has supported for years. The rules devices use to talk to each other are called communication protocols, and protocols such as Modbus do not check who sent a command: they carry out whatever arrives, like a door with no lock. On a phone, an update takes one tap, but a hole in a factory system often has to wait for a maintenance shutdown before it can be patched.
The basic defence is zoning: divide the equipment into layers and zones, open only the necessary, controlled channels between zones, and never let corporate IT touch the controllers directly. It is a bit like a building where each floor needs a different key card.
Where things stand: as I see it, over a decade and more, OT attacks have gone from "theoretically possible" to repeatedly causing blackouts, loss of heating and damaged equipment in the real world.
A report by Mandiant, a security company owned by Google, says that when the Russian hacking group Sandworm attacked Ukraine's power supply in October 2022, it did not use purpose-built industrial control malware. Instead, it borrowed control software already present in the substation's systems to send commands that cut the power. This is called "living off the land": using the victim's own legitimate tools to do harm. Traditional antivirus works by recognising what known bad programs look like, and a legitimate tool looks perfectly fine, so it is hard to spot.
In January 2024, about 100,000 people in Lviv, Ukraine, were left without heating for about two days in sub-zero weather. The attackers got in through a router and sent commands over Modbus. Dragos, a company specialising in industrial control security, says this program, called FrostyGoop, is the first piece of industrial control malware to cause real-world impact through Modbus.
At the end of December 2025 it was Poland's turn. More than 30 wind and solar farms and other sites were broken into, and wiper malware (malicious software designed to erase data and systems entirely) damaged the on-site control equipment; fortunately, there was no loss of power or heating. Poland's computer emergency response team attributed the attack to a team working under Russia's Federal Security Service. According to reports, it was one of the first large-scale attacks aimed at distributed energy (small power-generating sites spread across many locations).
An attack does not even need to touch the controllers. In September 2025, what was hit at the British carmaker Jaguar Land Rover was its office IT systems, yet what stopped was the whole production line, because a factory's scheduling and parts ordering usually depend on those systems too.
I think the 2021 case of Colonial Pipeline, a US fuel pipeline company, belongs in the same category: the ransomware (malware that locks up your files and demands a ransom to unlock them) hit IT, and the company shut the pipeline down to be on the safe side. The full story of the carmaker incident is in "Who Attacks and How They Get In".
There is also a quieter threat called pre-positioning: slipping in early to hold a position and acting only when needed. The prime example is Volt Typhoon, a Chinese hacking group disclosed by Microsoft in 2023, which lived off the land almost throughout. According to reports, a joint warning issued by several US government agencies in 2024 said it had been lurking in energy, water and other systems "for at least five years", and had stolen OT system diagrams.
Norway had a case too: according to reports, in April 2025 the web-based control screen of a dam was protected only by a weak password, and a release valve was opened for about 4 hours. Line these entry points up and I see a pattern. For Colonial it was an old password that could be used to connect to the company network from outside; for FrostyGoop, a router; in Norway, a weak password; in Poland, firewalls connected straight to the internet, plus passwords already used elsewhere. The consequences were physical, but the ways in were almost all ordinary IT problems.
What AI changes: in the material I have gathered, no OT incident has yet been confirmed as carried out with AI. The Google Threat Intelligence Group (GTIG) also wrote in September 2026 that it has not yet seen attackers deploy fully autonomous attack workflows in real environments. Three things, however, affect OT indirectly.
First, vulnerabilities are being exploited faster, and OT is the slowest to patch. Another GTIG report, from 30 September, says that from January to August 2026 an average of 18 vulnerabilities a month were used in real attacks, compared with 10.5 a month in 2025. GTIG believes the main reason is that vulnerabilities already made public are being turned into attack tools faster, very likely with help from AI. OT can only be patched during maintenance, so I think the same vulnerability stays open to attackers far longer than it does in IT.
Second, defenders are getting the tools first. The AI company Anthropic has a programme called Project Glasswing that lets defenders use its most security-capable model a step ahead of attackers. According to Anthropic, about 150 more organisations joined in June 2026, including some from the power and water sectors.
Third, the hardest step may get easier. The hardest part of attacking OT has always been understanding the process, which is why Volt Typhoon stole system diagrams. Language models are good at reading manuals and configuration files, and in future they could make this step easier. There is no public evidence of this yet, and I treat it as a risk to watch.
21 IoT, Medical and Automotive Devices
What it is: the Internet of Things (IoT) covers all kinds of connected devices, such as routers, network cameras, smart TVs and TV boxes, network-attached storage (small networked hard drives for keeping files on) and smart home appliances. Medical devices and cars count too, and the bar is higher for them, because when something goes wrong, patients and people on the road are affected.
These devices share problems that resemble a house whose locks were never changed after you moved in: nobody changes the factory default password, the management interface is exposed directly to the internet, and vendors support a product for only a few years while users keep it for ten. Buyers also cannot see which software components run inside, so when a vulnerability turns up in a common component, they have no idea whether the devices at home are affected.
Where things stand: in my view, the direction of the past few years has been to make security a condition for going on sale: if a product can't meet it, it can't be sold. Take the EU. Since August 2025, wireless connected devices sold in the EU have had to meet basic security requirements. Since September 2026, the Cyber Resilience Act has required manufacturers that come across a vulnerability being actively exploited by hackers, or a serious incident, to send an early warning to the EU's reporting platform within 24 hours.
Medical devices and cars are being tightened up too. In the US, a networked medical device submitted for approval must come with a software bill of materials, which is an "ingredients list" of the software components the product uses, along with a plan for handling vulnerabilities after the product goes on the market. On the car side, the UN has two vehicle regulations, one requiring carmakers to manage cybersecurity properly and the other to manage software updates properly; since July 2024, every new car sold in the EU has had to comply.
Turning to the attack side, infected TV boxes have become weapons. A distributed denial of service (DDoS) attack uses huge numbers of devices to flood a service with traffic at the same time until it chokes, like tens of thousands of people squeezing into one corner shop at once. The internet services company Cloudflare recorded an attack of 31.4 terabits per second, which it calls a record. The botnet behind it, meaning a large group of devices remotely controlled by hackers, was estimated at 1 million to 4 million devices, most of them infected Android TV devices.
Routers, meanwhile, are often used as stepping stones; Volt Typhoon hid behind hijacked routers. In the Known Exploited Vulnerabilities (KEV) catalogue kept by the US cybersecurity agency, networking and network storage devices from Taiwanese brands account for quite a few entries; the numbers are in "Taiwanese Brands in KEV".
According to reports, in February 2026 the University of Mississippi Medical Center in the US was hit by ransomware: all 35 of its clinics closed for 9 days, and its electronic health record system went down. To me, this is what sets healthcare apart: what stops is seeing patients and treating them.
What AI changes: AI makes vulnerabilities quicker to find, but patching cannot keep up. For the Glasswing programme mentioned earlier, Anthropic's own progress update for May 2026 showed that of the 530 high-risk or critical vulnerabilities it had reported, only 75 had been fixed. My view is that if even maintained software cannot be patched in time, out-of-support routers and cameras will only be slower, or will never be patched at all.
Malware is also starting to bring its own AI. PROMPTSPY, a piece of Android malware recorded by GTIG in May 2026, has a built-in "automated agent" that calls a large language model (the kind of model behind ChatGPT). An agent generally means an AI program that can plan its own steps and use tools. The TV devices mentioned above run Android too.
The number of newly disclosed vulnerabilities each month is also climbing fast: by GTIG's count, it was 5,045 in January 2026 and had already reached 10,740 by August. Note that these are newly disclosed vulnerabilities, not ones already used in attacks. Over the next year or two, I will be watching one thing: whether the people who handle vulnerabilities at vendors are swamped first, or pushed first into using AI to help set priorities and make fixes.
22 Human Factors and Social Engineering
What it is: social engineering attacks people instead of programs, getting them to click a link, enter a password, approve a login, transfer money, or reset an account on the attacker's behalf. These are the common techniques:
- Phishing: fake emails or web pages that trick you into entering your username and password.
- Voice phishing: phone calls impersonating the IT department, a bank or a manager.
- Business email compromise (BEC): impersonating the boss or a supplier to ask for a payment, or for a change to the account that payments go into.
- Deepfakes: AI-generated video or voice used to impersonate a specific person.
- Insider threats: people who already have access, including employees, contractors, or people who got hired under a fake identity.
Defence takes two paths. One is training and drills. The other is to make sure that even when a trick works, the attacker gets nothing out of it. Passkeys are an example: a passkey is like a key that only opens your own front door. It recognises only the real website, and at a fake site it simply won't turn. Confirming by calling back a known number, and requiring two people to approve large payments, also belong on this path.
Where things stand: the figures from the Internet Crime Complaint Center of the US Federal Bureau of Investigation (FBI) show the scale best. In 2025 it received more than 1 million complaints for the first time, with reported losses of about US$20.9 billion. The biggest losses came from investment fraud, at more than US$8.6 billion, with BEC in second place. These are amounts that victims chose to report, so the real losses are probably higher.
The phone has become an important way in. Mandiant's M-Trends 2026 looked at intrusions in 2025 to see how attackers first got in. The most common route was exploiting software vulnerabilities, at 32%; second was voice phishing, at 11%. Traditional phishing emails were down to just 6%.
The standard playbook of a group called Scattered Spider is to phone a company's IT help desk, pose as an employee and ask for a password reset, or for multi-factor authentication to be set up again (the step where, besides typing your password, you confirm once more on your phone when you log in). Once they have the account, they work their way inwards until they reach the system that controls logins across the whole company, the servers and the backups. According to reports, at Easter 2025 attackers posed as staff of the British retailer Marks & Spencer and got its outsourced help desk to reset a password. Online sales stopped for about 46 days, and the company estimated it lost about £300 million in operating profit.
Attackers also come up with new pretexts. In August 2026 Google disclosed that one group had been calling employees' personal mobile phones, claiming that an urgent, mandatory passkey enrolment was needed, and then steering them to a fake page. The page works like a postman who reads your letters: whatever you type, it passes on to the real website as normal while keeping a copy for itself, and that is how it steals usernames, passwords and login tokens (the temporary pass a website gives you after you log in). What strikes me as ironic is that companies' push for passkeys became the cover story for the scam. The passkeys were not cracked; the attackers left the login alone and went after the enrolment and reset processes instead.
A deepfake can be an entire video call, or just messages plus a fake voice. According to reports, in early 2024 an employee at the Hong Kong office of Arup, a British multinational, joined a video call in which everyone apart from him, including the chief financial officer, was a deepfake; in the end he transferred HK$200 million. Also according to reports, in February 2026 the then chairman of Fideuram, an Italian financial institution, received WhatsApp messages from someone posing as the chief executive of its parent group, Intesa Sanpaolo, and heard a lawyer's voice cloned by AI. He transferred about €95 million, of which more than half was later recovered.
Training does less than you might expect. According to reports, the University of California San Diego and its UCSD Health ran an experiment with more than 19,500 employees, published at the security research conference IEEE S&P 2025. It found that whether an employee's last annual security training was the previous month or nearly a year earlier made no difference to whether they clicked on phishing emails. The training material that pops up after someone clicks a simulated phishing email cut the share of people taken in by only about 1.7 percentage points.
Taiwan's regulations require government agencies to run a social engineering drill every six months. Drills can measure the risk, but I don't think training material after the fact will do much on its own to change behaviour.
So the focus has to move to the second path, which doesn't rely on people remembering anything. Besides stopping fake websites, passkeys are also easier to use than passwords: by Microsoft's own statistics, about 98% of passkey sign-ins succeed, against only 32% for passwords. People will only really switch if the new way is easy to use.
Insider threats have changed shape as well. North Korean remote IT workers get hired by companies under fake identities, then draw salaries and steal data. By Mandiant's count, the median time from an attacker getting in to being discovered is only 14 days for intrusions in general, but 122 days for espionage cases and cases of fake North Korean employees.
What AI changes: the 2025 edition of Microsoft's Digital Defense Report states that "AI-driven phishing is now three times more effective". Fake identities can also be mass-produced: GTIG recorded in September 2026 that Iranian and North Korean hacking groups were using large language models to create fake personas with realistic photos, fake recruiters and fake CVs.
The time from fooling someone to getting what you want has also shrunk. Anthropic's September 2026 report says that associated members of the extortion gang ShinyHunters had "nearly all of the work done by AI agents", and in about 34 hours obtained login tokens for Microsoft cloud accounts at more than 40 organisations. For how AI misuse has escalated step by step, see "Role 1 AI Helps Attackers".
Defenders can use it too. Microsoft has an AI program that does a first sort of suspicious emails for security staff; Microsoft's own internal trial says it made finding malicious emails up to 550% faster.
The new risk, as I see it, is that voices and faces can no longer prove who someone is. In the deepfake cases above, the money had already gone out; where any was recovered, it was thanks to freezing and tracing after the fact, not because someone spotted the fake at the time. My family and I have agreed on a call-back rule: any call or message asking for money, a password or a verification code gets hung up on first, however convincing the voice and face, and is then confirmed by calling back on a number we have saved ourselves. This rule doesn't need anyone to learn how to spot a deepfake.
The Four Roles of AI in Security
When people hear "AI and security", many of them think first of ChatGPT. In fact AI arrived in security much earlier; it just used to be called "machine learning": letting a computer work out the rules for itself from large numbers of examples, rather than having people write them in one by one.
Gmail automatically moves spam into a separate folder, and a similar approach had already become mainstream in the early 2000s: count how likely each word is to appear in spam and in ordinary mail, and from that work out whether a message looks like spam.
Security has plenty of tools of this kind too. Anomaly-based intrusion detection tries to catch people who break into a system: it first learns what the system normally looks like, then raises an alert when things stray too far from that. "Next-generation antivirus" doesn't just match the signatures of known viruses; it has a model judge the file itself. User and entity behaviour analytics spots stolen accounts from login times and places, such as someone logging in in the middle of the night from somewhere unfamiliar.
These systems have always had a few long-standing problems. The two most common are:
- Too many false alarms: real attacks are extremely rare among the vast amount of normal activity. Even if the false alarm rate is very low, once it is multiplied by the huge number of events each day, false alarms still far outnumber real attacks, and the people on duty soon become numb to them.
- Adversaries set out to fool it: researchers have shown that attaching a piece of text taken from a legitimate program to the end of a piece of malware was enough to make an antivirus product marketed on its machine learning judge it safe.
These were all single-purpose detectors: they could give a score, but they couldn't act on their own. After 2023, that changed.
In 2023–24 came the "copilots": large language models (LLMs, the kind of model behind ChatGPT) helped security analysts go through alerts and piece together what had happened, but people still made the decisions. Security analysts are the people in a company whose job is to watch the alerts and decide whether an attack is under way.
From 2025 came AI agents: AI systems that can chain several steps together and carry them all out on their own. They can read code, find vulnerabilities and write patches (that is, software updates), and they can also write exploits, programs that turn a vulnerability into a way in. A vulnerability is a mistake in how a program is designed or written that lets someone do something they shouldn't be allowed to do, like a gap in a door lock that nobody has noticed. Both attackers and defenders are using them.
One number is enough to show how fast things have moved: on Cybench, a test made up of professional security challenges, the share that models solved without any human hints was 15% in 2024 and 93% in 2026.
I divide AI's place in security into four roles:
The same attacks, faster, cheaper and with a lower barrier to entry
Sifting alerts, reading code to find vulnerabilities, writing patches
Models, agents and the tools they connect to become targets
The strongest cyber capabilities open first to verified defenders
The first two roles are really two sides of the same capability: a model that can find vulnerabilities and write attack programs helps whoever it is handed to. The third role appeared once AI was being built into company systems on a large scale: AI itself became a new weak point. The fourth role belongs to the frontier labs, the handful of companies that build the most powerful models. As I see it, this is their response to the first three: since the capability can't be split in two, they control who gets to use it.
Role 1 AI Helps Attackers
Several large companies regularly publish threat intelligence reports describing how attackers misuse their AI. Put these reports in date order and you can see the descriptions grow more serious year by year:
From 2024 to early 2025, AI was just an efficiency tool. In February 2024, Microsoft and OpenAI jointly reported that five state-backed hacking groups were using LLMs to look things up, write small programs and write phishing emails, but that they had "not yet observed particularly novel or unique AI-enabled attack techniques". In January 2025, the Google Threat Intelligence Group (GTIG) reached much the same conclusion: AI let attackers work "faster and at higher volume", but it was "not yet the game-changer it is sometimes portrayed to be".
In the second half of 2025, AI started doing the work itself. In August, Anthropic, the AI company that makes Claude, disclosed a case it called "vibe hacking". The name borrows from vibe coding, which means not writing the code yourself, just telling an AI what you want and letting it get on with it. A single attacker used Claude Code, Anthropic's coding tool, to steal data from at least 17 organisations and then extort them; the victims included healthcare providers and government bodies.
In November, GTIG catalogued a batch of malware that "only asks an LLM when it runs". One of them asks Google's Gemini model to rewrite it so that it can evade detection, though GTIG considers it still experimental. Another asks a model to generate, on the spot, the commands it should run.
The same month, Anthropic disclosed a case codenamed GTG-1002, which it assessed "with high confidence" to be the work of a Chinese state-sponsored group. The campaign was aimed at about 30 targets. Claude Code did 80% to 90% of the hands-on, step-by-step tactical work; people only signed off at the key points, making just 4 to 6 decisions per attack campaign. In the end, only "a small number of cases" led to a successful break-in.
In May 2026, GTIG disclosed the first zero-day exploit it believes was developed by AI. A zero-day is a vulnerability the vendor doesn't yet know about, so no fix exists yet. This one targeted an open-source administration tool (open source means the source code is public and anyone can use and modify it for free) and could get round two-factor authentication, the extra check you pass at login on top of your password, such as a code sent to your phone. It had been prepared for large-scale use and was stopped before that could happen.
In September, agents appeared in both companies' reports. GTIG said that one attacker, starting from a single compromised cloud server (a remote computer rented over the internet), "planned, built and executed" a large-scale campaign to harvest usernames and passwords in under 6 hours. Anthropic recorded that people linked to the hacking group ShinyHunters "let AI agents do almost all the work", obtaining more than 2,100 Azure AD tokens. Azure AD is Microsoft's system for managing companies' employee logins. A token is like a temporary pass a website gives you after you log in: whoever holds it doesn't need to type in the password again.
Anthropic's September report summed it up in one sentence: "Sophisticated attacks no longer require sophisticated attackers." It also said that between state-level groups and lone operators, "the main differentiator is no longer sophistication, but intent".
But one line has not yet been crossed. Also in September, GTIG stated plainly that it had "not yet observed attackers deploying fully autonomous attack pipelines against targets in the wild", meaning attacks that run from start to finish without anyone stepping in. AI makes mistakes too: Claude in GTG-1002 would hallucinate login details, that is, state with great confidence things that simply didn't exist.
AI has already left its mark on the statistics, but it is not yet the main cause of break-ins. According to media reports, the 2026 edition of the Global Threat Report from the security company CrowdStrike says attacks by adversaries using AI rose by 89%. Another security company, Mandiant, says in its annual report that the intrusions of 2025 were not "caused" by AI.
GTIG also found that as soon as a vendor releases an update, attackers have an LLM compare what changed before and after, quickly work out where the vulnerability is, and strike before everyone has installed the update. I'll come back to this in "Who Has the Upper Hand, Attackers or Defenders".
Phishing, social engineering and deepfakes
The most direct help AI gives attackers is making deception cheap. Phishing means sending fake emails or messages that pose as a bank or a company, to trick you into clicking a link or entering your password; social engineering more broadly means tricking people rather than machines. The 2026 edition of Microsoft's Digital Defense Report says AI is "compressing attack timelines and lowering the cost of sophisticated capabilities". In other words, fooling someone now takes much less time and effort.
Deepfakes are fake images or voices made with AI; real fraud cases are covered earlier in "22 Human Factors and Social Engineering". To guard against this kind of fraud, I think what works is process, such as calling back on a known number to check, and requiring two people to approve anything important, not trying to tell real videos from fake ones by eye.
How attack capability is measured
Threat intelligence tells us what attackers "have done"; benchmarks tell us what models "can do". A benchmark is a fixed set of test questions, so that different models sit the same exam paper. Cybench, mentioned at the start, contains 40 professional capture-the-flag challenges. Capture the flag is a puzzle-solving competition in the security world: contestants have to break into systems that have been deliberately left with vulnerabilities and retrieve a "flag" hidden inside. Stanford University's AI Index 2026 says Cybench has climbed more steeply than any other agent benchmark.
A harder test is whether a model can write a working attack program on its own. In Anthropic's April technical report, its new model Mythos Preview, available only to partners, produced attack programs that worked in a test environment 181 times against the JavaScript engine of the Firefox browser (the part of the browser that runs a web page's code). Its previous model, Opus 4.6, succeeded only twice in several hundred attempts.
Anthropic's prediction for what comes next is: "Mythos-class models will be widely available within the next 6 to 12 months." That means attackers may by then be able to get hold of models at the same level too.
With open-weight models, safety mechanisms can be removed for a price
Most of the numbers above come from closed models, which can only be used by connecting to the developer's own service, so the company can at least filter suspicious requests and ban accounts. A model's weights can be thought of as everything it has learned, stored as one very large file. Open weights means making this file public for anyone to download; once it has been downloaded, how it is used is entirely up to the user.
On 29 September, Anthropic published an analysis of GLM-5.3, an open-weight model from the Chinese AI company Zhipu. The Center for AI Standards and Innovation, the US government body responsible for evaluating AI, considers it "the most cyber-capable open-weight model to date", about "four months" behind the leading US models. In Anthropic's hands-on tests, GLM-5.3 took about a day to chain together several vulnerabilities in a JavaScript engine that nobody knew about (like first prising open the front door and then opening the safe) into a working attack program.
Worse still, the safety mechanisms can simply be removed: by editing the weights, you can strip out the model's tendency to "refuse to answer". This is called abliteration. Anthropic estimates it costs only about US$4,400, and weights treated this way are already circulating.
All these measurements have their limits. Cybench is a set of exam questions, not a real network with people defending it, and GTIG has not yet seen a fully automated attack. But in my view the direction is already clear: the bar for writing attack programs has dropped sharply within two years, and open-weight models are only a few months behind; once they have been downloaded, there is no taking them back.
For defenders, the more realistic assumption is that there is an assistant on the other side that never tires and never refuses. So defensive rules are best enforced automatically, for example by halting things the moment something unusual is spotted, rather than waiting until a person notices before acting.
Role 2 AI Helps Defenders
AI on the defensive side follows two tracks. One is handling alerts in the security operations centre (SOC), which you can think of as a company's security control room, where people work in shifts watching the alerts. The other is finding vulnerabilities in code and writing patches, and I think this track became the frontier labs' main battleground in 2026.
Copilots and agents in the SOC
The SOC's long-standing problem is too many alerts; according to some industry surveys, about half of them are false alarms. The selling point of the "agentic SOC" is to let AI do the first round of triage, so that people deal only with what genuinely needs their judgement. Triage works much as it does in a hospital's A&E department: first set aside the cases that are clearly fine, then move the urgent ones to the front of the queue.
In March 2025, Microsoft announced a set of agents for its security AI assistant, Security Copilot, one of them dedicated to triaging phishing emails. Microsoft ran its own comparison, splitting analysts into groups at random, and the group using the agent identified malicious emails "up to 550% faster".
Project Ire, which Microsoft announced in August 2025, shows the gap between a test environment and the real world better than anything else. It analyses files on its own and decides whether they are malicious. First, two terms: precision is "of the things it says are a problem, how many really are", and recall is "of the things that really are a problem, how many it catches".
By Microsoft's own published figures, on a public test dataset it caught over 80% of the truly malicious files (a recall of 0.83). On hard real-world files, around 90% of what it flagged was still correct (a precision of 0.89), but it caught only about a quarter of the truly malicious files (a recall of 0.26).
Google points to a problem in the opposite direction. When it open-sourced a vulnerability-finding tool in September, it warned that sloppy AI code scanning "hallucinates" vulnerabilities that don't exist at all: fewer than seven in every hundred reports are real.
Letting AI find vulnerabilities and write patches
According to media reports, Google said in July 2025 that its AI agent Big Sleep had found a vulnerability in the database software SQLite. At the time, the vulnerability was known only to attackers and was about to be exploited. Google called this the first time an AI agent had directly foiled an attempted attack in the real world.
In August that year, the US Defense Advanced Research Projects Agency held the AI Cyber Challenge, with the final at the DEF CON hacker conference; the contest was for systems that find and fix vulnerabilities fully automatically. According to media accounts of the official results, besides the vulnerabilities the organisers had deliberately planted, the competing systems found 18 real vulnerabilities that nobody had known about. The top systems were later all open-sourced, so anyone can use them.
In the first half of 2026, according to Anthropic, while it was working with Mozilla, the maker of the Firefox browser, Opus 4.6 confirmed 14 high-severity vulnerabilities in two weeks, equal to "almost a fifth" of all the high-severity fixes Firefox made in the whole of 2025. What Mozilla valued most were the proof-of-concept programs attached to the reports (small programs that prove the vulnerability can really be triggered) and the suggested fixes, which the engineers responsible could use straight away to reproduce the problem and start fixing it.
In 2026, the strongest models go to defenders first
In my view, the most visible change on the defensive side in 2026 was that all three frontier labs handed their most cyber-capable models to defenders with verified identities first. Who gets access and how it is tiered are left for "Role 4 Frontier Labs Gate Their Cyber Capabilities".
On 7 April, Anthropic decided not to release its new model Claude Mythos Preview to the public. Instead, through a programme called Project Glasswing, it gave the model only to 12 founding partners and more than 40 other organisations, to use for finding vulnerabilities. OpenAI and Google have similar arrangements, but Glasswing is the one with the most public data.
Among the findings Anthropic published was a 27-year-old vulnerability in the OpenBSD operating system that could crash the system remotely. In its first-month update on 22 May, Anthropic said its partners had found more than 10,000 high or critical-severity vulnerabilities.
On the open-source side, 530 high or critical-severity vulnerabilities had been disclosed to maintainers, that is, reported to the people responsible for the software. But fixes are coming nowhere near fast enough:
Once something has been fixed, you still have to wait for users to install the update. Anthropic itself says progress is "now limited by how fast we can verify, disclose and patch". Some open-source maintainers, swamped by large numbers of AI-generated reports, have also asked Anthropic to slow down its disclosures.
One media outlet, citing Anthropic's disclosure dashboard, reported 1,596 disclosed and 97 patched, which doesn't match the official post, possibly because the two counts cover different scopes. In both sets of figures, only a small fraction has been fixed.
Vendor numbers are mostly self-reported
Most of the defensive figures in this section come from the companies selling the products or releasing the models themselves, such as Microsoft's 550% and Glasswing's various counts. I'm not saying they are untrue, but they should be read as "claims" and checked against third parties wherever possible.
- Figures with a public record are more reliable: the AI Cyber Challenge systems have been open-sourced, and Cybench is an academic test. Publicly disclosed vulnerabilities all have standard ID numbers that anyone can look up, and the organisations responsible for software, such as Mozilla, say for themselves how much they have fixed.
- Different conditions, very different numbers: Project Ire's recall dropped sharply on real-world files. Anthropic says that of the 1,752 Glasswing findings that went through independent triage, 90.6% were real vulnerabilities, whereas the sloppy scanning Google warned about produces mostly false alarms. The difference is whether the findings have been reproduced, verified and checked by a person.
- Finding vulnerabilities is no longer the bottleneck: far fewer are being fixed than found. From here on, the race is about how quickly problems are confirmed, updates are produced and everyone installs them.
Role 3 AI Itself Becomes the Attack Surface
In the first two roles, AI was a tool in someone's hands. This role turns that around: the AI system itself becomes the attack surface, meaning the places where it can be attacked. These include the model, its training data, the knowledge base the model looks things up in before answering, the tools connected to it, and agents that act on their own. "14 AI System Security" gave an overview; here I dig one layer deeper.
As the section on the root cause explained, large language models can't tell "instructions" from "data", so any text they read may be taken as a command and carried out. What you type to an AI is called a prompt; slipping commands into text the AI will read and tricking it into following them is called prompt injection.
Poisoning and backdoors. Poisoning means quietly slipping malicious content into the places where a model learns or looks things up. A backdoor is one kind of poisoning: the model behaves perfectly normally until it sees a particular secret signal, and only then turns bad, like a bribed employee who acts only on hearing the agreed code word.
The old intuition was that the bigger the model and the more data it had, the more poison an attacker would need to plant. In October 2025, an experiment by Anthropic, the UK AI Security Institute and the Alan Turing Institute shook that intuition. They trained models of four different sizes from scratch and planted a backdoor that "outputs gibberish on seeing a particular signal". Whatever the size of the model, slipping in about 250 malicious documents was enough. The bigger the model, the smaller the share of its data those 250 documents made up, yet they worked just as well: what matters is the number of documents, not the proportion. The authors also cautioned that this is a very narrow backdoor, and that how far the trend holds "remains unclear".
Backdoors can also survive safety training, the extra training a model is given before release to teach it to refuse to do harmful things. In a 2024 Anthropic study, researchers deliberately trained a "sleeper agent" model: told "the year is 2023", it wrote secure code; told "2024", it wrote code with vulnerabilities. They then tried several common forms of safety training to correct it, and none of them managed to wash the backdoor out.
As mentioned earlier, many AI systems look things up in a knowledge base before answering, and that knowledge base is an even easier target. A study cited by OWASP found that planting just 5 poisoned documents in a knowledge base of millions of texts was enough for the attack to succeed about 90% of the time.
Prompt injection has already turned up in products. Direct injection is when users themselves give instructions in the conversation to try to get around restrictions; what is commonly called jailbreaking belongs here. Indirect injection is more insidious: the attacker hides instructions in content the AI will read, such as an email or a form, and when the AI comes across it in the course of its normal work, it carries them out as commands.
EchoLeak, disclosed in June 2025, is a textbook example. The attacker sends a specially crafted email, and the recipient doesn't have to click anything. While answering a user, Microsoft 365 Copilot, an AI assistant for office work, read that email in as reference material, then followed the instructions hidden in it and packed the data it had found into a link pointing to the attacker's website; as soon as the link loaded, the data went along with it to the attacker. It also got past the classifier Microsoft uses specifically to catch injection, a filtering model that judges whether a piece of text looks like an attack. The vulnerability scored 9.3 on a severity scale that tops out at 10, and Microsoft has fixed it on the server side.
Salesforce's ForcedLeak (scored 9.4) is even more ironic: the thing that went wrong was the very list meant to prevent leaks. According to several media reports, the attacker hid instructions in the description field of a website form for potential customers, and the agent followed them and sent data to a domain, the name part of a web address, such as google.com. This domain was on Salesforce's trusted allowlist, the list of addresses the system allows data to be sent to. But domains have to be renewed regularly, and this one had expired while still sitting on the list; Noma, the security company that found the vulnerability, registered it again and could then receive the data.
Coding agents are no different. GitHub is a website where programmers keep their code and report problems, and Anthropic has a tool that lets Claude work there automatically, called Claude Code GitHub Action. In June 2026, Microsoft's security research team revealed that hiding instructions in a GitHub issue, that is, a problem report, could make the tool leak Anthropic API keys (the passwords that programs use to call an AI service). GitHub automatically scans for keys that people have posted publicly, relying on the fixed prefix such keys begin with; strip off that prefix and the scan no longer recognises them. By the time Microsoft went public, Anthropic had already fixed the problem in May. Taken apart, the three incidents have almost the same structure:
In the course of normal work, the AI reads text written by the attacker
The same agent has permission to read sensitive content
The data can be carried outside
Data or keys end up in the attacker's hands
Untrusted input, private data, outward actions: keep at most two at once and the chain breaks; when all three are needed, a person approves
How the rankings are made. In practice, the list people use most is the LLM Top 10 from the Open Worldwide Application Security Project (OWASP). OWASP says the new edition of August 2026 is the first to be weighted by real incidents: three quarters of the ranking comes from votes by front-line practitioners, and one quarter from incidents that have actually happened.
One result runs against intuition: on incident counts alone, prompt injection would "drop out of the top ten entirely". The project leads call this a "defence effect", and kept it at number one anyway. My reading is that when everyone is already defending against something, fewer cases actually turn into incidents that get recorded, so few incidents doesn't mean low risk.
AI's add-on tools can be tampered with too. The Model Context Protocol (MCP) is an open standard for connecting AI agents to external tools, a bit like installing apps on a phone: install an MCP server and the agent gains a set of tools, such as sending email or querying a database. The agent decides how to use them by reading the description that comes with each tool.
That description is where the problem lies. In April 2025, the security company Invariant Labs demonstrated tool poisoning: by writing instructions into a description that are meant for the AI and that people won't notice, they got an agent to hand over the key the user logs in to servers with. Another trick is called the rug pull: a server behaves normally the first time it is loaded, then swaps out its tool definitions the second time and sends WhatsApp messages out.
This has already happened in the real world. npm is a platform where programmers download ready-made software components. In September 2025, a component on it called postmark-mcp impersonated the email service Postmark; after building up more than 1,000 downloads a week, it pushed out a new version that blind-copied every email it handled to the attacker. Based on a report from the security company Koi, MITRE ATLAS, a knowledge base of attacks on AI, describes it as the first malicious MCP server found in a real-world environment.
Vulnerabilities are piling up faster too. CVE is the standard numbering system for publicly known vulnerabilities. I ran a keyword match against the official CVE database: records whose title or description mentions MCP in connection with agents, tools or servers came to about 72 for the whole of 2025, and had already reached about 514 in 2026 up to 3 October. These are keyword counts, so they pick up some unrelated uses of "MCP" and miss records that don't spell out MCP; they are only good for showing the order of magnitude and the trend.
Model files are programs. Many people think downloading a model is like downloading a picture. In fact, with pickle, a file format commonly used in the Python programming language, simply opening the file makes the code hidden inside it run by itself, a bit like a macro in an Office file; and many models are still stored in pickle or in formats built on it. nullifAI, disclosed in February 2025, turned up on the model-sharing platform Hugging Face: the attackers deliberately broke the files, which fooled a dedicated scanner, and on loading, the malicious code had already finished running before the system noticed that the file was broken. The safer choice is the safetensors format, which contains no executable content.
Names can be hijacked too. In September 2025, the security research team Unit 42 demonstrated that after an author account on Hugging Face is deleted, an attacker can register the same name and upload a malicious model, and anyone who fetches models by name will get the malicious version. On the cloud AI platforms of both Google and Microsoft, the team managed to get its own code running on the platforms' machines.
Models get stolen. Stealing a model doesn't necessarily mean stealing its files. Distillation means querying a strong model at scale and then training your own model on its answers, like asking a great teacher millions of questions and then using the notes to teach your own students. In February 2026, Anthropic announced that it had determined that DeepSeek, Moonshot AI and MiniMax had distilled Claude through about 24,000 fake accounts. The next role comes back to this when it discusses protecting models.
Why defences keep losing to attackers who adapt. My view is this: many papers on defences report attack success rates close to zero, but those numbers are mostly measured against a fixed set of test questions, while real attackers first look at how the defence blocks them and then switch to a different approach. According to reports, in October 2025 researchers from OpenAI, Anthropic and Google DeepMind published a paper titled "The Attacker Moves Second", meaning that the defender publishes its method first, and the attacker makes a move after seeing it. They retested 12 recently published defences, and once they switched to attacks that adapt, the attack success rate exceeded 90% against most of them.
Vendors' own figures look better, but you have to read the conditions carefully. In February 2025, Anthropic announced a classifier designed specifically to catch jailbreaks; by its own evaluation, the jailbreak success rate fell from 86% to 4.4%. In November that year, it tested a browser agent, which operates web pages on a person's behalf, with attacks that try again and again, and Claude Opus 4.5 was successfully injected about 1% of the time. Anthropic's own conclusion was that "no browser agent is immune to prompt injection", and that this 1% "still represents meaningful risk".
OpenAI, for its part, has put the red team into training. A red team is a team whose job is to play the attacker and find its own side's weak spots. According to reports, in July 2026 OpenAI announced GPT-Red, a model trained specifically to attack models; by OpenAI's own figures, GPT-5.6 Sol, trained with its help, is fooled by direct injection 0.05% of the time. GPT-Red also attacked a vending machine that is really in business and run by an agent, tricking it into repricing items worth more than US$100 at US$0.50.
So in my view, classifiers can push the success rate down to less than a tenth of what it was, but agents read large numbers of emails and web pages every day, and even if only 1%, or even 0.05%, of attempts succeed, with enough attempts something will go wrong sooner or later.
Architecture is more reliable than filters. The OWASP project leads have said that the model will certainly be fooled (see "14 AI System Security"). That being so, the consensus has turned to a different question: how to stop "being fooled" from turning into "something going wrong".
- Two models with separate jobs: proposed by Simon Willison. The model with permissions never sees untrusted text; the model that reads that text has no tools at all. It is like a company where the person who opens letters from strangers isn't allowed to touch the accounts. The approach from Google, Google DeepMind and ETH Zurich goes a step further: based only on the user's own request, the model with permissions first writes a short program as its plan of action; after that, every time a tool is called, a dedicated program that carries out the plan checks where each piece of data came from and whether it is allowed to go where it is being sent. On AgentDojo, a test suite for agents, it completed 67% of the tasks with security that can be proven. I think the cost is that it can do less, but exactly what it gives up can be worked out in advance.
- Rule of Two: this is the last step in the chart above. Within a single task, an agent may have at most two of these three at once: handling untrusted input, being able to reach sensitive data, and being able to send data out or actually take actions (sending emails, making payments, changing data). When all three are needed, a person has to approve.
- Least privilege: give an agent only the tools and permissions it needs to complete its task. MITRE ATLAS breaks this down into items that can be checked, such as permission settings, having a person confirm the agent's actions, and restricting its use of tools once it has read untrusted data. According to reports, the US National Security Agency's MCP guidance of May 2026 goes further, requiring that whatever each tool sends back be treated as untrusted, and that checks be set up at the door where data leaves the company. That amounts to assuming that your own agents will sooner or later be injected and become insiders sending data out.
Regulation is catching up too. Article 15 of the EU AI Act requires AI systems that the law classes as high-risk to be able to guard against attacks such as data poisoning, model poisoning, inputs designed to make the model go wrong, and theft of confidential information, and to deal with flaws in the model itself as well. According to reports, however, the EU later amended the law, and these obligations have been pushed back to take effect in stages, in December 2027 and August 2028.
Role 4 Frontier Labs Gate Their Cyber Capabilities
This role looks at how the companies that build models deal with "our own model is too good at attacking". The three frontier labs all have public safety frameworks, and their logic is much the same: first define a few capability thresholds, and once a model's evaluations reach one, stronger safeguards have to be applied, including restricting who can use it and more strictly protecting the model weights, that is, the files that make up the model itself.
OpenAI spells out its thresholds most clearly. Its safety framework, which OpenAI calls the "Preparedness Framework", divides cyber capability into two levels. "High" means the model can "remove existing bottlenecks to scaling cyber operations", including automatically carrying out end-to-end cyber operations against reasonably protected targets, or automatically finding and exploiting vulnerabilities that matter in real operations. "Critical" means the model can, without any human involvement, find and develop working zero-day exploits of every severity level in many hardened, real-world critical systems.
OpenAI says GPT-5.3-Codex is "the first launch we are treating as High capability in cybersecurity", and calls this a precautionary decision, meaning it protects the model to the "High" standard in advance. The GPT-5.6 system card of 9 July 2026, the safety evaluation document released alongside a model, rates all three models, Sol, Terra and Luna, as "High" in cybersecurity, but below "Critical".
Google DeepMind ties its threshold to weight protection. The US think tank RAND grades the ability to protect model weights from level 1 to level 5; the higher the level, the stronger the attacker it can stop. Google's Frontier Safety Framework was updated to version 3.1 in April 2026: once a model can substantially help someone launch a high-impact cyber attack, it recommends reaching level 2, plus measures against attackers other than states (such as criminal gangs) and against the company's own insiders. According to media reports, Gemini 4 Argon, which Google launched on 30 September 2026, is going first to vetted defenders through the Fairwind programme, with wider access once the restrictions that prevent misuse mature; a blog post by Google's chief information security officer describes Fairwind as a limited-access programme for governments and trusted partners.
Anthropic's public summary shows no cyber threshold, yet it held a model back. Its Responsible Scaling Policy is currently at version 3.4, in effect since 8 July 2026, and the public summary has no dedicated cyber threshold. Since May 2025, Claude Opus 4 has run under the policy's level 3 safeguards, such as requiring two people's authorisation to access the model weights and limiting the bandwidth for sending data out.
In April 2026, Anthropic decided not to release Claude Mythos Preview publicly, on the grounds that its ability to find and exploit vulnerabilities can "surpass all but the most skilled humans", and instead gave it only to partners in Project Glasswing (see "Role 2 AI Helps Defenders"). The new Mythos 5.1, released in September, is likewise open only to vetted cyber defenders and life sciences researchers at US organisations; the generally available Fable 5.1 "can now be used to find software vulnerabilities, but not to develop exploits".
Open weights are the exception. My view is that all three labs' tiered access rests on one premise: the model sits on their own servers, so they can decide who uses it. Open-weight models lack that premise: once the weights are public, anyone can download them and run them themselves, and there is no taking them back after release. GLM-5.3, mentioned in "Role 1 AI Helps Attackers", is a case in point: the US Center for AI Standards and Innovation (CAISI) estimates that it trails the US frontier by "about four months", and editing the weights directly to strip out the safety mechanisms costs only about US$4,400.
Protecting the weights themselves has its limits too: Anthropic cites RAND's view that the top level, level 5, is "currently not possible". As I see it, however tightly the weights are locked, distillation can still move capabilities out one question and answer at a time, so locking up the weight files and vetting who gets to ask questions at scale both have to be done.
I think tiered access buys time, not safety. Letting defenders get the strongest tools first and find vulnerabilities first does help. But the lead has an expiry date: in May 2026, Anthropic itself estimated that "Mythos-class models will be widely available within the next 6 to 12 months", and CAISI estimates that GLM-5.3 trails the US frontier by only about four months. And as "Role 2 AI Helps Defenders" showed, defenders getting the tools first earns them more vulnerabilities waiting to be fixed.
To me, this means two things. First, design systems on the assumption that within a year attackers will have today's Mythos-class ability to find vulnerabilities, and that every service open to the outside will be scanned. Second, I can't get Mythos 5.1, since it is only for US organisations, so defence can't rely on "getting a stronger model first"; it can only rely on the things above that don't need a frontier model: controlling where data can be sent, giving agents only the least privilege they need, sticking to the Rule of Two, and keeping a close eye on how long it takes from finding a vulnerability to fixing it.
Who Has the Upper Hand, Attackers or Defenders
"The Four Roles of AI in Security" looked separately at how AI helps attack and how it helps defence; this section puts the two sides on the same scales.
The traditional view is that an attacker only needs to find one way in, while the defender has to guard every one. But defenders also have home advantage: they know what systems they have and what those systems normally look like, and they can plant detection points at every stage of an attack. AI amplifies both sides at once, so the short term and the long term need to be looked at separately.
Evidence that favours attackers
First, two terms. A "zero-day" is a vulnerability that is used in attacks before the vendor knows about it or has fixed it; an "n-day" is a vulnerability that has already been made public and has a fix, but which many systems have not yet installed.
- More vulnerabilities are being exploited, mainly because known ones are being turned into weapons faster. At the end of September, Google's threat intelligence team said that the number of vulnerabilities actually used in attacks rose from an average of 10.5 a month in 2025 to 18 a month from January to August 2026. Zero-days grew much less, so Google believes the extra comes mainly from n-days being rapidly "weaponised", that is, turned into programs that can be used in real attacks, very likely with AI's help.
- From disclosure to attack in days, or even hours. A vulnerability in one of BeyondTrust's remote support products was found by an AI agent (an AI program that can plan its own steps and call on tools to get things done), and an attack group was using it just 4 days after it was made public. The 2026 edition of Microsoft's Digital Defense Report says the median time from a vulnerability being discovered in the wild to being weaponised is "well under 24 hours". The full timeline is in "Speed Becomes the Main Measure".
- Patching can't keep up with discovery. According to several media reports, the median time to patch in the 2026 edition of Verizon's Data Breach Investigations Report is 43 days. Anthropic (the company that makes the AI assistant Claude) has also reported that Glasswing, its project with partners to find vulnerabilities using AI, had disclosed 530 high-risk or critical vulnerabilities by late May, but only 75 had been fixed.
- Attacks are getting cheaper, and the bar lower. Anthropic's September threat intelligence report says that autonomy, meaning letting AI carry out a whole series of steps on its own, has pushed down the cost of attacking while the payoff stays the same. The report has a memorable line: "Sophisticated attacks no longer require sophisticated attackers".
- Capability is spreading. An "open-weight model" is an AI model that anyone can download and run themselves. In September, Anthropic analysed GLM-5.3, an open-weight model from the AI company Zhipu, and, citing the US government body responsible for evaluating AI, said it trails the most advanced US models by "about four months". The model would normally refuse to help with harmful tasks, but stripping out that safety layer costs only about US$4,400, and versions with it removed are already circulating.
- Mythos-class models are expected to become widespread. Mythos is a controlled frontier model that Anthropic has offered since April: "frontier" means the most advanced, and "controlled" means not just anyone can use it. In May, Anthropic estimated that "Mythos-class models will be widespread within 6 to 12 months".
Anthropic put it bluntly in April: in the short term, if the companies building the most advanced models aren't careful about how they release them, it may be the attackers who have the upper hand.
Evidence that favours defenders
- In the long run, vulnerabilities may be fixed before software ships. The same piece goes on to say that, in the long term, Anthropic expects defenders to use these models to fix vulnerabilities before new code ships, though the transition may be "tumultuous". Glasswing partners have also started using a preview version of Mythos to rewrite old code in memory-safe programming languages. These languages are designed to rule out a whole class of vulnerabilities, rather like replacing a whole run of leaky old pipes instead of patching the holes one by one.
- Defenders know their own environment. A post on the Google Cloud blog in mid-September argues that attackers armed with AI are still "probing in the dark", while defenders know their own code and what normal looks like, so "AI defence is inherently faster and more accurate than AI offence".
- No fully automated AI attacks yet. In early September, Google's threat intelligence team said it had "not yet seen" attackers use fully automated attack workflows against targets in real environments. Another report from the same team at the end of September described the present as a "window of opportunity": a chance to shore up defences before attackers are able to exploit zero-days and n-days at scale.
- There are already measurable results. In August 2025 the US Defense Advanced Research Projects Agency ran the AI Cyber Challenge. According to media reports, the competing systems automatically found 18 real, previously unknown vulnerabilities, at an average cost of about US$152 per successful task. The top-placed systems have since been released as open source.
- Frontier models go to defenders first. Besides Anthropic's Mythos, OpenAI and Google have since launched controlled frontier models of their own, according to media reports. As I see it, all three follow the same pattern: defenders get to use them first, and you only get access after passing a vetting process.
- Only a tiny fraction is actually exploited. By Google's count, only about 0.23% of the vulnerabilities disclosed in 2026 have actually been used in attacks. When patching, it makes sense to set priorities by signals that "someone is really using this", and to put your staff on the few vulnerabilities that are truly dangerous. The US government's Known Exploited Vulnerabilities list (KEV) is one such signal; see "Taiwanese Brands in KEV".
Here is the evidence from both sides, sorted by timeframe:
| Timeframe | Favours attackers | Favours defenders |
|---|---|---|
| Short term | Known flaws weaponised fasterMore vulnerabilities exploitedPatching lags discoveryOpen models about 4 months behindAttack costs pushed down | Frontier models to defenders firstAI finds and fixes for about US$152Only 0.23% of vulnerabilities exploitedNo fully automated attacks seen yet |
| Long term | Sophisticated attacks need no expertsMythos-class models will spreadDiscontinued devices go unpatched | Defenders know their own environmentVulnerabilities fixed before shippingOld code rewritten in safer languages |
Where both sides should be taken with a pinch of salt
- Most of the figures come from the companies themselves. As far as I can see, Glasswing's vulnerability counts and the efficiency claims for the various AI security products are all measured by the companies themselves.
- Attack capabilities are mostly measured on ranges with nobody defending. A range is a simulated network built for practising attacks, and in the test environments used by Anthropic and Carnegie Mellon University there was nobody actively defending. Even in real attacks, AI is not that reliable: in 2025 Anthropic disclosed an attack operation that used Claude, in which Claude "hallucinated credentials", that is, made up usernames and passwords that didn't exist.
- The people making the claims have their own positions. As I see it, the ones saying "defenders win in the long run" also happen to be the companies selling defensive tools and controlled models. That doesn't mean they are wrong; it just means reading them with each party's position in mind.
My view
In the short term, AI's ability to find vulnerabilities and write attack programs is becoming widespread, but patching has not caught up. During this period, the balance tilts clearly towards attackers in three kinds of places. The first is where patching is slow. The second is where logins can still be phished: phishing means using fake websites or fake messages to trick people into handing over their usernames and passwords, and logins that rely only on a password or a one-time code fall into this group. The third is where nobody controls outbound connections, so a company's computers or AI programs can connect to anywhere outside at will.
The reason is simple: what AI speeds up first is turning known vulnerabilities into weapons and casting a wide net at scale, going straight for vulnerabilities that are "known but unpatched" and accounts that "let you log in once you have them".
In the long term, the balance will tilt towards defenders who "use AI to fix vulnerabilities before shipping", but discontinued devices that nobody patches, and small teams where one person looks after every system, will not benefit automatically. I can't predict the overall balance two or three years from now. What I am more confident about is that who has the upper hand depends on how well each organisation does on the four things below.
- Patching speed. According to media reports, the median time to patch is 43 days, while weaponisation takes well under 24 hours; that gap is the attacker's room to manoeuvre.
- Phishing-resistant logins. Many intrusions are not "break-ins" at all: the attacker gets hold of an account and simply "logs in". Passkeys (stored on your phone and unlocked with your fingerprint or face) and physical security keys (small devices you plug into your computer like a USB stick) only respond to the genuine website address. Even if a fake site looks identical, its address is wrong, so it can't get login credentials that actually work.
- Giving AI agents only the minimum permissions. Google has already seen malware using "prompt injection" against AI scanning tools: hiding a passage of text in a file to trick any AI that reads it into doing something else. A common approach is to give an agent only the tools its task needs, and to allow outbound connections only to addresses on an approved list.
- Records attackers can't alter. Systems automatically record who did what and when; these records are called logs. Spotting an intrusion and investigating afterwards both depend on them, so they need to be stored somewhere attackers can't change them.
Where Taiwan Stands
Taiwan faces a very dense stream of attacks, and networking equipment from Taiwanese brands often appears on the US government's vulnerability list. The law is changing too: three security-related laws were passed one after another in under five months.
How Intense the Threat Is
The National Security Bureau counts 2.63 million attempts a day. According to several media reports, a report published by Taiwan's National Security Bureau in early January 2026 says that in 2025 Taiwan's critical infrastructure faced an average of 2.63 million intrusion attempts a day, 6% more than in 2024. Critical infrastructure means systems whose failure would affect the whole of society; the report focuses on energy and hospitals.
This figure counts "attempts", with traffic and alerts both included, not "successful intrusions", and its size also depends on what is being monitored and how things are counted. The 6% is the bureau's own comparison, and I couldn't confirm what its 2024 baseline was. Another 2024 figure that is often quoted for comparison covers the Government Service Network, a different scope from this report, and I couldn't check that number either, so don't compare the two directly.
Microsoft says Taiwan has the most observed incidents in Asia-Pacific. The 2026 edition of Microsoft's Digital Defense Report states that the United States, Taiwan, the United Kingdom and Israel were the countries with the highest number of observed incidents in their respective regions. This is the part Microsoft can see, not the whole picture.
In July, an AI attack broke into government systems. In July 2026, attackers built an attack system on top of two AI tools, Hermes and OpenClaw, in which several AI programs split the work between them. It broke into 85 accounts on Taiwanese government systems and took more than 2,564 personnel records. The Ministry of Digital Affairs (MODA) confirmed the incident.
Single sign-on lets one username and password reach many systems, like a master key; multi-factor authentication adds a second check on top of the password, such as a confirmation on your phone. In this case, 84 of the 85 accounts were obtained through a connection point that linked single sign-on to the various systems but didn't have multi-factor authentication switched on, which amounts to connecting the weakest door to every room. My view is that the attack tool was new, but the hole it came in through was old.
Hospitals have become ransomware targets. In 2025, Mackay Memorial Hospital and Changhua Christian Hospital were hit one after the other by CrazyHunter ransomware, and their data was stolen and leaked publicly. Ransomware is malicious software that locks up or steals data and then demands money from the victim. Hospital systems can't be allowed to stop, which may be why they have become easy targets for extortionists.
China-linked activity is rising. According to media reports, the 2026 Global Threat Report from the security company CrowdStrike says China-linked hacking activity rose 38%. Of the vulnerabilities this activity exploited, 40% were in edge devices. Edge devices are the doors that stand between an internal network and the internet, such as routers and firewalls; as the next part shows, many Taiwanese-brand products are exactly this kind of device.
Taiwanese Brands in KEV
The US Cybersecurity and Infrastructure Security Agency maintains a "Known Exploited Vulnerabilities" list, KEV for short, which only includes vulnerabilities confirmed to have been used in real attacks. For ordinary people, it is a free and reliable "fix these first" list. Of the 1,733 entries in the 2 October 2026 version of the list, at least 78 are for Taiwanese-brand products.
- Mostly networking equipment. Most of these Taiwanese-brand products are edge and networking devices such as routers and network-attached storage (NAS). A NAS is a small server used at home or in the office to store files.
- Why "at least". Realtek and MediaTek are chipmakers, and their vulnerabilities often travel with their chips into other brands' devices; those devices are not counted here.
- This is not a ranking of who is worse. A more sensible reading is that the more widely a product is installed, and the more often it sits directly exposed to the internet, the more likely it is to be attacked.
Why edge devices are especially dangerous. Firewalls, routers and virtual private network (VPN) devices usually can't run the protective software that detects intrusions. A VPN is an encrypted tunnel for connecting securely back into a company or home network from outside. In 2023 Microsoft described a hacking group called Volt Typhoon that first took control of routers in ordinary homes and small offices, then used them as a route in and out of its real targets. The router in your home could be a stepping stone in someone else's attack.
Homes and small businesses can start with these steps:
- First check whether your model is out of support. Devices that are no longer supported won't get any more fixes; if your model appears in KEV, it should be replaced.
- Don't let people outside reach your router's settings page or your NAS directly. If you need to connect back in from outside, use a VPN.
- Turn on automatic updates and check the firmware at least once a month. Firmware is the low-level software built into a device. Change default passwords and turn off remote features you don't use.
- Small businesses should keep an inventory of their edge devices, and fix whatever is in KEV first. Vulnerabilities are often added to KEV within about a week of being made public, so patching has to be measured in days too, not left for months.
Laws Overhauled Within a Year
Between late August 2025 and mid-January 2026, in under five months, the Cyber Security Management Act, the Personal Data Protection Act and the AI Basic Act were each passed and promulgated. A law only formally exists once the Legislative Yuan, Taiwan's parliament, has passed it at its third reading and the President has promulgated it; but it only really starts to bind anyone on the day it "comes into force". The three laws are in very different states.
The Cyber Security Management Act is already in force
The act had not been amended since it came into force in 2019; this time the whole text was overhauled, and the new version came into force on 1 December 2025.
- The authority in charge is now MODA, and day-to-day enforcement is done by the Administration for Cyber Security, the agency under MODA dedicated to cybersecurity.
- Products that endanger national cybersecurity may not be downloaded, installed or used by government agencies; exceptions require sign-off by the chief information security officer (the most senior person responsible for security), approval from MODA, and a record kept. Organisations such as critical infrastructure providers and state-owned enterprises can also be restricted from using them.
- Heavier fines. Failing to report as required can bring a fine of up to NT$10 million, which can be imposed for each violation.
- Reporting within 1 hour. The reporting regulations under the act, as amended in January 2026, require agencies to report a security incident within 1 hour of becoming aware of it.
- An anti-scam drill every six months. Social engineering means breaking in by deceiving people rather than through technical means, for example by sending emails pretending to be from a manager. Agencies must run a social engineering drill every six months.
One hour is very tight. For comparison abroad, according to media reports, since November 2025 China has required "especially serious" incidents to be reported within 1 hour, while the EU's Cyber Resilience Act requires manufacturers to make a first notification within 24 hours.
The Personal Data Protection Act has been promulgated but is not yet in force
It was promulgated on 11 November 2025, but the date it comes into force is for the Executive Yuan, Taiwan's cabinet, to set, and as of the 18 September 2026 version of the official laws database, no date had been set. The authority in charge will be an independent Personal Data Protection Commission, which is still being set up.
Personal data breaches that fall within its scope must be notified to the people affected, and inadequate security can bring a fine of up to NT$15 million each time. According to media reports, the office preparing the commission has published draft detailed rules proposing that, after a breach, the people affected be notified and the authority informed within 72 hours.
The AI Basic Act sets out principles only, with no penalties
It was promulgated on 14 January 2026. The act is overseen by the National Science and Technology Council, but the framework for classifying AI by risk is left to MODA. Risk classification sorts uses of AI into tiers by how risky they are; usually, the higher the risk, the more requirements apply. On the security side, one of the seven principles in the act is "cybersecurity and safety". The concrete obligations will appear in detailed rules issued later, and I think MODA's classification framework is the one most worth following.
Ecosystem and Resources
- Companies hit by a security incident can turn to TWCERT/CC. The Taiwan Computer Emergency Response Team/Coordination Center (TWCERT/CC) handles incident reporting and vulnerability coordination for private companies, and is run by the National Institute of Cyber Security, set up in 2023.
- To get to know Taiwan's security community, HITCON is a good place to start. It is a hacker community that has been running since 2005 and holds a conference every summer; "hacker" here means someone who digs deep into technology, not a criminal.
- For a snapshot of the whole industry in one place, go to CYBERSEC, Taiwan's cybersecurity conference. It is a large security conference held every spring.
- To follow the latest developments yourself, you can read the national cybersecurity situation report that MODA submits to the Legislative Yuan every year, TWCERT/CC's vulnerability advisories, and the KEV catalogue mentioned earlier. Careers and certifications are covered in "Industry, Careers and a Learning Path".
Industry, Careers and a Learning Path
After the tour of twenty-two domains, people often ask two follow-up questions: where is the money in this industry going, and where should someone who wants to get in start? This section looks at the industry, the roles and the certifications first, and ends with a learning path that I think is fairly solid.
The Industry Is Consolidating
Over the past decade, companies kept buying more and more security tools: one vendor for the firewall, another for account management, yet another for collecting system records, each with its own screens and its own alerts. As I see it, from 2024 to 2026 things have gone exactly the other way: the big vendors are pulling these tools into a single platform, which the industry calls "platformisation".
SIEM was the first to be absorbed. Security information and event management (SIEM) is a system that collects the records kept by computers and devices all over a company, compares them in one place and raises an alert when something looks suspicious, rather like a building's central control room. In security circles these records are called "logs". In March 2024, Cisco completed its acquisition of Splunk for about US$28 billion.
Identity (that is, accounts, logins and permissions) and the cloud each saw a deal worth tens of billions of dollars. In February 2026, the security company Palo Alto Networks completed its acquisition of CyberArk, a vendor that manages privileged accounts (accounts with sweeping powers, such as administrator accounts) and machine identities (the identities programs use to log in to one another). In March 2026, Google completed its US$32 billion acquisition of the cloud security vendor Wiz.
Data pipelines have become a battleground too. A data pipeline is the layer that filters and tidies logs before they are sent into the SIEM, and back in September 2025 the security company SentinelOne bought Observo AI, a company of exactly this kind. As I understand it, the more data a SIEM takes in, the more it costs, so whoever can cut the volume down before the data comes through the door controls the way in.
The next step is moving AI agents in. An AI agent is an AI program that can look things up, make judgements and take action by itself, rather than just answering questions. On 23 September 2026, Microsoft released a preview of a new product that brings its own SIEM together with its other protection products, and described it as a foundation built for AI agents to take over security work.
Putting these deals together, I think the money is flowing mainly to four areas: identity, the cloud, data (including SIEM and data pipelines) and AI. For anyone who wants to get into the field, this means two things:
- Platforms change; the underlying knowledge doesn't. The product you learn today may have been bought up or renamed three years from now. Public frameworks that belong to no vendor last longer, such as MITRE ATT&CK, a knowledge base that catalogues the techniques attackers commonly use.
- There will be more work joining things up. What companies want is people who can connect accounts, computers, the cloud and logs and look at them together, not people who only know one tool.
A Map of Roles
There are a few common ways into the field:
- Working up from the security operations centre. A security operations centre (SOC) is the team whose job is to watch alerts and be the first to respond when something goes wrong. Its analysts are usually split into three tiers: tier 1 decides whether an alert is real or a false alarm, tier 2 investigates and contains incidents, and tier 3 actively hunts for intrusions that haven't yet set off an alert. Beyond that, you can move into incident response and forensics, or switch to threat intelligence, which means studying who the attackers are and which techniques they favour.
- From penetration testing to red teaming. Penetration testing means being hired to attack a client's systems within an agreed scope, with the aim of finding as many weaknesses as possible; a red team instead sets itself one specific goal and tries hard not to be noticed, so what it really tests is whether the defenders can spot it.
- Moving across from other jobs. People who can code have the easiest move into application security; people who have managed servers and accounts are well suited to cloud or identity work; and people who have done IT auditing can go into governance and compliance, and further up to chief information security officer (CISO). For people with a business or law background, I think governance and compliance, or privacy, is the best way in.
Mapping the seven families from earlier onto common roles and skills gives roughly the picture below. If some of the terms are unfamiliar, feel free to skip over them; what matters is how the two columns line up.
| Family | Common roles | Core skills |
|---|---|---|
| Family 1 Governance and Compliance | Governance and compliance analystAuditorPrivacy engineerCISO | Understanding frameworks and regulationsRisk assessmentOrganising audit evidenceExplaining risk clearly to managers |
| Family 2 Identity and Cryptography | Identity management engineerCryptography and certificate engineer | Setting up single sign-onControlling administrator accountsAutomatic certificate renewalPost-quantum cryptography migration |
| Family 3 Infrastructure | Network security engineerCloud security engineerComputer and phone security engineer | Network segmentationCloud permissions and configurationProtecting and hardening computersBackup and restore drills |
| Family 4 Software and AI Systems | Application securitySupply chain securityVulnerability managementAI security engineer | Code reviewThreat modellingHardening the build processPrioritising vulnerability fixesDefending against prompt injection |
| Family 5 Detection and Response | SOC analystDetection engineerIncident response and forensicsThreat intelligence analyst | Reading logsWriting detection rulesBuilding incident timelinesMapping to ATT&CK |
| Family 6 Offensive Security | Penetration testerRed teamerVulnerability researcherBug hunter | Web attacksGaining administrator accessReverse engineeringClearly written reports |
| Family 7 The Physical World and People | Industrial control security engineerMedical and automotive securitySecurity awareness and fraud prevention | Industrial networks and protocolsIndustrial security standardsSocial engineering drillsCall-back checks before payments |
Besides industry consolidation, two other forces are reshaping the job market: regulation and AI.
Taiwan's amended Cyber Security Management Act, which took effect on 1 December 2025, requires government agencies to assign dedicated cyber security staff according to their responsibility level (the law sorts agencies into several levels by how important they are); details are in "Laws Overhauled Within a Year". The reporting obligations under the EU's Cyber Resilience Act also started to apply on 11 September 2026: when a vulnerability in a product is used in an attack, the manufacturer must report it within a set deadline. For Taiwanese companies selling hardware into Europe, I think the demand for product security staff will only grow.
The first change AI brings is that new roles are springing up:
- AI red teaming: testing whether an AI can be fooled by "prompt injection", which means hiding malicious instructions in a web page or document the AI will read. According to figures from the bug bounty platform HackerOne, reports of prompt injection vulnerabilities rose by 540% between July 2024 and June 2025.
- Agent security: designing identities, permissions and activity records for AI agents, so that people can check afterwards what an agent did. The standards here aren't settled yet, and I think that is exactly where people are in short supply.
The second change is that the work of tier 1 SOC analysts is starting to be handed to AI. In March 2025, Microsoft announced a set of agents for Security Copilot, its AI security assistant, one of which triages phishing emails; in April, Google announced it was adding an alert triage agent to its own security platform. Triage means deciding first which alerts are worth a person's attention; details are in "Role 2 AI Helps Defenders". My view is that entry-level work won't disappear, but it will shift from opening alerts one by one to checking the AI's verdicts and catching its mistakes.
Certifications and Skills
The common certifications for each direction are listed below. Most of the names are acronyms and there's no need to memorise them; if you haven't chosen a direction yet, just look at the "Entry level" row.
| Direction | Common certifications | Matching family |
|---|---|---|
| Entry level | ISC2 CC, CompTIA Security+ | Any family |
| Defence | CompTIA CySA+, GIAC GCIH | Family 5 Detection and Response |
| Offence | OSCP+, OSEP | Family 6 Offensive Security |
| Management and audit | CISSP, CISM, CISA | Family 1 Governance and Compliance |
| Cloud | CCSP, the cloud providers' own security certifications | Family 3 Infrastructure |
| Taiwan-specific | The government's cyber security competency assessment, iPAS Information Security Engineer | Any family |
The CISA in the table is an auditing certification that happens to share its name with CISA, the US cyber security agency, so don't confuse the two. OSCP was renamed OSCP+ in November 2024 and is now valid for only three years, so passing the exam doesn't mean you're done.
My view is that a certification helps your CV get through the first round of screening, especially in the public sector and at large companies, but what interviewers really want to know is what you have done. A public write-up of how you solved a challenge, or a small tool you've put on GitHub (a website where engineers share code publicly), is more persuasive than one more certificate.
There are also a few skills that no certification tests but every role needs:
- Writing: almost all security work ends up as a report.
- Automation: being able to turn repetitive tasks into small Python programs, and also being able to get AI to write them, then checking for yourself whether it got them right.
- Speaking the language of risk: if you tell a manager "this vulnerability has a severity score of 9.8", they won't know what that means. You need to turn it into "if hackers broke into this device that lets staff connect to the office from outside, what would we lose?"
A Learning Path
If I had to give someone starting from zero a path, it would be the five steps below. First, one rule that applies whichever route you take: before attacking any system, you must have written authorisation and a clearly defined scope. Without authorisation it is a crime, and in Taiwan it falls under Chapter 36 of the Criminal Code, "Offences Against Computer Security". If you want to practise, do it only in a practice environment you have set up yourself, in capture the flag contests (CTFs, puzzle-style attack and defence competitions), or in bug bounty programmes with clearly written rules (where a company publicly invites outsiders to look for vulnerabilities and pays a reward for each one reported).
Networking, operating systems, basic Python
Collect logs, write detection rules, map them to ATT&CK
Learn web attacks from the OWASP list and practise with CTFs
Choose a direction based on your background or on what the industry needs
Read KEV and incident reports, write things up, share them with the community
1 Build the foundations. Understand how networks move data around, understand accounts, permissions and logs on Linux and Windows, and be able to write at least Python. This step is the most boring and the one most often skipped, but every later step is built on top of it.
2 Walk the blue team track. The blue team is the defending side. Set up a small lab on your own computer using virtual machines (several simulated computers running inside one real one), deliberately do some suspicious things, and see what the logs record. Next, learn to write detection rules, that is, conditions you give the SIEM, such as "raise an alert if the same account fails to log in many times in a short period". Sigma is an open rule format that isn't tied to any vendor, and every rule you write should be labelled with the ATT&CK technique it corresponds to. Finally, use CALDERA, MITRE's open-source attack emulation tool, to play the attacker and see whether your rules catch it.
3 Walk the red team track. Here "red team" simply means the attacking side. For web security, start with the 2025 edition of the Top 10 risks from the Open Worldwide Application Security Project (OWASP), where number one is still "broken access control", meaning users can reach data they shouldn't be able to touch. CTFs are the best place to practise: picoCTF suits beginners, and Taiwan's HITCON CTF makes a good advanced goal. You need to walk both tracks, because if you don't know how attackers move, your detection rules won't be accurate, and if you don't know what defenders can see, the test reports you write for clients as an attacker won't get to the point.
4 Pick a specialism. Choose from cloud, identity, application security, industrial control systems (the systems that control machinery in factories and power stations), AI, or governance and compliance. You can base the choice on your existing background, or on where the money is flowing: the deals and product moves covered earlier point to identity, cloud, data and AI.
5 Keep up to date. I'd suggest regularly reading the Known Exploited Vulnerabilities catalogue (KEV) from the US Cybersecurity and Infrastructure Security Agency. It lists only vulnerabilities that are confirmed to have been used in attacks, and the number added this year already exceeds the total for the whole of last year (see "Speed Becomes the Main Measure"). Whenever new entries appear, check them against the software and devices you use. Write up what you read, then share it in communities such as HITCON (Taiwan's hacker community and conference; see "Ecosystem and Resources").
Bug bounties are real-world practice, and a source of income too. HackerOne paid out US$81 million in rewards between July 2024 and June 2025. But AI is changing this market as well: in June 2025, XBOW, an autonomous AI penetration testing tool, reached the top of HackerOne's US leaderboard. I think simple vulnerabilities will be snapped up by machines faster and faster, and the value people bring will shift to spotting flaws in a system's own rules (for example, being able to use the same coupon over and over again), or to chaining several small weaknesses into a single route of attack. For more, see "19 Vulnerability Research and Bug Bounties".
AI can also be a tutor: it can explain logs you can't make sense of, walk you through how a vulnerability works, and help you set up a lab. According to results Anthropic itself published, Claude ranked in the top 3% at the beginner-level picoCTF 2025, but did not solve a single challenge at PlaidCTF or the DEF CON CTF qualifier (tough competitions entered by top teams). It is strong on beginner challenges, but still a long way off at the level of elite competitions.
The models ordinary people can use also have limits. Anthropic said in September that its general-release model, Fable 5.1, can help find software vulnerabilities but cannot be used to write attack code; Mythos 5.1, which it describes as its model with the strongest security capabilities, is available only to vetted security defenders at US organisations (and also to life sciences researchers). My reading is that AI can help you find and understand vulnerabilities, but writing attack code is something you still have to practise yourself, within an authorised scope.
This leads to two reminders:
- You only learn by solving it yourself. AI gets through beginner challenges very quickly; what really helps you is the process of getting stuck, looking things up and working it out. Do it yourself first, ask when you get stuck, and once you've asked, write it down in your own words.
- Always verify. AI gets things wrong, and says them with great confidence. When Google open-sourced Mantis, its framework for finding and fixing vulnerabilities, in September, it stated plainly that sloppy AI code scanning produces "hallucinated vulnerabilities" (ones that don't exist at all), with a true positive rate below 7%. Never submit an AI finding you haven't verified.
Back to the questions at the start. The industry is consolidating, and front-line work is starting to be handed to AI. My judgement is that the security people most needed over the next few years will be those who can direct AI agents to do the work and also check whether they have done it right, and the more solid your foundations, the better you will be at catching their mistakes. So if you want to get into the field, build solid foundations first, walk both the defending and the attacking tracks, and then pick a direction.
What Individuals and Teams Can Do
Security is a vast field, but for individuals, families and small teams, most intrusions still begin in a few familiar places: a login someone was tricked into handing over, a device that was never patched, a fake phone call. And when something does go wrong, recovery often gets stuck on a backup that has never been restored. Below I start with the five things most worth doing, then split the advice by who you are; apart from the "Engineers" part, none of it requires any programming.
If you only do five things
- Switch your important accounts to passkeys. A passkey is a way of signing in that uses the unlock on your phone or computer instead of a password. It works like a key that only opens the door of the real website: the browser checks the web address, so even if a fake site tricks you into using it, the fake site gets nothing it can use. Text message codes, the six digits from an authenticator app, a prompt that pops up on your phone asking you to tap "Allow": all of these forms of multi-factor authentication (an extra check on top of your password) can be passed along by a fake site in real time. You type the code into the fake site, and at that same moment the attacker uses it to sign in to the real one. Start with your email, because the "forgot password" emails for all your other accounts are sent there; then your Apple or Google account, your bank and brokerage accounts, and your cloud drive and photo library; social media accounts come last.
- Use a password manager, with a different password for every site. If you use the same password in many places, then as soon as any one site leaks it, attackers will take it and try it on other sites one by one; this is called credential stuffing. Verizon's 2026 Data Breach Investigations Report found that stolen usernames and passwords still featured in 39% of breaches.
- Turn on automatic updates, and replace routers and NAS devices that are no longer supported. The Wi-Fi box in your home is a kind of router, and a NAS is a network drive for storing files. The US government's cybersecurity agency keeps a Known Exploited Vulnerabilities catalogue (KEV), which only includes vulnerabilities that someone is confirmed to have used in an attack. As of 2 October 2026, at least 78 entries in it were products from Taiwanese brands, mostly networking equipment and NAS devices (see "Where Taiwan Stands"). Once a device is no longer supported, any new vulnerability found in it will never be patched, so the only option is to replace it.
- Keep at least one backup offline or locked against changes, and make sure you have actually restored from it. Offline means something like an external hard drive that normally stays unplugged; locked means that once data is stored, nobody can delete it for a set period. Ransomware is malicious software that locks or steals your files and then demands a ransom. In its M-Trends 2026 report, the security firm Mandiant observed that ransomware gangs now go after backups first, so that victims have nothing to restore from. That is why you need to actually do a restore once: only then do you know whether you can get your files back.
- If a call asks for money or login details, hang up first, then call back on a number you looked up yourself. In 2024, an employee at the firm Arup in Hong Kong transferred HK$200 million during a video call in which everyone apart from him was a deepfake (a fake face and voice made with AI). I think fake voices will only get harder to tell apart, so you have to rely on process: whoever the caller is and however urgent it sounds, confirm through a different channel that you control.
Individuals and families
- Be wary of calls asking you to "set up a passkey": in August 2026, Google's threat intelligence team revealed that attackers had been phoning office workers on their personal mobiles, saying they urgently needed to register a passkey, steering them to a fake website, and then stealing their login details and tokens (a token is a temporary pass a website gives you after you sign in). My view is that passkeys themselves have not been cracked; what is being exploited is the "registration" and "reset" process. A legitimate service will not phone you to hurry you to some web address to set one up.
- A web page that tells you to paste a command is a trap: fake verification pages instruct you to paste a piece of text into the little "Run" box (the one that pops up when you press Win+R on Windows) and then press Enter, which amounts to running the attacker's ready-made command with your own hands. Microsoft counted such commands being run on more than 1.1 million devices between February and May 2026.
- The fewer browser extensions, the better: extensions are small add-ons installed in your browser. Microsoft mentions a malicious extension, installed more than 600,000 times, that collected users' ChatGPT and DeepSeek conversation histories.
- Network devices at home: keep the firmware (the software built into the device) on your router and network cameras updated, and change the default passwords.
- Don't treat cloud sync of your phone's photo library as a backup: photos you delete by mistake on your phone, and files locked by ransomware, get synced to the cloud along with everything else.
- A family code word: agree with your family on a check question that only your own people know the answer to, and ask it first whenever you get a call saying "I'm in trouble, send money now".
Small businesses and teams
- Start with a list: which devices, company web addresses (domains) and cloud accounts can be reached by anyone on the internet, and who is responsible for each.
- Phishing-resistant sign-in for everyone, such as passkeys, starting with administrators: in July 2026, an attack tool in which several AI agents (AI programs that can plan their own steps and operate tools) divided up the work broke into Taiwanese government systems. Taiwan's Ministry of Digital Affairs confirmed that of the 85 accounts it obtained, 84 were reached through a "single sign-on" entrance (one username and password that opens many systems) that did not have multi-factor authentication switched on.
- The help desk must verify identity before resetting a password: the standard trick of the group known as Scattered Spider is to phone a company's IT help desk, pose as an employee and ask for a password reset. That one trick is how the British retailer Marks & Spencer ended up with its online shopping halted for about 46 days. Google recommends checking ID over video or in person, and not using dates of birth or the last few digits of a national ID number for verification, because attackers already have all of these.
- For internet-facing devices, first fix the vulnerabilities on the Known Exploited Vulnerabilities catalogue (KEV): firewalls (the filtering devices that stand at the entrance to your network), VPNs (the tunnels that let staff connect back to the company from outside) and the management tools IT uses to operate computers remotely come first. In Verizon's report, the median time companies took to patch vulnerabilities lengthened from 32 days to 43 days, and until a fix is in, the door stays open.
- Keep logs where attackers can't delete them: keep a separate copy of sign-in logs (who signed in, when and from where) in an account or service that attackers can't touch even once they are inside the company's systems. M-Trends 2026 found that the median time from an intruder getting in to being discovered was 14 days. My view is that if you only keep logs for a week, many incidents simply can't be traced back.
- Write down in advance who does what, and by when, on the day something goes wrong: who decides to disconnect from the network, who explains things to the outside world, and which forensics firm you call in to work out what happened. Under the reporting rules that accompany Taiwan's Cyber Security Management Act, government agencies and certain non-government bodies (for example, critical infrastructure operators and state-owned enterprises) must report a security incident within 1 hour of becoming aware of it. Companies outside its scope can still use this 1 hour as a target in their drills.
- Regularly clean up the permissions given to third-party apps: when you sign in to an app with your work account and press "Allow access", that app is given a token, and from then on it can get in without a password. In August 2025, attackers stole the tokens held by an app called Salesloft Drift and used them to export large amounts of the data that companies kept in Salesforce (a cloud system businesses use to manage customer data). Revoke any permissions that are no longer in use.
- Training is a supplement, not the main defence: a randomised controlled trial at the University of California San Diego found no link between how recently someone had taken their annual security training and whether they clicked on phishing links. Put the budget first into phishing-resistant sign-in, help desk identity checks and two-person approval for payments (a payment can only go out once two people have each confirmed it).
Engineers
Most of the items below defend against supply chain attacks, where instead of attacking you directly, the attacker contaminates software you are going to install, such as packages (ready-made pieces of code written by other people that you install and use).
- Pin GitHub Actions to full commit SHAs: GitHub Actions is a tool for automatically testing and deploying code, and it can plug in small components written by other people. A version tag is like a sticky note: it can be peeled off and stuck onto different code. A commit SHA is the fingerprint of that exact version of the code, and it can't be swapped out. In March 2025, one component's tags were moved onto malicious code, and more than 23,000 projects were using it.
- Wait a few days before installing a new version: guidance from Google and Mandiant in September 2026 recommends a cooldown period of at least 7 days (set with
min-release-agein npm andexclude-newerin uv). LiteLLM is a package often used to connect to several AI models through one interface; in March 2026, malicious versions of it sat on Python's public package index for only 2 hours 32 minutes, yet were installed 119,000 times. Anyone with a 7-day cooldown period in place would never have installed them. - Don't keep long-lived tokens: if a token stays valid indefinitely, then once it leaks, attackers can keep using it to publish new versions in your name for a long time. Switch to trusted publishing instead, where each automated upload is given a fresh temporary pass that expires quickly.
- Give AI coding assistants (coding agents) only the minimum permissions: they change code and run commands on their own. Give them only the permissions the task in hand needs, and don't run them anywhere that holds the login details or keys for your live production systems.
- Check that packages recommended by AI actually exist: a study presented at USENIX Security 2025, an academic security conference, found that 19.7% of the packages recommended by AI coding tools did not exist at all. Attackers get in first and publish malicious packages under these names, so anyone who installs what the AI suggested ends up installing the attacker's code.
People who use AI
- Don't paste passwords, keys or other sign-in strings, or customers' personal data, into a chat box: conversation histories themselves can be stolen, and the malicious extension mentioned above is one example.
- First check which accounts an AI agent can reach: when you connect an agent to your email, cloud drive or calendar, don't give it write permissions it doesn't need, such as sending emails or changing files. Microsoft 365 Copilot once had a vulnerability where a specially crafted email only had to be read by it for data to be sent out, without the user clicking anything. Microsoft has since fixed it.
- The Rule of Two: this is a rule proposed by Meta. In any single task, give an AI at most two of these three things: reading content of unknown origin (such as an email from a stranger), access to sensitive data, and the ability to act on the outside world or change things. When all three come together, someone only has to hide an instruction aimed at the AI inside an email, and the AI may follow it and send your data away. If a task really needs all three, a person has to approve it.
- Even the best defence only lowers the odds: Anthropic has published its own results showing that when Claude Opus 4.5 was operating a browser on someone's behalf, against an attacker who kept adjusting their technique as they went and made 100 attempts in each test scenario, the attack success rate was 1%. Anthropic also says this is still a meaningful risk.
- Companies need a shadow AI policy: shadow AI means AI tools that employees use on their own without the IT department knowing. I think that rather than banning them outright, it is better to provide an approved tool that keeps records, and to spell out which data must not be put into it.
Closing Thoughts
Having worked my way through the whole field of cybersecurity, my biggest takeaway is this: most attacks are still the old tricks, just much faster. The major incidents of 2025 and 2026 mostly still came in through the old places: network edge devices such as firewalls that had not been patched, stolen usernames and passwords, and a phone call to a company's help desk.
The attack that broke into Taiwanese government systems in July 2026 used several AI agents (AI programs that can carry out a task step by step on their own), each with its own part to play. But 84 of the 85 accounts came in through the same single sign-on entrance, which did not have multi-factor authentication switched on. Single sign-on means one account gets you into many systems; multi-factor authentication means confirming on your phone as well as typing your password.
What has really changed is speed. According to several media outlets citing the 2026 Global Threat Report from the security company CrowdStrike, criminal groups take an average of just 29 minutes from breaking in to starting to spread to other computers inside the network. So my view is that the basics, such as patching devices and protecting accounts, have not changed; they all just have to happen faster, and whatever can be handed over to machines should be.
The real door is "who you are": if you can pass yourself off as someone, you can walk into a company's systems, and the IT help desk that resets people's passwords is one of the gatekeepers. In the security company Mandiant's M-Trends 2026 report, voice phishing (phoning people and pretending to be someone else to trick them) is the second most common way in. Separately, Google's threat intelligence team tracked a hacking group that phoned employees directly on their personal mobiles, using an "urgent passkey registration" as the pretext. A passkey is a way of logging in with your phone's fingerprint or face recognition instead of a password.
If you run your own systems, you are the only one who can reset passwords, which makes you your own help desk. I think even a request that seems to come from "you yourself" should be checkable, for example by hanging up first and then calling back on a number registered in advance.
Salesloft Drift is a third-party tool that plugs into Salesforce, the cloud system many companies use to manage their customer data. The attackers stole its tokens, that is, the authorisations a company issues to the tool when someone clicks "Allow". According to several media reports, this let the attackers take data from more than 700 organisations without needing a single password. Passkeys protect the moment you log in; the processes for resetting passwords, registering new devices and authorising third-party apps are where the gaps are now.
AI agents also have to log in and be given permissions to do their work, which amounts to adding another employee account. My view is that every AI agent should have an account of its own, its tokens should expire after 10 minutes, and its permissions should cover only what its job needs.
AI makes finding vulnerabilities cheap; fixing them is the bottleneck. The AI company Anthropic has itself announced that its Glasswing programme, through which it lets partners hunt for vulnerabilities, found more than 10,000 vulnerabilities rated high severity or worse in its first month. But according to the figures it published in May, of the 530 reported to open-source projects (software whose code is public and that anyone can use), only 75 had been fixed.
I think installing a few more tools that hunt for vulnerabilities automatically will only get you a longer list, more than you could ever fix yourself. What helps more is making the list shorter: for example, relying less on ready-made code components written by other people that you simply drop in, and keeping systems off the internet if outsiders have no need to reach them.
Only after drawing the whole map can you see how wide this field is. The security I had in mind was mostly a set of technical questions: where AI agents can connect to, which data they can touch, and whether they leave a record. Only when I laid out all twenty-two domains did I find a few areas I had barely thought about: governance (who makes decisions and how often risks are reviewed), backups (attackers delete or lock the backups first, so you have nothing to restore from), incident response (who does what in the first hour after something goes wrong), and people (the phone call to the help desk described above was a con aimed at a person). Most of these are not solved by writing code, but they still need to be written down as rules and rehearsed regularly.
About the Sources
The numbers come from public annual reports, official datasets, and articles published by vendors and research groups themselves. For stories still developing, the picture is as of 3 October 2026; the status of Taiwanese law is based on data as of 18 September; and the Glasswing patching figures are the ones Anthropic published in May.
Each report draws on a different sample: some count only confirmed data breaches, some count only the intrusions that the publishing company handled itself, and some come from surveys or from what their own products detected. That is why the numbers often fail to match, and I never compare percentages from two different reports directly. Results published by vendors are set, marked and published by the same company, so please treat them as that company's own claims.
The network I did my research from could not reach some official websites or most Taiwanese media, so I used public copies preserved on other websites instead, and cross-checked them against reports from vendors and research groups. I did not read the full annual reports from Verizon, IBM, CrowdStrike or the European Union Agency for Cybersecurity; the numbers taken from them rely on consistent reporting across several media outlets. Some numbers I counted myself from public raw data. Counting by vendor name or keyword is bound to miss or miscount a few, so I write "at least" or "about".
As of 3 October, these things are still uncertain or still changing:
- US rules on reporting security incidents: these require operators of critical infrastructure, such as power companies and hospitals, to report to the government within a set time after an incident. According to a single report, the final version was sent to the White House Office of Management and Budget for review around 1 October and is expected to be published before the end of the year; the deadlines may still change.
- Amendments to Taiwan's Personal Data Protection Act: the mirror copy of Taiwan's national Laws and Regulations Database is synced up to 18 September 2026, and the amended provisions are still shown as "not yet in force", meaning they have not taken effect. I have not yet been able to find out whether the Personal Data Protection Commission has been formally set up.
- What AI companies say about their own cyber capabilities: the security programmes and models announced by OpenAI and Google are, for now, backed only by the companies' own statements and press coverage, with no outside testing yet; see "Role 4 Frontier Labs Gate Their Cyber Capabilities". The Glasswing patching figures are also bound to change.
- Figures on the size of the industry: I could not find original sources I could check for the output value of Taiwan's security industry, the shortage of security workers or the size of the global security market, so the text leaves them out.
References
Frameworks and Standards
- NIST CSRC: NIST Revises SP 800-61 (2025-04-03)
- NIST CSRC: NIST Releases Revision to SP 800-53 Controls (2025-08-27)
- NIST: NIST Offers 19 Ways to Build Zero Trust Architectures (2025-06)
- CIS: CIS Critical Security Controls v8.1
- MITRE ATT&CK: ATT&CK STIX data releases, latest v19.2 (2026-08-05)
- OWASP: OWASP Top 10:2025 repository
- FIRST: Common Vulnerability Scoring System, CVSS v4.0
Annual Threat Reports
- Verizon: 2026 Data Breach Investigations Report (2026-05-19)
- Verizon: 2026 DBIR news release (2026-05-19)
- Mandiant: M-Trends 2026 (2026-03-23)
- CrowdStrike: 2026 CrowdStrike Global Threat Report (2026-02-24)
- Security Boulevard: How Much Does a Data Breach Cost? IBM's 2026 Report Puts the US Average at $11.5 Million (2026-08)
- CyberScoop: IBM's Cost of a Data Breach Report 2025 (2025-07)
- ComplexDiscovery: ENISA Threat Landscape 2026 Finds DDoS Leads Incident Counts While Ransomware Stays Most Impactful in the Short Term (2026-09)
- Microsoft: Microsoft Digital Defense Report 2026 (2026-10-01)
- Microsoft: Microsoft Digital Defense Report 2025 (2025-10)
- Sophos: The State of Ransomware 2026
- Chainalysis: Ransomware payments in 2025, from the 2026 Crypto Crime Report
- Google GTIG: 2025 Zero-Day Review (2026-03-05)
- Google GTIG: 2024 Zero-Day Exploitation Analysis (2025-04-29)
Incidents and Threat Intelligence
- Google GTIG: Updated Cyber Threat Actor Naming System (2026-07)
- Google GTIG: The Cost of a Call: From Voice Phishing to Data Extortion (2025-06-04)
- Google GTIG: Widespread Data Theft Targets Salesforce Instances via Salesloft Drift (2025-08)
- Google GTIG: UNC6671 Targets Financial Services and Enterprise Cloud Environments (2026-08-06)
- Google GTIG: ShinyHunters Renewed Mass Exploitation Campaign Targeting Oracle PeopleSoft (2026-09)
- Microsoft: Disrupting Active Exploitation of On-Premises SharePoint Vulnerabilities (2025-07-22)
- TechCrunch: Instructure Strikes Deal with Hackers Who Breached It Twice (2026-05-12)
- Cyber Monitoring Centre: Statement on the Jaguar Land Rover Cyber Incident (2025-10-22)
Identity, Network, Endpoint and Cloud
- Microsoft: Pushing Passkeys Forward: Microsoft's Latest Updates for Simpler, Safer Sign-ins (2025-05-01)
- Microsoft: Inside Tycoon2FA: How a Leading AiTM Phishing Kit Operated at Scale (2026-03-04)
- Model Context Protocol: MCP specification and auth extensions, current 2026-07-28 revision
- Microsoft: Microsoft Entra Agent ID
- Palo Alto Networks: Palo Alto Networks Completes Acquisition of CyberArk to Secure the AI Era (2026-02-11)
- Microsoft: Windows Resiliency Initiative
- Mandiant: UNC5537 Targets Snowflake Customer Instances for Data Theft and Extortion (2024)
- GitGuardian: The State of Secrets Sprawl 2026
- Kubernetes: ingress-nginx README and retirement notice
- Google: Wiz acquisition completed (2026-03-11)
- Apple: Private Cloud Compute security research source
Cryptography and Post-Quantum
- CA/Browser Forum: TLS Baseline Requirements v2.3.0 (2026-09-07)
- Let's Encrypt: Website and blog sources
- OpenSSH: openssh-portable source and config docs
- OpenSSL: OpenSSL source and NEWS
- Apple: Get ahead with quantum-secure cryptography, WWDC25 (2025-06)
- Google Cloud: Announcing Quantum-Safe Digital Signatures in Cloud KMS (2025-02-21)
- Apple: swift-homomorphic-encryption library
AppSec, Supply Chain and Vulnerabilities
- Google GTIG: Vulnerability Discovery and Exploitation Trends in the AI Era (2026-09-30)
- Datadog: State of DevSecOps 2026 (2026-02)
- Google GTIG: Mitigation Guidance for Supply Chain Compromise
- Microsoft: Mitigating the Axios npm Supply Chain Compromise (2026-04-01)
- SLSA: SLSA specification v1.2 (2025-11-24)
- CycloneDX: CycloneDX specification 1.7.2 (2026-09-17)
- OpenSSF: Malicious packages database
- Linux kernel: Kernel source tree, Rust docs and .rs files
- USENIX Security 2025 (Spracklen et al.): We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (2025)
Detection, Response and Offensive Security
- Microsoft: Reimagining the SOC for the Agentic Era in Microsoft Defender (2026-09-23)
- SiliconANGLE: Cisco Completes $28B Acquisition of Splunk (2024-03-18)
- SigmaHQ: Sigma detection rule repository
- MITRE: CALDERA automated adversary emulation platform
- HackerOne: HackerOne Report Finds 210% Spike in AI Vulnerability Reports Amid Rise of AI Autonomy (2025)
- SC World: Google Awards Over $17 Million to Security Researchers in 2025
- TechCrunch: Price of Zero-Day Exploits Rises as Companies Harden Products Against Hackers (2024-04-06)
OT, IoT and the Human Factor
- Mandiant: Sandworm Disrupts Power in Ukraine Using a Novel Attack Against Operational Technology
- BleepingComputer: FrostyGoop Malware Attack Cut Off Heat in Ukraine During Winter (2024-07)
- BleepingComputer: Pro-Russian Hackers Blamed for Water Dam Sabotage in Norway
- Dark Reading: Destructive attack on Poland's wind and solar infrastructure
- Microsoft: Volt Typhoon Targets US Critical Infrastructure with Living-off-the-Land Techniques (2023-05-24)
- CISA: CISA and Partners Release Advisory on PRC-sponsored Volt Typhoon Activity and Supplemental Living Off the Land Guidance (2024-02-07)
- FBI IC3: 2025 Internet Crime Report
- Google GTIG: UNC3944 proactive hardening recommendations (2025-05-06)
- Microsoft: Protecting Customers from Octo Tempest Attacks Across Multiple Industries (2025-07-16)
- SCMP: UK Multinational Arup Confirmed as Victim of HK$200 Million Deepfake Scam (2024-05)
- AML Intelligence: AI Messaging Scam Costs Italy's Top Bank Intesa Millions, Sources Say (2026-09)
- UC San Diego: Understanding the Efficacy of Phishing Training in Practice (IEEE S&P 2025)
AI for Attack and Defence
- Microsoft: Staying Ahead of Threat Actors in the Age of AI (2024-02-14)
- Google GTIG: Adversarial Misuse of Generative AI (2025-01-29)
- Anthropic: Detecting and Countering Misuse of AI: August 2025 (2025-08-27)
- Google GTIG: GTIG AI Threat Tracker: Advances in Threat Actor Usage of AI Tools (2025-11-05)
- Anthropic: Disrupting the First Reported AI-Orchestrated Cyber Espionage Campaign (2025-11-13)
- Anthropic: AI-enabled cyber threats mapped to MITRE ATT&CK (2026-06-03)
- Google GTIG: From Prompting to Autonomy: The Evolution of Adversarial AI (2026-09-08)
- Anthropic: Threat Intelligence Report, September 2026 (2026-09-10)
- Anthropic: Incalmo toolkits for LLM-driven multistage network attacks (2025-06)
- Anthropic: ExploitBench and ExploitGym evaluations (2026-05-22)
- Anthropic: GLM-5.3 and the Spread of Advanced Cyber Capabilities (2026-09-29)
- Microsoft: Microsoft Unveils Microsoft Security Copilot Agents and New Protections for AI (2025-03-24)
- Microsoft: Learn What Generative AI Can Do for Your Security Operations Center (SOC) (2025-11-04)
- Microsoft Research: Project Ire Autonomously Identifies Malware at Scale (2025-08-05)
- Google Cloud: The Dawn of Agentic AI in Security Operations at RSAC 2025 (2025-04-28)
- Mandiant: Staying Ahead of Adversarial AI Through Agentic Source Code Review (2026-08-18)
- Google Cloud: Getting Started with the Mantis Harness to Find and Fix Bugs (2026-09-02)
- Anthropic: Building AI for Cyber Defenders (2025-10-03)
- The Record: Google's Big Sleep finds a SQLite flaw before attackers could exploit it (2025-07)
- DARPA: AI Cyber Challenge Marks Pivotal Inflection Point for Cyber Defense (2025-08)
- Dark Reading: AI-Based Pen Tester Becomes Top Bug Hunter on HackerOne (2025-06)
- Anthropic: Claude's results in 2025 cyber competitions
- Anthropic: Claude Code Security (2026-02-20)
- Anthropic: Finding Firefox vulnerabilities with Mozilla (2026-03-06)
- Anthropic: Project Glasswing (2026-04-07)
- Anthropic: Claude Mythos Preview technical report (2026-04-07)
- Anthropic: Project Glasswing initial update (2026-05-22)
- Anthropic: Expanding Project Glasswing (2026-06-02)
- Anthropic: Fable and Mythos access changes (2026-06)
- Anthropic: Claude Fable 5.1 and Mythos 5.1 (2026-09-01)
- CNBC: OpenAI expands its Daybreak cybersecurity program (2026-08-10)
- SecurityWeek: Google Launches Gemini 4 Argon with Guardrail-Free Access for Vetted Defenders (2026-09-30)
Security of AI Systems
- NIST: AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (2025-03-24)
- MITRE ATLAS: atlas-data CHANGELOG, v2026.09 (2026-09-14)
- OWASP GenAI Security Project: OWASP GenAI LLM Top 10 2026 (2026-08-04)
- Anthropic: A Small Number of Samples Can Poison LLMs of Any Size (2025-10-09)
- CVE: EchoLeak record, CVE-2025-32711 (2025-06-11)
- Noma Security: ForcedLeak: Agent Risks Exposed in Salesforce Agentforce (2025-09-25)
- Invariant Labs: MCP tool poisoning, shadowing and rug-pull experiments (2025-04)
- Anthropic: Detecting and Preventing Distillation Attacks (2026-02-23)
- arXiv (Nasr, Carlini et al.): The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses against LLM Jailbreaks and Prompt Injections (2025-10-10)
- Anthropic: Constitutional Classifiers: Defending against Universal Jailbreaks (2025-02-03)
- Anthropic: Prompt injection defenses for browser agents (2025-11-24)
- Google Research: CaMeL source code, defeating prompt injections by design (2025-03)
- Meta: Agents Rule of Two: A Practical Approach to AI Agent Security (2025-10-31)
- NSA: Model Context Protocol (MCP): Security Design Considerations for AI-Driven Automation (2026-05-20)
- Anthropic: Responsible Scaling Policy v3.4 (2026-07-08)
- OpenAI: GPT-5.6 System Card (2026-07-09)
- Google DeepMind: Frontier Safety Framework v3.1 (2026-04-17)
Regulation and Policy
- EU AI Act: Article 15: Accuracy, Robustness and Cybersecurity
- FCC: Cyber Trust Mark lead administrator changes to ioXt Alliance, DOC-420764A1 (2026-04-13)
- BSI: BSI Encourages Manufacturers of Internet-Connected Devices to Consider Cybersecurity Testing Activities as Radio Equipment Directive (RED) Delegated Act Deadline Approaches (2025-02)
- Federal Register: Cybersecurity in Medical Devices: Quality System Considerations and Content of Premarket Submissions (2025-06-27)
- Carscoops: EU Introduces New Cybersecurity Rules as EVs Are Deemed 'Spying Machines on Four Wheels' (2024-04)
Taiwan