AI-Discovered Flaws Are Twice as Likely to Enable Remote Code Execution, Google Finds
Google's Threat Intelligence Group (GTIG) published a report on September 30, 2026 titled "Vulnerability Discovery and Exploitation Trends in the AI Era," authored by Robin Grunewald, Supriya Mazumdar, and Kelli Vanderlee. Drawing on disclosure and exploitation data from January 2025 through August 2026, the report concludes that AI is not just accelerating vulnerability discovery — it is changing the type and risk profile of what gets found. Chief among the findings: 50% of AI-discovered vulnerabilities result in remote code execution (RCE), compared to just 26% of vulnerabilities found through conventional methods.
Details
| Attribute | Value |
|---|---|
| Report | Vulnerability Discovery and Exploitation Trends in the AI Era |
| Publisher | Google Threat Intelligence Group (GTIG) |
| Authors | Robin Grunewald, Supriya Mazumdar, Kelli Vanderlee |
| Published | September 30, 2026 |
| Coverage period | January 2025 – August 2026 |
| Monthly disclosures (Jan → Aug 2026) | 5,045 → 10,477 (July) → 10,740 (August, peak) |
| High-risk disclosure growth | 167% (131 in January to 350 in August) |
| AI-discovered RCE rate | 50% (vs. 26% for non-AI discoveries) |
| Exploited vulnerabilities, Jan–Aug 2026 | 141 (already exceeding all of 2025's 127) |
| Cumulative AI/LLM-stack CVEs tracked | 2,076 (Jan 2025–Aug 2026); 1,500+ in 2026 alone |
What Google Found
Disclosure Volume Is Accelerating
GTIG recorded monthly vulnerability disclosures nearly doubling in 2026, climbing from 5,045 in January to 10,477 in July and peaking at 10,740 in August. The report cautions that raw volume can be misleading — automated CVE assignment across open-source ecosystems inflates the count, and CVEs whose description merely mentions the Linux kernel alone generated roughly 5,000 entries between January and August with no in-the-wild zero-day exploitation observed among them. Even accounting for that noise, GTIG found high-risk disclosures (rated by GTIG's own methodology, not CVSS) grew 167%, from 131 in January to 350 in August.
AI-Discovered Vulnerabilities Skew More Severe
The report's central finding is a shift in what kind of bugs are being found. AI-discovered vulnerabilities were rated 39% low-risk and 58% medium-risk, versus 69% low-risk and 28% medium-risk for non-AI discoveries — and both categories converge near 3–4% high-risk. The more significant gap is in exploitability outcome: AI agents are proving disproportionately good at surfacing memory corruption and logic bypass flaws in core C/C++ libraries, runtimes, and hypervisors through fuzzing harnesses and memory-state modeling — the exact bug classes most likely to yield RCE. That is why 50% of AI-found vulnerabilities enable RCE versus 26% of traditionally discovered ones.
The AI Stack Itself Is a Growing Attack Surface
GTIG tracked 2,076 cumulative AI-related CVE disclosures between January 2025 and August 2026, with more than 1,500 of those landing in 2026 alone. The largest single category was AI orchestration and agent frameworks — tools like Flowise, Langflow, and LangChain — accounting for 782 CVEs, largely RCE via untrusted workflow serialization. Other categories included AI web apps and portals such as Open-WebUI and AnythingLLM (230 CVEs, SSRF and stored XSS), inference and serving infrastructure like vLLM, Ollama, and Triton (212 CVEs, unauthenticated APIs and model deserialization flaws), ML frameworks and model hubs including PyTorch and Hugging Face (99 CVEs, memory-safety violations), and frontier model providers — Anthropic, Gemini, OpenAI (97 CVEs, RCE and sandbox escapes). Only a handful of the 2,076 have been confirmed exploited in the wild so far, including flaws in LiteLLM and Langflow.
Case Study: CVE-2026-1731 and the Hacktron Agent
GTIG's headline example is CVE-2026-1731, an unauthenticated OS command injection flaw in BeyondTrust Privileged Remote Access (PRA) and Remote Support, discovered autonomously by the third-party AI research agent Hacktron AI. Following public disclosure, one threat cluster began exploiting the flaw within four days; five additional threat clusters followed within seven days. Post-exploitation activity observed by GTIG included privilege escalation, data exfiltration, and deployment of secondary payloads including SNOWLIGHT, SPARKRAT, and cryptomining malware. Google frames the case as proof that defensive AI tooling surfaces exactly the kind of high-impact vulnerabilities adversaries are actively hunting for — and that once found, they get weaponized fast.
Exploitation and Zero-Day Trends
Overall exploitation activity climbed from an average of 10.5 vulnerabilities per month in 2025 to 18 per month in 2026 — GTIG counted 141 distinct exploited vulnerabilities in the first eight months of 2026 alone, already surpassing 2025's full-year total of 127. Still, only 0.23% of all disclosed vulnerabilities (roughly 1 in 431) were confirmed exploited in the wild. High-risk vulnerability exploitation more than doubled, from 28 in all of 2025 to 75 in the January–August 2026 window. Zero-day exploitation ticked up more modestly — from 8 per month in 2025 to 11 per month in 2026, with an August spike to 22 — but zero-days still accounted for 62% of all confirmed in-the-wild exploitation, underscoring how disproportionately impactful they remain relative to their raw numbers.
Impact Assessment
| Impact Area | Description |
|---|---|
| Vulnerability management | Mass-patch-everything models break down as disclosure volume outpaces triage capacity |
| Time-to-exploit | n-day exploitation windows are compressing as AI automates analysis and weaponization |
| AI/LLM infrastructure | Orchestration frameworks (Langflow, Flowise, LangChain) carry the largest AI-related CVE burden |
| Third-party AI research agents | Tools like Hacktron surface high-impact bugs that adversaries weaponize within days of disclosure |
| Zero-day exploitation | Zero-days remain a small share of disclosures but drive the majority of confirmed in-the-wild exploitation |
| Bug severity mix | AI discovery methods skew toward memory-corruption and logic flaws that are more likely to yield RCE |
Recommendations
For Security Teams
Move away from mass-patching toward threat-intelligence-driven triage. Given that only a fraction of a percent of disclosed vulnerabilities are ever exploited, prioritize patching based on exploitation signals, exposure, and RCE potential rather than CVSS score alone. Treat any AI-discovered vulnerability disclosure affecting your stack as higher-urgency by default, given its elevated likelihood of enabling RCE.
For Software Vendors and AppSec Teams
Adopt pre-release AI-assisted code review — Google specifically points to tools like CodeMender — as standard practice to catch memory-corruption and logic flaws before public disclosure, reducing the volume of post-release findings entirely. Fuzzing harnesses and memory-state modeling, the same techniques driving AI-discovery gains for defenders, should be integrated into CI pipelines.
For AI/LLM Platform Operators
Treat orchestration and agent frameworks (Langflow, Flowise, LangChain, and similar tooling) as high-priority attack surface — they accounted for the largest share of AI-related CVEs in this dataset. Sandbox autonomous agentic workloads, restrict untrusted workflow serialization/deserialization, and audit inference-serving infrastructure (vLLM, Ollama, Triton) for unauthenticated API exposure.
For IT Administrators
Patch BeyondTrust PRA/Remote Support against CVE-2026-1731 immediately if not already done — exploitation began within four days of disclosure and multiple threat clusters are actively using it to deploy SNOWLIGHT, SPARKRAT, and cryptominers. More broadly, assume the window between disclosure and mass exploitation is shrinking and adjust patch SLAs accordingly.
Key Takeaways
- AI-discovered vulnerabilities are 50% likely to enable remote code execution, nearly double the 26% rate for traditionally discovered flaws.
- Monthly vulnerability disclosures nearly doubled in 2026, from 5,045 in January to a peak of 10,740 in August, though automated CVE assignment inflates part of that growth.
- High-risk disclosures grew 167%, and confirmed exploited vulnerabilities (141) already exceeded all of 2025's total (127) within the first eight months of 2026.
- The AI/LLM technology stack itself is a fast-growing attack surface, with 2,076 cumulative CVEs tracked and orchestration frameworks like Langflow and LangChain the largest single category.
- CVE-2026-1731, found by the third-party Hacktron AI agent in BeyondTrust PRA/Remote Support, was exploited by multiple threat clusters within a week of disclosure, delivering SNOWLIGHT, SPARKRAT, and cryptomining payloads.
- Zero-days remain rare in raw numbers but account for 62% of confirmed in-the-wild exploitation — reinforcing that severity, not volume, should drive patch prioritization.