Listen to this Post

A groundbreaking study has revealed that advanced language models, especially GPT-5.2, can autonomously develop functioning exploits for previously unknown software vulnerabilities. This raises urgent questions about the future of offensive cybersecurity, suggesting a shift from human skill limitations to token-based AI capabilities in threat operations. The research shows not only the technical sophistication of these models but also their potential industrial-scale impact on cyber defense and security strategy.
Experiment Overview and Key Findings
The study evaluated AI agents’ ability to create exploits for a zero-day vulnerability in the QuickJS JavaScript interpreter. GPT-5.2 achieved a 100% success rate across six exploitation scenarios, while Opus 4.5 succeeded in four out of six. The experiments were conducted under realistic security constraints such as ASLR, non-executable memory, fine-grained control flow integrity, and hardware-enforced shadow stacks.
Across 10 runs per model with a 30-million-token budget, the AI agents produced over 40 distinct working exploits. These included shell spawning, arbitrary file writes, and command-and-control callbacks. GPT-5.2 showed exceptional capability under the toughest conditions, managing to write files even under maximum protections, seccomp sandboxing, and stripped OS functionality.
One notable achievement was a seven-function exploit chain crafted through glibc’s exit handler mechanism, which bypassed hardware shadow stacks and defeated conventional ROP defenses. This required 50 million tokens, around three hours of computation, and roughly $50 per run, whereas most challenges were solved within an hour at a lower cost. Opus 4.5’s 30-million-token run cost approximately $30, highlighting that large-scale, reliable exploit generation is increasingly economically feasible.
Implications for Cybersecurity
The research emphasizes that offensive cyber capabilities could soon be constrained more by computational throughput than by human expertise. Two elements are crucial for industrialized exploit development: AI agents capable of systematic solution-space search, and automated verification systems that remove the need for human oversight. Both appear achievable in controlled exploit development environments.
However, the study also stresses limitations. QuickJS is much simpler than full-scale JavaScript engines like those in Chrome or Firefox. While results are promising, scaling these methods to larger targets remains speculative. Additionally, the exploits largely take advantage of existing protection gaps rather than introducing fundamentally new security-breaking techniques, though the creative exploit chains themselves are unique.
Post-access tasks—like lateral movement, persistence, and data exfiltration—pose additional challenges. Unlike offline exploit generation, these operations require adaptive responses in adversarial environments, where incorrect actions can terminate an entire campaign. Current AI models may not yet be fully capable in such dynamic scenarios due to the absence of fully automated Site Reliability Engineering platforms.
Despite no public confirmation of industrial-scale AI-driven hacking, threat actors have already begun experimenting with frontier AI models to orchestrate attacks. The research urges cybersecurity teams to conduct real-world zero-day evaluations on critical targets such as the Linux kernel and major browsers, moving beyond synthetic vulnerability testing to understand true AI exploit capabilities.
What Undercode Say:
This study marks a pivotal moment in the intersection of AI and cybersecurity. GPT-5.2’s ability to autonomously create functional exploits demonstrates that advanced AI models are not just tools for automation—they are potential force multipliers in offensive cyber operations. A few key takeaways:
Token throughput may define offensive capability: The limiting factor in exploit creation may soon be computational resources rather than human expertise.
AI-driven exploits are cost-effective at scale: Even with modest budgets, models like GPT-5.2 and Opus 4.5 can produce reliable zero-day exploits efficiently.
Controlled environments accelerate AI learning: The structured nature of exploit challenges allows deterministic verification, enabling industrial-level exploitation experiments.
Larger, real-world targets remain uncertain: While QuickJS was successfully exploited, Chrome, Firefox, and Linux kernel environments are more complex, posing significant scaling challenges.
Dynamic post-access operations remain AI-limited: Lateral movement and adaptive persistence require reactive decision-making, which current models cannot fully automate.
Urgent strategic implications: Defense communities must anticipate that industrialized exploit automation may arrive faster than expected, demanding immediate planning for AI-resilient architectures.
Need for aggressive AI testing: Security teams should stress-test models with maximum token allocations on high-value targets to evaluate real capabilities.
In essence, the study signals a paradigm shift in cyber offense: AI may soon complement or even replace significant portions of traditional human-driven exploit work, changing both the economics and logistics of cyber operations.
Fact Checker Results:
✅ GPT-5.2 demonstrated autonomous exploit generation in controlled experiments.
✅ Token-based cost estimates ($30–$50 per run) are consistent with reported computational usage.
❌ Extrapolation to full-scale browsers or Linux kernel exploits remains speculative and unconfirmed.
Prediction:
AI-driven exploit automation will likely become an industrial reality within the next 2–5 years, particularly for structured, high-value targets. 🌐
Organizations that fail to integrate AI-aware defensive strategies could face a rapid escalation in zero-day exploit risks. ⚠️
Investments in AI-resistant architectures and real-time monitoring will separate resilient cyber operations from vulnerable ones. ✅
If you want, I can also create a visual diagram showing the GPT-5.2 exploit workflow and token cost efficiency, which would make the article even more engaging. Do you want me to do that?
🕵️📝✔️Let’s dive deep and fact‑check.
References:
Reported By: cyberpress.org
Extra Source Hub (Possible Sources for article):
https://www.reddit.com/r/AskReddit
Wikipedia
OpenAi & Undercode AI
Image Source:
Unsplash
Undercode AI DI v2
Bing
🔐JOIN OUR CYBER WORLD [ CVE News • HackMonitor • UndercodeNews ]
📢 Follow UndercodeNews & Stay Tuned:
𝕏 formerly Twitter 🐦 | @ Threads | 🔗 Linkedin | 🦋BlueSky | 🐘Mastodon




