Google’s PageBreak AI Agent Finds 500 Flaws in Its Web Apps
The situation illustrates a trend toward using AI and deterministic validation to identify flaws and exploitability, and provide a risk assessment.
The situation illustrates a trend toward using AI and deterministic validation to identify flaws and exploitability, and provide a risk assessment.
A Google internal AI agent has discovered more than 500 cross-site scripting (XXS) flaws on the company’s own Web applications, by using a combination of AI and validation to provide an overview of vulnerabilities that can likely be exploited and thus should be addressed.
PageBreak is an agent of Google’s Product Security team that the company developed to test the security of its first-party Web applications and address any security challenges, the company quietly revealed in a blog post last month. The agent began in pilot form last November and moved to a full-fledged product in January, with a mission “to autonomously scale vulnerability discovery while minimizing manual toil,” Google information security engineer Michał Bentkowski wrote in the post.
Among the flaws in Google Web apps identified by PageBreak are a cache poisoning in apis.google.com; an XSS flaw in admin.google.com; and insecure external handshakes in browser extensions. The company outlined the three flaws, which have been fixed, in a companion blog post.
Google did not immediately respond to a request for the status of fixes for all the flaws identified by PageBreak. However, Google eventually intends to connect PageBreak with CodeMender, its automated vulnerability-fixing system, so that once the former identifies a verified flaw, the latter can generate a proposed fix for product engineers to validate and apply.
Perhaps more important, however, is PageBreak’s change in approach to using large language models (LLMs) in application security by giving them access to source code and security tooling, and letting them find vulnerabilities faster than human researchers, Bentkowski said.
So far, there have been challenges to this approach, such as “distinguishing a genuine, exploitable flaw from a convincing hallucination,” which often increases the burden on product teams rather than reducing it, he explained.
How PageBreak Addresses Patch Prioritization
The thinking behind the development of PageBreak is to go beyond identifying possible vulnerabilities to use autonomous agents and deterministic exploit validation to establish which flaws are truly worth addressing, he said. PageBreak does this by identifying a potential weakness, attempting to exploit it in a running environment, and reporting the result only when the vulnerability can be demonstrated.
“The core of this deterministic approach lies in its suite of specialized, non-AI-written validators,” Bentkowski explained in the post. “When the agent identifies a potential flaw, it passes the hypothesis to a validator which then executes a real payload to confirm the exploit.”
The validation logic and interface vary based on the vulnerability class (such as XSS, SQL injection, remote code execution) and the application surface (such as HTTP or gRPC).
“This approach results in a near-zero false-positive rate, ensuring that we avoid overloading product teams with unverified vulnerability reports,” he said.
The idea is to streamline the work of security teams that want to use AI to help investigate flaws but not be overwhelmed by the amount of false positives they have to sift through, notes Rickard Carlsson, CEO of Detectify.
“For the last couple of years, AI security tools have produced more potential vulnerabilities than teams can realistically investigate,” he says. “The important part of Google’s approach is that the agent doesn’t get the final word. A separate, non-AI validator has to execute a real payload against the running application before something is treated as a confirmed finding.”
Going on the Cyber Offensive With AI
Google based PageBreak’s usage primarily on Gemini models, including Gemini 3.1 Pro and Gemini 3.5 Flash. However, the company said it can work with other models as well.
For defenders, the tool demonstrates a potential direction for how they might use AI to identify security gaps and flaws in their environments and applications. Indeed, for Carlsson, the takeaway is “you shouldn’t let an agent grade its own homework.”
Although AI can find vulnerabilities at a scale humans can’t match, it’s only useful if an organization “can reliably separate exploitable vulnerabilities from convincing noise,” he says.
“That requires deterministic validation against what is actually running,” Carlsson says. “Get that right, and AI becomes much more useful defensively because teams can automate not just finding potential weaknesses, but proving which ones actually matter.”
PageBreak illustrates a broader trend toward offensive testing becoming “more continuous, autonomous, and evidence-driven,” observes Darin Fredde, senior director of technical marketing engineering at Ridge Security.
“Finding something is no longer enough,” he says. “We increasingly need to prove it, fix it, and verify that the fix worked.”