Palo Alto Networks is betting that one AI model is not enough for enterprise vulnerability discovery. In testing disclosed with the Sept. 22 launch of Unit 42 Continuous Frontier AI Defense, the company says no single model found more than 40% of vulnerabilities across the enterprise codebases and live environments it evaluated.
Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6-Cyber overlapped on less than 10% of the exposures they identified, according to Palo Alto.
For CISOs and security architects, the results shift the evaluation from model rankings to demonstrated coverage. Palo Alto’s multi-model design targets the blind spots it found, but its public materials do not show how much of the full vulnerability set the combined system detects.
Palo Alto builds around single-model blind spots
Palo Alto’s technical disclosure says its evaluation covered enterprise codebases and live environments, but it does not publish the number of applications or codebases tested, individual model detection rates, or the benchmark dataset. It also does not explain the denominator behind the less-than-10% overlap figure.
Continuous Frontier AI Defense uses GPT-5.6-Cyber, Claude Mythos 5, and open-weight models in a proprietary harness that routes security tasks to different models. The service builds on Frontier AI Defense, launched April 17 with a point-in-time exposure analysis and security blueprint. In August, Unit 42 added the OpenAI and Anthropic models to that earlier analysis.
Palo Alto says it developed and validated the approach over six months of internal work and more than 100 Unit 42 customer engagements. In its own deployment, high- and critical-severity findings per product were 3.2 times the level produced by legacy testing, while mean time to remediate fell 51%. Its technical post says 37% of identified exposures were high or critical, while two out of three validated exposures in third-party applications had no known CVE.
Those are company-reported results, not independent benchmarks. A recent benchmark of AI vulnerability repair found that some AI-generated patches could stop an exploit while still breaking legitimate software behavior, underscoring how strongly security results can depend on validation methodology.
Palo Alto also says its system can chain individually minor weaknesses into viable attack paths. In one financial-sector engagement, it linked an identity re-verification failure, a skipped one-time-password check, and a session-routing flaw into a path that could have enabled account takeover and payment fraud without victim action. That approach reflects a broader shift toward AI-assisted exploit chaining.
What eWeek Found: Palo Alto quantifies model gaps, not multi-model coverage
Palo Alto’s September service extends April’s point-in-time exposure analysis into continuous testing built around multiple models. Its results showed large coverage gaps between individual models, providing the technical rationale for the multi-model design.
The company has not disclosed the combined detection rate of its multi-model system, the total vulnerability set used for testing, or false-positive and false-negative rates. Axios separately reported Palo Alto’s findings and strategy but did not independently reproduce the benchmark.
Model count alone does not establish security coverage. Procurement teams should ask vendors for validated detection rates, false-positive and false-negative data, attack-path testing methodology, and evidence that additional models close material gaps rather than simply generate more findings. IBM is also using OpenAI models to identify and validate exploitable software flaws, giving buyers another example of frontier models being paired with enterprise security workflows.
Want to learn more AI tips, tricks, and prompting techniques? Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work.
Learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. eWeek readers get free 7-day access. Check out all the lessons here →


