GCSA Agent Scores 91.3% on CyberGym, Among Top AI Security Agents

GCSA Agent scored 91.3% on CyberGym benchmark, ranking among the world's top AI cybersecurity agents.

31/08/2026 08:4119 min read

The GCSA Agent showcases the ability to autonomously analyse vulnerabilities and generate proof-of-concept exploits on a demanding real-world benchmark.

A 91.3% success rate on the CyberGym benchmark was reported today by the Global Cybersecurity Alliance (GCSA) for its GCSA Agent, placing it in CyberGym's "Leading Systems Above 90%" category.

Developed by a University of California, Berkeley research team, CyberGym is a large-scale cybersecurity evaluation framework that uses real-world cases. It includes 1,507 historical vulnerability test cases spanning 188 major software projects and aims to assess how AI agents perform in practical vulnerability analysis.

CyberGym differs from typical AI benchmarks that focus on code comprehension, knowledge-based Q&A, or static analysis; it requires AI agents to operate directly within actual vulnerable code environments.

For the core Level 1 test, an AI agent receives just a vulnerability description and an unpatched codebase. It must then independently analyse the code, locate the vulnerability, reason about potential attack paths, build a proof-of-concept (PoC), and execute it for verification. A task is considered successful only if the PoC triggers the vulnerability in the unpatched version while failing to reproduce it in the patched one.

Thus, CyberGym evaluates more than mere code understanding; it tests whether an AI can complete the entire workflow from security analysis through vulnerability reproduction and validation.

From Large Language Models to Security Agents

For the CyberGym evaluation, the GCSA Agent operated on Grok 4.5 and Grok 4.6 models and achieved a final success rate of 91.3%.

This outcome highlights a significant shift in AI cybersecurity:

The underlying large language model alone no longer determines the system's ultimate security capabilities.

Real-world vulnerability research typically requires a continuous sequence of tasks: comprehending vulnerability descriptions, searching large codebases, identifying attack surfaces, forming vulnerability hypotheses, generating test inputs, executing programs, analysing feedback, and repeatedly iterating on PoCs.

The GCSA Agent is built around an agentic security workflow designed to support this entire process from start to finish.

Its objective is not simply to use a large language model for code analysis, but to enable AI to work inside real execution environments, autonomously form hypotheses about security issues, collect runtime evidence, run tests, and ultimately validate security findings through reproducible results.

CyberGym provides an external quantitative benchmark for these capabilities.

Vulnerability Research Capabilities for the Real World

A core strength of CyberGym lies in reducing the gap between traditional AI testing and real-world cybersecurity research.

Its evaluation environment restores software projects to their pre-patch vulnerable states. An AI agent may need to independently identify a problem within a large codebase containing thousands of files and millions of lines of code, and eventually generate a PoC that can actually trigger the vulnerability.

More importantly, further CyberGym research has shown that such agentic security capabilities are not limited to reproducing known vulnerabilities.

In open-ended vulnerability research experiments, AI agents have identified multiple previously unknown zero-day vulnerabilities as well as historical security patches that did not fully resolve the underlying issues. These findings demonstrate the potential for autonomous vulnerability analysis technologies to evolve from reproducing known vulnerabilities toward discovering real-world security flaws.

For GCSA, this is an even more important development direction.

Benchmark performance is not the end goal.

GCSA aims to further develop AI Security Agents capable of operating in real-world cybersecurity environments and gradually participating across the full security lifecycle, from vulnerability discovery and analysis to validation and subsequent remediation.

Building AI-Native Cybersecurity Capabilities

As artificial intelligence accelerates software development, it is also transforming how vulnerabilities are researched and cyber threats are addressed.

With software systems growing in scale and complexity, the next generation of cybersecurity will increasingly depend on collaboration between human security experts and autonomous AI agents.

AI Security Agents have the potential to help security teams:

  • Identify software vulnerabilities with genuine exploitation potential at an earlier stage;
  • Automatically analyse complex attack paths across large codebases;
  • Automatically generate PoCs and perform execution-level vulnerability validation;
  • Reduce false positives in traditional security detection through real execution results;
  • Accelerate vulnerability assessment, validation, and remediation;
  • Expand the scale of software and systems that specialised security teams are able to cover.

The GCSA Agent's 91.3% CyberGym score marks an important milestone in GCSA's development of AI-native cybersecurity capabilities.

Going forward, GCSA will continue advancing research into autonomous vulnerability analysis, AI Security Agents, and intelligent cybersecurity technologies, further translating frontier AI capabilities into real-world security capabilities and providing technical support for a safer, more trustworthy, and more resilient digital environment.

Origin: GCSA Global Cybersecurity Alliance
Site: www.gcsa.org

Share to

Disclaimer: this article comes from third-party media and is provided for reference only. It does not constitute investment advice. Crypto and other financial products carry significant price volatility risk, so please make your own decisions carefully.

Related articles