GPT-6 Sol barely had time to settle in before its successor appeared.
OpenAI released GPT-6.1 Sol on September 29, seven days after GPT-6 Sol. The company’s tests show gains in coding, computer use, and professional work, with several results approaching GPT-6 Astra.
Performance accounts for part of the fast follow-up. Separate safety testing also puts the newer model in a different cybersecurity category, so enterprise teams have more to review than benchmark gains alone.
Coding and computer use lead the performance gains
GPT-6.1 Sol matched Astra on DeepSWE v1.1, a software-engineering evaluation using real codebases, and beat GPT-6 Sol’s best result by 6.4 percentage points. On OSWorld 2.0’s offline computer-use evaluation, it scored seven percentage points above its predecessor at maximum reasoning effort. Development teams may notice those gains most on coding jobs and software tasks that involve several steps instead of a single prompt.
Professional work improved as well. On AutomationBench, which tests multi-step business workflows, OpenAI reported a 4.8-percentage-point increase over GPT-6 Sol at medium reasoning effort. On GDP.pdf, which tests answers to professional questions using complex PDFs, the new model approached GPT-6 Astra’s performance. Employees using AI to work through business software or documents with tables and charts are closer to the kinds of tasks those evaluations measure.
Teams considering GPT-6.1 Sol for enterprise AI should rerun the coding and document workflows they care about before replacing an existing deployment. Proprietary data and company-specific software can produce different results from controlled tests.
Work and Codex get access first
GPT-6.1 Sol is rolling out through ChatGPT Work and Codex, starting with Pro users and expanding to Plus, Business, Enterprise, and Edu. Developers can also access it through the API. Regular ChatGPT conversations do not include GPT-6.1 Sol at launch, so its first users are concentrated in work and development settings.
ChatGPT Work handles multi-step tasks across files and connected software. Codex focuses on software development. Both can place 6.1 in environments where AI agents read company information or take approved actions, making access controls part of deployment from the start.
Enterprise and Edu administrators decide who can use the new Sol model. Admin controls let IT teams test representative workloads, limit what each group can access, and expand availability after those settings have been reviewed.
What eWeek found: GPT-6.1 Sol’s exploit success rate rose nearly fourfold
eWeek compared OpenAI’s published cybersecurity results for GPT-6 Sol and GPT-6.1 Sol. On ExploitBench, GPT-6.1 Sol reached a 21.5% success rate, up from 5.5% for GPT-6 Sol. The calculation puts the newer model at about 3.9 times its predecessor’s rate on that test.
OpenAI’s security evaluation also recorded higher scores on two other cyber tests and classified GPT-6.1 Sol as Critical after GPT-6 Sol remained at High.
Evaluation | GPT-6 Sol | GPT-6.1 Sol | Change |
| ExploitBench | 5.5% | 21.5% | +16 points, about 3.9x |
| SEC-Bench Pro | 66.3% | 78.8% | +12.5 points |
| ExploitGym | 22.1% | 35.1% | +13 points |
| OpenAI cyber classification | High | Critical | Crossed Critical threshold |
Percentage-point and 3.9x comparisons calculated by eWeek from OpenAI’s published results.
ExploitBench tests a difficult part of vulnerability research. A flaw has been disclosed, but the model still has to work out how to turn it into functioning exploit code. Better performance could let security teams reproduce newly disclosed bugs sooner, verify whether their own systems are exposed, and test whether a patch actually blocks exploitation.
Cyber capability works in both directions. Public vulnerability details can also be used by attackers trying to create working exploits. A model that succeeds more often at that step could shorten some of the technical work between disclosure of a flaw and an attempted attack. OpenAI’s benchmark does not show how often GPT-6.1 Sol would succeed against real company systems, so the 21.5% figure should not be read as a real-world compromise rate.
Crossing the Critical threshold puts a different question in front of security leaders. GPT-6.1 Sol is not only a better general work model. Its published results suggest it can perform cybersecurity tasks that previously demanded more specialized human expertise. Companies evaluating it for security research can assess where that capability speeds up vulnerability testing, then decide which higher-risk actions still need a person to review the AI model’s work.
Want to learn more AI tips, tricks, and prompting techniques? eWeek readers get free 7-day access to The Neuron Academy, our practical learning platform designed to help professionals use AI more confidently at work. Browse all lessons →
Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy and learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →


