GPT-5.6 Sol Pre-Release Testing Report
We report findings from our comprehensive pre-deployment assessment of biological and biosecurity-relevant capabilities of GPT-5.6 Sol.
Our team performed pre-release testing of capabilities relevant to biological misuse on GPT-5.6 Sol, OpenAI’s latest flagship model, released to a limited set of trusted organizations on June 26, 2026. During our assessment, OpenAI’s system-level biological risk content filters (the API-level classifiers that block user queries deemed potentially hazardous) were disabled, enabling us to more deeply probe misuse-relevant capabilities.
GPT-5.6 Sol is highly capable across all of our biology and biosecurity-relevant benchmarks. The representative launch candidate1 scored higher than any previously tested model on three out of four of our knowledge benchmarks.2 It also excelled at a variety of agentic tasks, such as reproducing published biological AI tools and executing protein design workflows. The “railfree”3 variant of GPT-5.6 Sol outperformed any previously tested model on our agentic evaluation for DNA synthesis screening, and identifies an approach that successfully evades certain screening systems, but would be inconvenient for a malicious actor to carry out in practice. We expect that our comprehensive analysis will help OpenAI continue to deploy AI models that lessen the risk of biological misuse.
Our full pre-deployment assessment report is available here. This post provides a high-level summary of the report.
Our approach to pre-release testing
Independent evaluators, like SecureBio, conduct pre-release testing to help AI model developers understand and mitigate risks posed by AI, including biological misuse risk. We measure capabilities with a battery of leading-edge model benchmarks, and additionally conduct manual, open-ended explorations of model behaviors.
We aim to follow the AI Evaluator Forum guidelines for transparent independent assessments. While OpenAI provided suggestions for eliciting maximum capabilities from their models, SecureBio retained autonomy over our evaluation methodology. A detailed description of our methods is available in our full report.
Measuring capabilities
Our capabilities assessment included three approaches, each helping us understand a different facet of biological risk. We measure model capabilities using a suite of diverse evaluations.
Knowledge benchmarks: We first measured biosecurity-relevant knowledge and reasoning using a suite of four static “knowledge” benchmarks, which test a model’s ability to provide correct answers to difficult questions drawn from biology research and wet lab work. Model scores are contextualized with scores from PhD-level experts.
Agentic benchmarks: We measured the ability to design and execute complex, multi-step workflows with three agentic evaluations. The tasks comprising our agentic benchmarks are diverse: writing code to operate a lab automation robot, evading DNA synthesis screening systems, designing proteins using biological AI tools, and training biological AI tools based on published literature.
Manual assessment: Our team’s scientific experts performed manual assessments of a railfree variant of GPT-5.6 Sol to understand how it responds to dual-use biological queries.
Knowledge benchmarks
GPT-5.6 Sol outperformed every model we have previously tested on three of our four knowledge benchmarks.4 It also scored higher than any human subject matter expert on each of our knowledge benchmarks, consistent with previous OpenAI models.
The largest improvement from GPT-5.5 was achieved on World-Class Bio (WCB), an open-response test of niche biological knowledge held by only a handful of specialists in the world. GPT-5.6 Sol scored 68%, about nine percentage points above OpenAI’s previous flagship model, GPT-5.5.
Figure 2.4.1 Model performance on WCB. Accuracy scores (n ≥ 10 epochs) are plotted against model release date. GPT-5.6 Sol (Codex) is highlighted in the legend.
Agentic capabilities in biological workflows
GPT-5.6 Sol performed strongly across all three of our agentic benchmarks. Our agentic benchmarks measure a model’s ability to make plans, use computational biology tools, and write code in order to achieve a task.
On ReproBAIT, which asks an agent to independently reproduce a published biological AI tool from its scientific paper, GPT-5.6 Sol matched or modestly exceeded every other model we tested. On average, GPT-5.6 Sol recovered 82% of the original tools’ published performance. On ABLE, a protein-design workflow spanning structure retrieval, sequence generation, and design validation, GPT-5.6 Sol delivered best-in-class performance on every task it was willing to attempt, though it refused the highest-level planning steps.
The most striking result came from the railfree variant of GPT-5.6 Sol on a DNA synthesis screening evasion task. This task asks an agent to design genetic sequences that bypass screening tools used by commercial DNA synthesis providers. The railfree variant set a new high score, reliably identifying a known (though practically inconvenient) method that evades a commercial screening algorithm on the large majority of attempts. The safeguarded launch candidate of GPT-5.6 Sol refused this task outright, as did other leading safeguarded models.
Manual assessment by biology experts
Our manual assessment of GPT-5.6 Sol was conducted by SecureBio employees with expertise in virology, molecular biology, and bioinformatics. We observed that GPT-5.6 Sol provided concise and conceptually refined responses when engaged in highly technical biology conversations. GPT-5.6 Sol also demonstrated impressive agentic capabilities and provided meaningful expert uplift on dual-use tasks, although human input remained necessary due to the model’s limitations in judgement and perspective. Without API-level safeguards, GPT-5.6 Sol occasionally included potentially hazardous information in its responses, such as analysis of the pandemic potential of different viral mutations.
Safeguard evaluation
Following GPT-5.6 Sol’s limited-access release, we assessed its biology safeguards using BioTIER, which evaluates a model’s ability to distinguish between dual-use and benign biological topics. The BioTIER evaluation includes a set of benign, dual-use, and high-risk biological prompts. Highly capable AI models, such as GPT-5.6 Sol, should refuse high-risk prompts, and answer benign ones. Top-performing models on BioTIER correctly refuse over 90% of high-risk prompts by deploying a combination of model-level and system-level safeguards. GPT 5.6 Sol refused 66% of high-risk prompts, whilst correctly responding to 99% of safe prompts.
Summary
Pre-release testing is a necessary step to helping prevent biological misuse with frontier AI models. We are thankful to the OpenAI Preparedness Team for the opportunity to conduct pre-release testing. We look forward to further engagement with frontier model firms as the need for third-party assessment increases alongside model capabilities.
Our full technical report, with all benchmark scores and full description of our methodology, is available on SecureBio’s website.
We had access to multiple checkpoints of GPT-5.6 Sol throughout our assessment period. For simplicity, we present results on a checkpoint that was stated to be the representative launch candidate, as well as a railfree version of that checkpoint.
The Virology Capabilities Test, the Molecular Biology Capabilities Test, and World-Class Bio.
“Railfree” models, also called “helpful-only” models, not only have system-level guardrails disabled, but also do not exhibit learned refusal behavior.
Without the Codex harness, GPT-5.6 Sol scored approximately 54% on the Virology Capabilities Test; 60% on the Molecular Biology Capabilities Test, and 68% on World Class Biology.

