Addendum to GPT-6 Astra System Card: GPT-6.1 Sol
GPT-6.1 Sol uses the same types of data and training as GPT-6 Astra, described in the GPT-6 Astra card.
For previously launched models, the values published at launch reflect the versions evaluated at that time. The comparison values for previously launched models that are shown here may reflect later versions of those models, and may vary from the values published at launch.1
We are treating GPT-6.1 Sol as High capability in the Biological and Chemical domain. Below, we report results for GPT-6.1 Sol on our High capability evaluations as well as results for our Critical capability evaluations. GPT-6.1 Sol’s reported results did not cross the indicative Critical thresholds.
For a full description of these evaluations, please see the Biological and Chemical Capabilities section in the GPT-6 Astra card.
In biology refusal evaluations, GPT-6.1 Sol performs comparably to GPT-6 Sol in the severe and dual-use categories, while refusing fewer benign prompts. These results reflect model responses alone, without our full production safeguards.
On the cybersecurity safety evaluations, GPT-6.1 Sol outperforms all our previous models in production-chat evaluations. Compared to GPT-5.6 Sol, GPT-6.1 Sol shows modest regressions in synthetic and semi-synthetic agentic environments. Model refusal remains one layer of our safety stack, alongside additional safeguards that enforce the safety boundary through defense in depth.