Estimating worst case frontier risks of open weight LLMs
What changed
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity.
Why it matters
A concrete addition to Practical AI: Tools, Models & Frameworks: it changes what's available to builders today rather than being general commentary.
How it compares
Related prior coverage to compare against:
- gpt-oss-120b & gpt-oss-20b Model Card
- gpt-oss-safeguard technical report
- Introducing gpt-oss
Sources
- Estimating worst case frontier risks of open weight LLMs (openai-blog)primary