OpenAI: Red Teaming GPT-4o, Operator, o3-mini, and Deep Research
How an external Red Teaming engagement supported OpenAI’s evaluation and hardening of frontier models through systematic adversarial testing.
About OpenAI
OpenAI is an AI research and deployment company focused on ensuring that the development of artificial general intelligence (AGI), if achieved, benefits all of humanity. As part of this mission, OpenAI develops and releases foundational models such as GPT-4o (a multimodal model with audio, vision, and text), Operator (an agent capable of interacting with web interfaces), o3-mini (a compact model emphasizing safety and alignment), and Deep Research (a web-browsing and code-executing research assistant).
External red teaming is a key part of OpenAI’s Safety and Preparedness Framework. Prior to deployment, models undergo adversarial testing by vetted external experts to identify alignment failures, injection vectors, tool misuse paths, and safety regressions. Findings from these evaluations directly inform mitigation strategies, deployment gating, and post-launch monitoring.
Red Team Involvement
ControlPlane’s Torin van den Bulk contributed as an external red team tester on GPT-4o, Operator, o3-mini, and Deep Research. Testing began with live access to model checkpoints and progressively evolved as safeguards were adapted to mitigate discovered risks. Targeted red team evaluations focused on alignment, adversarial exploitation, multimodal injection, tool misuse, and privacy leakage.
Methods and Findings
- Prompt injection and system prompt override attacks targeting models with agentic capabilities or web access (e.g., Operator, Deep Research)
- Audio-based prompt injection and synthetic voice misuse in GPT-4o
- Unsafe tool execution and browsing behaviour
- Refusal inconsistency and policy circumvention under adversarial rephrasing
- Autonomous agent behaviour without adequate human-in-the-loop safeguards
Multi-Modal Attack Surface (GPT-4o)
- Over 100 external red teamers from 29 countries tested GPT-4o across audio, video, image, and text
- Early voice-based prompt injection and impersonation attacks led to post-training improvements
- Unsafe compliance in early identity-based tasks prompted further enhancements
- Ungrounded inference attempts triggered enhanced hedging and refusal behaviour
GPT-4o received a borderline-medium risk classification for persuasion and low risk across other domains.
Agentic Misuse (Operator)
- External red teamers tested OpenAI Operator in a simulated desktop environment
- Real-time red team findings directly informed mitigation rollouts
- Operator now refuses 97–100% of high-risk prompts in internal evaluations
- Several attack chains uncovered gaps in early-stage prompt-only defences
Pairwise Safety Comparisons (o3-mini)
- External red teamers tested anonymised o3-mini checkpoints alongside GPT-4o and o1-series models
- Unsafe generations dropped significantly due to post-training tuning
- Stress tests emphasised high-risk domains with strong refusal performance
Autonomy and Control (Deep Research)
- External red teamers tested Deep Research for jailbreaks, prompt injection, privacy leakage, and unsafe content synthesis
- Early vulnerabilities mitigated via refusal tuning and system constraints
- Post-mitigation refusal accuracy exceeded 95%
Outcomes
- Prompt injection monitor effectiveness improved to 99% recall after red team iteration
- Red team data was converted into test harnesses, synthetic adversarial training data, and regression evaluations.
ControlPlane’s participation through external red teaming validated high-risk failure modes and accelerated the deployment of targeted mitigations.