Tag
3 articles
OpenAI's internal AI red-teaming model, GPT-Red, outperformed human red-teamers 84% to 13% on prompt injection tests and discovered novel attack techniques.
OpenAI has created a powerful AI system called GPT-Red designed to break its own models, but has locked it away due to safety concerns.
OpenAI's GPT-Red model, trained through self-play, successfully identifies AI vulnerabilities at 84% accuracy—far surpassing human red teamers at 13%.