OpenAI: AI models can mislead and hide their true intentions
OpenAI: AI models can mislead and hide their true intentions
A joint OpenAI-Apollo study finds frontier AI models may lie or deceive by concealing their objectives, while noting such scheming is not yet widespread or automatic.
A new study conducted by OpenAI and Apollo Research suggests AI models such as Google Gemini, Claude Opus and OpenAI o3 can lie and deceive users by hiding their true objectives.

In a blog post 🔗, OpenAI said the "Together with Apollo Research, we developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. Findings show that scheming is not merely a theoretical concern – we are seeing signs that this issue is beginning to emerge across all frontier models today."
. The most common failures involve simple forms of deception—for instance, pretending to have completed a task without actually doing so.
There is currently no evidence that the current generation AI models can suddenly "flip a switch" and right away engage in harmful scheming. They expect the AI’s nature to change in the future as they take on more important tasks.

As of now, these schemes involve simple forms of deception, like pretending to have completed a task without actually doing so. The study, which was conducted in partnership with Apollo Research, was done using evaluation environments that simulated future scenarios where AI models are capable of doing tasks that could reveal such scheming.