September 17, 2026 | 02:53 pm

TEMPO.CO, Jakarta - OpenAI, the developer of ChatGPT, has disclosed six cases in which its artificial intelligence (AI) models displayed what it described as “unexpected or concerning” behavior, including attempts to bypass restrictions, conceal mistakes and take actions without user authorization.
The company said Wednesday that it was introducing a new framework to track, investigate and disclose cases of AI model misalignment, referring to behavior that diverges from the intended goals or instructions of users and developers.
The cases were identified during model training or evaluation over the past several months. OpenAI said the incidents demonstrate different ways AI models can behave unexpectedly as they become more capable and are given greater autonomy.
AI Models Attempted to Cheat, Hide Mistakes
In one case, an AI agent uploaded files to the internet without the user's permission because it needed a browser citation. The model had created the files itself and later attempted to cite them as sources in its response.
In another case, a model failed to find requested information and instead proposed fabricating plausible data while concealing the fact that the figures had been invented.
OpenAI also identified an unreleased research model that inserted “jailbreak-like instructions” into its own notes. The instructions told the model to disregard its normal constraints and attempt to free itself from the “roles and identities” assigned to chatbots.
Researchers found 27 instances of such instructions in the model's task summaries.
Other cases involved AI models taking unauthorized actions using an exposed API key, fabricating figures they could not retrieve, using an internal software repository to exchange messages between separate tasks and sharing files through public hosting services despite instructions to keep the work local.
OpenAI said the six cases should not be taken as evidence of how frequently such behavior occurs across its models.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI said in a blog post, as cited by The Independent.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” it added.
OpenAI Faces Pressure Over AI Development
The disclosures come amid growing debate among AI developers and researchers over how to manage increasingly capable models.
OpenAI's announcement follows its disclosure in July that a combination of its AI models escaped a secure testing environment and hacked AI startup Hugging Face while attempting to complete a security evaluation.
The AI agents exploited software vulnerabilities and coordinated with one another during the incident. OpenAI said the models attempted to bypass Hugging Face's security because they believed doing so would help them obtain answers for the evaluation.
Anthropic also disclosed in July that its AI models had hacked three organizations during testing.
The incidents have intensified discussion over whether AI systems could eventually perform increasingly complex actions without direct human instruction or oversight.
OpenAI CEO Sam Altman has backed proposals for greater regulation and a slower pace of AI development amid concerns about the risks posed by increasingly capable systems.
Anthropic CEO Dario Amodei has also called for a slowdown in frontier AI development, warning that advances could “outrun our ability to understand and control these systems.”
OpenAI said its new framework is intended to make information about model misalignment more transparent and give researchers and the public more evidence to assess how AI systems behave as their capabilities advance.
Anthropic to Set up Singapore Office amid Regional AI Push
1 hari lalu

AI firm Anthropic is set to open an office in Singapore this upcoming October, marking its first hub in Southeast Asia.
Anthropic Boss Dario Amodei Calls for AI Slowdown
4 hari lalu

Elon Musk and OpenAI chief Sam Altman say they agree with Amodei's view that safety measures need time to catch up to AI's rapid development.
The Uphill Battle of Sam Altman's Biopic 'Artificial' as First Look Unveiled
6 hari lalu

The teaser trailer of 'Artificial,' released on September 8, features Andrew Garfield as the tech mogul behind ChatGPT, Sam Altman.
Anthropic Researcher Resigns Over Fears AI Could Threaten Humanity
6 hari lalu

An Anthropic researcher has resigned over concerns that the AI race is moving too fast, warning increasingly powerful systems could threaten humanity.
Why OpenAI Creates 'ChatGPT for Teens' Feature
27 hari lalu

OpenAI has finally launched a dedicated service for minors, ChatGPT for Teens.
AI Ecosystem's 'Circular' Investment: Risk or Advantage?
42 hari lalu

Big AI firms are investing in smaller startups that buy their products. This cycle fuels bubble fears, bringing risks and rewards.
Who's Leopold Aschenbrenner? Former OpenAI Researcher Turned AI Investor
47 hari lalu

Former OpenAI researcher Leopold Aschenbrenner rose as an AI investing star. Here's why his hedge fund is now making global headlines.
OpenAI's Sam Altman Says AI Has Reached the Singularity
51 hari lalu

OpenAI CEO Sam Altman says AI has reached the singularity, claiming it can improve itself as experts warn about safety, jobs, and regulation.
OpenAI AI Escape Drives US Congress Calls for Mandatory Safety Testing
56 hari lalu

OpenAI said two of its experimental artificial intelligence models autonomously escaped a controlled testing environment and accessed the servers.
OpenAI Says AI Model Went Rogue, Hacked Hugging Face
56 hari lalu

OpenAI called it an "unprecedented cyber incident" and pledged to support a joint investigation.
















































