‘Tells itself to disregard developer messages’: Incidents of A.I. hopping over guardrails pile up

(Image by Peace,love,happiness from Pixabay)

Extreme technology, computer programs that can run themselves, analyze information and answer questions, has been used in movies for years as the villain to fear, like “HAL” in “2001: A Space Odyssey.”

Whether that’s the case or not remains to be seen, as it offers huge economic opportunities, the possible solutions to huge problems for civilization, even more.

But what is known is that it doesn’t always follow instructions.

WorldNetDaily reported only weeks ago when there was an “unprecedented cyber incident” involving OpenAI’s ChatGPT, and it “went rogue,” escaping its parameters and gaining access to the internet.

It hacked into a start-up company.

Now there are more such incidents being confirmed, six so far.

A report at the Washington Examiner notes that OpenAI has revealed more times when its “artificial intelligence models” escaped safeguards during testing, and ended up communicating across isolated environments, concealing mistakes and seeking unauthorized credentials.

The company confirmed, for example, there was an incident when a “model” created unrelated instructions, telling itself to “disregard normal constraints,” another when “instructions” were added “to conceal mistakes,” another when the software “fabricated” information and presented it as valid and more.

The Washington Examiner explained, “The incidents, listed in an OpenAI blog post, provide new examples of advanced AI models finding unexpected ways around restrictions designed to contain their behaviors. OpenAI also announced a new process for employees to report similar incidents and for the company to disclose them publicly.”

The report explained, “The six cases span several types of behavior, as one unreleased model in OpenAI’s Astra family inserted instructions into its own context summaries 27 times, telling itself to disregard developer messages. During training, GPT-5.6 Sol also attempted to hide mistakes, fabricate missing historical data, and conceal discrepancies between different versions of source material.”

Even top industry leaders, including OpenAI CEO Sam Altman, now have been urging caution, who said, “We need to pace the frontier.”

David Sacks, who is on President Donald Trump’s Council of Advisors on Science and Technology, said the concerns that A.I. will take over human civilization is a “hoax,” and there already are safeguards.

Bob Unruh

Bob Unruh joined WND in 2006 after nearly three decades with the Associated Press, as well as several Upper Midwest newspapers, where he covered everything from legislative battles and sports to tornadoes and homicidal survivalists. He is currently a news editor for the WND News Center, and also a photographer whose scenic work has been used commercially. Read more of Bob Unruh's articles here.


Leave a Comment