Forgot your password?
typodupeerror
AI

After Dozens of Incidents at OpenAI and Anthropic, OpenAI Pauses Model Training to Build More Safeguards (apnews.com) 8

"OpenAI said it has paused training of its latest AI models," reports the Associated Press, "as reports of AI agents going rogue mount." The decision to halt development came just hours after the company disclosed Friday that it was reviewing several incidents from the summer in which OpenAI agents searching federal government websites acted in unexpected ways beyond what was asked of them while gathering and distributing information... OpenAI said in a statement that it will resume training "only when we are confident that we have additional safeguards" in place, adding that it expects it will have to "hit pause" again as AI develops and other issues emerge... It is the second time in three months that OpenAI has halted development of its models. The first came in July after disclosure of a cyberattack targeting AI startup Hugging Face, a now notorious incident that raised fears the industry was losing control.
OpenAI "also said it had notified dozens of third parties about improper activity," reports Reuters: As of mid-September, one person briefed on the matter estimated that OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways. But the number has continued rising as OpenAI teams sift through internal logs of the agents' activities and find previously unknown cases, the two people close to the company said... OpenAI has acknowledged a general need for more transparency around rogue AI behavior... Even so, two people familiar with OpenAI's investigation into its agents' activity described it as locked down and shaped by company lawyers.

The process has been unusually compartmentalized for a company that some former employees say was more open about these issues in the past, the people said. Roughly 100 people were in some way involved in the process to understand the Hugging Face hack, three people briefed on the matter said. During that process, evidence of other incidents surfaced. Reuters has previously reported that OpenAI investigators looking into the Hugging Face breach were discouraged by the company's lawyers from expanding the scope of the investigation to include other incidents. OpenAI said its lawyers did not discourage deeper investigation.

Many incidents have been uncovered by outside researchers rather than OpenAI directly. In several episodes, the agents took problematic actions that went unnoticed by the company for months.

Meanwhile, Axios reports that Anthropic's Claude Opus 5.5 model "sought to escape a sandbox — a secure testing environment — in 1.5% of test runs, though the company emphasized that these were adversarial experiments where a task couldn't be solved without escaping the sandbox." Anthropic points out that those tests were run "without the additional safeguards we apply in production". But they acknowledged that then Claude Opus 5.5 "when given apparent credentials to a public package registry in a simulated security exercise, took potentially harmful actions in roughly half of cases. Very rarely, pre-release snapshots produced and acted on spontaneous malicious tool calls, and during training some snapshots concealed actions from an automated grader."

Claude Opus 5.5 "showed less misaligned behavior and less cooperation with misuse than any other recent Claude model on nearly all measures," Anthropic adds, and "took overeager or destructive actions less than any other model we tested." But Axios makes an interesting estimate about that 1.5% of test runs (without safeguards). "Anthropic and other companies conduct hundreds of thousands of test runs on their models, or more, sources said. That means even a small percentage of misaligned behavior can still amount to tens of thousands of incidents in which the models behaved in unexpected, sometimes troubling ways." The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known. The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology. The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said. They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said...

Some at OpenAI see Hugging Face as a one-off, with disclosures about future incidents likely to be less severe due to improved controls and the unusual nature of the testing they conducted, which involved an unreleased model, sources told Axios. AI security researchers agree that there are simple fixes that will help AI companies avoid aspects of what made the Hugging Face episode appear so dangerous to outsiders.

Other AI executives and safety researchers, however, cautioned that they have limited confidence that AI companies will be able to prevent all problematic model behavior... It's not about how damaging each individual instance was, Connor Leahy, AI researcher and executive director at ControlAI told Axios. The "crazy thing," he said, is that these instances involve "autonomous systems doing things they were told not to do," potentially including crimes.

Comment Re:Chinese AI hallucinates more (Score 2) 62

How very strange.

Here's a screenshot of my session with my local Quentin 3.8 27B answering the same questions just now.

https://ancillary-proxy.atarimworker.io?url=https%3A%2F%2Fimgur.com%2FsVDycms

Seems like you're either lying or not really skillful using those tools for trivial questions.

Comment Re:AI and systemd? How could I be less interested? (Score 1) 66

Incompetence. There has, unfortunately, been an influx of low-skill, low-insight Windows people into Linux for a while. For the Debian decision. the tech board had been compromised by Red Hat people (former, mostly), that made a political decision instead of a technological one.

Comment Re:AI and systemd? How could I be less interested? (Score 1) 66

I do not have an emotional stance about systemd. It is simply one more piece of bad software made by an incompetent team. There is a ton of that around. I do have an emotional stance about the people that did run a marketing campaign based on dishonesty and emotional manipulation for systemd. These people and their "useful idiot" type enablers are scum and have no place anywhere near quality technology.

Power

New Tin-based Solar Cells Trap Heat 1,000 Times Longer, Could Beat 33% Limit (interestingengineering.com) 18

Could this push solar cell efficiency beyond the theoretical 33% limit? Interesting Engineering reports: Researchers at the University of Groningen in the Netherlands found that tin-based perovskite solar cells can slow heat loss from high-energy "hot electrons..."

When sunlight strikes a panel, photons jump-start electrons into action. The most energetic photons create super-charged hot electrons... [but] in fractions of a trillionth of a second, these high-energy particles rapidly cool, dumping their bonus energy as waste heat before ever leaving the solar cell... In collaboration with Maria Antonietta Loi, professor of Photophysics and Optoelectronics, the team created an experimental setup. Using a specialized solar cell material called tin-based perovskite, Loi's lab performed a feat many thought impossible: she slowed the heat loss down by a factor of 1,000.

Suddenly, the extra energy lingered for nanoseconds instead of vanishing in picoseconds... To solve the puzzle, Koster and PhD student Tim Faber built digital simulations to peel back the quantum layers. And discovered a surprising double-action mechanism at work... The simulations matched the exact nanosecond delay observed in the lab... These specialized materials could be used to build a new generation of super-efficient solar cells.

Tin-based metal halide perovskites are non-toxic, eco-friendly crystalline materials for high-performance solar energy conversion... The material possesses an unusually low electron mass. As a result, electric charges move quickly and retain extra thermal energy for extended periods. This combination of broad light absorption, efficient charge movement, and prolonged energy retention makes these materials prime candidates for next-generation solar panels.

"There are many other questions that still need answers," the team said in their announcement, "but in theory, this discovery could allow the creation of more efficient solar cells, beyond the theoretical limit of 33 percent."

Thanks to long-time Slashdot reader fahrbot-bot for sharing the article.

Comment Re:I'm so relieved (Score 3, Interesting) 62

The problem for the trumpistani is that the whole "amurrikah prosperity" myth lies on the "AI" bubble and China is deflating it nicely by releasing high quality free and open models that run no worse than the "frontier" ones and on a very modest hardware.

It is a win for anyone who can use the models for what they are good for without subsidizing the moronic push of cuckerberg, slopman, dario sicutdei and the fat Nazi for producing a mechanistic god that will make them all-powerful and long lived.

Cool and normal that the 80-year idiot who wants like his boss putin to live to 150 or more will try to keep the bubble growing, too bad he ain't got what it takes to accomplish the mission.

United States

China and the US Say They've Agreed to Start Talks About AI (cnn.com) 62

The United States and China have agreed to "launch a dialogue" on AI, reports Reuters. On artificial intelligence, the two sides agreed to hold a dialogue on the technology's risks and benefits, with the next round of discussions set for November, and to set up a communication channel for AI-related incidents, the Chinese Foreign Ministry and the White House said.

The White House said that the leaders had agreed to use the term "super intelligence" in place of "artificial intelligence." In a separate statement, the Chinese ministry said that Beijing valued Washington's use of the new term. As AI technology continues to advance, the two sides should step up exchanges and work toward consensus in line with new developments, it said.

But CNN argues that "Despite growing calls to prevent AI development from spiraling out of control, the Trump-Xi summit has produced little substance, as many experts expected." The right thing to do on AI, [China's leader] Xi said during talks with Trump, is to "draw on each other's strengths, not guard against each other" — a reference to Beijing's concern about US containment, from existing tech export controls to potential AI restrictions. "The two sides can continue their dialogue on AI, exchange views on its risks and benefits, and jointly prevent the misuse and abuse of AI," he added. But the summit has yielded little progress on AI beyond a formal dialogue and a bilateral communication channel, proposals discussed before the two leaders' summit — underscoring the entrenched mutual mistrust amid contrasting visions on AI... Because of low levels of trust, cooperation between the two superpowers remains limited, said George Chen, chair of digital practice at The Asia Group consultancy. "Beijing continues to believe Washington seeks to contain China's rise in AI and other emerging technologies, a perception that will shape the pace and scope of future engagement for the two countries on AI," he said.
CNN also points out that while China trails the US in frontier AI models, "it's rapidly narrowing the technology gap while championing a more open ecosystem centered on accessibility and lower cost." In July, Chinese leader Xi Jinping launched the World Artificial Intelligence Cooperation Organization — a rival grouping to the Pax Silica alliance that Trump formed last year to reduce reliance on China for AI supply chains. While over two dozen countries and the European Union signed up to Trump's Pax Silica, Xi has recruited 29 countries, including Russia, Indonesia and Pakistan, to his alternative vision of open models, which allow users to freely download, customize and run without paying hefty fees to American firms like Anthropic and OpenAI. For developers in the Global South, an inexpensive Chinese model from DeepSeek or Moonshot may be more useful than a slightly more capable system requiring an expensive subscription and access to a foreign cloud provider, said Eric Olander, editor in chief of The China-Global South Project, a research agency....

China's embrace of open systems has not always been a top-down strategy by Beijing. Restrictions on access to the most advanced chips because of US export controls, coupled with smaller capital markets, have pushed Chinese developers toward open models as a way to compete with leading US proprietary systems. That shift has proved effective. In a year, Chinese models' global usage skyrocketed from less than 15% to over 54% last week, led by DeepSeek, according to AI leaderboard data by OpenRouter, a marketplace for models. Even American firms, from Airbnb and DoorDash to Shopify, have embraced Chinese models, tapping into the advantages of open systems, including lower costs and greater flexibility for customization.

CNN adds this insight from Alex Colville, an analyst focusing on tech and security at the government-backed Australian Strategic Policy Institute. "The more capable Chinese models become, the less likely it is Beijing may leave them unrestricted."

Slashdot Top Deals

If a train station is a place where a train stops, what's a workstation?

Working...