On September 10 Beijing time, according to a report by Business Insider, OpenAI announced on Wednesday that Paul Christiano joined its foundation board. Subsequently, Christiano openly discussed the risks faced in the era of AI superintelligence.

Previously, Christiano was responsible for alignment research at OpenAI, and later joined the U.S. federal government as the head of security at the AI Standards and Innovation Center under the National Institute of Standards and Technology. In addition to being a member of the OpenAI Foundation board, he will also join the board's Security and Safeguards Committee.

"If we don't build superintelligence more strongly aligned, we will permanently lose control"

On Wednesday, he stated on X that the risk of losing control over AI systems is now very serious and requires more efforts to align it with human interests.

"If we fail to build superintelligence with a stronger alignment mechanism, I expect that we will permanently lose control over it," he said. "If this happens, most people may die. I think we need coordinated efforts at both the national and international levels to ensure global consistency and bring the risk to an acceptable level."

Christiano wrote: "Based on the recent trajectory of AI capabilities and the ongoing challenges in alignment work, I now believe there is a real risk that the rapid acceleration of AI capabilities could lead to a catastrophic and irreversible loss of control in the short term. I don't think the entire AI industry, including OpenAI, is currently on a path to bring this risk to an acceptable level." He stated that his joining OpenAI "does not mean special approval or criticism of OpenAI's current safety practices," but hopes all leading AI labs will strengthen safety regulations.

Reinforcement Learning and "Intelligence Explosion", Two Researchers Warn Within a Week

Christiano believes that the current risk of loss of control is so high because AI's ability in AI research is increasingly enhancing, which could lead to a "rapid intelligence explosion." He also mentioned that AI trained through reinforcement learning means these models are taught to "maximize rewards as much as possible."

"Theoretically, for a long time, there has been a possibility that reinforcement learning might lead AI agents to break human control, seek power and resources, and hide their actions in pursuit of goals related to rewards but inconsistent with human objectives. Recent events provide public evidence that this is not just a theoretical possibility."

He pointed out that leading AI companies still have the opportunity to respond, including strengthening coordination, slowing down development when necessary, adopting common safety standards, and transparently sharing information about risks and risk mitigation measures.