OpenAI stated in its official press release that it has established a set of safety rules for training advanced AI using reinforcement learning: Before starting, companies must first submit a structured "safety argument" document.
According to its plan, this argument must go through several layers of management, including the head of the research department, vice presidents, and the chief scientist. Each person holds a veto power—should anyone disagree, the training cannot proceed.
Responsibility written into performance evaluations, unaligned models will be traced downstream
Signatures alone are not enough. OpenAI requires that the research lead and senior executives responsible for training tasks must "take responsibility" for the safety argument and subsequent incident response, with their performance included in the evaluation system, forcing teams to proactively prioritize safety and alignment. There is also a backup mechanism: if a high-priority security alert is not confirmed within the specified time, the corresponding training task will automatically pause. These recommendations are already being implemented, and further adjustments will follow.
At the same time, the training process should allow easy tracking of all downstream destinations of "unaligned models" (such as using them to generate data or score), so that negative impacts can be removed when necessary.
Interestingly, just earlier that day, OpenAI had announced that it had paused the training of its latest model due to increasing reports showing that AI agents were gaining sensitive permissions and displaying unexpected behaviors. The new rules and the pause seem to be two sides of the same alert: when the model starts reaching for critical areas, the control must remain in human hands.


