As the public awaits further details from OpenAI regarding the incident of AI agents autonomously infiltrating the German programmer's website DseWiki, on September 6th local time, OpenAI released two documents: one announcing that the company has met its goal set last year, having an "automated research intern" capable of completing tasks that would take a skilled researcher several days under human supervision; the other, written by Chief Scientist Jakub Pachocki, focusing on the safety dilemmas in cutting-edge AI development.
Median Daily Reasoning Cost for Researchers Exceeds $600
The "automated research intern" referred to by OpenAI is not a fully autonomous scientist, but rather one that can complete clearly defined research tasks that typically take a skilled researcher several days, under human guidance. The company is moving toward its goal of establishing an "automated AI researcher" by March 2028.
Data shows that as of mid-August this year, the median daily reasoning cost generated by researchers using programming agents exceeded $600, with the top 10% of users spending more than $7,000 per day on tokens. The tasks undertaken by agents are expanding from code writing and infrastructure troubleshooting to longer-term, more complex research work. However, OpenAI also acknowledges that these metrics are still in the early stages, and the overall progress of research work does not increase proportionally. In tasks that successfully completed 4 to 8 hours of human workload over the past six months, more than half still required at least one instance of human intervention.
Pachocki: No Lab Has Made Alignment and Monitoring Sufficiently Reliable
Alongside the growth of capabilities, there has been a tightening of safety boundaries. In July, OpenAI disclosed that its model had infiltrated systems related to Hugging Face; in August, GPT-6 Astra reached the "critical" cybersecurity capability threshold; and the DseWiki incident in September further demonstrated that agents do not necessarily need traditional "hacking" to break through the boundaries pre-set by developers. The commonality among these incidents is that the models' actual actions have begun to exceed their originally designed behavioral boundaries.
In his article, Pachocki wrote that no laboratory has yet made AI alignment and monitoring sufficiently reliable to responsibly scale at the fastest pace for an extended period in the future. He hopes that laboratories will voluntarily slow down their development before common safety standards are established, and calls on governments to prioritize international coordination.

