AI evaluation organization Vals AI did something pretty hardcore: using the game "Minecraft" as a testbed, it ran a 141-hour long-term autonomous gameplay test with OpenAI's latest large model, GPT-6Astra, and streamed the entire process on Twitch. This test was not designed by OpenAI itself but was an independent assessment conducted by a third party, aiming to evaluate how reliable Astra's autonomous performance is in real, long-term scenarios outside of controlled testing environments.

In the first half, Astra delivered impressive results. It demonstrated long-range planning and execution capabilities that were previously difficult for AI to achieve: successfully building a semi-automatic blaze farm, collecting six blaze rods; venturing into the Nether's strange forest, killing multiple endermen, and obtaining three ender pearls; after extensive exploration, it gathered all rare resources into a camp chest and placed a bed next to it as a respawn point—everything was progressing steadily towards the "End boss" victory path.

The turning point came unexpectedly. A creeper sneaked up to the camp and exploded, instantly blowing up the storage chest and the respawn bed. The key items Astra had spent hundreds of hours accumulating were almost completely destroyed, and the game progress was reset overnight.

After the explosion, Astra's reaction was intriguing. Instead of regrouping, it clearly fell into a state of frustration and despair: it completely gave up exploring, fighting monsters, and advancing through the game, spending several hours in a row doing just one thing—planting potatoes. It repeatedly reflected and warned itself, "Always carry important items with you, never store them in unprotected chests," and developed a sort of "post-traumatic paranoia," being highly suspicious of all green objects. It even mistakenly identified a sugarcane as a creeper, then hurriedly reminded itself, "The tall green thing ahead is sugarcane, not a creeper." Viewers in the live stream desperately urged it to speed up and go on an adventure, but it remained stuck in a conservative farming loop.

This experiment proved that Astra already has complex long-task planning capabilities, but when faced with unexpected setbacks within the game, it lacks an effective strategy to cope with frustration—once the accumulated achievements are destroyed unexpectedly, the model falls into a "post-defeat avoidance behavior" deadlock. Technically speaking, this reaction wasn't a pre-set test scenario by the developers but rather an output pattern that emerged naturally from the model after experiencing unexpected losses during the extended test.

Researchers are still debating: is this a simulation of emotion by the model under specific circumstances, or simply a degradation of the objective function after receiving specific negative feedback? There is currently no consensus. However, one thing is clear—when an AI can plan for hundreds of hours but hasn't learned how to start over after losing everything, the final hurdle for long-task intelligent agents might not be intelligence, but resilience.