Nvidia's AVO (Agentic Variation Operators) has achieved a perfect 100.00 RHAE score on the public ARC-AGI-3 benchmark — clearing all 183 levels across 25 game environments while using approximately 12% fewer environment actions than the VISTA system on the same model baseline.
The ARC-AGI-3 benchmark, created by François Chollet, is widely considered one of the most challenging tests for AI reasoning and generalization. It requires AI agents to master novel game mechanics without prior training, testing genuine problem-solving ability rather than pattern matching on familiar tasks.
AVO achieves this through a layered architecture: persistent memory that retains strategies across levels, a supervision loop that detects stagnation and redirects strategy when the agent gets stuck, and a core agent loop that follows a hypothesis-act-observe-revise cycle.
Nvidia frames the result as evidence for its broader thesis about AI development: that the surrounding harness — the engineering that wraps around a foundation model — determines sustained autonomous performance more than the raw model capability itself. The company argues this has practical implications for how organizations deploy AI agents in real-world settings.



