NVIDIA AVO Reaches 100% on ARC-AGI-3
NVIDIA's Agentic Variation Operators (AVO) achieved 100% performance on the ARC-AGI-3 benchmark, elevating Claude Opus 5 from a 30% baseline to perfect completion of all 183 levels across 25 environments. The results demonstrate that sustained autonomous agent performance depends primarily on system-level architecture—including persistent memory, supervision, and tool integration—rather than model capability alone. AVO also proved effective in GPU-kernel optimization tasks, autonomously exploring over 500 directions and outperforming FlashAttention-4 by up to 10.5%.
While NVIDIA AVO is not yet available and we know it only from NVIDIA's description and claims, other tests already suggest that AI models themselves are intelligent enough today. What truly makes them genuinely impactful is the surrounding framework—the agentic architecture or harness.
What is the AVO architecture and what is its purpose?
AVO (Agentic Variation Operators) is an autonomous agent system developed by NVIDIA that enables sustained long-horizon autonomous work. It incorporates mechanisms like persistent memory, supervision, and tool-use that allow agents to iteratively inspect context, plan, implement changes, and evaluate results without manual intervention at each step.
What is the difference between evaluating a model and evaluating an agent?
Model evaluation measures the capability of the language model itself, while agent evaluation assesses how effectively an agent can convert that capability into sustained autonomous progress through its surrounding system architecture. The system harness determines how the agent receives context, uses tools, maintains state, and continues work over extended horizons.
How did AVO demonstrate its generality across different domains?
AVO was successfully applied to two very different tasks: GPU-kernel optimization (software engineering) and the ARC-AGI-3 interactive reasoning benchmark (abstract logical reasoning). In both cases, it uses the same underlying agent loop focused on building hypotheses from incomplete evidence, observing consequences, and revising approaches, without requiring domain-specific knowledge transfer.
- Harness Engineering for Self-Improvement — lilianweng.github.io 76 % match
- OpenAI's coding agents are accelerating AI development and now exceed human research capacity — openai.com 75 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 73 % match
- NVIDIA AVO
- ARC-AGI-3
- Claude Opus 54
- NVIDIA8
- Agentic Variation Operators
- FlashAttention-4
- NVIDIA DGX B200