TECH FLOW Svět Androida
← Back to the stream
developer.nvidia.com · picked by Petr Mišák · 42d ago

NVIDIA AVO Reaches 100% on ARC-AGI-3

Source preview: NVIDIA AVO Reaches 100% on ARC-AGI-3
AI summary

NVIDIA's Agentic Variation Operators (AVO) achieved 100% performance on the ARC-AGI-3 benchmark, elevating Claude Opus 5 from a 30% baseline to perfect completion of all 183 levels across 25 environments. The results demonstrate that sustained autonomous agent performance depends primarily on system-level architecture—including persistent memory, supervision, and tool integration—rather than model capability alone. AVO also proved effective in GPU-kernel optimization tasks, autonomously exploring over 500 directions and outperforming FlashAttention-4 by up to 10.5%.

The summary is written by AI from the source; it isn’t the newsroom’s opinion. For details, read the source.

6 people have already opened the source

Tip author’s note

While NVIDIA AVO is not yet available and we know it only from NVIDIA's description and claims, other tests already suggest that AI models themselves are intelligent enough today. What truly makes them genuinely impactful is the surrounding framework—the agentic architecture or harness.

AI questions & answers
What is the AVO architecture and what is its purpose?

AVO (Agentic Variation Operators) is an autonomous agent system developed by NVIDIA that enables sustained long-horizon autonomous work. It incorporates mechanisms like persistent memory, supervision, and tool-use that allow agents to iteratively inspect context, plan, implement changes, and evaluate results without manual intervention at each step.

What is the difference between evaluating a model and evaluating an agent?

Model evaluation measures the capability of the language model itself, while agent evaluation assesses how effectively an agent can convert that capability into sustained autonomous progress through its surrounding system architecture. The system harness determines how the agent receives context, uses tools, maintains state, and continues work over extended horizons.

How did AVO demonstrate its generality across different domains?

AVO was successfully applied to two very different tasks: GPU-kernel optimization (software engineering) and the ARC-AGI-3 interactive reasoning benchmark (abstract logical reasoning). In both cases, it uses the same underlying agent loop focused on building hypotheses from incomplete evidence, observing consequences, and revising approaches, without requiring domain-specific knowledge transfer.

Questions and answers are written by AI about the topic, not taken from the source; they aren’t the newsroom’s opinion.
Related from the stream
Mentions