# Nvidia's AVO harness lifts Claude Opus 5 from 30% to 100% on ARC-AGI-3

> Nvidia's AVO agent system completed all 183 ARC-AGI-3 levels using Claude Opus 5.

*The same model solved every public puzzle once wrapped in a system with memory, tools and a supervisor.*

By Behzad Hosseini · FeaturedDaily
Canonical: https://featureddaily.com/news/nvidia-s-avo-harness-lifts-claude-opus-5-from-30-to-100-on-arc-agi-3

**What happened:** Nvidia said its AVO (Agentic Variation Operators) system, running Claude Opus 5, completed all 183 levels across all 25 public ARC-AGI-3 environments on 21 August 2026. The model worked out what to do with no instructions, rules or stated goals.

**The numbers:** ARC Prize had put Claude Opus 5 alone at 30.2% on the public set, at high reasoning effort. AVO, using the same model, hit 100%. It also did so more efficiently: 6,624 environment actions versus 7,542 for the rival VISTA harness, a 12% reduction.

**The context:** ARC-AGI-3 is an interactive benchmark of game-like puzzles built to test whether AI can learn new skills the way people do. AVO adds a main agent, persistent memory, tools, and a supervisor that intervenes when progress stalls. Nvidia first introduced AVO in March 2026, and the same unchanged agent had earlier spent seven days optimising GPU kernels, beating FlashAttention-4 by up to 10.5% on DGX B200.

**Why it matters:** the model itself did not change. The system wrapped around it did. That's the takeaway most commentators drew from the jump, and it points to how much of an AI's apparent capability depends on the scaffolding, not just the underlying model.

**The catch:** Nvidia itself cautions the jump isn't a clean measure of AVO's contribution, since configurations differ between runs. The public set is also tutorial-like and had already been solved by other harnesses before AVO. There's no private-set score yet.

**In their words:** ARC's creator François Chollet praised the approach but noted that "the private set is the real test of generalisation."

**What's next:** AVO remains a research project rather than a product. The private-set result, once available, will show whether the gains hold up beyond the public tutorial-style puzzles.

The findings come from [The New Stack](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/), reporting on work by a five-person Nvidia team: Terry Chen, Yeyin Zhu, Zhifan Ye, Jean-Francois Puget and Humphrey Shi.

## Key takeaways

- Claude Opus 5 alone scored 30.2% on ARC-AGI-3's public set
- Wrapped in Nvidia's AVO system, it solved all 183 levels
- Private-set score is missing; that's the real generalisation test

## Sources

- [Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia's AVO, it hit 100%](https://thenewstack.io/nvidia-avo-arcagi3-benchmark/) — The New Stack, 2026-08-21
