CHAPTERS
- 0:00 – 0:06
Sholto introduces Claude Opus 4.5 as Anthropic’s top model
Sholto, an engineer and researcher at Anthropic, frames Claude Opus 4.5 as the company’s best model to date. He highlights strong performance across coding, agentic work, and everyday productivity tasks like spreadsheets.
- •Speaker introduction and context (Anthropic engineer/researcher)
- •Positioning Opus 4.5 as “best model yet”
- •Claims leadership in coding and agentic tasks
- •Emphasis on practical day-to-day work (e.g., spreadsheets)
- 0:06 – 0:12
Reliability and intuition: “it just gets it”
He notes that Opus 4.5 is difficult to demo purely via benchmarks because its advantage is qualitative. The model feels more trustworthy and aligned with intent, leading to smoother collaboration.
- •Hard-to-measure advantage: intuitive understanding
- •Greater user trust in outputs
- •Qualitative improvement over prior experiences
- •Implied reduction in friction during real work
- 0:12 – 0:18
Fewer interventions: longer autonomous runs on tasks
Sholto describes needing to step in less often while the model works. He also shares anecdotal team feedback where Opus 4.5 solved bugs that previous models (e.g., Sonnet) struggled to find.
- •Increasing time between human interventions
- •Better at debugging and root-cause discovery
- •Anecdotal comparison vs. Sonnet
- •Team-reported wins in real engineering workflows
- 0:18 – 0:24
Efficiency via deliberation: knowing when to think before acting
Opus 4.5 is described as more efficient not only by speed but by making fewer wrong turns. It reportedly pauses to “think” before making changes, improving the likelihood of correct edits.
- •Efficiency framed as fewer mistakes, not just faster output
- •Better judgment about when to deliberate
- •More accurate code or file changes
- •Reduced rework from incorrect modifications
- 0:24 – 0:30
Benchmark highlight: top score on a two-hour take-home engineering task
He cites a demanding two-hour engineering take-home where Opus 4.5 outperformed every human score recorded. This is used as evidence of strong end-to-end engineering capability under time constraints.
- •Reference to a two-hour intensive engineering take-home
- •Model scored higher than any human recorded
- •Claimed robustness on complex, realistic tasks
- •Performance framed as exceptional among both models and humans
- 0:30
Improved frontend and vision for better computer use
Sholto calls out specific capability gains: significantly better frontend work and substantially better vision. These improvements make Opus 4.5 more effective at using computers and interacting with visual interfaces.
- •Notable upgrade in frontend development capability
- •Major improvement in vision performance
- •Better at “using computers” via visual/UI understanding
- •Broader applicability to workflows involving screens and interfaces
Availability: launching today across major cloud platforms
He closes with a release and distribution update: Opus 4.5 is available immediately. For the first time, it’s offered on every major cloud platform, and viewers are invited to share feedback.
- •Opus 4.5 available “today”
- •First time on every major cloud platform
- •Emphasis on accessibility and deployment options
- •Call for user feedback
