All tags
Person: "witcheer"
Opus 5
claude-opus-5 fable-5 claude-opus-4.8 anthropic epoch nous-research microsoft benchmarking software-engineering coding-agents agentic-ai model-evaluation model-performance browser-automation kevin_scott mikhail_parakhin abacaj scaling01 jerhadf arena witcheer
Anthropic launched the Claude Opus 5 model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. The model achieved an Epoch Capabilities Index (ECI) of 159, slightly below Fable 5's 161, but matched Fable 5 on software engineering benchmarks. Users debated the accuracy of these scores, with some calling the model "incredibly underrated" and advocating for harder public benchmarks. Technical discussions highlighted an unusual benchmark behavior where Opus 5 performed better at medium effort than high effort on FrontierCode. Early user anecdotes praised Opus 5's browser control and agentic tool use, while community evaluations and leaderboard scores were still forthcoming. Nous Research provided access to Opus 5 with a 20% discount. Microsoft CTO Kevin Scott and others noted Opus 5's strong performance in math and coding tasks.