All tags
Person: "mikhail_parakhin"
Opus 5
claude-opus-5 fable-5 claude-opus-4.8 anthropic epoch nous-research microsoft benchmarking software-engineering coding-agents agentic-ai model-evaluation model-performance browser-automation kevin_scott mikhail_parakhin abacaj scaling01 jerhadf arena witcheer
Anthropic launched the Claude Opus 5 model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. The model achieved an Epoch Capabilities Index (ECI) of 159, slightly below Fable 5's 161, but matched Fable 5 on software engineering benchmarks. Users debated the accuracy of these scores, with some calling the model "incredibly underrated" and advocating for harder public benchmarks. Technical discussions highlighted an unusual benchmark behavior where Opus 5 performed better at medium effort than high effort on FrontierCode. Early user anecdotes praised Opus 5's browser control and agentic tool use, while community evaluations and leaderboard scores were still forthcoming. Nous Research provided access to Opus 5 with a 20% discount. Microsoft CTO Kevin Scott and others noted Opus 5's strong performance in math and coding tasks.
not much happened today
claude gpt-5.2-pro dgm-h rllm anthropic meta-ai-fair agent-frameworks workflow-automation multi-agent-systems reinforcement-learning reward-models self-improving-agents benchmark-generation operational-efficiency closed-loop-feedback jenny_zhang jase_weston mikhail_parakhin jeremyphoward
Anthropic introduced Claude Cowork and Claude Code enabling desktop control of mouse, keyboard, and screen in a macOS research preview, expanding agent capabilities beyond APIs and browsers. The agent ecosystem is evolving towards long-running, parallel, tool-rich workflows with projects like Hermes Agent, T3 Code, Command Center, and Parchi enhancing multi-agent orchestration and autonomous task management. Operational challenges such as fragility and inefficiency in subagents, including GPT-5.2 Pro and Claude browser/computer use, highlight the need for closed-loop feedback systems. Research from Meta AI advances self-improving agents with Hyperagents / DGM-H enabling meta-level procedural improvements, and unifies reinforcement learning post-training with RLLM (RL + LM-as-RM) to improve reward modeling across task types. Additionally, WebArena-Infinity drastically reduces browser environment construction costs, accelerating benchmark and environment generation.