All tags
Topic: "browser-automation"
Opus 5
claude-opus-5 fable-5 claude-opus-4.8 anthropic epoch nous-research microsoft benchmarking software-engineering coding-agents agentic-ai model-evaluation model-performance browser-automation kevin_scott mikhail_parakhin abacaj scaling01 jerhadf arena witcheer
Anthropic launched the Claude Opus 5 model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. The model achieved an Epoch Capabilities Index (ECI) of 159, slightly below Fable 5's 161, but matched Fable 5 on software engineering benchmarks. Users debated the accuracy of these scores, with some calling the model "incredibly underrated" and advocating for harder public benchmarks. Technical discussions highlighted an unusual benchmark behavior where Opus 5 performed better at medium effort than high effort on FrontierCode. Early user anecdotes praised Opus 5's browser control and agentic tool use, while community evaluations and leaderboard scores were still forthcoming. Nous Research provided access to Opus 5 with a 20% discount. Microsoft CTO Kevin Scott and others noted Opus 5's strong performance in math and coding tasks.
GPT-Realtime-2, -Translate, and -Whisper: new SOTA realtime voice APIs
gpt-realtime-2 gpt-5.5 codex openai anthropic goodfireai scale-ai voice-models streaming-translation transcription benchmarking context-windows browser-automation cybersecurity interpretability neural-geometry manifolds ai-safety rlhf micahcarroll milesbrundage ryanpgreenblatt
OpenAI released GPT-Realtime-2, a voice model with GPT-5-class reasoning, tool use, interruption handling, and extended context windows up to 128K tokens, achieving top scores on Big Bench Audio and Conversational Dynamics benchmarks. They also launched a Chrome plugin for Codex enabling browser control and multitasking, and introduced GPT-5.5 with Trusted Access for Cyber for secure defensive workflows and red teaming. Anthropic introduced Natural Language Autoencoders for interpreting model activations as human-readable text, aiding interpretability and debugging, while Goodfire proposed a neural geometry research agenda focusing on manifolds as primitives for neural network behavior. Anthropic also announced The Anthropic Institute to advance AI safety and economic resilience research.
not much happened today
OpenAI expanded its Agents SDK by separating the agent harness from compute/storage, enabling long-running, durable agents with features like file/computer use, skills, memory, and compaction. The harness is now open-source and supports execution via partner sandboxes, fostering a new ecosystem with integrations from Cloudflare, Modal, Vercel, and others. Cloudflare launched Project Think, a next-gen Agents SDK with durable execution and sandboxed code, alongside Agent Lee, a prompt-driven UI agent using sandboxed TypeScript, and introduced real-time voice pipelines and browser automation tools. Hermes Agent focuses on persistent skill formation by learning from completed workflows, positioning itself as a professional agent distinct from GUI-first assistants like OpenClaw. "Hermes autonomously backfills tracking data, updates cron jobs, and saves workflows as reusable skills," highlighting its advanced workflow management capabilities.