All tags
Topic: "context-compaction"
not much happened today
gpt-5.6-luna gpt-5.6-terra gpt-5.6-sol arc-agi-3 inkling-small inkling gemini-robotics-2 openai thinking-machines lmsys modal unsloth artificial-analysis google price-optimization agent-systems memory-retention context-compaction multimodality mixture-of-experts model-compression benchmarking open-weights multimodal-models model-efficiency model-deployment embodied-ai robotics long-context sama fchollet kimmonismus gneubig scaling01 mervenoyann
OpenAI aggressively cut prices for GPT-5.6 Luna by 80% and Terra by 20%, introducing a faster Sol Fast tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The ARC-AGI-3 debate highlighted that the complete agent system, including memory retention and tool orchestration, is critical beyond just the base model. Thinking Machines released Inkling-Small, an open-weights, multimodal MoE model with 276B parameters (12B active), delivering performance comparable to the original Inkling at a quarter of the size, supporting audio, images, and Python-based image inspection. Benchmarks show Inkling-Small excels in coding and multimodality tasks, with 1M-context support and broad open inference stack adoption. The news also mentions Google's Gemini Robotics 2 advancing embodied AI from tabletop to full-body control.
Claude Opus 4.5: 3rd new SOTA coding model in past week, 1/3 the price of Opus
claude-opus-4.5 gemini-3-pro gpt-5.1-codex-max opus-4.1 sonnet-4.5 anthropic amazon google anthropic coding agents tool-use token-efficiency benchmarking api model-pricing model-performance effort-control context-compaction programmatic-tool-calling alexalbert__ btibor91 scaling01 klieret
Anthropic launched Claude Opus 4.5, a new flagship model excelling in coding, agents, and tooling with a significant 3x price cut compared to Opus 4.1 and improved token efficiency using 76% fewer output tokens. Opus 4.5 achieved a new SOTA on SWE-bench Verified with 80.9% accuracy, surpassing previous models like Gemini 3 Pro and GPT-5.1-Codex-Max. The update includes advanced API features such as effort control, context compaction, and programmatic tool calling, improving tool accuracy and reducing token usage. Claude Code is now bundled with Claude Desktop, and new integrations like Claude for Chrome and Excel are rolling out. Benchmarks show Opus 4.5 breaking the 80% barrier on SWE-bench Verified and strong performance on ARC-AGI-2 and BrowseComp-Plus.