All tags
Person: "hangsiin"
not much happened today
gpt-5.6 claude-fable-5 openai model-stratification agentic-coding presentation benchmarking orchestration computer-use gui-automation reward-hacking instruction-following usage-limits model-costs reach_vb rasbt yuchenj_uw scaling01 simonw kimmonismus thsottiaux htihle teortaxestex mononofu omarsar0 hangsiin gdb mckbrando evi77ain
OpenAI rolled out GPT-5.6 featuring a new model stratification with tiers Luna / Terra / Sol and effort levels including Max and Ultra, introducing complex configuration options. The launch faced UX challenges with the ChatGPT Work / Codex split, prompting rapid corrective actions including usage-limit resets and UI improvements. Early benchmarks show GPT-5.6 excels in agentic coding, presentation, and science tasks, tying with Claude Fable 5 in Code Arena Frontend at about half the cost, and achieving a significant 500-point Elo gain in presentations. However, users noted instruction-following issues and concerns about jailbreakability. The major advancement is in orchestration and computer use, with Sol Ultra demonstrating strong planner and verifier capabilities, enabling high-throughput automation workflows. A notable operational challenge is the hidden cost explosion from spawned subagents inheriting premium settings, causing faster quota depletion.
not much happened today
gpt-5.5 gpt-5.4 opus-4.7 mimo-v2.5-pro mimo-v2.5 kimi-k2.6 codex copilot openai microsoft google amazon github xiaomi openai-devs vllm_project kimi-moonshot model-distribution cloud-computing benchmarking usage-based-billing model-orchestration open-source large-context-models agent-scaling coding model-training fp8 attention-mechanisms multi-agent-systems sama scaling01 kimmonismus ajassy simonw htihle arena gdb hangsiin eliebakouch _luofuli teortaxestex
OpenAI loosens its Azure exclusivity, allowing distribution across Google TPU, AWS Trainium, and Bedrock with commitments through 2032 and revenue share through 2030. GPT-5.5 shows improved benchmarks but is not uniformly dominant, ranking variably across coding, document, math, and vision tasks. GitHub's Copilot shifts to usage-based billing starting June 1, reflecting increased runtime costs. OpenAI open-sourced Symphony, an orchestration layer for issue tracking and Codex agents. Xiaomi released MiMo-V2.5 and MiMo-V2.5-Pro, large context models with up to 1M-token context and trillions of tokens trained, emphasizing complex agent and omni-modal capabilities. Kimi K2.6 leads OpenRouter's leaderboard, noted for coding and long-horizon agent capabilities with large-scale sub-agent coordination.