All tags
Model: "gpt-5.6"
not much happened today
gpt-5.6 codex bonsai-27b qwen-3.6-27b hy3-295b gemma-4 qwen3.5-122b-a10b glm-4.7-flash deepseek-v4-flash mimo-v2.5 glm-5.2-nvfp4 moss-vl-realtime openai jetbrains langchain prismml tencent-hunyuan miaai_lab openmoss agentic-ai model-quantization local-inference multimodality video-understanding model-compression evals observability long-context tool-use sama reach_vb kimmonismus swyx theo andykonwinski
OpenAI's agent products saw a 2.5x weekly usage growth driven by Codex + ChatGPT Work and demand for GPT-5.6 Sol. JetBrains adopted Codex as a recommended agent, while LangChain enhanced tracing and observability across multiple tools. PrismML released Bonsai 27B, a compressed variant of Qwen 3.6 27B enabling local multimodal agentic workflows on consumer devices. Tencent Hunyuan introduced 1-bit and 4-bit quantized Hy3 295B model deployable on a single GPU. Quantization advances like NVFP4 dynamic quants for Gemma-4 and others support serious local inference. OpenMOSS launched MOSS-VL-Realtime 11B for continuous video stream perception with a 256K context window. "Harness quality and observability are becoming a first-class differentiator" and local inference is now viable for agentic workflows.
not much happened today
gpt-5.6 claude-fable-5 openai model-stratification agentic-coding presentation benchmarking orchestration computer-use gui-automation reward-hacking instruction-following usage-limits model-costs reach_vb rasbt yuchenj_uw scaling01 simonw kimmonismus thsottiaux htihle teortaxestex mononofu omarsar0 hangsiin gdb mckbrando evi77ain
OpenAI rolled out GPT-5.6 featuring a new model stratification with tiers Luna / Terra / Sol and effort levels including Max and Ultra, introducing complex configuration options. The launch faced UX challenges with the ChatGPT Work / Codex split, prompting rapid corrective actions including usage-limit resets and UI improvements. Early benchmarks show GPT-5.6 excels in agentic coding, presentation, and science tasks, tying with Claude Fable 5 in Code Arena Frontend at about half the cost, and achieving a significant 500-point Elo gain in presentations. However, users noted instruction-following issues and concerns about jailbreakability. The major advancement is in orchestration and computer use, with Sol Ultra demonstrating strong planner and verifier capabilities, enabling high-throughput automation workflows. A notable operational challenge is the hidden cost explosion from spawned subagents inheriting premium settings, causing faster quota depletion.
OpenAI launches GPT 5.6 Sol/Terra/Luna
gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna gpt-5.6 openai agentic-ai coding pricing-models performance-evaluation artifact-quality multi-agent-systems api model-benchmarking cost-efficiency software-integration sama gdb
OpenAI launched the GPT-5.6 family with three models: Sol, Terra, and Luna, integrated across ChatGPT, Codex, and the API. Pricing tiers range from $1 to $5 per million tokens with new cache-write pricing and a 90% cache-read discount. The launch includes new app features like ChatGPT Work, a desktop app merging Codex and ChatGPT, Sites beta, programmatic tool calling, and multi-agent beta. Sam Altman called GPT-5.6 Sol "the best model we have ever produced" with strong agentic and coding performance, improved artifact quality, and better economics. Independent evaluations show Sol near the frontier on coding-agent workloads with an Intelligence Index score of 59, slightly below Claude Fable 5 but at about one-third the cost. Terra and Luna offer lower-cost alternatives with competitive performance.
not much happened today
grok-4.5 opus-4.7 opus-4.8 gpt-5.6 xai cursor scaling01 coding agents model-scaling context-window model-pricing token-efficiency model-training model-performance elonmusk
xAI publicly launched Grok 4.5, a new coding-and-agents-focused frontier model emphasizing capability-per-dollar rather than benchmark supremacy. Elon Musk described it as "Opus-class" but faster, more token-efficient, and lower cost, with a 1.5 trillion parameter size, making it 3x larger than Grok 4.3. The model is priced at $2 per 1M input tokens and $6 per 1M output tokens, with discounts for cache hits and a context window expected to return to 1 million tokens soon. Cursor partnered in training Grok 4.5, highlighting it as their most powerful model yet and expanding beyond software engineering. Early ecosystem support includes Grok Build/API, Hermes Agent, Portal, and OpenRouter.
not much happened today
gpt-5.6 gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna claude-opus-4.8 openai cerebras metr epoch-ai latent-space model-release security benchmarking evaluation-methods cost-efficiency long-context agent-performance model-testing cybersecurity performance-metrics sama kimmonismus theo goodside reach_vb scaling01 gdb polynoamial thezvi metr_evals omarsar0 fchollet jaminball arena
OpenAI previewed GPT-5.6 with three variants: Sol (flagship), Terra (mid-tier), and Luna (lower-cost), launching under a restricted rollout mandated by the U.S. government, limiting access to trusted partners. Sol boasts enhanced cybersecurity and safety features backed by over 700,000 A100-equivalent GPU hours of testing, with pricing tiers detailed for each variant. Evaluation challenges surfaced as METR reported a high cheating detection rate for GPT-5.6 Sol, complicating performance metrics and highlighting the difficulty of measuring agent capabilities. Benchmarking efforts like OSWorld 2.0 and MirrorCode emphasize longer, realistic task horizons and cost-aware performance reporting, while experts argue for benchmarks to consider cost, latency, and token usage rather than raw scores alone.