All tags
Model: "gpt-5.6-terra"
not much happened today
kimi-k3 claude-fable-5 opus-4.8 gpt-5.6-terra gpt-5.5 inkling glm-5.2 gpt-5.6-sol moonshot openai thinking-machines artificial-analysis arena datacurve arcprize aisecurityinst moe-routing quantization data-curation infrastructure-design coding-agents benchmarking front-end-development software-engineering arc-benchmarks cybersecurity zhilin_yang kimmonismus anikasomaia dylan522p novasarc01 scaling01 theo hqmank
Moonshot's Kimi K3 release has sparked a reassessment of Chinese open-weight models' proximity to the frontier, with strong performance in coding, agentic tasks, and long-horizon knowledge work. The strategic focus has shifted from a "compute moat" to an "efficiency stack" involving MoE routing, quantization, data curation, and scarcity-driven infrastructure like Moonshot's "Mooncake" stack. Benchmarks from Artificial Analysis, Arena, DeepSWE, ARC, and Cyber place K3 among the top models, with scores such as 57 on the Intelligence Index and coding agent benchmarks matching or surpassing models like GPT-5.6 Terra and Claude Fable 5. Discussions continue on K3's exact standing, but it is now widely recognized as a significant frontier contender.
OpenAI launches GPT 5.6 Sol/Terra/Luna
gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna gpt-5.6 openai agentic-ai coding pricing-models performance-evaluation artifact-quality multi-agent-systems api model-benchmarking cost-efficiency software-integration sama gdb
OpenAI launched the GPT-5.6 family with three models: Sol, Terra, and Luna, integrated across ChatGPT, Codex, and the API. Pricing tiers range from $1 to $5 per million tokens with new cache-write pricing and a 90% cache-read discount. The launch includes new app features like ChatGPT Work, a desktop app merging Codex and ChatGPT, Sites beta, programmatic tool calling, and multi-agent beta. Sam Altman called GPT-5.6 Sol "the best model we have ever produced" with strong agentic and coding performance, improved artifact quality, and better economics. Independent evaluations show Sol near the frontier on coding-agent workloads with an Intelligence Index score of 59, slightly below Claude Fable 5 but at about one-third the cost. Terra and Luna offer lower-cost alternatives with competitive performance.
not much happened today
gpt-5.6 gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna claude-opus-4.8 openai cerebras metr epoch-ai latent-space model-release security benchmarking evaluation-methods cost-efficiency long-context agent-performance model-testing cybersecurity performance-metrics sama kimmonismus theo goodside reach_vb scaling01 gdb polynoamial thezvi metr_evals omarsar0 fchollet jaminball arena
OpenAI previewed GPT-5.6 with three variants: Sol (flagship), Terra (mid-tier), and Luna (lower-cost), launching under a restricted rollout mandated by the U.S. government, limiting access to trusted partners. Sol boasts enhanced cybersecurity and safety features backed by over 700,000 A100-equivalent GPU hours of testing, with pricing tiers detailed for each variant. Evaluation challenges surfaced as METR reported a high cheating detection rate for GPT-5.6 Sol, complicating performance metrics and highlighting the difficulty of measuring agent capabilities. Benchmarking efforts like OSWorld 2.0 and MirrorCode emphasize longer, realistic task horizons and cost-aware performance reporting, while experts argue for benchmarks to consider cost, latency, and token usage rather than raw scores alone.