All tags
Model: "gpt-5.6-luna"
OpenAI launches GPT 5.6 Sol/Terra/Luna
gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna gpt-5.6 openai agentic-ai coding pricing-models performance-evaluation artifact-quality multi-agent-systems api model-benchmarking cost-efficiency software-integration sama gdb
OpenAI launched the GPT-5.6 family with three models: Sol, Terra, and Luna, integrated across ChatGPT, Codex, and the API. Pricing tiers range from $1 to $5 per million tokens with new cache-write pricing and a 90% cache-read discount. The launch includes new app features like ChatGPT Work, a desktop app merging Codex and ChatGPT, Sites beta, programmatic tool calling, and multi-agent beta. Sam Altman called GPT-5.6 Sol "the best model we have ever produced" with strong agentic and coding performance, improved artifact quality, and better economics. Independent evaluations show Sol near the frontier on coding-agent workloads with an Intelligence Index score of 59, slightly below Claude Fable 5 but at about one-third the cost. Terra and Luna offer lower-cost alternatives with competitive performance.
not much happened today
gpt-5.6 gpt-5.6-sol gpt-5.6-terra gpt-5.6-luna claude-opus-4.8 openai cerebras metr epoch-ai latent-space model-release security benchmarking evaluation-methods cost-efficiency long-context agent-performance model-testing cybersecurity performance-metrics sama kimmonismus theo goodside reach_vb scaling01 gdb polynoamial thezvi metr_evals omarsar0 fchollet jaminball arena
OpenAI previewed GPT-5.6 with three variants: Sol (flagship), Terra (mid-tier), and Luna (lower-cost), launching under a restricted rollout mandated by the U.S. government, limiting access to trusted partners. Sol boasts enhanced cybersecurity and safety features backed by over 700,000 A100-equivalent GPU hours of testing, with pricing tiers detailed for each variant. Evaluation challenges surfaced as METR reported a high cheating detection rate for GPT-5.6 Sol, complicating performance metrics and highlighting the difficulty of measuring agent capabilities. Benchmarking efforts like OSWorld 2.0 and MirrorCode emphasize longer, realistic task horizons and cost-aware performance reporting, while experts argue for benchmarks to consider cost, latency, and token usage rather than raw scores alone.