All tags
Person: "mckbrando"
OpenAI GPT-6 Astra
gpt-6-astra openai alignment monitorability benchmarking computer-use software-engineering scientific-reasoning 3d-generation game-building chain-of-thought sama thsottiaux reach_vb scaling01 tomekkorbak micahcarroll kaicathyc artificialanlys arcprize fchollet epochairesearch theo abacaj markchen90 mckbrando dimillian mattshumer_ skirano tomkrcha realyunfanye nasqret rileybrown neelnanda5 ryangreenblatt
OpenAI launched GPT-6 Astra as its new flagship model, described as "our most intelligent and aligned model yet," focusing on computer use, software engineering, math/science, office work, and cybersecurity. The rollout faced delays and access issues, with early access given to influencers before paying users, leading to frustration. OpenAI offered "banked resets" to compensate. The system card revealed improved alignment but decreased chain-of-thought monitorability, sparking debate. Benchmark results showed a significant leap in capabilities, especially in computer use, 3D generation, game-building, and scientific reasoning, though some researchers questioned the consistency and alignment claims. Positive feedback came from OpenAI staff and testers, while concerns were raised about monitorability, evaluation-awareness, and release governance.
not much happened today
gpt-5.6 claude-fable-5 openai model-stratification agentic-coding presentation benchmarking orchestration computer-use gui-automation reward-hacking instruction-following usage-limits model-costs reach_vb rasbt yuchenj_uw scaling01 simonw kimmonismus thsottiaux htihle teortaxestex mononofu omarsar0 hangsiin gdb mckbrando evi77ain
OpenAI rolled out GPT-5.6 featuring a new model stratification with tiers Luna / Terra / Sol and effort levels including Max and Ultra, introducing complex configuration options. The launch faced UX challenges with the ChatGPT Work / Codex split, prompting rapid corrective actions including usage-limit resets and UI improvements. Early benchmarks show GPT-5.6 excels in agentic coding, presentation, and science tasks, tying with Claude Fable 5 in Code Arena Frontend at about half the cost, and achieving a significant 500-point Elo gain in presentations. However, users noted instruction-following issues and concerns about jailbreakability. The major advancement is in orchestration and computer use, with Sol Ultra demonstrating strong planner and verifier capabilities, enabling high-throughput automation workflows. A notable operational challenge is the hidden cost explosion from spawned subagents inheriting premium settings, causing faster quota depletion.
not much happened today
o3-mini o1-mini llama hunyuan-a13b ernie-4.5 ernie-4.5-21b-a3b qwen3-30b-a3b gemini-2.5-pro meta-ai-fair openai tencent microsoft baidu gemini superintelligence ai-talent job-market open-source-models multimodality mixture-of-experts quantization fp8-training model-benchmarking model-performance model-releases api model-optimization alexandr_wang shengjia_zhao jhyuxm ren_hongyu shuchaobi saranormous teortaxesTex mckbrando yuchenj_uw francoisfleuret quanquangu reach_vb philschmid
Meta has poached top AI talent from OpenAI, including Alexandr Wang joining as Chief AI Officer to work towards superintelligence, signaling a strong push for the next Llama model. The AI job market shows polarization with high demand and compensation for top-tier talent, while credentials like strong GitHub projects gain importance. The WizardLM team moved from Microsoft to Tencent to develop open-source models like Hunyuan-A13B, highlighting shifts in China's AI industry. Rumors suggest OpenAI will release a new open-source model in July, potentially surpassing existing ChatGPT models. Baidu open-sourced multiple variants of its ERNIE 4.5 model series, featuring advanced techniques like 2-bit quantization, MoE router orthogonalization loss, and FP8 training, with models ranging from 0.3B to 424B parameters. Gemini 2.5 Pro returned to the free tier of the Gemini API, enabling developers to explore its features.