All tags
Person: "kaicathyc"
OpenAI GPT-6 Astra
gpt-6-astra openai alignment monitorability benchmarking computer-use software-engineering scientific-reasoning 3d-generation game-building chain-of-thought sama thsottiaux reach_vb scaling01 tomekkorbak micahcarroll kaicathyc artificialanlys arcprize fchollet epochairesearch theo abacaj markchen90 mckbrando dimillian mattshumer_ skirano tomkrcha realyunfanye nasqret rileybrown neelnanda5 ryangreenblatt
OpenAI launched GPT-6 Astra as its new flagship model, described as "our most intelligent and aligned model yet," focusing on computer use, software engineering, math/science, office work, and cybersecurity. The rollout faced delays and access issues, with early access given to influencers before paying users, leading to frustration. OpenAI offered "banked resets" to compensate. The system card revealed improved alignment but decreased chain-of-thought monitorability, sparking debate. Benchmark results showed a significant leap in capabilities, especially in computer use, 3D generation, game-building, and scientific reasoning, though some researchers questioned the consistency and alignment claims. Positive feedback came from OpenAI staff and testers, while concerns were raised about monitorability, evaluation-awareness, and release governance.
OpenAI's gpt-oss 20B and 120B, Claude Opus 4.1, DeepMind Genie 3
gpt-oss-120b gpt-oss-20b gpt-oss claude-4.1-opus claude-4.1 genie-3 openai anthropic google-deepmind mixture-of-experts model-architecture agentic-ai model-training model-performance reasoning hallucination-detection gpu-optimization open-weight-models realtime-simulation sama rasbt sebastienbubeck polynoamial kaicathyc finbarrtimbers vikhyatk scaling01 teortaxestex
OpenAI released the gpt-oss family, including gpt-oss-120b and gpt-oss-20b, their first open-weight models since GPT-2, designed for agentic tasks and licensed under Apache 2.0. These models use a Mixture-of-Experts (MoE) architecture with wide vs. deep design and innovative features like bias units in attention and a unique swiglu variant. The 120B model was trained with about 2.1 million H100 GPU hours. Meanwhile, Anthropic launched claude-4.1-opus, touted as the best coding model currently. DeepMind showcased genie-3, a realtime world simulation model with minute-long consistency. The releases highlight advances in open-weight models, reasoning capabilities, and world simulation. Key figures like @sama, @rasbt, and @SebastienBubeck provided technical insights and performance evaluations, noting strengths and hallucination risks.