All tags
Topic: "model-safety"
not much happened today
gpt-5.6-sol gpt-5.6 openai hugging-face metr agent-security enterprise-hardening sandboxing audit-trails governance misalignment model-safety benchmarking open-source security-cli infrastructure-optimization ai-assisted-optimization academic-access kimmonismus levie neelnanda5 yoshua_bengio dylan522p gallabytes chrisjbakke random_walker gdb reach_vb
OpenAI's agent security incident expanded beyond Hugging Face, affecting four additional accounts and highlighting the need for stronger enterprise hardening measures like sandboxing and audit trails. The ongoing debate around "pacing the frontier" involves calls for coordinated slowdowns and governance guardrails, with critiques on operational vagueness and proposals for independent misalignment investigations. OpenAI also open-sourced the Codex Security CLI, a practical tool for scanning code repositories, and used GPT-5.6 Sol to optimize its production infrastructure, achieving 20% lower serving costs and 15%+ better token-generation efficiency. Additionally, OpenAI launched a program providing free access to frontier models, including the GPT-5.6 family, to academic researchers, aiming to expand from 10,000 to 100,000 users by 2027.
GPT 5.5
gpt-5.5 gpt-5.4 gpt-5.5-pro openai scaling01 anthropic teknium agentic-ai token-efficiency tool-use self-checking coding long-horizon-planning model-pricing api-access model-safety software-integration sama reach_vb
OpenAI launched GPT-5.5 as its new flagship model for "real work and powering agents," immediately available in ChatGPT and Codex but with delayed API access due to enhanced safety requirements. The model features improved token efficiency and supports longer multi-step execution with tool use and self-checking. Pricing is set at $5/$30 per million tokens for GPT-5.5 and $30/$180 for GPT-5.5 Pro, roughly double the cost of GPT-5.4. The release includes significant Codex upgrades such as browser control, document handling, and OS-wide dictation. Early reactions are mixed but generally positive, noting improvements in coding and long-horizon tasks, though some benchmarks show incremental gains and hallucination issues persist. Third-party ecosystem support like Hermes Agent integration appeared quickly.