All tags
Person: "kylebrussell"
Claude Fable 5.1 and Claude Mythos 5.1
claude-fable-5.1 claude-mythos-5.1 astra anthropic openai nous-research perplexity-ai coding model-architecture safety enterprise-ai benchmarking cache-optimization cybersecurity recurrent-depth chain-of-thought model-transparency sama alexalbert__ eliebakouch ethancaballero valsai stevendillmann scaling01 artificialanlys theo teknuim gregkamradt kylebrussell boazbaraktcs kimmonismus
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1, which share base weights but differ in safeguards and routing, showing improved coding performance and usability with a 75% cache-read price cut to $0.25/MTok. Benchmarks highlight strong coding/science results, though Fable 5.1 costs about 20% more per task than its predecessor. Adoption revealed aggressive safety triggers framed as Enterprise Frontier Safeguards for enterprise deployments. Meanwhile, OpenAI previewed Astra, its first model reaching the Critical cybersecurity preparedness level, demonstrating advanced cyber capabilities and employing a recurrent depth/looped transformer architecture, sparking debate on its impact on chain-of-thought reasoning and model transparency. Sam Altman noted safety work slowed Astra's deployment, indicating future models may prioritize safeguards over speed.
not much happened today
gpt-5.3-codex claude-opus-4.6 openai anthropic cursor_ai github microsoft builder-tooling cybersecurity api-access model-rollout agentic-ai long-context serving-economics throughput-latency token-efficiency workflow-design sama pierceboggan kylebrussell natolambert omarsar0 sam_altman
OpenAI launched GPT-5.3-Codex with a Super Bowl ad emphasizing "You can just build things" as a product strategy, focusing on builder tooling over chat interfaces. The model is rolling out across Cursor, VS Code, and GitHub with phased API access and is flagged as their first "high cybersecurity capability" model. Sam Altman reported over 1M Codex app downloads in the first week and strong weekly user growth. Meanwhile, Anthropic's Claude Opus 4.6 is recognized as a leading "agentic generalist" model, topping text and code leaderboards but noted for high token usage. Discussions around serving economics and "fast mode" behavior highlight practical deployment considerations. Additionally, Recursive Language Models (RLMs) introduce a novel approach using a second programmatic context space to extend long-context capabilities.
nothing much happened today
o1 chatgpt-4o llama-3-1-405b openai lmsys scale-ai cognition langchain qdrant rohanpaul_ai reinforcement-learning model-merging embedding-models toxicity-detection image-editing dependency-management automated-code-review visual-search benchmarking denny_zhou svpino alexandr_wang cwolferesearch rohanpaul_ai _akhaliq kylebrussell
OpenAI's o1 model faces skepticism about open-source replication due to its extreme restrictions and unique training advances like RL on CoT. ChatGPT-4o shows significant performance improvements across benchmarks. Llama-3.1-405b fp8 and bf16 versions perform similarly with cost benefits for fp8. A new open-source benchmark "Humanity's Last Exam" offers $500K in prizes to challenge LLMs. Model merging benefits from neural network sparsity and linear mode connectivity. Embedding-based toxic prompt detection achieves high accuracy with low compute. InstantDrag enables fast, optimization-free drag-based image editing. LangChain v0.3 releases with improved dependency management. Automated code review tool CodeRabbit adapts to team coding styles. Visual search advances integrate multimodal data for better product search. Experts predict AI will be default software by 2030.