All tags
Model: "inkling"
not much happened today
gpt-5.6-luna gpt-5.6-terra gpt-5.6-sol arc-agi-3 inkling-small inkling gemini-robotics-2 openai thinking-machines lmsys modal unsloth artificial-analysis google price-optimization agent-systems memory-retention context-compaction multimodality mixture-of-experts model-compression benchmarking open-weights multimodal-models model-efficiency model-deployment embodied-ai robotics long-context sama fchollet kimmonismus gneubig scaling01 mervenoyann
OpenAI aggressively cut prices for GPT-5.6 Luna by 80% and Terra by 20%, introducing a faster Sol Fast tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The ARC-AGI-3 debate highlighted that the complete agent system, including memory retention and tool orchestration, is critical beyond just the base model. Thinking Machines released Inkling-Small, an open-weights, multimodal MoE model with 276B parameters (12B active), delivering performance comparable to the original Inkling at a quarter of the size, supporting audio, images, and Python-based image inspection. Benchmarks show Inkling-Small excels in coding and multimodality tasks, with 1M-context support and broad open inference stack adoption. The news also mentions Google's Gemini Robotics 2 advancing embodied AI from tabletop to full-body control.
not much happened today
kimi-k3 claude-fable-5 opus-4.8 gpt-5.6-terra gpt-5.5 inkling glm-5.2 gpt-5.6-sol moonshot openai thinking-machines artificial-analysis arena datacurve arcprize aisecurityinst moe-routing quantization data-curation infrastructure-design coding-agents benchmarking front-end-development software-engineering arc-benchmarks cybersecurity zhilin_yang kimmonismus anikasomaia dylan522p novasarc01 scaling01 theo hqmank
Moonshot's Kimi K3 release has sparked a reassessment of Chinese open-weight models' proximity to the frontier, with strong performance in coding, agentic tasks, and long-horizon knowledge work. The strategic focus has shifted from a "compute moat" to an "efficiency stack" involving MoE routing, quantization, data curation, and scarcity-driven infrastructure like Moonshot's "Mooncake" stack. Benchmarks from Artificial Analysis, Arena, DeepSWE, ARC, and Cyber place K3 among the top models, with scores such as 57 on the Intelligence Index and coding agent benchmarks matching or surpassing models like GPT-5.6 Terra and Claude Fable 5. Discussions continue on K3's exact standing, but it is now widely recognized as a significant frontier contender.
not much happened today
inkling thinking-machines-lab huggingface vllm_project lmsysorg modal baseten databricks mixture-of-experts multimodality foundation-models model-licensing context-window open-weights model-release miramurati soumithchintala johnschulman2 lilianweng natolambert artificialanlys scaling01
Thinking Machines Lab launched Inkling, its first fully released open-weights foundation model family, featuring 975B parameters with 41B active parameters in a Mixture-of-Experts architecture. Inkling supports multimodality with text, image, and audio inputs and text output, is Apache 2.0 licensed, and offers up to 1M context window. The model is available on platforms like Tinker, Hugging Face, and partners, with broad ecosystem support from vLLM, SGLang, Modal, Baseten, and Databricks. Key figures such as Mira Murati, Soumith Chintala, John Schulman, and Lilian Weng highlighted its open weights, customization, and practical use focus. Independent commentators noted it as the strongest U.S.-based open-weight release to date, though still behind top Chinese open-weight and best closed models on some benchmarks.