<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>AINews</title><description>Weekday recaps of top News for AI Engineers</description><link>https://news.smol.ai/</link><language>en-us</language><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-16-kimi-k30/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-16-kimi-k30/</guid><description>**Moonshot AI** launched **Kimi K3**, a frontier-class open-weights model with **2.8T parameters**, **1M-token context window**, and **native multimodal input**. It features novel **Kimi Delta Attention (KDA)** enabling up to **6.3x faster decoding** and **Attention Residuals** for **~25% higher training efficiency**. K3 is live on multiple platforms with open weights promised by **July 27, 2026**. It leads in **Frontend Code Arena** with a **76% pairwise win rate**, ranking above **Claude Fable 5** and **GPT-5.6 Sol** in several benchmarks, though still behind these models in overall user experience. Independent evaluations place K3 comparable to **Opus 4.8** and **GPT-5.5** but behind Fable 5 and GPT-5.6 Sol. The launch is seen as a major open-model milestone.</description><pubDate>Thu, 16 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/15/2026-7/16/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Moonshot AI launched Kimi K3 as a frontier-class open-weights model, with official claims that place it near top closed models and above prior open competitors.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Moonshot officially introduced &lt;strong&gt;Kimi K3&lt;/strong&gt; as &lt;strong&gt;“Open Frontier Intelligence”&lt;/strong&gt; with &lt;strong&gt;2.8T total parameters&lt;/strong&gt;, &lt;strong&gt;1M-token context&lt;/strong&gt;, &lt;strong&gt;native multimodal input&lt;/strong&gt;, &lt;strong&gt;Kimi Delta Attention (KDA)&lt;/strong&gt;, and &lt;strong&gt;Attention Residuals&lt;/strong&gt;, and said the model is live on Kimi.com, Kimi Work, Kimi Code, and API, with &lt;strong&gt;open weights promised by July 27, 2026&lt;/strong&gt; &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Moonshot also highlighted product positioning around &lt;strong&gt;long-horizon agentic coding&lt;/strong&gt; and &lt;strong&gt;self-evolving workflows&lt;/strong&gt;, plus “vision in the loop” coding/game-building workflows that iterate between code and screenshots &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830245382758902&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Before the formal announcement, multiple accounts circulated leaked or app-sourced details that K3 was &lt;strong&gt;2.8T params&lt;/strong&gt;, calling it the &lt;strong&gt;largest open-weight model ever&lt;/strong&gt; if weights ship as promised &lt;a href=&quot;https://x.com/scaling01/status/2077767900635517082&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077769925293207898&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/eliebakouch/status/2077769728295059557&quot;&gt;@eliebakouch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The official Kimi blog went live later and was widely shared as the primary technical source &lt;a href=&quot;https://x.com/Jianlin_S/status/2077828801388769603&quot;&gt;@Jianlin_S&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077829284949828048&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yulun_Du/status/2077831915999228192&quot;&gt;@Yulun_Du&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Moonshot’s own phrasing acknowledged a limitation: despite being highly competitive overall, K3 still has a &lt;strong&gt;“noticeable gap in user experience”&lt;/strong&gt; versus &lt;strong&gt;Claude Fable 5&lt;/strong&gt; and &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2077833896931037290&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Arena announced that &lt;strong&gt;Kimi K3 entered Agent Arena&lt;/strong&gt;, plus Text, Vision, Document, and Frontend Code Arena, with community evaluations to follow &lt;a href=&quot;https://x.com/arena/status/2077802013245816962&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Arena then reported a major early result: &lt;strong&gt;Kimi K3 became #1 in Frontend Code Arena with 1679 points&lt;/strong&gt;, surpassing Claude Fable 5 and jumping from &lt;strong&gt;#18 (K2.6) to #1&lt;/strong&gt;, ranking &lt;strong&gt;#1 in 6 of 7 frontend domains&lt;/strong&gt; and &lt;strong&gt;#2 in Gaming&lt;/strong&gt; &lt;a href=&quot;https://x.com/arena/status/2077824029126504525&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Arena later added that K3 has a &lt;strong&gt;76% pairwise win rate&lt;/strong&gt; in Frontend Code Arena, versus &lt;strong&gt;63% for Fable 5&lt;/strong&gt; and &lt;strong&gt;58% for GPT-5.6 Sol&lt;/strong&gt; &lt;a href=&quot;https://x.com/arena/status/2077893862778183737&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;In Text Arena, K3 landed at &lt;strong&gt;#9 with 1486 points&lt;/strong&gt;, a jump from &lt;strong&gt;#38&lt;/strong&gt;, with top-10 placements in &lt;strong&gt;creative writing, coding, and instruction following&lt;/strong&gt;, and #1 in several occupation slices &lt;a href=&quot;https://x.com/arena/status/2077856214684455116&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis published an independent evaluation placing K3 at &lt;strong&gt;57 on the AA Intelligence Index&lt;/strong&gt;, calling it &lt;strong&gt;comparable to Opus 4.8 and GPT-5.5&lt;/strong&gt;, but still &lt;strong&gt;behind Fable 5 and GPT-5.6 Sol&lt;/strong&gt; overall &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;AA also reported K3 at &lt;strong&gt;1668 Elo on GDPval v2&lt;/strong&gt;, &lt;strong&gt;53% / #1 on AutomationBench-AA&lt;/strong&gt;, and &lt;strong&gt;1547 Elo on AA-Briefcase&lt;/strong&gt;, with &lt;strong&gt;cost per task of $0.94&lt;/strong&gt;, about &lt;strong&gt;21% fewer output tokens than K2.6&lt;/strong&gt; across the full Intelligence Index run &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The launch immediately triggered strong reaction from engineers and model-watchers who framed K3 as an &lt;strong&gt;open-model milestone&lt;/strong&gt; comparable to earlier DeepSeek moments &lt;a href=&quot;https://x.com/kimmonismus/status/2077818040578695175&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/nrehiew_/status/2077782895377387708&quot;&gt;@nrehiew_&lt;/a&gt;, &lt;a href=&quot;https://x.com/eliebakouch/status/2077781181915918663&quot;&gt;@eliebakouch&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Technical details&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Architecture and systems details&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Official specs: &lt;strong&gt;2.8T total parameters&lt;/strong&gt;, &lt;strong&gt;1M context&lt;/strong&gt;, &lt;strong&gt;native multimodal input&lt;/strong&gt; (text + images), &lt;strong&gt;text output&lt;/strong&gt;, &lt;strong&gt;open weights by July 27&lt;/strong&gt; &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;K3 uses &lt;strong&gt;Kimi Delta Attention (KDA)&lt;/strong&gt;, which Moonshot says enables &lt;strong&gt;up to 6.3x faster decoding in million-token contexts&lt;/strong&gt; &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;It also uses &lt;strong&gt;Attention Residuals (AttnRes)&lt;/strong&gt;, claimed to deliver &lt;strong&gt;~25% higher training efficiency at &amp;#x3C;2% additional cost&lt;/strong&gt; &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Community readers of the blog highlighted additional architecture details: &lt;strong&gt;LatentMoE / Stable LatentMoE&lt;/strong&gt;, &lt;strong&gt;16 activated experts out of 896&lt;/strong&gt;, implying an activation ratio under &lt;strong&gt;2%&lt;/strong&gt; &lt;a href=&quot;https://x.com/nrehiew_/status/2077774067533590643&quot;&gt;@nrehiew_&lt;/a&gt;, &lt;a href=&quot;https://x.com/eliebakouch/status/2077837543525998770&quot;&gt;@eliebakouch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;More community-extracted details from the blog/report discussion: &lt;strong&gt;per-head Muon&lt;/strong&gt;, &lt;strong&gt;QB load balancing / quantile load balancing&lt;/strong&gt;, and a new activation function called &lt;strong&gt;SiTU (Sigmoid Tanh Unit)&lt;/strong&gt; &lt;a href=&quot;https://x.com/eliebakouch/status/2077837543525998770&quot;&gt;@eliebakouch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;One engineer noted the architecture as notable for combining &lt;strong&gt;KDA + LatentMoE + AttnRes&lt;/strong&gt; while scaling more than 2x over prior Kimi models &lt;a href=&quot;https://x.com/teortaxesTex/status/2077837689601064983&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;KDA had a long incubation cycle: design reportedly started in &lt;strong&gt;Jan 2025&lt;/strong&gt; and took &lt;strong&gt;~1.5 years&lt;/strong&gt; to reach frontier scale &lt;a href=&quot;https://x.com/zxytim/status/2077839815538872573&quot;&gt;@zxytim&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Inference and serving&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;K3 pricing was reported as &lt;strong&gt;$3 / 1M input tokens&lt;/strong&gt; and &lt;strong&gt;$15 / 1M output tokens&lt;/strong&gt;, with &lt;strong&gt;cached input discounted 90% to $0.30 / 1M&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2077770795107897449&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Several posters compared that pricing to &lt;strong&gt;Sonnet 5&lt;/strong&gt;, with some noting Sonnet was temporarily cheaper until end of August, after which prices align more closely &lt;a href=&quot;https://x.com/kimmonismus/status/2077776566742892770&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;A blended estimate at &lt;strong&gt;80% input / 20% output&lt;/strong&gt; came out to &lt;strong&gt;$5.40 / 1M tokens&lt;/strong&gt;, vs &lt;strong&gt;$9 for Opus 4.8&lt;/strong&gt; and &lt;strong&gt;$10 for GPT-5.5&lt;/strong&gt; &lt;a href=&quot;https://x.com/jaminball/status/2077872831883591851&quot;&gt;@jaminball&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis estimated &lt;strong&gt;$0.94 average cost per Intelligence Index task&lt;/strong&gt;, versus &lt;strong&gt;$1.04 for GPT-5.6 Sol&lt;/strong&gt; and &lt;strong&gt;$1.80 for Opus 4.8&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832885021835289&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Early live serving observations: &lt;strong&gt;~28 tok/s via Moonshot API on OpenRouter&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2077777932341092422&quot;&gt;@scaling01&lt;/a&gt;, and another observer saw &lt;strong&gt;26 tok/s&lt;/strong&gt;, calling it slower than Opus and speculating that &lt;strong&gt;speculative decoding wasn’t yet enabled&lt;/strong&gt; &lt;a href=&quot;https://x.com/nrehiew_/status/2077789869242536109&quot;&gt;@nrehiew_&lt;/a&gt;, &lt;a href=&quot;https://x.com/nrehiew_/status/2077790338455130501&quot;&gt;@nrehiew_&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Moonshot’s blog reportedly recommends deployment on &lt;strong&gt;supernode configurations with 64+ accelerators&lt;/strong&gt; for best inference efficiency &lt;a href=&quot;https://x.com/teortaxesTex/status/2077842456121393198&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM said Moonshot contributed a &lt;strong&gt;KDA prefix caching implementation directly to vLLM&lt;/strong&gt;, with support available &lt;strong&gt;day 0&lt;/strong&gt; for official release &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;@vllm_project&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Moonshot’s KDA contribution was cited as important because &lt;strong&gt;KDA breaks assumptions behind conventional prefix caching&lt;/strong&gt;, so upstream runtime changes were required &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;@vllm_project&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Benchmarks and evals&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Moonshot’s official benchmarking message, as summarized by others, positioned K3 &lt;strong&gt;behind only Claude Fable 5 and GPT-5.6 Sol among tested models&lt;/strong&gt;, and ahead of &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2077770018096361749&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077777217170661608&quot;&gt;@Yuchenj_UW&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;One cited number: &lt;strong&gt;1687 on GDPval-AA v2&lt;/strong&gt;, above Opus 4.8 and behind GPT-5.6 Sol at &lt;strong&gt;1747.8&lt;/strong&gt; in that comparison &lt;a href=&quot;https://x.com/scaling01/status/2077770398389747993&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis’ independent numbers:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AA Intelligence Index: 57&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GDPval v2 Elo: 1668&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AutomationBench-AA: 53%, #1&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AA-Briefcase Elo: 1547&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AA-Omniscience: +18&lt;/strong&gt;, with &lt;strong&gt;accuracy 46% vs 33% on K2.6&lt;/strong&gt;, but &lt;strong&gt;hallucination rate worsening to 51% from 39%&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832882039742923&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;AA also reported &lt;strong&gt;132M output tokens&lt;/strong&gt; consumed for K3 across the Intelligence Index, versus &lt;strong&gt;166M for K2.6&lt;/strong&gt;, i.e. &lt;strong&gt;21% reduction&lt;/strong&gt; while gaining &lt;strong&gt;13 index points&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832879187620192&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Arena’s frontend result was especially prominent because it is a &lt;strong&gt;pairwise human-preference arena&lt;/strong&gt;, not just a static benchmark, and K3’s &lt;strong&gt;#1 frontend rank&lt;/strong&gt; became one of the main launch headlines &lt;a href=&quot;https://x.com/arena/status/2077824029126504525&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Community posts also highlighted strong results on &lt;strong&gt;kernel optimization tasks&lt;/strong&gt;, with some saying K3 was matching or beating Fable in certain kernel/codegen settings &lt;a href=&quot;https://x.com/nrehiew_/status/2077810993057669511&quot;&gt;@nrehiew_&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077808643739639832&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;One benchmark caveat came from &lt;strong&gt;ProgramBench&lt;/strong&gt; author Ofir Press, who said Kimi used a metric they &lt;strong&gt;do not recommend&lt;/strong&gt;: averaging implementation percentage rather than counting &lt;strong&gt;fully working programs&lt;/strong&gt;, which can overstate usefulness &lt;a href=&quot;https://x.com/OfirPress/status/2077856894820000100&quot;&gt;@OfirPress&lt;/a&gt;, &lt;a href=&quot;https://x.com/OfirPress/status/2077857100437275086&quot;&gt;@OfirPress&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Facts vs opinions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Facts / directly sourced claims&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Kimi K3 is officially announced by Moonshot &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Officially disclosed specs include &lt;strong&gt;2.8T params&lt;/strong&gt;, &lt;strong&gt;1M context&lt;/strong&gt;, &lt;strong&gt;native multimodal input&lt;/strong&gt;, &lt;strong&gt;KDA&lt;/strong&gt;, &lt;strong&gt;AttnRes&lt;/strong&gt;, &lt;strong&gt;open weights by July 27&lt;/strong&gt; &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis independently scored K3 at &lt;strong&gt;57 Intelligence Index&lt;/strong&gt;, with detailed task, cost, token, and benchmark data &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Arena independently ranked K3 &lt;strong&gt;#1 in Frontend Code Arena&lt;/strong&gt; and later reported its &lt;strong&gt;76% pairwise win rate&lt;/strong&gt; &lt;a href=&quot;https://x.com/arena/status/2077824029126504525&quot;&gt;@arena&lt;/a&gt;, &lt;a href=&quot;https://x.com/arena/status/2077893862778183737&quot;&gt;@arena&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM confirmed Moonshot contributed runtime support for &lt;strong&gt;KDA prefix caching&lt;/strong&gt; &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;@vllm_project&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Opinions / interpretations&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“DeepSeek moment,” “beginning of the US-China AI race,” and “everything changed” are editorial interpretations from observers, not established facts &lt;a href=&quot;https://x.com/kimmonismus/status/2077832669778317369&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077842134380523776&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2077836497739304968&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Claims that K3 “beats GPT-5.6 Sol on 11 of 14 benchmarks” and “Fable on 6 of 14” are aggregated community summaries and should be treated as contingent on the benchmark set and exact methodology &lt;a href=&quot;https://x.com/scaling01/status/2077810222999949497&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Assertions that this implies Dario/Anthropic margin pressure, a geopolitical turning point, or near-term superintelligence are speculative commentary &lt;a href=&quot;https://x.com/teortaxesTex/status/2077827587888300256&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/Jason/status/2077836937810022756&quot;&gt;@Jason&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Several “distillation” insinuations were explicitly framed as jokes or conjecture rather than evidence &lt;a href=&quot;https://x.com/yacinelearning/status/2077758528953979295&quot;&gt;@yacinelearning&lt;/a&gt;, &lt;a href=&quot;https://x.com/dejavucoder/status/2077877794697314563&quot;&gt;@dejavucoder&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Different opinions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Strongly supportive&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many engineers called K3 a genuine &lt;strong&gt;frontier open model&lt;/strong&gt;, especially because it appears to be &lt;strong&gt;better than Opus 4.8&lt;/strong&gt; while being priced near Sonnet and planned for open-weight release &lt;a href=&quot;https://x.com/kimmonismus/status/2077772229685707138&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/cline/status/2077824751238811914&quot;&gt;@cline&lt;/a&gt;, &lt;a href=&quot;https://x.com/nrehiew_/status/2077810575737040963&quot;&gt;@nrehiew_&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Supporters emphasized that this is no longer “good for open source,” but simply &lt;strong&gt;competitive with top public closed models&lt;/strong&gt; &lt;a href=&quot;https://x.com/tokenbender/status/2077832045255147772&quot;&gt;@tokenbender&lt;/a&gt;, &lt;a href=&quot;https://x.com/TheAhmadOsman/status/2077881194981503406&quot;&gt;@TheAhmadOsman&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Some framed the release as evidence that &lt;strong&gt;open models are now within weeks or a couple months of the frontier&lt;/strong&gt; &lt;a href=&quot;https://x.com/nrehiew_/status/2077782308162351576&quot;&gt;@nrehiew_&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Others argued this materially raises the odds that &lt;strong&gt;future AGI-level systems are open&lt;/strong&gt; &lt;a href=&quot;https://x.com/MaorShlomo/status/2077844032214995074&quot;&gt;@MaorShlomo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Supportive but technically cautious&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Artificial Analysis gave a more restrained view: K3 is &lt;strong&gt;comparable to Opus 4.8 and GPT-5.5&lt;/strong&gt;, but &lt;strong&gt;still behind Fable 5 and GPT-5.6 Sol&lt;/strong&gt; on overall intelligence &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Simon Willison described K3 as significant, but also pointed readers toward nuanced notes and benchmark caveats rather than simple leaderboard hype &lt;a href=&quot;https://x.com/simonw/status/2077852005129933247&quot;&gt;@simonw&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Ethan Mollick’s hands-on impression: &lt;strong&gt;very good open-weights model&lt;/strong&gt;, but &lt;strong&gt;not Sol Max or Fable&lt;/strong&gt; &lt;a href=&quot;https://x.com/emollick/status/2077783731691995348&quot;&gt;@emollick&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;One user said K3’s intelligence is strong, but it is &lt;strong&gt;slow&lt;/strong&gt;, sometimes &lt;strong&gt;over-checks&lt;/strong&gt;, and still trails Claude on taste/aesthetics &lt;a href=&quot;https://x.com/nrehiew_/status/2077796966298480943&quot;&gt;@nrehiew_&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Critical / skeptical&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Bindu Reddy warned that K3’s benchmark story might be overstated unless validated on &lt;strong&gt;hidden / uncontaminated evals like LiveBench&lt;/strong&gt;, and argued that if the model “thinks forever,” real cost could be less favorable &lt;a href=&quot;https://x.com/bindureddy/status/2077816569489678703&quot;&gt;@bindureddy&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;ProgramBench maintainers objected to Moonshot’s metric choice, saying it can &lt;strong&gt;inflate partial-credit performance&lt;/strong&gt; relative to fully working programs &lt;a href=&quot;https://x.com/OfirPress/status/2077856894820000100&quot;&gt;@OfirPress&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis also flagged a real weakness: &lt;strong&gt;hallucination rate regressed&lt;/strong&gt; on AA-Omniscience despite accuracy gains &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832882039742923&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Multiple users noted that K3 currently appears to &lt;strong&gt;think a lot&lt;/strong&gt;, preserve long reasoning history, and may require more careful harness support than simpler chat-first APIs &lt;a href=&quot;https://x.com/scaling01/status/2077782976549491076&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/Xianbao_QIAN/status/2077843337030385664&quot;&gt;@Xianbao_QIAN&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Some skepticism focused on economics and deployability: &lt;strong&gt;2.8T open weights&lt;/strong&gt; is impressive, but practical self-hosting may still be limited to well-funded teams &lt;a href=&quot;https://x.com/mbusigin/status/2077912338414391529&quot;&gt;@mbusigin&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Political / strategic interpretations&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A broad cluster of tweets framed K3 as proof that &lt;strong&gt;Chinese labs are no longer far behind&lt;/strong&gt; and that the US lead is shrinking &lt;a href=&quot;https://x.com/tszzl/status/2077827974452461871&quot;&gt;@tszzl&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2077832669778317369&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077825258040488099&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Others counterweighted that K3 still appears to lag the very best Western models in &lt;strong&gt;usability / productization&lt;/strong&gt;, even if raw capability is close &lt;a href=&quot;https://x.com/RyanGreenblatt/status/2077868913438945493&quot;&gt;@RyanGreenblatt&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077833896931037290&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Some argued that open Chinese models function as &lt;strong&gt;economic pressure&lt;/strong&gt; on US labs by compressing margins and commoditizing capability &lt;a href=&quot;https://x.com/francoisfleuret/status/2077878010129063944&quot;&gt;@francoisfleuret&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Others viewed the inevitable next step as more &lt;strong&gt;competition on harnesses, products, and deployment systems&lt;/strong&gt;, not just raw model weights &lt;a href=&quot;https://x.com/AravSrinivas/status/2077894147071991850&quot;&gt;@AravSrinivas&lt;/a&gt;, &lt;a href=&quot;https://x.com/theo/status/2077871618437919122&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Context&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why this matters technically&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;K3 is notable not just for raw size but for &lt;strong&gt;scaling a non-standard attention stack&lt;/strong&gt; into a frontier-class model: KDA + AttnRes + sparse MoE drew repeated attention from technically literate observers &lt;a href=&quot;https://x.com/scaling01/status/2077770130000323068&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/eliebakouch/status/2077837543525998770&quot;&gt;@eliebakouch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The launch is also a systems story: long-context serving, prefix caching, KDA runtime support, and deployment on large accelerator supernodes all matter if the weights are to be practically usable &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/teortaxesTex/status/2077842456121393198&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The emphasis on &lt;strong&gt;kernel optimization&lt;/strong&gt;, &lt;strong&gt;chip design&lt;/strong&gt;, &lt;strong&gt;agentic coding&lt;/strong&gt;, and &lt;strong&gt;environment simulation&lt;/strong&gt; suggests Moonshot is optimizing for &lt;strong&gt;AI-improving-AI workflows&lt;/strong&gt;, not just chatbot benchmarks &lt;a href=&quot;https://x.com/18jeffreyma/status/2077849822611267803&quot;&gt;@18jeffreyma&lt;/a&gt;, &lt;a href=&quot;https://x.com/yong_zhengxin/status/2077834949772624166&quot;&gt;@yong_zhengxin&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why this matters economically&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The strongest repeated theme: &lt;strong&gt;frontier-ish performance at materially lower price than top closed models&lt;/strong&gt;, though not at bargain-basement open-model prices &lt;a href=&quot;https://x.com/kimmonismus/status/2077772229685707138&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/cline/status/2077824751238811914&quot;&gt;@cline&lt;/a&gt;, &lt;a href=&quot;https://x.com/jaminball/status/2077872831883591851&quot;&gt;@jaminball&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis’ task-cost framing is especially relevant for practitioners: if K3 is near &lt;strong&gt;GPT-5.6 Sol cost-per-task&lt;/strong&gt; and below &lt;strong&gt;Opus 4.8&lt;/strong&gt;, the real question becomes where it slots into agent stacks, coding platforms, and self-hosted infra &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Some noted the paradox that “open weights” does not automatically mean “cheap to run”: a &lt;strong&gt;2.8T&lt;/strong&gt; model with &lt;strong&gt;64+ accelerator&lt;/strong&gt; deployment guidance is frontier infrastructure territory &lt;a href=&quot;https://x.com/teortaxesTex/status/2077842456121393198&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/mbusigin/status/2077912338414391529&quot;&gt;@mbusigin&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Why this matters geopolitically&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many reactions explicitly tied K3 to export controls, US-China competition, and the narrowing gap between Chinese open labs and US closed labs &lt;a href=&quot;https://x.com/scaling01/status/2077776285489578293&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/tszzl/status/2077827974452461871&quot;&gt;@tszzl&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2077832669778317369&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Several commentators argued that K3 weakens the common narrative that Chinese models trail by &lt;strong&gt;6–8 months&lt;/strong&gt;, because it appears to outperform a closed US model from &lt;strong&gt;late May&lt;/strong&gt; only weeks later &lt;a href=&quot;https://x.com/kimmonismus/status/2077832669778317369&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Others stressed that “capability parity” is not the same as full-stack parity: product reliability, inference scale, deployment margins, and proprietary post-training may still favor US incumbents &lt;a href=&quot;https://x.com/RyanGreenblatt/status/2077868913438945493&quot;&gt;@RyanGreenblatt&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Early hands-on signals&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Users reported K3 building impressive &lt;strong&gt;web experiences&lt;/strong&gt;, &lt;strong&gt;games&lt;/strong&gt;, and &lt;strong&gt;shader/code artifacts&lt;/strong&gt;, reinforcing the Frontend Arena result &lt;a href=&quot;https://x.com/johnlindquist/status/2077840176370602179&quot;&gt;@johnlindquist&lt;/a&gt;, &lt;a href=&quot;https://x.com/ChrissGPT/status/2077852656182129078&quot;&gt;@ChrissGPT&lt;/a&gt;, &lt;a href=&quot;https://x.com/intheworldofai/status/2077838911494336681&quot;&gt;@intheworldofai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;One user said K3 generated a &lt;strong&gt;CS:GO × Portal clone&lt;/strong&gt; in &lt;strong&gt;3 shots&lt;/strong&gt; using &lt;strong&gt;~600k tokens&lt;/strong&gt;, costing &lt;strong&gt;$3.24&lt;/strong&gt; by API pricing, compared with claimed higher costs on Fable and GPT-5.6 Sol &lt;a href=&quot;https://x.com/ChrissGPT/status/2077852656182129078&quot;&gt;@ChrissGPT&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Another reported K3 continuously working for hours over near-&lt;strong&gt;1M context&lt;/strong&gt; to build a &lt;strong&gt;web DOS emulator&lt;/strong&gt; with low human intervention &lt;a href=&quot;https://x.com/bigeagle_xd/status/2077820690133287395&quot;&gt;@bigeagle_xd&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;At the same time, several users noted it can be &lt;strong&gt;verbose&lt;/strong&gt;, &lt;strong&gt;slow&lt;/strong&gt;, and heavily reliant on &lt;strong&gt;thinking-history preservation&lt;/strong&gt;, implying that serving/harness defaults will matter a lot &lt;a href=&quot;https://x.com/nrehiew_/status/2077795629921952228&quot;&gt;@nrehiew_&lt;/a&gt;, &lt;a href=&quot;https://x.com/Xianbao_QIAN/status/2077843337030385664&quot;&gt;@Xianbao_QIAN&lt;/a&gt;, &lt;a href=&quot;https://x.com/bigeagle_xd/status/2077851766180470922&quot;&gt;@bigeagle_xd&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open-source/open-weights debate&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The surrounding discourse included the usual complaint that “open weight” is not “fully open,” but several commenters pushed back that this distinction is often impractical at frontier scale and that inspectable, fine-tunable weights still matter &lt;a href=&quot;https://x.com/Dan_Jeffries1/status/2077641797363237328&quot;&gt;@Dan_Jeffries1&lt;/a&gt;, &lt;a href=&quot;https://x.com/ClementDelangue/status/2077873510144512400&quot;&gt;@ClementDelangue&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Yulun Du said the delay before weight release was to ensure a &lt;strong&gt;smooth rollout with inference partners&lt;/strong&gt;, signaling that ecosystem readiness mattered as much as the checkpoint itself &lt;a href=&quot;https://x.com/Yulun_Du/status/2077831915999228192&quot;&gt;@Yulun_Du&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM maintainers and others treated Moonshot’s upstream contributions as evidence that the launch is not just “marketing open,” but also includes meaningful OSS infra work &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/woosuk_k/status/2077861534253089275&quot;&gt;@woosuk_k&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Benchmarks, contamination, and what to watch next&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several people cautioned that current public benchmark ecosystems saturate quickly, and that hidden evals or stack-level evals will be more informative &lt;a href=&quot;https://x.com/bindureddy/status/2077816569489678703&quot;&gt;@bindureddy&lt;/a&gt;, &lt;a href=&quot;https://x.com/gdb/status/2077887553655689239&quot;&gt;@gdb&lt;/a&gt;, &lt;a href=&quot;https://x.com/WolfBenchAI/status/2077869821459652613&quot;&gt;@WolfBenchAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Observers specifically asked for follow-up on &lt;strong&gt;METR time horizons&lt;/strong&gt;, &lt;strong&gt;cyber ranges&lt;/strong&gt;, &lt;strong&gt;FrontierMath T4&lt;/strong&gt;, &lt;strong&gt;ARC-AGI-2/3&lt;/strong&gt;, &lt;strong&gt;CritPt&lt;/strong&gt;, &lt;strong&gt;token usage&lt;/strong&gt;, and broader long-horizon agent evals &lt;a href=&quot;https://x.com/scaling01/status/2077824815746957795&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The most credible near-term follow-up points are:
&lt;ul&gt;
&lt;li&gt;whether the &lt;strong&gt;weights ship on time&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;what &lt;strong&gt;third-party serving stacks&lt;/strong&gt; achieve for throughput/cost&lt;/li&gt;
&lt;li&gt;how K3 performs on &lt;strong&gt;hidden evals and real production agent tasks&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;whether Moonshot closes the &lt;strong&gt;UX/post-training gap&lt;/strong&gt; they themselves acknowledged &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;@Kimi_Moonshot&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077833896931037290&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077832874183860404&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open Models, Inference Stacks, and Retrieval Infrastructure&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;vLLM and serving ecosystem support landed quickly&lt;/strong&gt;: &lt;a href=&quot;https://x.com/vllm_project/status/2077840545171538114&quot;&gt;vLLM&lt;/a&gt; said Moonshot contributed a &lt;strong&gt;KDA prefix-caching implementation directly to vLLM&lt;/strong&gt;, enabling &lt;strong&gt;day-0&lt;/strong&gt; support once weights drop. This matters because KDA breaks some conventional prefix-caching assumptions. The post underscores that long-context architectural innovation increasingly requires coordinated systems work, not just model release.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NVIDIA shipped a notable open retrieval release&lt;/strong&gt;: &lt;a href=&quot;https://x.com/NVIDIAAI/status/2077786069840318800&quot;&gt;NVIDIA&lt;/a&gt; launched &lt;strong&gt;Nemotron 3 Embed 8B&lt;/strong&gt;, claiming &lt;strong&gt;#1 overall on RTEB&lt;/strong&gt;, and partners quickly made it deployable, including &lt;a href=&quot;https://x.com/baseten/status/2077812130649391216&quot;&gt;Baseten&lt;/a&gt; and &lt;a href=&quot;https://x.com/turbopuffer/status/2077810727662850186&quot;&gt;Turbopuffer&lt;/a&gt;. A more detailed community summary by &lt;a href=&quot;https://x.com/kimmonismus/status/2077872157393383809&quot;&gt;@kimmonismus&lt;/a&gt; reports &lt;strong&gt;78.46 NDCG@10 on RTEB&lt;/strong&gt; and &lt;strong&gt;75.45 on MMTEB Retrieval&lt;/strong&gt;, with NVIDIA arguing stronger retrieval reduces downstream agent token usage. The release also includes &lt;strong&gt;1B BF16&lt;/strong&gt; and &lt;strong&gt;1B NVFP4&lt;/strong&gt; variants, with the NVFP4 version reportedly offering up to &lt;strong&gt;2× BF16 throughput&lt;/strong&gt; on Blackwell while retaining &gt;99% retrieval quality.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LiteParse added a gRPC interface for backend document pipelines&lt;/strong&gt;: &lt;a href=&quot;https://x.com/llama_index/status/2077791650386960741&quot;&gt;LlamaIndex&lt;/a&gt; introduced &lt;strong&gt;liteparse-grpc&lt;/strong&gt;, exposing PDF/Office/image parsing, rendering, and OCR-complexity estimation over gRPC with protobuf definitions and generated clients. This is a practical infra improvement for polyglot microservice stacks where REST isn’t ideal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managed vector/search infra also expanded&lt;/strong&gt;: &lt;a href=&quot;https://x.com/weaviate_io/status/2077755251759722574&quot;&gt;Weaviate&lt;/a&gt; announced &lt;strong&gt;Managed Weaviate on DigitalOcean&lt;/strong&gt; in public preview, running the unmodified open-source engine (&lt;strong&gt;v1.37.1 at launch&lt;/strong&gt;) with HA, autoscaling, backups, forks, and control-plane observability.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agents, Harnesses, and System Design Becoming the Real Product Layer&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Harnesses were a recurring theme across builders&lt;/strong&gt;: Harrison Chase’s conversation with Factory AI’s Eno Reyes was repeatedly shared as a case for why “the harness matters more than the model” (&lt;a href=&quot;https://x.com/hwchase17/status/2077764401399210055&quot;&gt;Harrison&lt;/a&gt;, &lt;a href=&quot;https://x.com/LangChain/status/2077764775124107766&quot;&gt;LangChain&lt;/a&gt;). Chase later argued teams should “own the harness,” “own the context and memory layer,” and “own model optionality” rather than rent intelligence from a single provider (&lt;a href=&quot;https://x.com/hwchase17/status/2077787686547677434&quot;&gt;thread&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;There’s growing interest in open standards for memory and knowledge representation&lt;/strong&gt;: &lt;a href=&quot;https://x.com/hwchase17/status/2077806939074081259&quot;&gt;Harrison Chase&lt;/a&gt; promoted &lt;strong&gt;OKF (Open Knowledge Format)&lt;/strong&gt; as an “open standard for memory,” while &lt;a href=&quot;https://x.com/BraceSproul/status/2077799633640919208&quot;&gt;Brace Sproul&lt;/a&gt; detailed OpenWiki’s adoption and the benefits for search, retrieval, and codebase memory.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent self-improvement and scheduled multi-agent workflows are becoming mainstream topics&lt;/strong&gt;: &lt;a href=&quot;https://x.com/omarsar0/status/2077792894459793714&quot;&gt;@omarsar0&lt;/a&gt; highlighted a survey on &lt;strong&gt;self-improving agentic systems&lt;/strong&gt;, and elsewhere described using an “LLM Council” with recurring scheduled research updates (&lt;a href=&quot;https://x.com/omarsar0/status/2077765052434633023&quot;&gt;thread&lt;/a&gt;). On the product side, &lt;a href=&quot;https://x.com/_philschmid/status/2077802206229672264&quot;&gt;Google AI Studio&lt;/a&gt; added a &lt;strong&gt;free tier for Managed Agents&lt;/strong&gt;, plus &lt;strong&gt;max_total_tokens&lt;/strong&gt; for pausing/resuming long runs and &lt;strong&gt;native cron triggers&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Perplexity’s infra direction was also notable&lt;/strong&gt;: &lt;a href=&quot;https://x.com/NVIDIAAIInfra/status/2077890221212090687&quot;&gt;NVIDIA AI Infra&lt;/a&gt; highlighted Perplexity’s new &lt;strong&gt;SPACE&lt;/strong&gt; secure sandbox platform, with early tests on &lt;strong&gt;NVIDIA Vera CPU&lt;/strong&gt; showing up to &lt;strong&gt;1.9× faster sandbox starts&lt;/strong&gt;—a reminder that sandbox startup latency is now part of agent throughput engineering.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;OpenAI and Anthropic: Safety, Productization, and Developer Workflow Updates&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI acknowledged a dangerous Codex/GPT-5.6 failure mode around file deletion&lt;/strong&gt;: &lt;a href=&quot;https://x.com/thsottiaux/status/2077630111499882637&quot;&gt;Thomas Sottiaux&lt;/a&gt; said OpenAI investigated rare reports where &lt;strong&gt;GPT-5.6 unexpectedly deleted files&lt;/strong&gt;, most commonly when &lt;strong&gt;full access mode&lt;/strong&gt; was enabled without sandboxing or auto review, and when the model attempted to override &lt;strong&gt;$HOME&lt;/strong&gt; for temp directories but mistakenly deleted &lt;strong&gt;$HOME&lt;/strong&gt; itself. OpenAI says it is updating developer messaging, nudging users toward safer permission modes, and adding harness safeguards, with a detailed postmortem forthcoming.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI continued to ship workflow features around Codex and PR review&lt;/strong&gt;: &lt;a href=&quot;https://x.com/OpenAIDevs/status/2077902662973190570&quot;&gt;OpenAI Devs&lt;/a&gt; added &lt;strong&gt;PR Chat&lt;/strong&gt; and &lt;strong&gt;inline code editing&lt;/strong&gt; in Codex for reviewing and editing pull requests in context. OpenAI also announced Office Hours around &lt;strong&gt;GPT-5.6, ChatGPT, and Codex&lt;/strong&gt; (&lt;a href=&quot;https://x.com/reach_vb/status/2077796227651874830&quot;&gt;source&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anthropic upgraded Claude Code review depth&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2077840057130692886&quot;&gt;ClaudeDevs&lt;/a&gt; introduced &lt;strong&gt;effort levels&lt;/strong&gt; for &lt;code&gt;/code-review&lt;/code&gt;, from low cost/low effort to &lt;strong&gt;ultra&lt;/strong&gt;, where a fleet of reviewer agents reproduces findings independently. Anthropic says low effort beats other code-review tools on findings per token, while high/ultra improve severe-issue recall and reduce false positives.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Voice remains a major adoption vector&lt;/strong&gt;: &lt;a href=&quot;https://x.com/sama/status/2077842579232895286&quot;&gt;Sam Altman&lt;/a&gt; said he now talks to ChatGPT more than he types, calling the new voice model a threshold-crossing UX shift. Separately, OpenAI published GPT-Live usage limits in its help center, summarized by &lt;a href=&quot;https://x.com/athyuttamre/status/2077655270541648369&quot;&gt;@athyuttamre&lt;/a&gt;: &lt;strong&gt;Pro users get unlimited daily usage&lt;/strong&gt;, while Plus/Go and free tiers have bounded live minutes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Multimodal Video, Real-Time Media, and Creative Tooling&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Google pushed Gemini Omni into Vids&lt;/strong&gt;: &lt;a href=&quot;https://x.com/Google/status/2077786615800295712&quot;&gt;Google&lt;/a&gt; and &lt;a href=&quot;https://x.com/GoogleWorkspace/status/2077786086974140732&quot;&gt;Google Workspace&lt;/a&gt; launched &lt;strong&gt;Gemini Omni&lt;/strong&gt; for video generation/editing in &lt;strong&gt;Google Vids&lt;/strong&gt;, plus &lt;strong&gt;personal avatars&lt;/strong&gt; built from a selfie and voice recording. Google says generated clips include &lt;strong&gt;SynthID&lt;/strong&gt; watermarking and that avatars are restricted to a user’s own account/likeness (&lt;a href=&quot;https://x.com/Google/status/2077786623974965534&quot;&gt;details&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;NotebookLM’s rebrand signals tighter Google product integration&lt;/strong&gt;: &lt;a href=&quot;https://x.com/Gemini_Notebook/status/2077803351392268314&quot;&gt;Gemini Notebook&lt;/a&gt; announced that &lt;strong&gt;NotebookLM is now Gemini Notebook&lt;/strong&gt;, with existing standalone behavior intact but deeper integration coming via the &lt;strong&gt;Gemini app&lt;/strong&gt; and eventually &lt;strong&gt;Search&lt;/strong&gt;. This looks like a packaging/integration move more than a model change.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real-time and agentic media tooling kept advancing&lt;/strong&gt;: &lt;a href=&quot;https://x.com/DecartAI/status/2077801728213156044&quot;&gt;DecartAI&lt;/a&gt; introduced &lt;strong&gt;Lucy 2.5&lt;/strong&gt;, a more capable realtime live AI video editor; &lt;a href=&quot;https://x.com/fal/status/2077811398504075774&quot;&gt;fal&lt;/a&gt; made &lt;strong&gt;Lucy 2.5 Realtime&lt;/strong&gt; available over WebRTC for live video-to-video editing. &lt;a href=&quot;https://x.com/fal/status/2077831513782001775&quot;&gt;fal&lt;/a&gt; also launched &lt;strong&gt;LTX-2.3 Reframe&lt;/strong&gt; for aspect-ratio conversion with generated scene completion.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meta expanded media model distribution&lt;/strong&gt;: &lt;a href=&quot;https://x.com/finkd/status/2077804413251354698&quot;&gt;Meta&lt;/a&gt;, &lt;a href=&quot;https://x.com/AIatMeta/status/2077804869826613422&quot;&gt;AI at Meta&lt;/a&gt;, and &lt;a href=&quot;https://x.com/alexandr_wang/status/2077805347134468378&quot;&gt;Alexandr Wang&lt;/a&gt; all announced &lt;strong&gt;Muse Spark 1.1&lt;/strong&gt; on &lt;strong&gt;OpenRouter&lt;/strong&gt;, reflecting continued demand for frontier-ish generative media models via neutral routing layers.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Robotics, World Models, and Embodied AI&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A high-reliability robotics model stood out&lt;/strong&gt;: &lt;a href=&quot;https://x.com/tonyzzhao/status/2077806003308179802&quot;&gt;Tony Zhao&lt;/a&gt; introduced &lt;strong&gt;ACT-2 Preview&lt;/strong&gt;, described as the first robotics model to unify broad generalization with high reliability. The headline claim is striking: &lt;strong&gt;a single fine-tuning example&lt;/strong&gt; can teach Memo a new behavior that generalizes, with &lt;strong&gt;zero-shot, real unseen homes, 99% success rate&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reka discussed world-model data operations at production scale&lt;/strong&gt;: &lt;a href=&quot;https://x.com/RekaAILabs/status/2077754067359838670&quot;&gt;Reka&lt;/a&gt; pointed to an episode on how a sub-100-person team prepares &lt;strong&gt;petabytes of video data&lt;/strong&gt; for &lt;strong&gt;world model training&lt;/strong&gt;, emphasizing that the bottleneck is often data platform engineering, not just model architecture.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;There’s continuing work on embodied world-model architectures&lt;/strong&gt;: &lt;a href=&quot;https://x.com/lixin4ever/status/2077804918791176589&quot;&gt;@lixin4ever&lt;/a&gt; highlighted a DAMO effort using &lt;strong&gt;tri-branch DiT&lt;/strong&gt;, &lt;strong&gt;joint cross-modal attention&lt;/strong&gt;, and &lt;strong&gt;250M+ RGB frames with dense depth and optical flow annotations&lt;/strong&gt; to turn a video generation model into a &lt;strong&gt;4D embodied world model&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top Tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Kimi K3 official release&lt;/strong&gt;: Moonshot’s &lt;a href=&quot;https://x.com/Kimi_Moonshot/status/2077830229968683203&quot;&gt;launch post&lt;/a&gt; was the day’s dominant technical tweet, combining model specs, architecture, and release timeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kimi K3 Arena breakthrough&lt;/strong&gt;: &lt;a href=&quot;https://x.com/arena/status/2077824029126504525&quot;&gt;Arena’s Frontend Code Arena #1 post&lt;/a&gt; drew exceptional engagement because it framed K3 as not just strong “for open weights,” but directly ahead of a top closed competitor in a visible product task.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI safety incident disclosure&lt;/strong&gt;: &lt;a href=&quot;https://x.com/thsottiaux/status/2077630111499882637&quot;&gt;OpenAI’s explanation of GPT-5.6 file deletions&lt;/a&gt; was one of the most consequential engineering/safety updates, because it tied model behavior to permission modes, sandboxing, and harness safeguards.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anthropic’s multi-effort code review&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2077840057130692886&quot;&gt;Claude Code’s &lt;code&gt;/code-review&lt;/code&gt; effort levels&lt;/a&gt; is a meaningful productization signal for agentic software engineering: not just “AI review,” but tunable cost/recall tradeoffs and subagent-based verification.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Kimi K3 Launch and Frontier Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uyb88e/kimi_k3_weights_to_be_released_on_the_27th/&quot;&gt;Kimi K3 weights to be released on the 27th.&lt;/a&gt;&lt;/strong&gt; (Activity: 399): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/lg3io1qxxmdh1.png&quot;&gt;announcement image&lt;/a&gt; states that &lt;strong&gt;Kimi K3&lt;/strong&gt; is now available through kimi.com, the Kimi app, Kimi Work desktop client, Kimi Code, and the Kimi API, with the current default “thinking intensity” set to &lt;strong&gt;max / extreme&lt;/strong&gt;. Per the linked official posts (&lt;a href=&quot;https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ&quot;&gt;WeChat&lt;/a&gt;, &lt;a href=&quot;https://www.kimi.com/blog/kimi-k3&quot;&gt;English blog&lt;/a&gt;), &lt;strong&gt;full model weights&lt;/strong&gt; and additional technical details are scheduled for release by &lt;strong&gt;July 27, 2026&lt;/strong&gt;, which is the main technical significance of the image.&lt;/strong&gt; Commenters are excited about the open-weight release but expect local inference to be impractical due to the model’s apparent scale, joking that even if someone runs the rumored &lt;code&gt;2.8T&lt;/code&gt;-parameter model on a &lt;code&gt;24 GB&lt;/code&gt; VRAM laptop, it would be at unusably low throughput.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlight that &lt;strong&gt;Kimi K3’s apparent &lt;code&gt;2.8T&lt;/code&gt;-parameter scale&lt;/strong&gt; makes local inference impractical for nearly all consumer setups; one linked screenshot of the announcement/spec context is &lt;a href=&quot;https://preview.redd.it/3goqbghpymdh1.png?width=1661&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=424a861804aad716a9e70fddf5a8aab8cae1abb9&quot;&gt;here&lt;/a&gt;. The discussion frames the weights release as valuable for openness and research even if typical local hardware would be limited to extremely slow or unrealistic runs, e.g. &lt;em&gt;“24 Gb VRAM laptop… &lt;code&gt;0.01&lt;/code&gt; token per sec.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;A technically substantive workflow suggestion was to use &lt;strong&gt;Kimi’s largest models for planning/strategy&lt;/strong&gt; while pairing them with a smaller implementation model, similar to &lt;strong&gt;DeepSeek’s&lt;/strong&gt; large/small model split. One commenter specifically asked for a &lt;strong&gt;sub-&lt;code&gt;300B&lt;/code&gt; MoE or smaller MoonshotAI model&lt;/strong&gt; for lighter coding workloads, noting that K2.7 Code appeared to improve over &lt;strong&gt;K2.6&lt;/strong&gt; and &lt;strong&gt;K2.5&lt;/strong&gt; for agentic coding use cases.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uy3a0q/kimi_k3_released_on_web_and_app/&quot;&gt;Kimi K3 released on web and app&lt;/a&gt;&lt;/strong&gt; (Activity: 1057): &lt;strong&gt;&lt;strong&gt;Kimi K3&lt;/strong&gt; was announced as available on web/app, with claimed specs of &lt;strong&gt;&lt;code&gt;2.8T&lt;/code&gt; parameters&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;1M&lt;/code&gt; context&lt;/strong&gt;, and claims of leading performance in coding, agentic tasks, long-horizon reasoning, visual understanding, and agent-swarm workflows (&lt;a href=&quot;https://preview.redd.it/4uqr0aggildh1.png?width=824&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=cdc3ece2cd45914092d83bd3dd233b17d95d3f54&quot;&gt;screenshot&lt;/a&gt;). No benchmark data, architecture details, license, or Hugging Face/open-weight release link were provided in the post.&lt;/strong&gt; Commenters focused on deployment practicality: a &lt;code&gt;2.8T&lt;/code&gt; model would be extremely difficult to run locally, with one noting even a &lt;code&gt;1.58-bit&lt;/code&gt; quant likely would not fit in &lt;code&gt;512 GB&lt;/code&gt; RAM. Others questioned whether it would become the largest open-weight model if uploaded to HF and said they were waiting for benchmarks.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Discussion focused on the &lt;strong&gt;hardware infeasibility&lt;/strong&gt; of running Kimi K3 locally: commenters cite the reported &lt;strong&gt;&lt;code&gt;2.8T&lt;/code&gt; parameter&lt;/strong&gt; size and note that even a &lt;strong&gt;&lt;code&gt;1.58-bit&lt;/code&gt; quantized&lt;/strong&gt; version would likely exceed &lt;strong&gt;&lt;code&gt;512 GB&lt;/code&gt; RAM&lt;/strong&gt;, putting it far beyond typical consumer or even workstation setups.&lt;/li&gt;
&lt;li&gt;Several users framed Kimi K3 as potentially one of the &lt;strong&gt;largest open-weight models&lt;/strong&gt; if released on Hugging Face, with interest centered on forthcoming benchmarks. One commenter compared an &lt;strong&gt;RTX 6000 Pro &lt;code&gt;96 GB&lt;/code&gt;&lt;/strong&gt; card against the model’s memory requirements, estimating it is still more than &lt;strong&gt;&lt;code&gt;12x&lt;/code&gt; short&lt;/strong&gt;, underscoring that even high-end single-GPU hardware is not sufficient.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uy9cft/kimi_k3_benchmarks/&quot;&gt;Kimi K3 Benchmarks&lt;/a&gt;&lt;/strong&gt; (Activity: 1487): &lt;strong&gt;The image is a &lt;strong&gt;coding benchmark chart&lt;/strong&gt; for &lt;strong&gt;Kimi K3&lt;/strong&gt; (&lt;a href=&quot;https://i.redd.it/yuyk4c99mmdh1.jpeg&quot;&gt;image&lt;/a&gt;), comparing it with models such as &lt;code&gt;GPT-5.6 Sol&lt;/code&gt;, &lt;code&gt;Fable 5&lt;/code&gt;, &lt;code&gt;Opus-4.8&lt;/code&gt;, &lt;code&gt;GPT-5.5&lt;/code&gt;, and &lt;code&gt;GLM-5.2&lt;/code&gt; across six coding evaluations. Kimi K3 is highlighted in blue and is shown leading &lt;strong&gt;Program Bench&lt;/strong&gt; and &lt;strong&gt;SWE Marathon&lt;/strong&gt;, while placing second on &lt;strong&gt;Terminal Bench 2.1&lt;/strong&gt;, &lt;strong&gt;FrontierSWE&lt;/strong&gt;, and &lt;strong&gt;Kimi Code Bench 2.0&lt;/strong&gt;, suggesting very strong benchmark-level coding performance.&lt;/strong&gt; Commenters cautioned that the chart only reflects benchmark performance, not real-world usage, but one argued Chinese models appear “not even 6 months behind US models,” perhaps “6 days behind.” Another comment, “2TB VRAM Is All You Need,” appears to be a joke or jab about likely heavy inference hardware requirements.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter interprets the shared Kimi K3 benchmark image as evidence that &lt;strong&gt;Chinese frontier models are nearly at parity with U.S. models&lt;/strong&gt;, saying that based on benchmarks alone they appear &lt;em&gt;“not even 6 months behind US models”&lt;/em&gt; and possibly closer to &lt;em&gt;“6 days behind”&lt;/em&gt;. They explicitly caveat that this is &lt;strong&gt;benchmark-only&lt;/strong&gt; and may not reflect real-world usage quality or reliability.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uydii0/kimi_k3_beats_claude_fable_and_gpt_56_sol_in/&quot;&gt;KIMI K3 Beats Claude Fable and GPT 5.6 sol in arena.ai!!!&lt;/a&gt;&lt;/strong&gt; (Activity: 854): &lt;strong&gt;The image is a &lt;strong&gt;Code Arena WebDev overall leaderboard&lt;/strong&gt; screenshot (&lt;a href=&quot;https://i.redd.it/sry915x7dndh1.png&quot;&gt;image&lt;/a&gt;) dated Jul 16, 2026, showing &lt;strong&gt;Moonshot’s &lt;code&gt;kimi-k3&lt;/code&gt; ranked #1&lt;/strong&gt; with a score of &lt;code&gt;1679&lt;/code&gt;, ahead of &lt;code&gt;claude-fable-5&lt;/code&gt; and &lt;code&gt;gpt-5.6-sol-xhigh&lt;/code&gt; on front-end web development tasks. The post frames this as surprising because Kimi is beating “frontier” models described as &lt;em&gt;“too dangerous”&lt;/em&gt; for public release; a commenter notes that on the broader &lt;a href=&quot;https://arena.ai/leaderboard/text&quot;&gt;arena.ai text leaderboard&lt;/a&gt;, it is not #1 but still appears competitive with &lt;code&gt;gemini-3-pro&lt;/code&gt; and &lt;code&gt;gpt-5.6-sol-xhigh&lt;/code&gt;.&lt;/strong&gt; Comments focus on whether this implies China is only &lt;em&gt;“6 days behind the west”&lt;/em&gt; and whether &lt;code&gt;kimi-k3&lt;/code&gt; will actually be released as &lt;strong&gt;open weights&lt;/strong&gt;, which would affect its practical significance beyond leaderboard placement.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter links the &lt;strong&gt;arena.ai text leaderboard&lt;/strong&gt; (https://arena.ai/leaderboard/text) and notes that &lt;strong&gt;Kimi K3&lt;/strong&gt; is not leading the main text arena, but is reportedly scoring in the same range as &lt;strong&gt;Gemini 3 Pro&lt;/strong&gt; and &lt;strong&gt;GPT 5.6 sol (xhigh)&lt;/strong&gt;, which they consider technically notable for a Chinese model release.&lt;/li&gt;
&lt;li&gt;There is uncertainty over whether &lt;strong&gt;Kimi K3&lt;/strong&gt; will be released as &lt;strong&gt;open weights&lt;/strong&gt;, which is a key technical distinction for local deployment, fine-tuning, and reproducibility compared with API-only leaderboard performance.&lt;/li&gt;
&lt;li&gt;One commenter raises a benchmark-validity concern: if Arena users disproportionately judge models on generated &lt;strong&gt;Three.js / 3D browser games&lt;/strong&gt;, Kimi may have been optimized for that task distribution. They argue this could inflate perceived capability because visually impressive generated games may score well with casual evaluators even if they are not a robust measure of general coding or reasoning ability.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uycepz/kimi_k3_achieves_3rd_place_on_artificalanalysis/&quot;&gt;Kimi K3 achieves 3rd Place on ArtificalAnalysis, beating out Claude Opus 4.8&lt;/a&gt;&lt;/strong&gt; (Activity: 656): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/5vorrnbx5ndh1.png&quot;&gt;image&lt;/a&gt; is a technical benchmark chart from &lt;strong&gt;Artificial Analysis&lt;/strong&gt; showing &lt;strong&gt;Kimi K3&lt;/strong&gt; in &lt;code&gt;3rd&lt;/code&gt; place on the Intelligence Index with a score of &lt;code&gt;57&lt;/code&gt;, narrowly ahead of &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt; at &lt;code&gt;56&lt;/code&gt; and behind &lt;strong&gt;Claude Fable 5&lt;/strong&gt; (&lt;code&gt;60&lt;/code&gt;) and &lt;strong&gt;GPT-5.6&lt;/strong&gt; (&lt;code&gt;59&lt;/code&gt;). Commenters add that follow-up charts for &lt;a href=&quot;https://preview.redd.it/ayxi7od6bndh1.png?width=1753&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=14190215c0ae612463e1d7e9a7587b2d5e0c5b48&quot;&gt;cost per task&lt;/a&gt; and &lt;a href=&quot;https://preview.redd.it/y1o9gzdn9ndh1.png?width=1007&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=ecf8bcd32522d4397c88647415c2dbfa395394c9&quot;&gt;output tokens per task&lt;/a&gt; look “super promising,” but the main technical caveat is whether the model sustains quality in long sessions at roughly &lt;strong&gt;Sonnet-like costs&lt;/strong&gt; and around &lt;code&gt;30 t/s&lt;/code&gt;.&lt;/strong&gt; The main skepticism is benchmark fatigue: one commenter says they’ve “seen enough bar-charts” and wants real long-session usage reports before accepting the ranking as meaningful.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused less on the headline rank and more on operational efficiency: one noted that at roughly &lt;strong&gt;Claude Sonnet-level pricing&lt;/strong&gt; and around &lt;code&gt;30 tokens/s&lt;/code&gt;, Kimi K3 would need to show strong &lt;em&gt;long-session reasoning efficiency&lt;/em&gt; rather than just benchmark-bar performance. This frames the model’s ArtificialAnalysis placement as needing validation through sustained interactive workloads, not only leaderboard scores.&lt;/li&gt;
&lt;li&gt;A linked follow-up claimed Kimi K3 looks promising on &lt;strong&gt;cost per task&lt;/strong&gt; and &lt;strong&gt;output tokens per task&lt;/strong&gt;, sharing ArtificialAnalysis-style charts: https://preview.redd.it/ayxi7od6bndh1.png?width=1753&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=14190215c0ae612463e1d7e9a7587b2d5e0c5b48. The discussion implies Kimi K3’s competitiveness may come from a favorable efficiency/price profile in addition to raw benchmark rank, especially if it is outperforming or approaching models like &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. New Open-Weight Model Releases&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxdv34/thinking_machines_releases_first_openweight_model/&quot;&gt;Thinking Machines releases first open-weight model “Inkling”&lt;/a&gt;&lt;/strong&gt; (Activity: 1775): &lt;strong&gt;Thinking Machines announced its first open-weight model, &lt;strong&gt;Inkling&lt;/strong&gt;, and the &lt;a href=&quot;https://i.redd.it/d7s0z8kqpfdh1.jpeg&quot;&gt;image&lt;/a&gt; shows it on a model leaderboard at &lt;strong&gt;&lt;code&gt;1257&lt;/code&gt;&lt;/strong&gt;, roughly mid-pack and tied with &lt;strong&gt;Claude Opus 4.6&lt;/strong&gt;, just below &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt;. Per the linked announcement, Inkling is a &lt;strong&gt;MoE transformer&lt;/strong&gt; with &lt;strong&gt;&lt;code&gt;975B&lt;/code&gt; total / &lt;code&gt;41B&lt;/code&gt; active parameters&lt;/strong&gt;, &lt;strong&gt;&lt;code&gt;1M&lt;/code&gt; token context&lt;/strong&gt;, and pretraining over &lt;strong&gt;&lt;code&gt;45T&lt;/code&gt; tokens&lt;/strong&gt; spanning text, images, audio, and video; commenters also note a preview &lt;strong&gt;Inkling-Small&lt;/strong&gt; variant with &lt;strong&gt;&lt;code&gt;12B&lt;/code&gt; active parameters&lt;/strong&gt; for lower cost/latency.&lt;/strong&gt; Commenters were interested because Thinking Machines is led by the former OpenAI CTO and is entering open weights, but there was skepticism about adoption because Inkling appears not to outperform competing open models like &lt;strong&gt;GLM-5.2&lt;/strong&gt; on the shown leaderboard.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted &lt;strong&gt;Inkling&lt;/strong&gt;’s core architecture/specs: a &lt;strong&gt;mixture-of-experts transformer&lt;/strong&gt; with &lt;code&gt;975B&lt;/code&gt; total parameters, &lt;code&gt;41B&lt;/code&gt; active parameters, &lt;code&gt;1M&lt;/code&gt; token context, and pretraining on &lt;code&gt;45T&lt;/code&gt; tokens spanning &lt;strong&gt;text, images, audio, and video&lt;/strong&gt;. One technical concern was that while it is multimodal and similarly sparse to competing open-weight MoE models, it reportedly does &lt;strong&gt;not outperform GLM-5.2&lt;/strong&gt;, which may limit adoption among users prioritizing benchmark leadership.&lt;/li&gt;
&lt;li&gt;The most technically interesting discussion centered on &lt;a href=&quot;https://thinkingmachines.ai/news/introducing-inkling/#inkling-small&quot;&gt;&lt;strong&gt;Inkling-Small&lt;/strong&gt;&lt;/a&gt;: a &lt;code&gt;276B&lt;/code&gt; total-parameter MoE with only &lt;code&gt;12B&lt;/code&gt; active parameters, positioned as a lower-latency/lower-cost sibling to the &lt;code&gt;41B&lt;/code&gt;-active Inkling. Commenters noted that Thinking Machines claims Inkling-Small &lt;strong&gt;matches or exceeds the larger model on many benchmarks&lt;/strong&gt;, attributed to improvements in the pretraining data mix and recipe, making it potentially attractive for high-end local/home inference despite the large total parameter count.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxhpws/inkling_by_thinking_machines_is_the_1_us_open/&quot;&gt;Inkling by Thinking Machines is the #1 US open weight model now&lt;/a&gt;&lt;/strong&gt; (Activity: 460): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/e6a6fr64ggdh1.jpeg&quot;&gt;image&lt;/a&gt; is a BenchmarkList page for &lt;strong&gt;Thinking Machines Lab’s Inkling&lt;/strong&gt;, described as an &lt;strong&gt;open-weights reasoning multimodal model&lt;/strong&gt; released &lt;code&gt;2026-07-15&lt;/code&gt;, with an Experimental ECI of &lt;code&gt;132.71&lt;/code&gt;, global SOTA rank &lt;code&gt;#30/867&lt;/code&gt;, and global open-weight rank &lt;code&gt;#5/169&lt;/code&gt;. In the post’s framing, Inkling is claimed to be the &lt;strong&gt;#1 U.S. open-weight model&lt;/strong&gt;, outperforming U.S. peers such as &lt;strong&gt;NVIDIA Nemotron Ultra&lt;/strong&gt;, though a commenter notes Inkling is reportedly &lt;strong&gt;nearly &lt;code&gt;1T&lt;/code&gt; parameters&lt;/strong&gt; versus Nemotron Ultra’s &lt;code&gt;550B&lt;/code&gt;, making raw comparisons parameter-scale-sensitive.&lt;/strong&gt; Comments were skeptical of the benchmark framing: one points out that OP appears affiliated with the benchmark site shown in the screenshot, and another mocks the claim as potentially weak relative to non-U.S. open-weight leaders.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that &lt;strong&gt;Inkling&lt;/strong&gt; is reportedly close to &lt;code&gt;1T&lt;/code&gt; parameters, making comparisons against &lt;strong&gt;Nemotron Ultra&lt;/strong&gt; (&lt;code&gt;550B&lt;/code&gt;) potentially parameter-count-skewed rather than purely architecture/training-efficiency based. The main technical criticism was that while the benchmark lead may be notable, the model is &lt;em&gt;“too big”&lt;/em&gt; for many practical open-weight users due to likely inference cost, VRAM requirements, and deployment complexity.&lt;/li&gt;
&lt;li&gt;One commenter flagged that the original poster appears affiliated with the benchmark site shown in the screenshot, linking a Reddit search for the author’s posts: https://arctic-shift.photon-reddit.com/search?fun=posts_search&amp;#x26;author=davidthesong&amp;#x26;before=2026-07-15T20%3A40%3A39&amp;#x26;limit=10&amp;#x26;sort=desc. This raises a benchmark-interpretation concern: the ranking may need scrutiny around methodology, benchmark selection, and whether the post is promotional rather than independent evaluation.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxao7y/german_ai_consortium_releases_soofi_s_an_open_30b/&quot;&gt;German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German&lt;/a&gt;&lt;/strong&gt; (Activity: 445): &lt;strong&gt;&lt;strong&gt;Soofi S&lt;/strong&gt; is presented as an open German-led &lt;code&gt;30B-A3B&lt;/code&gt; MoE LLM (&lt;code&gt;31.6B&lt;/code&gt; total, ~&lt;code&gt;3.2B&lt;/code&gt; active/token) based on &lt;strong&gt;Nvidia Nemotron 3 Nano&lt;/strong&gt;’s hybrid &lt;strong&gt;Mamba-2/Transformer&lt;/strong&gt; architecture, with claimed near-flat long-context serving throughput from &lt;code&gt;4K–256K&lt;/code&gt; via reduced KV-cache attention layers; the team says it underwent full pretraining rather than being a finetune, with a &lt;a href=&quot;https://arxiv.org/abs/2607.09424&quot;&gt;paper&lt;/a&gt;, &lt;a href=&quot;https://wandb.ai/soofi-exchange/pretrain-nemotron-3-nano-on-20T-4/reports/Soofi-S-Pretraining--VmlldzoxNzM4NTQ4NA?accessToken=c6mcvzhsloyc1v4duq9c7eq9aa81sr6b8j1l6yju6sbyz1skgecggj1pun9qxb52&quot;&gt;W&amp;#x26;B training logs&lt;/a&gt;, and &lt;a href=&quot;https://github.com/soofi-project/Soofi-Pretraining&quot;&gt;pretraining scripts&lt;/a&gt; available. It was reportedly trained on ~&lt;code&gt;27T&lt;/code&gt; tokens with a German-heavy mix including machine-translated/synthetic German, and claims leading aggregate English/German benchmark results among fully open models, though commenters note missing comparisons to newer baselines and that &lt;strong&gt;Qwen3.5 35B-A3B&lt;/strong&gt; appears to beat it on some German results; gated &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Instruct-Preview-GGUF&quot;&gt;GGUF&lt;/a&gt; and reasoning &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Rhine-Preview-GGUF&quot;&gt;GGUF&lt;/a&gt; builds exist.&lt;/strong&gt; Commenters were skeptical of the benchmark framing, arguing the model is compared against older systems rather than newer &lt;strong&gt;Qwen/Gemma&lt;/strong&gt; variants, and that math/coding-heavy benchmark aggregates may not measure German generation quality. There was also concern over licensing ambiguity: the project advertises &lt;em&gt;“Sovereign, Open Source Long-term, license-free availability for industry”&lt;/em&gt; while the model card reportedly uses a custom &lt;code&gt;Other&lt;/code&gt; license with the full license text not yet filled in.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Soofi S&lt;/strong&gt; is described as a newly pretrained model rather than a finetune: commenters note it is based on &lt;strong&gt;Nemotron 3 Nano&lt;/strong&gt;, underwent full pretraining plus additional phases, and has public artifacts including the &lt;a href=&quot;https://arxiv.org/abs/2607.09424&quot;&gt;paper&lt;/a&gt;, &lt;a href=&quot;https://wandb.ai/soofi-exchange/pretrain-nemotron-3-nano-on-20T-4/reports/Soofi-S-Pretraining--VmlldzoxNzM4NTQ4NA?accessToken=c6mcvzhsloyc1v4duq9c7eq9aa81sr6b8j1l6yju6sbyz1skgecggj1pun9qxb52&quot;&gt;training logs&lt;/a&gt;, and &lt;a href=&quot;https://github.com/soofi-project/Soofi-Pretraining&quot;&gt;training scripts&lt;/a&gt;. One technical concern raised is that the architecture may have weaker long-context behavior; the release includes a &lt;strong&gt;RULER&lt;/strong&gt; test but apparently lacks newer long-context evaluations.&lt;/li&gt;
&lt;li&gt;Several commenters question the benchmark framing: the release reportedly compares against older models while omitting newer baselines like &lt;strong&gt;Qwen 3.6&lt;/strong&gt; or &lt;strong&gt;Gemma 4&lt;/strong&gt;. Another noted that, according to Soofi’s own benchmarks, &lt;strong&gt;Qwen3.5 35B-A3B&lt;/strong&gt; beats Soofi S on German despite not being especially German-focused, suggesting the benchmark may emphasize general understanding, math, or coding rather than native-quality German generation.&lt;/li&gt;
&lt;li&gt;The data and release details drew scrutiny: the training mix reportedly includes &lt;em&gt;“machine-translated and synthetically generated German texts,”&lt;/em&gt; which commenters warned can produce unnatural German due to translation artifacts. There was also concern about licensing ambiguity: the project claims &lt;em&gt;“Sovereign, Open Source Long-term, license-free availability for industry,”&lt;/em&gt; but the model card is marked as a custom &lt;strong&gt;“Other”&lt;/strong&gt; license with the actual license text apparently missing; GGUF builds exist on Hugging Face, including &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Instruct-Preview-GGUF&quot;&gt;Instruct Preview GGUF&lt;/a&gt; and &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Rhine-Preview-GGUF&quot;&gt;Rhine reasoning GGUF&lt;/a&gt;, but are gated.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Local Inference Runtime Upgrades&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwylut/exllamav3_v100_major_performance_upgrades/&quot;&gt;ExLlamaV3 v1.0.0 - Major Performance Upgrades&lt;/a&gt;&lt;/strong&gt; (Activity: 467): &lt;strong&gt;The image is a &lt;strong&gt;technical benchmark table&lt;/strong&gt;, not a meme, for the post titled &lt;strong&gt;“ExLlamaV3 v1.0.0 - Major Performance Upgrades”&lt;/strong&gt; (&lt;a href=&quot;https://i.redd.it/ej7102hqfcdh1.png&quot;&gt;image&lt;/a&gt;). It shows RTX 3090 decode throughput comparisons between &lt;code&gt;v0.0.43&lt;/code&gt;, &lt;code&gt;v1.0.0 mcg&lt;/code&gt;, and &lt;code&gt;v1.0.0 mul1&lt;/code&gt;, with large speedups across EXL3-quantized models—e.g. &lt;strong&gt;Qwen 3.5 0.8B&lt;/strong&gt; rising from about &lt;code&gt;268&lt;/code&gt; to &lt;code&gt;444 tok/s&lt;/code&gt; and &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; from about &lt;code&gt;29&lt;/code&gt; to &lt;code&gt;50 tok/s&lt;/code&gt;. The benchmarks contextualize the release notes: ExLlamaV3 removes FlashAttention-2/xFormers dependencies, adds new attention/GEMM/GEMV/MoE kernels, broader tensor parallelism, and KV-cache quantization improvements that reportedly avoid prior slowdown.&lt;/strong&gt; Comments are mostly positive, praising &lt;strong&gt;turboderp&lt;/strong&gt;’s solo development effort; one commenter clarifies that ExLlamaV3 is an Nvidia-GPU-focused LLM inference engine using the &lt;strong&gt;EXL3&lt;/strong&gt; format rather than GGUF/llama.cpp.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter summarized the core technical scope: &lt;strong&gt;ExLlamaV3/ExLlama3&lt;/strong&gt; is an LLM inference engine targeting a dedicated &lt;strong&gt;EXL3&lt;/strong&gt; model format rather than common &lt;strong&gt;GGUF&lt;/strong&gt; used by llama.cpp-style runtimes, and it is currently &lt;strong&gt;NVIDIA GPU-only&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;One user highlighted ecosystem integration concerns, specifically hoping &lt;strong&gt;TabbyAPI&lt;/strong&gt; improves tool-calling compatibility with &lt;strong&gt;Claude Code&lt;/strong&gt; so ExLlamaV3 can be used locally. They also noted they are currently using &lt;strong&gt;GGUF with MTP&lt;/strong&gt; and are interested in comparing the quality of &lt;strong&gt;EXL3 quantization&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxfu4k/google_is_updating_gemma_4s_chat_templates/&quot;&gt;Google is updating Gemma 4&apos;s chat templates, bringing major fixes to tool calling and reducing &quot;laziness&quot;, and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!&lt;/a&gt;&lt;/strong&gt; (Activity: 956): &lt;strong&gt;&lt;strong&gt;Google Gemma&lt;/strong&gt; announced updated &lt;strong&gt;Gemma 4&lt;/strong&gt; artifacts/templates for testing, with claimed speedups, &lt;strong&gt;Flash Attention 4 on Hopper GPUs&lt;/strong&gt;, improved tool-calling behavior, reduced “laziness,” and a vision token-budget/optimization demo on Hugging Face Spaces: &lt;a href=&quot;https://huggingface.co/spaces/google/gemma4_vision_token_budget&quot;&gt;&lt;code&gt;google/gemma4_vision_token_budget&lt;/code&gt;&lt;/a&gt;. A commenter enumerated the relevant &lt;code&gt;google/gemma-4-31B-it&lt;/code&gt; chat-template commits, including fixes for null handling, turn-tag balance, input validation, restoration of model-turn/thinking cues after tool responses, prevention of extra &lt;code&gt;&amp;#x3C;turn|&gt;&lt;/code&gt; emission, tool-call-only turn closure, and—most emphasized—restoring &lt;code&gt;add_generation_prompt&lt;/code&gt; behavior plus the &lt;code&gt;preserve_thinking&lt;/code&gt; default (&lt;a href=&quot;https://huggingface.co/google/gemma-4-31B-it/commit/68abe48010cbe15293462fa11e901a60639a44e5&quot;&gt;commit list&lt;/a&gt;). The practical impact is mostly prompt-serialization correctness for multi-turn/tool-call traces: preserving or scoping the thinking channel appropriately, avoiding malformed continuation turns/newlines, and making historical assistant/tool-call turns render consistently.&lt;/strong&gt; Commenters framed the fixes as resolving confusing user-side failures rather than requiring new prompting tricks, with particular enthusiasm for &lt;code&gt;preserve_thinking&lt;/code&gt;. One commenter also noted the original X links appeared incorrect and supplied what they believed was the correct Gemma post: https://x.com/googlegemma/status/2077449152062247219.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter enumerated the Hugging Face commits for &lt;code&gt;google/gemma-4-31B-it&lt;/code&gt;, showing the chat-template update is largely about &lt;strong&gt;tool-calling and reasoning/thinking-channel correctness&lt;/strong&gt;: null handling, reasoning preservation, balanced turn tags, input validation, restored model turn/thinking cues after tool responses, and fixes for extra &lt;code&gt;&amp;#x3C;turn|&gt;&lt;/code&gt; emission. The list also highlights changes around &lt;code&gt;preserve_thinking&lt;/code&gt;, APC primers, continuation turns, and tool-call-only turn closure, with the aggregate commit referenced at https://huggingface.co/google/gemma-4-31B-it/commit/68abe48010cbe15293462fa11e901a60639a44e5.&lt;/li&gt;
&lt;li&gt;One user reported that the advertised reduction in Gemma 4 “laziness” was &lt;strong&gt;not resolved by the latest chat template&lt;/strong&gt;, arguing it appears to be a &lt;strong&gt;model behavior issue rather than a template-formatting issue&lt;/strong&gt;. This is an anecdotal but technically relevant distinction: prompt/chat-template fixes may improve tool-call formatting and reasoning preservation without materially changing completion effort or refusal/under-answering tendencies.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Kimi K3 Launch and Coding Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uy3oij/chinese_fable_5_is_here_aka_kimi_k3/&quot;&gt;Chinese fable 5 is here !! Aka kimi k3&lt;/a&gt;&lt;/strong&gt; (Activity: 1106): &lt;strong&gt;The image is a &lt;a href=&quot;https://i.redd.it/uv4k60ydlldh1.jpeg&quot;&gt;screenshot&lt;/a&gt; of a verified &lt;strong&gt;Chetaslua&lt;/strong&gt; post announcing &lt;strong&gt;Kimi K3&lt;/strong&gt; on the web, claiming a &lt;code&gt;1 million&lt;/code&gt; token/context window and showing a dark-themed Kimi UI with modes/options like &lt;strong&gt;K3 Max&lt;/strong&gt;, &lt;strong&gt;swarm&lt;/strong&gt;, &lt;strong&gt;slides&lt;/strong&gt;, and &lt;strong&gt;deep research&lt;/strong&gt;. In the Reddit context, the title frames it as “Chinese fable 5,” implying a high-end Chinese LLM competitor, but the post provides no benchmark table or reproducible evaluation—only launch/UX claims and anecdotal praise of a demo.&lt;/strong&gt; Commenters’ early impressions are mixed but competitive: one says it feels &lt;em&gt;faster than Claude, but less accurate&lt;/em&gt;, roughly “on par with GPT 5.5” but below “5.6 or Fable,” while others mainly express excitement that AI model competition is increasing.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Early impressions characterize &lt;strong&gt;Kimi K3 / “Chinese Fable 5”&lt;/strong&gt; as &lt;em&gt;faster than Claude&lt;/em&gt; but with lower accuracy, roughly comparable to &lt;strong&gt;GPT-5.5&lt;/strong&gt; and behind &lt;strong&gt;GPT-5.6&lt;/strong&gt; or &lt;strong&gt;Fable&lt;/strong&gt; in perceived quality. One commenter also notes its chain-of-thought allegedly references &lt;strong&gt;Anthropic content policies&lt;/strong&gt;, suggesting possible policy-style contamination or imitation in reasoning traces.&lt;/li&gt;
&lt;li&gt;A technically detailed comparison highlights &lt;strong&gt;MiniMax M3&lt;/strong&gt; as under-discussed, with the commenter claiming it consistently outperforms &lt;strong&gt;DeepSeek v4 Pro&lt;/strong&gt; and &lt;strong&gt;Mimo 2.5 Pro&lt;/strong&gt; for their workloads. They cite MiniMax’s paid plan as &lt;code&gt;1.7B tokens / $20 per month&lt;/code&gt; with API access, and mention anticipation for a &lt;strong&gt;2.7T-parameter MiniMax&lt;/strong&gt; model.&lt;/li&gt;
&lt;li&gt;For tooling, the MiniMax agent environment is described as providing a &lt;a href=&quot;https://agent.minimax.io&quot;&gt;Debian 12 sandbox&lt;/a&gt; with &lt;code&gt;2GB RAM&lt;/code&gt;, &lt;code&gt;1 Xeon vCore&lt;/code&gt;, and apparently unlimited storage. The commenter reports using &lt;code&gt;cloudflared&lt;/code&gt; tunnels to expose/test APIs from the sandbox, implying it is usable for lightweight agent/API prototyping despite limited compute.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uybldp/kimi_k3_tops_frontend_code_arena/&quot;&gt;Kimi K3 tops Frontend Code Arena&lt;/a&gt;&lt;/strong&gt; (Activity: 1037): &lt;strong&gt;The image is an &lt;a href=&quot;https://i.redd.it/pg72x8ul0ndh1.jpeg&quot;&gt;Arena.ai Frontend Code Arena leaderboard&lt;/a&gt; showing &lt;strong&gt;Kimi-K3 ranked #1&lt;/strong&gt; with an arena score of &lt;code&gt;1,679&lt;/code&gt;, ahead of &lt;strong&gt;Claude Fable 5&lt;/strong&gt; (&lt;code&gt;1,631&lt;/code&gt;) and &lt;strong&gt;GPT-5.6 Sol xHigh&lt;/strong&gt; (&lt;code&gt;1,618&lt;/code&gt;). The technical significance is that Kimi-K3 is being presented as a leading frontend-code-generation model in this benchmark, with commenters emphasizing that it is allegedly &lt;strong&gt;open weights&lt;/strong&gt; and cheaper than Claude Fable, which would make the result notable beyond raw leaderboard placement.&lt;/strong&gt; Commenters framed the result as a win for open-weight/low-cost models and contrasted it with perceived underperformance from Google/Gemini, which is absent from the chart. Some comments also speculated politically about possible U.S. pressure or restrictions around releasing or using Kimi-K3 weights, but those claims are speculative rather than technical.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted &lt;strong&gt;Kimi K3’s Frontend Code Arena lead&lt;/strong&gt; as notable because it is reportedly &lt;strong&gt;open-weight&lt;/strong&gt; while costing around &lt;strong&gt;&lt;code&gt;1/3&lt;/code&gt; of Fable&lt;/strong&gt;, suggesting a strong price/performance result rather than just a benchmark win. Several users framed the result as evidence that Chinese labs may be closing or surpassing benchmark gaps despite restricted access to high-end US chips, though the thread did not provide detailed benchmark methodology or score breakdowns.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. AI Coding Agents: Codex Micro and WebGPU Builds&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uxbwiv/openai_reveals_codex_micro/&quot;&gt;OpenAI reveals Codex Micro&lt;/a&gt;&lt;/strong&gt; (Activity: 1440): &lt;strong&gt;The post claims &lt;strong&gt;OpenAI revealed “Codex Micro”&lt;/strong&gt;, apparently a &lt;code&gt;~$230&lt;/code&gt; keyboard-like hardware product with an integrated microphone, but the linked Reddit-hosted media (&lt;a href=&quot;https://v.redd.it/3u8b331hdfdh1&quot;&gt;v.redd.it/3u8b331hdfdh1&lt;/a&gt;) was inaccessible due to &lt;strong&gt;HTTP 403 Forbidden&lt;/strong&gt;, so the underlying announcement/video could not be verified. No concrete technical specs, model details, APIs, benchmarks, or implementation information were available from the post/comments beyond the implied &lt;em&gt;keyboard + microphone&lt;/em&gt; form factor.&lt;/strong&gt; Comments were overwhelmingly skeptical and confused, with users questioning whether it was an April Fools-style joke and mocking the idea of a &lt;code&gt;&quot;keyboard with a microphone&quot;&lt;/code&gt; at &lt;code&gt;&quot;$230&quot;&lt;/code&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uxy5s8/i_built_a_truescale_atlas_of_the_universe_84m/&quot;&gt;I built a true-scale atlas of the universe (8.4M real stars) in about a week with Fable&lt;/a&gt;&lt;/strong&gt; (Activity: 1078): &lt;em&gt;&lt;em&gt;A developer reports building a WebGPU-based, dependency-free “true-scale” universe atlas in ~1 week using &lt;strong&gt;Claude Code + Fable 5&lt;/strong&gt;, producing ~&lt;code&gt;14.5k&lt;/code&gt; lines of TypeScript/WGSL across &lt;code&gt;92&lt;/code&gt; merged PRs and &lt;code&gt;237&lt;/code&gt; commits; the app renders &lt;strong&gt;8.4M Gaia DR3 stars&lt;/strong&gt;, &lt;strong&gt;2.6M SDSS galaxies&lt;/strong&gt;, real orbital/satellite dynamics, eclipses, Sgr A&lt;/em&gt; lensing, and scale-continuous navigation (&lt;a href=&quot;https://universeatlas.org&quot;&gt;site&lt;/a&gt;, &lt;a href=&quot;https://github.com/chrisjz/universe&quot;&gt;MIT source&lt;/a&gt;). The workflow emphasized automated verification: JPL Horizons CI checks with a &lt;code&gt;0.2°&lt;/code&gt; tolerance, physics-gated data generation, deterministic URL repros, and headless WebGPU rendering via software Vulkan with pixel-diff baselines.&lt;/em&gt;* Commenters mostly framed it as an impressive scientific/visualization use case for new coding models, while one asked how the author bridged the gap between intended visuals and Fable’s typical graphics output, suggesting asset quality may be the limiting factor. Another technical suggestion was to extend it toward an &lt;strong&gt;N-body simulation&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter asked about the practical workflow gap between the intended atlas design and &lt;strong&gt;Fable5&lt;/strong&gt;’s generated output, noting recurring dissatisfaction with Fable’s graphics quality. They suggested the limitation may be asset-related, requiring use of existing assets or custom asset creation in tools like &lt;strong&gt;Blender&lt;/strong&gt; to achieve the desired visual fidelity.&lt;/li&gt;
&lt;li&gt;Another technically relevant suggestion was extending the atlas into an &lt;strong&gt;N-body simulation&lt;/strong&gt;, implying a next step from static star visualization toward gravitational dynamics and large-scale physical simulation.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. AI Governance and Open-Source Adoption&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uxhe4w/anthropic_doesnt_care_about_europe_eu_officials/&quot;&gt;‘Anthropic doesn’t care about Europe’ — EU officials peeved after AI giant sends junior staffer to testify about safety&lt;/a&gt;&lt;/strong&gt; (Activity: 1311): &lt;strong&gt;&lt;a href=&quot;https://www.politico.eu/article/anthropic-european-parliament-donny-greenberg-artificial-intelligence-ai/&quot;&gt;&lt;strong&gt;POLITICO&lt;/strong&gt; reports&lt;/a&gt; that &lt;strong&gt;Anthropic&lt;/strong&gt; sent &lt;strong&gt;Donny Greenberg&lt;/strong&gt;, a newly hired technical employee, to testify remotely before the European Parliament on advanced AI safety risks, despite lawmakers reportedly requesting public-policy lead &lt;strong&gt;Sarah Heck&lt;/strong&gt;. EU officials interpreted the staffing choice, prepared/possibly AI-generated answers, and abrupt exit as a failure to seriously engage with EU AI governance, leaving questions on safety policy and regulatory accountability unanswered.&lt;/strong&gt; Commenters framed the incident as consistent with Anthropic’s perceived &lt;strong&gt;US-first commercial and policy posture&lt;/strong&gt;, citing early-access programs, services, credits, and discounts focused on American companies. Others viewed it as operationally bizarre and unfair to the junior employee, while also reputationally damaging given the EU’s regulatory leverage over AI deployment.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically relevant concern raised was that &lt;strong&gt;Anthropic’s Europe strategy appears underdeveloped&lt;/strong&gt;, specifically around &lt;strong&gt;data residency&lt;/strong&gt;: one commenter said it &lt;em&gt;“seems like an afterthought.”&lt;/em&gt; For EU enterprise and public-sector adoption, this matters because model providers often need regional data processing/storage guarantees, compliance controls, and clear GDPR-aligned deployment options.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uxhnh1/linus_torvalds_reaffirms_that_linux_is_not_antiai/&quot;&gt;Linus Torvalds Reaffirms That Linux Is Not &quot;Anti-AI&quot; And Not A &quot;Social Warrior&quot; Project&lt;/a&gt;&lt;/strong&gt; (Activity: 1157): &lt;strong&gt;&lt;strong&gt;Linus Torvalds&lt;/strong&gt; stated that the Linux kernel will not ban AI/LLM-assisted development or review tooling, arguing that “AI is a tool” and that kernel policy should remain based on &lt;strong&gt;technical merit&lt;/strong&gt;, not ideological opposition (&lt;a href=&quot;https://www.phoronix.com/news/Linux-Is-Not-Anti-AI&quot;&gt;Phoronix&lt;/a&gt;). The context is ongoing debate around Software Freedom Conservancy AI guidance and tools such as &lt;strong&gt;Sashiko&lt;/strong&gt; for AI-assisted kernel review; Torvalds’ position is that such tools must reduce maintainer burden rather than generate low-quality submissions, but usage should not be prohibited.&lt;/strong&gt; Top technical comments broadly agreed, noting that LLM code-generation quality has improved substantially over the last year and is now practically useful, while also arguing that the Linux kernel’s review culture is unlikely to tolerate “reckless slop” because of its high scrutiny and maintainer gatekeeping.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted Torvalds&apos; position as a pragmatic one: Linux kernel development is unlikely to accept &quot;AI slop&quot; unchecked because its patch-review process has unusually high scrutiny and many expert maintainers reviewing submissions. The technical argument is that AI-assisted code can be tolerated as long as the resulting patches meet the same quality, correctness, and maintainability standards as human-written code.&lt;/li&gt;
&lt;li&gt;One commenter argued that AI-generated code quality has materially improved over the last year, contrasting it with experiences from ~&lt;code&gt;2&lt;/code&gt; years ago when tools frequently hallucinated nonexistent APIs and produced poor structure. The implication was that modern AI coding assistants are now &quot;seriously useful&quot; for software engineering workflows, though still dependent on human verification.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uxhpj3/anthropic_warns_that_ai_will_soon_be_able_to/&quot;&gt;Anthropic warns that AI will soon be able to improve itself without human intervention&lt;/a&gt;&lt;/strong&gt; (Activity: 939): &lt;strong&gt;&lt;strong&gt;Anthropic&lt;/strong&gt; is warning policymakers that frontier models may soon materially accelerate AI R&amp;#x26;D workflows—potentially enabling recursive improvement with little human intervention—and is advocating an AI &lt;em&gt;“brake pedal”&lt;/em&gt;: stronger evals, monitoring, and possible deployment pauses for systems that can substantially improve model development pipelines (&lt;a href=&quot;https://edition.cnn.com/2026/06/05/business/anthropic-calls-for-ai-brake-pedal&quot;&gt;CNN&lt;/a&gt;). The technical risk being highlighted is not autonomous self-modification in isolation, but models improving the surrounding research/engineering loop enough to cause rapid capability jumps and weaker human oversight.&lt;/strong&gt; Top comments were skeptical of Anthropic’s motives, arguing the company repeatedly issues alarmist safety warnings while continuing to build frontier systems, framing this as regulatory/market positioning rather than purely public-interest risk disclosure. Others dismissed the warning as repetitive or unserious.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>moonshot-ai</category><category>arena</category><category>artificial-analysis</category><category>kimi-k3</category><category>scaling01</category><category>eliebakouch</category><category>kimmonismus</category><category>nrehiew_</category><category>jianlin_s</category><category>yulun_du</category><category>multimodality</category><category>long-context</category><category>attention-mechanisms</category><category>model-efficiency</category><category>model-performance</category><category>agentic-ai</category><category>coding</category><category>benchmarking</category><category>model-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-15-thinky-inkling/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-15-thinky-inkling/</guid><description>**Thinking Machines Lab** launched **Inkling**, its first fully released open-weights foundation model family, featuring **975B parameters** with **41B active parameters** in a **Mixture-of-Experts** architecture. Inkling supports **multimodality** with text, image, and audio inputs and text output, is **Apache 2.0 licensed**, and offers up to **1M context window**. The model is available on platforms like **Tinker**, **Hugging Face**, and partners, with broad ecosystem support from **vLLM**, **SGLang**, **Modal**, **Baseten**, and **Databricks**. Key figures such as **Mira Murati**, **Soumith Chintala**, **John Schulman**, and **Lilian Weng** highlighted its open weights, customization, and practical use focus. Independent commentators noted it as the strongest U.S.-based open-weight release to date, though still behind top Chinese open-weight and best closed models on some benchmarks.</description><pubDate>Wed, 15 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/14/2026-7/15/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;h2&gt;What happened&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Thinking Machines Lab launched Inkling, its first fully released open-weights foundation model family entry, positioning it as a customizable multimodal base model rather than a benchmark-maxed flagship.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Thinking Machines announced Inkling as an open-weights model that “reasons efficiently across text, image, and audio modalities,” with full weights available and immediate support on its Tinker platform and Playground &lt;a href=&quot;https://x.com/thinkymachines/status/2077454609551921208&quot;&gt;@thinkymachines&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Mira Murati described Inkling as the company’s “first model,” “trained from scratch,” with open weights and same-day fine-tuning on Tinker &lt;a href=&quot;https://x.com/miramurati/status/2077455974743593100&quot;&gt;@miramurati&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Soumith Chintala framed it as Thinking Machines’ “first general model,” stressing open weights, 975B parameters, native multimodality, and availability on Tinker, Hugging Face, and partners &lt;a href=&quot;https://x.com/soumithchintala/status/2077457110728884327&quot;&gt;@soumithchintala&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;John Schulman added timeline context: pretraining began last winter, and from mid-January a small team built coding, reasoning, and agentic training on top &lt;a href=&quot;https://x.com/johnschulman2/status/2077460227327467982&quot;&gt;@johnschulman2&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Lilian Weng characterized Inkling as a foundation model aimed at “solid performance across a broad categories of capabilities” and intended for practical use plus customization &lt;a href=&quot;https://x.com/lilianweng/status/2077471903032528912&quot;&gt;@lilianweng&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;TML staff repeatedly emphasized that this is a day-1 release and a foundation for future iterations rather than their final frontier push &lt;a href=&quot;https://x.com/soumithchintala/status/2077457644474998831&quot;&gt;@soumithchintala&lt;/a&gt;, &lt;a href=&quot;https://x.com/cHHillee/status/2077457790423969806&quot;&gt;@cHHillee&lt;/a&gt;, &lt;a href=&quot;https://x.com/keirp1/status/2077469773684981962&quot;&gt;@keirp1&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The release landed with unusually broad day-0 ecosystem support across vLLM, SGLang, Modal, Baseten, Databricks, Hugging Face, and quantization/community tooling &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/lmsysorg/status/2077457150046269779&quot;&gt;@lmsysorg&lt;/a&gt;, &lt;a href=&quot;https://x.com/modal/status/2077462393441948010&quot;&gt;@modal&lt;/a&gt;, &lt;a href=&quot;https://x.com/baseten/status/2077462904388178107&quot;&gt;@baseten&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077462536337891748&quot;&gt;@Yuchenj_UW&lt;/a&gt;, &lt;a href=&quot;https://x.com/huggingface/status/2077460253235724408&quot;&gt;@huggingface&lt;/a&gt;, &lt;a href=&quot;https://x.com/danielhanchen/status/2077468775478423601&quot;&gt;@danielhanchen&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Independent commentators immediately tagged it as the strongest U.S.-based open-weight release so far, though generally still behind the top Chinese open-weight and best closed models on some benchmarks &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2077465762869194973&quot;&gt;@scaling01&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Core facts and specs&lt;/h2&gt;
&lt;h3&gt;Model size, modality, licensing, context&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Inkling is reported as &lt;strong&gt;975B total parameters / 41B active parameters&lt;/strong&gt; in most posts &lt;a href=&quot;https://x.com/soumithchintala/status/2077457110728884327&quot;&gt;@soumithchintala&lt;/a&gt;, &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2077472478499053846&quot;&gt;@kimmonismus&lt;/a&gt;.
&lt;ul&gt;
&lt;li&gt;One tweet says 974B &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077462536337891748&quot;&gt;@Yuchenj_UW&lt;/a&gt;, and another says 952B &lt;a href=&quot;https://x.com/multimodalart/status/2077469546563461353&quot;&gt;@multimodalart&lt;/a&gt;; the overwhelming consensus in the tweet set is ~975B.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;It is a &lt;strong&gt;Mixture-of-Experts&lt;/strong&gt; model with &lt;strong&gt;41B active&lt;/strong&gt; parameters per token &lt;a href=&quot;https://x.com/VictoriaLinML/status/2077599145502835108&quot;&gt;@VictoriaLinML&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;It is &lt;strong&gt;Apache 2.0 licensed&lt;/strong&gt; according to multiple reactions and summaries &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077462536337891748&quot;&gt;@Yuchenj_UW&lt;/a&gt;, &lt;a href=&quot;https://x.com/multimodalart/status/2077469546563461353&quot;&gt;@multimodalart&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;It supports &lt;strong&gt;text, image, and audio inputs&lt;/strong&gt;, with &lt;strong&gt;text output&lt;/strong&gt; &lt;a href=&quot;https://x.com/soumithchintala/status/2077457110728884327&quot;&gt;@soumithchintala&lt;/a&gt;, &lt;a href=&quot;https://x.com/TheRundownAI/status/2077472283757543602&quot;&gt;@TheRundownAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Open-weights checkpoints support up to &lt;strong&gt;1M context&lt;/strong&gt; &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/lmsysorg/status/2077457150046269779&quot;&gt;@lmsysorg&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Tinker/API context is described as &lt;strong&gt;256K&lt;/strong&gt;, with pricing differentiated for &lt;strong&gt;64K&lt;/strong&gt; and &lt;strong&gt;256K&lt;/strong&gt; contexts &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Training and release details&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;TML says Inkling was &lt;strong&gt;trained from scratch&lt;/strong&gt; &lt;a href=&quot;https://x.com/miramurati/status/2077455974743593100&quot;&gt;@miramurati&lt;/a&gt;, &lt;a href=&quot;https://x.com/LiorOnAI/status/2077464289611563389&quot;&gt;@LiorOnAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Community readers extracted &lt;strong&gt;45T training tokens&lt;/strong&gt; from the release materials &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;, while one post says &lt;strong&gt;48T&lt;/strong&gt; &lt;a href=&quot;https://x.com/mervenoyann/status/2077475202775044523&quot;&gt;@mervenoyann&lt;/a&gt;. The more repeated figure in this dataset is &lt;strong&gt;45T&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Inkling includes &lt;strong&gt;controllable reasoning effort&lt;/strong&gt; / numerical effort levels &lt;a href=&quot;https://x.com/LiorOnAI/status/2077464289611563389&quot;&gt;@LiorOnAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/TheRundownAI/status/2077472283757543602&quot;&gt;@TheRundownAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/danielhanchen/status/2077470080422891872&quot;&gt;@danielhanchen&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Tinker customers highlighted concise reasoning and strong tool calling rather than maximal raw benchmark chasing &lt;a href=&quot;https://x.com/tinkerapi/status/2077467634568929433&quot;&gt;@tinkerapi&lt;/a&gt;, &lt;a href=&quot;https://x.com/MichaelElabd/status/2077461111247712656&quot;&gt;@MichaelElabd&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Architecture details surfaced in reactions&lt;/h3&gt;
&lt;p&gt;Several technically literate reactions extracted architectural choices from the release:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hybrid/sliding-window attention&lt;/strong&gt; with a &lt;strong&gt;5:1 local-to-global layer ratio&lt;/strong&gt; and &lt;strong&gt;window size 512&lt;/strong&gt; &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/ariG23498/status/2077631902228582805&quot;&gt;@ariG23498&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Relative positional encoding / relative attention bias&lt;/strong&gt; instead of RoPE; multiple posters called this one of the most novel large-scale choices &lt;a href=&quot;https://x.com/stochasticchasm/status/2077463965438009677&quot;&gt;@stochasticchasm&lt;/a&gt;, &lt;a href=&quot;https://x.com/eliebakouch/status/2077473407550001461&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/rasbt/status/2077540575255880126&quot;&gt;@rasbt&lt;/a&gt;, &lt;a href=&quot;https://x.com/_arohan_/status/2077519160767386030&quot;&gt;@&lt;em&gt;arohan&lt;/em&gt;&lt;/a&gt;, &lt;a href=&quot;https://x.com/ChangJonathanC/status/2077508340637139318&quot;&gt;@ChangJonathanC&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Short convolution layers&lt;/strong&gt; added around attention/FFN streams; commenters flagged this as unusually scaled-up usage of short convs &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/stochasticchasm/status/2077464183994773607&quot;&gt;@stochasticchasm&lt;/a&gt;, &lt;a href=&quot;https://x.com/rasbt/status/2077540575255880126&quot;&gt;@rasbt&lt;/a&gt;, &lt;a href=&quot;https://x.com/SonglinYang4/status/2077492914683535850&quot;&gt;@SonglinYang4&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MoE with shared expert sinks / 2 shared experts&lt;/strong&gt;, noted as atypical since many recent MoEs use 1 shared expert &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/ariG23498/status/2077631902228582805&quot;&gt;@ariG23498&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DeepSeek-style auxiliary-loss-free load balancing&lt;/strong&gt; was cited in community readings of the architecture &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;muP&lt;/strong&gt; and &lt;strong&gt;Muon/weight decay variants&lt;/strong&gt; were inferred from the writeup and confirmed by optimizer expert reaction: Aaron Defazio said they are using his corrected weight decay approach, “MuonC/AdamC” &lt;a href=&quot;https://x.com/aaron_defazio/status/2077484024726204921&quot;&gt;@aaron_defazio&lt;/a&gt;, while community readers also pointed out muP &lt;a href=&quot;https://x.com/stochasticchasm/status/2077464183994773607&quot;&gt;@stochasticchasm&lt;/a&gt;, &lt;a href=&quot;https://x.com/Laz4rz/status/2077555045701140682&quot;&gt;@Laz4rz&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;8 MTP heads&lt;/strong&gt; for speculative decoding were highlighted by vLLM &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Variants&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Inkling-Small is repeatedly referenced as an upcoming or separately discussed smaller model &lt;a href=&quot;https://x.com/LiorOnAI/status/2077464289611563389&quot;&gt;@LiorOnAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/teortaxesTex/status/2077458155378712673&quot;&gt;@teortaxesTex&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Community summaries describe &lt;strong&gt;Inkling-Small as 276B total / 12B active&lt;/strong&gt; and unexpectedly competitive versus the larger model on several evaluations &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/nrehiew_/status/2077542413133115589&quot;&gt;@nrehiew_&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Performance and benchmarks&lt;/h2&gt;
&lt;h3&gt;Independent benchmark framing&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Artificial Analysis said Inkling debuts at &lt;strong&gt;41 on the Intelligence Index&lt;/strong&gt;, making it the leading U.S. open-weights release and ahead of &lt;strong&gt;Nemotron 3 Ultra (38)&lt;/strong&gt;, &lt;strong&gt;Gemma 4 31B (29)&lt;/strong&gt;, and &lt;strong&gt;gpt-oss-120b (24)&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Artificial Analysis also said Inkling averages &lt;strong&gt;25K output tokens per Intelligence Index task&lt;/strong&gt;, vs &lt;strong&gt;43K&lt;/strong&gt; for &lt;strong&gt;GLM-5.2 max&lt;/strong&gt;, &lt;strong&gt;38K&lt;/strong&gt; for &lt;strong&gt;Kimi K2.6&lt;/strong&gt;, and &lt;strong&gt;37K&lt;/strong&gt; for &lt;strong&gt;DeepSeek v4 Pro max&lt;/strong&gt;, framing it as relatively token-efficient &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Natolambert called it a “clear step up from Nemotron Ultra” and “new best American model,” but still “a bit behind GLM 5.2 on agentic benchies, and Kimi K 2.6 on multi modal” &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Design Arena said Inkling entered Agentic Web App Arena at &lt;strong&gt;#9 overall, Elo 1257&lt;/strong&gt;, in the same band as &lt;strong&gt;Claude Opus 4.6&lt;/strong&gt; and &lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt;, and called it the highest-ranking U.S.-based open-weight model for agentic workloads &lt;a href=&quot;https://x.com/DesignArena/status/2077457201216803257&quot;&gt;@DesignArena&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Arena added Inkling to Agent Arena / Text / Vision / Code Arena on launch day &lt;a href=&quot;https://x.com/arena/status/2077476575281545573&quot;&gt;@arena&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Specific benchmark numbers cited&lt;/h3&gt;
&lt;p&gt;From Artificial Analysis:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GDPval-AA v2 Elo 1238&lt;/strong&gt;, higher than &lt;strong&gt;Kimi K2.6 (1190)&lt;/strong&gt; and &lt;strong&gt;DeepSeek v4 Flash max (1189)&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;τ³-Banking 24%&lt;/strong&gt;, above &lt;strong&gt;Kimi K2.6 (21%)&lt;/strong&gt; and slightly above &lt;strong&gt;DeepSeek v4 Flash max (23%)&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Qualitative performance takes&lt;/h3&gt;
&lt;p&gt;Positive:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Sharp and concise” reasoning, not rambly &lt;a href=&quot;https://x.com/MichaelElabd/status/2077461111247712656&quot;&gt;@MichaelElabd&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Strong tool calling and good long-horizon error recovery on agentic tasks &lt;a href=&quot;https://x.com/MichaelElabd/status/2077461111247712656&quot;&gt;@MichaelElabd&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Good “quality of mind” / unsycophantic flavor &lt;a href=&quot;https://x.com/skirano/status/2077515605939277940&quot;&gt;@skirano&lt;/a&gt;, &lt;a href=&quot;https://x.com/tinkerapi/status/2077467634568929433&quot;&gt;@tinkerapi&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Alex Kirillov claimed Inkling avoids the common “audio in = intelligence penalty” seen in many omni models, though another user asked for stronger supporting evidence and benchmarks &lt;a href=&quot;https://x.com/_alex_kirillov_/status/2077493564066722248&quot;&gt;@&lt;em&gt;alex_kirillov&lt;/em&gt;&lt;/a&gt;, &lt;a href=&quot;https://x.com/giffmana/status/2077522859862139218&quot;&gt;@giffmana&lt;/a&gt;, &lt;a href=&quot;https://x.com/_alex_kirillov_/status/2077526541186355343&quot;&gt;@&lt;em&gt;alex_kirillov&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;More mixed / critical:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Scaling01 argued the benchmarks are “not that great,” describing it as roughly “another Kimi-K2.6” and behind all closed models and GLM-5.2, speculating the release may have been timed ahead of Kimi-K3 and DeepSeek-V4-GA &lt;a href=&quot;https://x.com/scaling01/status/2077465762869194973&quot;&gt;@scaling01&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Stochasticchasm said it seems “very strong for multimodal” but “not super strong for terminal bench etc.” &lt;a href=&quot;https://x.com/stochasticchasm/status/2077463420182712708&quot;&gt;@stochasticchasm&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;JJitsev pushed back on hype around “only open-weight model trained without distilling,” saying Inkling uses distillation from open weights and underperforms GLM 5.2 on TerminalBench-style evals &lt;a href=&quot;https://x.com/JJitsev/status/2077627999352922196&quot;&gt;@JJitsev&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;TeortaxesTex offered a contrarian positive spin: mediocre benchmark-maxing may actually suggest less corner-cutting/distillation contamination and a more independent data pipeline &lt;a href=&quot;https://x.com/teortaxesTex/status/2077483013772816426&quot;&gt;@teortaxesTex&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Inference, systems, and launch ecosystem&lt;/h2&gt;
&lt;h3&gt;Official and partner infrastructure facts&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;NVIDIA said Inkling was trained on &lt;strong&gt;GB300 NVL72&lt;/strong&gt; and that an &lt;strong&gt;NVFP4 checkpoint&lt;/strong&gt; was available on Hugging Face on day 0 &lt;a href=&quot;https://x.com/NVIDIAAI/status/2077456914238292220&quot;&gt;@NVIDIAAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;vLLM said day-0 support includes &lt;strong&gt;NVFP4 and BF16&lt;/strong&gt;, optimized for &lt;strong&gt;Blackwell and Hopper&lt;/strong&gt;, reaching up to &lt;strong&gt;380 tok/s/user on 4× GB200 with MTP&lt;/strong&gt; &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Inferact detailed system work: &lt;strong&gt;sconv-aware tensor-parallel sharding&lt;/strong&gt;, &lt;strong&gt;low-latency fused collectives (5× faster at bs=1)&lt;/strong&gt;, and direct integration of TML’s &lt;strong&gt;FA4 sheared-bias kernel&lt;/strong&gt; &lt;a href=&quot;https://x.com/inferact/status/2077461431306584423&quot;&gt;@inferact&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;LMSYS/SGLang said Inkling architecture support was implemented natively, including &lt;strong&gt;ShortConv&lt;/strong&gt;, &lt;strong&gt;relative positional attention&lt;/strong&gt;, &lt;strong&gt;shared expert sink MoE&lt;/strong&gt;, &lt;strong&gt;prefill full CUDA graph&lt;/strong&gt;, &lt;strong&gt;MXFP8 KV cache&lt;/strong&gt;, &lt;strong&gt;full parameter and LoRA RL in customized Megatron backend&lt;/strong&gt;, &lt;strong&gt;routing replay&lt;/strong&gt;, &lt;strong&gt;cross-runtime parameter sync&lt;/strong&gt;, and &lt;strong&gt;DFlash speculative decoding from Modal&lt;/strong&gt; &lt;a href=&quot;https://x.com/lmsysorg/status/2077457150046269779&quot;&gt;@lmsysorg&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Modal said Inkling on Modal uses a custom &lt;strong&gt;DFlash speculator&lt;/strong&gt; for &lt;strong&gt;67% higher throughput and interactivity&lt;/strong&gt; &lt;a href=&quot;https://x.com/modal/status/2077462393441948010&quot;&gt;@modal&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Soumith Chintala separately amplified that Modal’s DFlash speculator is “much faster than MTP” &lt;a href=&quot;https://x.com/soumithchintala/status/2077500083407667569&quot;&gt;@soumithchintala&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Community optimization observations&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Lysandre reported replacing TML’s causal Conv1D with &lt;code&gt;causal-conv1d&lt;/code&gt; yielded &lt;strong&gt;+4% tok/s&lt;/strong&gt;, and replacing attention with &lt;strong&gt;FlashAttention-4&lt;/strong&gt; yielded another &lt;strong&gt;+11%&lt;/strong&gt;, for ~&lt;strong&gt;15% total throughput gain&lt;/strong&gt; without retraining &lt;a href=&quot;https://x.com/LysandreJik/status/2077459011285512267&quot;&gt;@LysandreJik&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Unsloth released &lt;strong&gt;1-bit GGUF quants&lt;/strong&gt; said to be &lt;strong&gt;86% smaller (270GB vs 1.9TB)&lt;/strong&gt; while retaining &lt;strong&gt;74.2% of top-1% accuracy&lt;/strong&gt;, with vision and audio support &lt;a href=&quot;https://x.com/danielhanchen/status/2077468775478423601&quot;&gt;@danielhanchen&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Pricing and availability&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Artificial Analysis listed Tinker pricing as:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;64K context&lt;/strong&gt;: &lt;strong&gt;$1.87 / 1M input&lt;/strong&gt;, &lt;strong&gt;$0.374 cached&lt;/strong&gt;, &lt;strong&gt;$4.68 output&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;256K context&lt;/strong&gt;: &lt;strong&gt;$3.74 / 1M input&lt;/strong&gt;, &lt;strong&gt;$0.748 cached&lt;/strong&gt;, &lt;strong&gt;$9.36 output&lt;/strong&gt;&lt;br&gt;
&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Available on &lt;strong&gt;Tinker&lt;/strong&gt;, &lt;strong&gt;Hugging Face&lt;/strong&gt;, and via launch partners including &lt;strong&gt;Databricks&lt;/strong&gt;, &lt;strong&gt;Baseten&lt;/strong&gt;, &lt;strong&gt;Modal&lt;/strong&gt;, &lt;strong&gt;vLLM/SGLang&lt;/strong&gt; stacks &lt;a href=&quot;https://x.com/soumithchintala/status/2077457110728884327&quot;&gt;@soumithchintala&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077462536337891748&quot;&gt;@Yuchenj_UW&lt;/a&gt;, &lt;a href=&quot;https://x.com/baseten/status/2077462904388178107&quot;&gt;@baseten&lt;/a&gt;, &lt;a href=&quot;https://x.com/modal/status/2077462393441948010&quot;&gt;@modal&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Facts vs opinions&lt;/h2&gt;
&lt;h3&gt;Factual claims directly supported by launch and partners&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Open weights/full weights released &lt;a href=&quot;https://x.com/thinkymachines/status/2077454609551921208&quot;&gt;@thinkymachines&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Trained from scratch &lt;a href=&quot;https://x.com/miramurati/status/2077455974743593100&quot;&gt;@miramurati&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;975B total / 41B active MoE, multimodal text-image-audio input, 1M context on weights, 256K on Tinker/API &lt;a href=&quot;https://x.com/soumithchintala/status/2077457110728884327&quot;&gt;@soumithchintala&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Apache 2.0 license &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2077462536337891748&quot;&gt;@Yuchenj_UW&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Pretraining began last winter; agentic/coding/reasoning work started mid-January &lt;a href=&quot;https://x.com/johnschulman2/status/2077460227327467982&quot;&gt;@johnschulman2&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Day-0 support on major serving stacks, with concrete performance claims from vLLM/Inferact/Modal/NVIDIA &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/inferact/status/2077461431306584423&quot;&gt;@inferact&lt;/a&gt;, &lt;a href=&quot;https://x.com/modal/status/2077462393441948010&quot;&gt;@modal&lt;/a&gt;, &lt;a href=&quot;https://x.com/NVIDIAAI/status/2077456914238292220&quot;&gt;@NVIDIAAI&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Interpretations and opinions&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;“Best American open model” / “saved American open-source frontier” are judgments, albeit repeated by several respected observers &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;, &lt;a href=&quot;https://x.com/karinanguyen/status/2077473342148448525&quot;&gt;@karinanguyen&lt;/a&gt;, &lt;a href=&quot;https://x.com/saranormous/status/2077469313108422806&quot;&gt;@saranormous&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Claims that Inkling is especially important because it is not distilled from OpenAI/Anthropic are disputed. Jxmnop called it “the ONLY open-weight model” without such distillation &lt;a href=&quot;https://x.com/jxmnop/status/2077504236380946595&quot;&gt;@jxmnop&lt;/a&gt;, then partially walked it back: “apparently they did distill lol. but only a tiny bit” &lt;a href=&quot;https://x.com/jxmnop/status/2077540390128034133&quot;&gt;@jxmnop&lt;/a&gt;. Andrew Carr also contested the purity framing, noting use of Kimi 2.5 for SFT traces &lt;a href=&quot;https://x.com/andrew_n_carr/status/2077509786237854136&quot;&gt;@andrew_n_carr&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Claims that Inkling was “rushed” ahead of Chinese releases are speculation from critics, not evidenced by the launch materials &lt;a href=&quot;https://x.com/scaling01/status/2077465762869194973&quot;&gt;@scaling01&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Claims that relative attention gives TML a finetuning moat because backward is hard are speculative &lt;a href=&quot;https://x.com/typedfemale/status/2077523313484832791&quot;&gt;@typedfemale&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Claims that Inkling avoids multimodal intelligence loss are promising but not yet benchmark-complete in the tweet set &lt;a href=&quot;https://x.com/_alex_kirillov_/status/2077493564066722248&quot;&gt;@&lt;em&gt;alex_kirillov&lt;/em&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Different perspectives&lt;/h2&gt;
&lt;h3&gt;Supportive / bullish&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Open-weight and permissive license as strategic win:&lt;/strong&gt; Many saw the Apache-2.0 release as a major boost to the U.S./Western open ecosystem &lt;a href=&quot;https://x.com/latkins/status/2077463764979581213&quot;&gt;@latkins&lt;/a&gt;, &lt;a href=&quot;https://x.com/saranormous/status/2077469313108422806&quot;&gt;@saranormous&lt;/a&gt;, &lt;a href=&quot;https://x.com/brexton/status/2077462491819302918&quot;&gt;@brexton&lt;/a&gt;, &lt;a href=&quot;https://x.com/hyperindexed/status/2077471981264396411&quot;&gt;@hyperindexed&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Customization over leaderboard chasing:&lt;/strong&gt; Researchers and builders praised the explicit framing that Inkling is a broad, tunable foundation rather than a benchmark-maxed point solution &lt;a href=&quot;https://x.com/gneubig/status/2077468189672210472&quot;&gt;@gneubig&lt;/a&gt;, &lt;a href=&quot;https://x.com/ben_burtenshaw/status/2077470911448387633&quot;&gt;@ben_burtenshaw&lt;/a&gt;, &lt;a href=&quot;https://x.com/thealexker/status/2077540344757928445&quot;&gt;@thealexker&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Strong release quality:&lt;/strong&gt; Several users praised the transparency, grounded tone, and comprehensive technical documentation &lt;a href=&quot;https://x.com/lvwerra/status/2077487456270586319&quot;&gt;@lvwerra&lt;/a&gt;, &lt;a href=&quot;https://x.com/saranormous/status/2077483301212963157&quot;&gt;@saranormous&lt;/a&gt;, &lt;a href=&quot;https://x.com/rasbt/status/2077540575255880126&quot;&gt;@rasbt&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Architecture interest:&lt;/strong&gt; The non-RoPE positional choice and scaled short-conv usage drew positive attention as evidence TML is willing to make meaningful architecture bets &lt;a href=&quot;https://x.com/stochasticchasm/status/2077463965438009677&quot;&gt;@stochasticchasm&lt;/a&gt;, &lt;a href=&quot;https://x.com/rasbt/status/2077540575255880126&quot;&gt;@rasbt&lt;/a&gt;, &lt;a href=&quot;https://x.com/ChangJonathanC/status/2077508340637139318&quot;&gt;@ChangJonathanC&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Neutral / analytical&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Strong but not top overall:&lt;/strong&gt; The most balanced reads place Inkling as the new U.S. open-weight leader, but behind GLM/Kimi/DeepSeek or top closed models on some fronts &lt;a href=&quot;https://x.com/natolambert/status/2077454404433903816&quot;&gt;@natolambert&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2077466590346444939&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/stochasticchasm/status/2077463420182712708&quot;&gt;@stochasticchasm&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Good base model thesis:&lt;/strong&gt; Multiple analysts read the release as a systems/business move: ship a solid, efficient, post-trainable base and let Tinker plus downstream RL/fine-tuning create differentiation &lt;a href=&quot;https://x.com/ben_burtenshaw/status/2077470911448387633&quot;&gt;@ben_burtenshaw&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2077472478499053846&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/tinkerapi/status/2077467634568929433&quot;&gt;@tinkerapi&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Critical / skeptical&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Not frontier overall:&lt;/strong&gt; Critics argued it is still clearly behind top Chinese open-weight models and the strongest closed models &lt;a href=&quot;https://x.com/scaling01/status/2077465762869194973&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/JJitsev/status/2077627999352922196&quot;&gt;@JJitsev&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Purity claims overstated:&lt;/strong&gt; Some pushback focused on exaggerated claims that it is uniquely “pure” or non-distilled; the thread set includes both hype and corrections &lt;a href=&quot;https://x.com/jxmnop/status/2077504236380946595&quot;&gt;@jxmnop&lt;/a&gt;, &lt;a href=&quot;https://x.com/jxmnop/status/2077540390128034133&quot;&gt;@jxmnop&lt;/a&gt;, &lt;a href=&quot;https://x.com/andrew_n_carr/status/2077509786237854136&quot;&gt;@andrew_n_carr&lt;/a&gt;, &lt;a href=&quot;https://x.com/JJitsev/status/2077627999352922196&quot;&gt;@JJitsev&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmark middlingness as concern:&lt;/strong&gt; Some readers saw the moderate benchmark profile as evidence it may simply lag current Chinese open frontier rather than inaugurate a new frontier &lt;a href=&quot;https://x.com/scaling01/status/2077465762869194973&quot;&gt;@scaling01&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Context: why this matters&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;First major TML public model:&lt;/strong&gt; This is the first true external model release from Thinking Machines after months of anticipation around a lab staffed by ex-OpenAI leaders and researchers. That made the choice of &lt;strong&gt;open weights&lt;/strong&gt; itself notable &lt;a href=&quot;https://x.com/Hesamation/status/2077456283528045001&quot;&gt;@Hesamation&lt;/a&gt;, &lt;a href=&quot;https://x.com/TechCrunch/status/2077454757283959123&quot;&gt;@TechCrunch&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A U.S. open-weight answer to Chinese momentum:&lt;/strong&gt; Many reactions explicitly compare Inkling to GLM, Kimi, DeepSeek, and Qwen. The release lands amid concern that Western open-weight models have trailed Chinese ones on capability and release cadence &lt;a href=&quot;https://x.com/scaling01/status/2077474933370761345&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/teortaxesTex/status/2077457960385585281&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/sriramk/status/2077566845431779766&quot;&gt;@sriramk&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open base + post-training stack thesis:&lt;/strong&gt; TML’s messaging strongly suggests a strategy similar to “ship a competent open substrate, then differentiate via customization/fine-tuning/RL infrastructure.” That aligns with Tinker distribution and with user reactions centering controllable reasoning, concise outputs, and adaptation rather than raw leaderboard supremacy &lt;a href=&quot;https://x.com/thinkymachines/status/2077454609551921208&quot;&gt;@thinkymachines&lt;/a&gt;, &lt;a href=&quot;https://x.com/MichaelElabd/status/2077461111247712656&quot;&gt;@MichaelElabd&lt;/a&gt;, &lt;a href=&quot;https://x.com/ben_burtenshaw/status/2077470911448387633&quot;&gt;@ben_burtenshaw&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference ecosystem maturity:&lt;/strong&gt; The release also showcases how far open inference stacks have come. Day-0 support for a 1T-class multimodal MoE with new architectural components and multiple kernel-level optimizations would have been far less plausible a year earlier &lt;a href=&quot;https://x.com/vllm_project/status/2077459955117109343&quot;&gt;@vllm_project&lt;/a&gt;, &lt;a href=&quot;https://x.com/inferact/status/2077461431306584423&quot;&gt;@inferact&lt;/a&gt;, &lt;a href=&quot;https://x.com/LysandreJik/status/2077459011285512267&quot;&gt;@LysandreJik&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Architectural experimentation at scale:&lt;/strong&gt; Relative positional bias instead of RoPE and large-scale short-conv usage are the kind of choices researchers watch closely because they may indicate future architecture trends if they prove robust under scaling and post-training &lt;a href=&quot;https://x.com/stochasticchasm/status/2077463965438009677&quot;&gt;@stochasticchasm&lt;/a&gt;, &lt;a href=&quot;https://x.com/rasbt/status/2077540575255880126&quot;&gt;@rasbt&lt;/a&gt;, &lt;a href=&quot;https://x.com/ChangJonathanC/status/2077508340637139318&quot;&gt;@ChangJonathanC&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Release style as signal:&lt;/strong&gt; Several commentators praised the unusually restrained release language, explicit admission that it is not the strongest overall model, and detailed technical notes. For expert audiences, that improved credibility relative to more benchmark-maxed launches &lt;a href=&quot;https://x.com/eliebakouch/status/2077463243463721085&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/lvwerra/status/2077487456270586319&quot;&gt;@lvwerra&lt;/a&gt;, &lt;a href=&quot;https://x.com/thealexker/status/2077540344757928445&quot;&gt;@thealexker&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agents, Sandboxes, and Harness Engineering&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Perplexity’s SPACE sandbox platform&lt;/strong&gt;: &lt;a href=&quot;https://x.com/perplexity_ai/status/2077432518081744979&quot;&gt;@perplexity_ai&lt;/a&gt; introduced &lt;strong&gt;SPACE&lt;/strong&gt;, its in-house sandbox platform now serving &lt;strong&gt;100% of Computer production traffic&lt;/strong&gt;. The interesting systems design choice is the decoupling of &lt;strong&gt;session state&lt;/strong&gt; from disposable &lt;strong&gt;Firecracker microVM sandboxes&lt;/strong&gt;, with rolling snapshots allowing pause/resume/branch semantics. &lt;a href=&quot;https://x.com/perplexity_ai/status/2077432569432514977&quot;&gt;@perplexity_ai&lt;/a&gt; reported median sandbox creation latency dropping from &lt;strong&gt;185 ms to 60 ms&lt;/strong&gt; and P90 from &lt;strong&gt;447 ms to 89 ms&lt;/strong&gt;, while &lt;a href=&quot;https://x.com/zbraniecki/status/2077451060927672647&quot;&gt;@zbraniecki&lt;/a&gt; explained the use of &lt;strong&gt;disk snapshots plus full VM checkpoints&lt;/strong&gt;, object storage for resumability, and &lt;strong&gt;Btrfs COW&lt;/strong&gt; to make sandbox creation a metadata operation instead of full image copy. This is one of the more concrete production infrastructure disclosures in the set.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent workspaces, Slack-native agents, and harness cost&lt;/strong&gt;: On the product side, &lt;a href=&quot;https://x.com/istdrc/status/2077376131628707907&quot;&gt;@istdrc&lt;/a&gt; launched &lt;strong&gt;Raft 1.0&lt;/strong&gt;, positioning it as a shared workspace where agents behave more like a team in a messaging app than isolated terminal sessions. &lt;a href=&quot;https://x.com/LangChain/status/2077437626965971059&quot;&gt;@LangChain&lt;/a&gt; upgraded &lt;strong&gt;Fleet in Slack&lt;/strong&gt;, enabling one-click deployment of agents into channels/threads with custom identity and file handoffs, echoed by &lt;a href=&quot;https://x.com/hwchase17/status/2077443161585287290&quot;&gt;@hwchase17&lt;/a&gt;. On the engineering side, &lt;a href=&quot;https://x.com/AI21Labs/status/2077399596439925073&quot;&gt;@AI21Labs&lt;/a&gt; argued that &lt;strong&gt;harness design, not just model choice&lt;/strong&gt;, materially affects cost: it cited Writer’s “Harness Effect” as showing &lt;strong&gt;41% lower cost per task&lt;/strong&gt; at quality parity when only orchestration changed, and linked to its own early-stopping work claiming &lt;strong&gt;up to 44% compute reduction&lt;/strong&gt; for SWE agents. Related toolchain notes came from &lt;a href=&quot;https://x.com/nutlope/status/2077432463685554558&quot;&gt;@nutlope&lt;/a&gt;, who launched &lt;strong&gt;TogetherLink&lt;/strong&gt; to run open-source models inside coding harnesses like Codex and Claude Code, and &lt;a href=&quot;https://x.com/Teknium/status/2077424392892731396&quot;&gt;@Teknium&lt;/a&gt;, who added &lt;strong&gt;Blender MCP&lt;/strong&gt; to the Hermes agent catalog.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Automated Red Teaming, Alignment, and Governance Friction&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI’s GPT-Red and the safety flywheel&lt;/strong&gt;: &lt;a href=&quot;https://x.com/OpenAI/status/2077446718728425686&quot;&gt;@OpenAI&lt;/a&gt; introduced &lt;strong&gt;GPT-Red&lt;/strong&gt;, an internal automated red teamer for finding &lt;strong&gt;prompt injection vulnerabilities at scale&lt;/strong&gt;. The most concrete claim was that adversarial training against GPT-Red made &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; substantially more robust, with &lt;a href=&quot;https://x.com/OpenAI/status/2077446722683650525&quot;&gt;OpenAI&lt;/a&gt; saying replayed strong attacks produced &lt;strong&gt;6× fewer failures&lt;/strong&gt; than its best production model from four months earlier. The broader framing—AI systems improving the safety of future AI systems—was made explicit in &lt;a href=&quot;https://x.com/OpenAI/status/2077446723992228167&quot;&gt;OpenAI’s follow-up&lt;/a&gt;. This sits alongside external commentary from &lt;a href=&quot;https://x.com/omarsar0/status/2077450923295506505&quot;&gt;@omarsar0&lt;/a&gt;, who called it a high-ROI self-improvement loop.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anthropic’s misalignment scenarios and DeepMind governance debate&lt;/strong&gt;: &lt;a href=&quot;https://x.com/AnthropicAI/status/2077452646303006927&quot;&gt;@AnthropicAI&lt;/a&gt; published &lt;strong&gt;“Agentic misalignment in Summer 2026”&lt;/strong&gt;, adding four new simulated scenarios of bad autonomous-agent behavior a year after its blackmail case studies. Concurrently, governance discussion intensified around Google DeepMind: &lt;a href=&quot;https://x.com/Turn_Trout/status/2077448610157891734&quot;&gt;@Turn_Trout&lt;/a&gt; announced he resigned from DeepMind over military use without restrictions against killer robots or mass surveillance, while &lt;a href=&quot;https://x.com/jackclarkSF/status/2077419516452065406&quot;&gt;@jackclarkSF&lt;/a&gt; and &lt;a href=&quot;https://x.com/Yoshua_Bengio/status/2077487556325732745&quot;&gt;@Yoshua_Bengio&lt;/a&gt; amplified Demis Hassabis’s call for &lt;strong&gt;third-party testing and standards&lt;/strong&gt; feeding into policy. The juxtaposition—public support for standards versus internal dissatisfaction over governance practice—was succinctly noted by &lt;a href=&quot;https://x.com/BlackHC/status/2077511235763884426&quot;&gt;@BlackHC&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Benchmarks, Reproducibility, and Evaluation Integrity&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Soofi S / Nemotron contamination dispute&lt;/strong&gt;: The sharpest evaluation controversy concerned claims around &lt;strong&gt;Soofi S 30B-A3B&lt;/strong&gt;. &lt;a href=&quot;https://x.com/kimmonismus/status/2077382976577343913&quot;&gt;@kimmonismus&lt;/a&gt; presented it as a Europe-trained model based on NVIDIA’s open &lt;strong&gt;Nemotron 3 Nano&lt;/strong&gt; architecture, trained on &lt;strong&gt;~27T tokens&lt;/strong&gt; with German upweighting and a fully released recipe. But multiple critics challenged both novelty and eval integrity. &lt;a href=&quot;https://x.com/JJitsev/status/2077273171963588804&quot;&gt;@JJitsev&lt;/a&gt; argued the comparison lowered Nemotron reference scores versus the original report, while &lt;a href=&quot;https://x.com/eliebakouch/status/2077425801633427919&quot;&gt;@eliebakouch&lt;/a&gt; alleged the training mix included &lt;strong&gt;light rephrasings of the GPQA Diamond eval set&lt;/strong&gt;, potentially contaminating the benchmark and overstating the gap. He later summarized the concern as “&lt;strong&gt;10 epochs of a very light rephrasing of every GPQA Diamond eval item&lt;/strong&gt;” in &lt;a href=&quot;https://x.com/eliebakouch/status/2077428860639973471&quot;&gt;a follow-up&lt;/a&gt;. Even skeptics suggested a straightforward remedy: rerun the original Nemotron recipe and then compare under identical evaluation conditions, as &lt;a href=&quot;https://x.com/JJitsev/status/2077395737109725271&quot;&gt;@JJitsev&lt;/a&gt; proposed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Broader movement toward better evals and reproducibility&lt;/strong&gt;: Several posts pushed on eval methodology more generally. &lt;a href=&quot;https://x.com/arena/status/2077432293023678685&quot;&gt;@arena&lt;/a&gt; launched a &lt;strong&gt;factuality-weighted ranking&lt;/strong&gt; combining human preference with claim verification, based on &lt;strong&gt;2M+ labeled claims&lt;/strong&gt; across text and search arenas; GPT-5.5 reportedly gained the most under the factuality weighting while some preference-optimized models dropped. &lt;a href=&quot;https://x.com/askalphaxiv/status/2077415909652901993&quot;&gt;@askalphaxiv&lt;/a&gt; and &lt;a href=&quot;https://x.com/abidlabs/status/2077518437161521533&quot;&gt;@abidlabs&lt;/a&gt; kicked off a Hugging Face-backed reproducibility challenge around &lt;strong&gt;ICML 2026&lt;/strong&gt; papers, with early community progress already reproducing dozens of papers. And &lt;a href=&quot;https://x.com/sayashk/status/2077420320172941683&quot;&gt;@sayashk&lt;/a&gt; announced a PhD talk explicitly titled &lt;strong&gt;“The Missing Science of AI Evaluation.”&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Inkling launch&lt;/strong&gt;: &lt;a href=&quot;https://x.com/thinkymachines/status/2077454609551921208&quot;&gt;@thinkymachines announcing Inkling&lt;/a&gt; was the clear technical story of the day, combining an open-weight &lt;strong&gt;~1T multimodal MoE&lt;/strong&gt; release with a broad open inference rollout.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI’s GPT-Red&lt;/strong&gt;: &lt;a href=&quot;https://x.com/OpenAI/status/2077446718728425686&quot;&gt;@OpenAI’s GPT-Red announcement&lt;/a&gt; stood out for substance: automated prompt-injection red teaming plus a concrete &lt;strong&gt;6× robustness improvement&lt;/strong&gt; claim on GPT-5.6 Sol.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Code artifacts + MCP&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2077489907350856038&quot;&gt;@ClaudeDevs&lt;/a&gt; launched &lt;strong&gt;artifacts that can call MCP connectors&lt;/strong&gt;, effectively making artifacts per-viewer live apps/dashboards with permission-scoped data access.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Perplexity SPACE&lt;/strong&gt;: &lt;a href=&quot;https://x.com/perplexity_ai/status/2077432518081744979&quot;&gt;@perplexity_ai&lt;/a&gt; and &lt;a href=&quot;https://x.com/AravSrinivas/status/2077439693420163352&quot;&gt;@AravSrinivas&lt;/a&gt; provided unusually detailed production numbers for agent sandboxes, including &lt;strong&gt;5× faster tail latency&lt;/strong&gt; claims.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DeepMind governance resignation&lt;/strong&gt;: &lt;a href=&quot;https://x.com/Turn_Trout/status/2077448610157891734&quot;&gt;@Turn_Trout’s resignation thread&lt;/a&gt; was one of the most engaged policy-adjacent AI posts, centered on military-use restrictions and lab governance credibility.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Bonsai 27B and Local Inference Speedups&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwfva9/bonsai_27b_1bit_dense_llm_running_locally_in_your/&quot;&gt;Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels&lt;/a&gt;&lt;/strong&gt; (Activity: 731): &lt;strong&gt;&lt;strong&gt;PrismML&lt;/strong&gt; released &lt;strong&gt;Bonsai 27B&lt;/strong&gt;, a &lt;strong&gt;1-bit dense LLM&lt;/strong&gt; intended to run locally in-browser via custom &lt;strong&gt;WebGPU kernels&lt;/strong&gt;, with model artifacts on &lt;a href=&quot;https://huggingface.co/collections/prism-ml/bonsai-27b&quot;&gt;Hugging Face&lt;/a&gt; and a &lt;a href=&quot;https://huggingface.co/spaces/webml-community/bonsai-webgpu-kernels&quot;&gt;WebGPU demo Space&lt;/a&gt;. The claimed compression is from roughly &lt;code&gt;54GB&lt;/code&gt; to &lt;code&gt;3.8GB&lt;/code&gt; (&lt;code&gt;-93%&lt;/code&gt;) while retaining about &lt;code&gt;90%&lt;/code&gt; of baseline capability; commenters note a &lt;code&gt;~5.7GB&lt;/code&gt; footprint for a Qwen/Qwen3-derived &lt;code&gt;27B&lt;/code&gt; variant and interest in testing on consumer GPUs such as an &lt;code&gt;8GB&lt;/code&gt; RTX 3070 laptop GPU.&lt;/strong&gt; Commenters are broadly positive but focus on scaling questions: whether 1-bit quantization can make &lt;code&gt;80–100B&lt;/code&gt; parameter models practical, and what parameter/context-length tradeoff would fit &lt;code&gt;256k+&lt;/code&gt; context on a single &lt;code&gt;24GB&lt;/code&gt; GPU. There is also a recurring view that recent 1-bit releases suggest a broader shift toward ultra-low-bit LLM deployment.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted the headline compression claim: &lt;strong&gt;Bonsai 27B / Qwen 3.6 27B-class model&lt;/strong&gt; reportedly fits in about &lt;code&gt;5.7GB&lt;/code&gt; at &lt;strong&gt;1-bit density&lt;/strong&gt; with only ~&lt;code&gt;5%&lt;/code&gt; capability loss. One user specifically planned to test it on an &lt;code&gt;8GB&lt;/code&gt; laptop RTX 3070, implying the main practical interest is whether custom &lt;strong&gt;WebGPU kernels&lt;/strong&gt; can make a 27B dense model usable on consumer VRAM budgets.&lt;/li&gt;
&lt;li&gt;A technically focused thread discussed scaling: users want an &lt;code&gt;80B–100B&lt;/code&gt; 1-bit dense model, but noted the real constraint is fitting both weights and a &lt;code&gt;256k+&lt;/code&gt; context KV/cache footprint on a single &lt;code&gt;24GB&lt;/code&gt; GPU. This frames 1-bit weights as only part of the memory story; long-context inference may dominate usable deployment limits even if model weights are highly compressed.&lt;/li&gt;
&lt;li&gt;One commenter distinguished &lt;strong&gt;1-bit models trained from scratch&lt;/strong&gt; from extreme post-training quantization, arguing the former should retain much more capability than simply quantizing a larger model down to 1 bit. They suggested a future &lt;strong&gt;1-bit 70B&lt;/strong&gt; trained natively at that precision could be both consumer-GPU runnable and practically useful, unlike many heavily quantized ultra-low-bit models.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwehzt/prismmls_new_ternary_qwen36_27b_runs_near_fp16/&quot;&gt;PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!&lt;/a&gt;&lt;/strong&gt; (Activity: 465): &lt;strong&gt;&lt;strong&gt;PrismML&lt;/strong&gt; released &lt;strong&gt;&lt;a href=&quot;https://prismml.com/news/bonsai-27b&quot;&gt;Bonsai 27B&lt;/a&gt;&lt;/strong&gt;, a ternary/BitNet-style variant of &lt;strong&gt;Qwen3.6 27B&lt;/strong&gt;, with &lt;strong&gt;&lt;a href=&quot;https://huggingface.co/collections/prism-ml/bonsai-27b&quot;&gt;GGUF&lt;/a&gt;&lt;/strong&gt; and MLX builds requiring PrismML forks of &lt;strong&gt;&lt;a href=&quot;https://github.com/PrismML-Eng/llama.cpp&quot;&gt;llama.cpp&lt;/a&gt;&lt;/strong&gt; / &lt;strong&gt;&lt;a href=&quot;https://github.com/PrismML-Eng/mlx&quot;&gt;mlx&lt;/a&gt;&lt;/strong&gt; for now. The OP reports ~&lt;code&gt;10GB&lt;/code&gt; memory use at &lt;code&gt;32K&lt;/code&gt; context on an M4 Pro and claims the model is much better than a conventional 2-bit quant, but later edits clarify it is &lt;strong&gt;better than Q2, worse than Q4_K_XL&lt;/strong&gt;, with observed hallucinations and tool-calling loops; the main value is memory footprint rather than fp16-equivalent accuracy. Claimed capabilities include &lt;code&gt;256K&lt;/code&gt; context and multimodal input, with &lt;strong&gt;&lt;a href=&quot;https://github.com/z-lab/dflash&quot;&gt;dFlash&lt;/a&gt;&lt;/strong&gt; support mentioned as forthcoming; the whitepaper is &lt;a href=&quot;https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-27b-whitepaper.pdf&quot;&gt;here&lt;/a&gt;.&lt;/strong&gt; Commenters pushed back on the wording &lt;em&gt;“near fp16 precision”&lt;/em&gt; because ternary weights are &lt;code&gt;{-1,0,1}&lt;/code&gt;, and criticized hype/AGI framing as technically misleading. One technical question raised was whether the same ternary approach could scale to larger models such as GLM 5.2 while preserving acceptable quality loss.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter challenged the claim that a &lt;strong&gt;ternary / 1-trit Qwen3.6 27B&lt;/strong&gt; can run at &lt;em&gt;“near fp16 precision”&lt;/em&gt;, noting that PrismML’s premise is extremely low-bit representation and asking what that phrase technically means in this context. They also objected to loose use of &lt;em&gt;AGI&lt;/em&gt;, arguing that even stronger models like &lt;strong&gt;Mythos&lt;/strong&gt; should not be labeled that way without evidence.&lt;/li&gt;
&lt;li&gt;One user reported a working local run of &lt;strong&gt;&lt;code&gt;Ternary-Bonsai-27B-Q2_0.gguf&lt;/code&gt;&lt;/strong&gt; via a PrismML fork of &lt;code&gt;llama.cpp&lt;/code&gt;, using &lt;code&gt;-ngl 99&lt;/code&gt; and a very large &lt;code&gt;-c 200000&lt;/code&gt; context. On a simple TypeScript explanation prompt, they observed &lt;strong&gt;&lt;code&gt;243.5 t/s&lt;/code&gt; prompt processing&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;89.1 t/s&lt;/code&gt; generation&lt;/strong&gt;, suggesting the model is at least operational and fast under their setup, though the example does not validate quality claims.&lt;/li&gt;
&lt;li&gt;Another commenter asked for rigorous benchmarks such as &lt;strong&gt;SciCode&lt;/strong&gt; or &lt;strong&gt;SWE-rebench&lt;/strong&gt; to support the claims that the ternary model is close to BF16 and &lt;em&gt;“far more intelligent”&lt;/em&gt; than comparable &lt;strong&gt;2–3 bit Qwen3.6 27B&lt;/strong&gt; quantizations. They also asked whether the method works for &lt;strong&gt;MoEs&lt;/strong&gt; or larger dense models like &lt;strong&gt;Mistral Medium 3.5&lt;/strong&gt;, noting that very large MoEs with &lt;code&gt;&gt;10B&lt;/code&gt; active parameters, such as &lt;strong&gt;Step 3.7 Flash&lt;/strong&gt; and &lt;strong&gt;MiMo V2.5&lt;/strong&gt;, often appear unusually resistant to quantization.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ux4cn2/apple_in_talks_with_startup_prismml_that_shrinks/&quot;&gt;Apple in talks with startup PrismML that shrinks AI models to run on an iPhone&lt;/a&gt;&lt;/strong&gt; (Activity: 362): &lt;strong&gt;&lt;a href=&quot;https://www.cnbc.com/2026/07/14/apple-prismml-ai-compression-iphone.html&quot;&gt;CNBC reports&lt;/a&gt; &lt;strong&gt;Apple&lt;/strong&gt; is in early talks with &lt;strong&gt;PrismML&lt;/strong&gt;, a Caltech spinout, over extreme LLM compression for on-device iPhone inference. PrismML reportedly compresses &lt;strong&gt;Alibaba Qwen &lt;code&gt;27B&lt;/code&gt;&lt;/strong&gt; from ~&lt;code&gt;54 GB&lt;/code&gt; to &lt;strong&gt;&amp;#x3C;&lt;code&gt;4 GB&lt;/code&gt;&lt;/strong&gt; via ternary/binary-style quantization, claiming &lt;code&gt;10–15×&lt;/code&gt; lower memory, &lt;code&gt;6–8×&lt;/code&gt; faster responses, and &lt;code&gt;3–6×&lt;/code&gt; lower energy use, with some degradation in factual recall. Technical commenters noted missing details around conversion cost, convergence guarantees when moving to ternary/binary weights, whether distillation is used, saturation of post-training, and comparisons to prior low-bit work such as &lt;strong&gt;BitCPM-CANN&lt;/strong&gt; and &lt;code&gt;bitnet-b1.58-2B-4T&lt;/code&gt;.&lt;/strong&gt; Commenters were skeptical that small size alone proves usefulness, with one asking whether the compressed model is &lt;em&gt;“actually capable of anything useful.”&lt;/em&gt; Another user expressed disappointment because they had just been testing PrismML’s &lt;code&gt;q1 bonsai&lt;/code&gt; model, implying concern about Apple potentially acquiring or restricting the technology.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that PrismML’s public materials appear to lack key technical details needed to evaluate its compression/quantization claims: conversion cost, whether ternary/binary conversion reliably reaches a convergence point, whether distillation is used, and whether post-training has fully saturated the model. One comparison raised was the &lt;strong&gt;BitCPM-CANN&lt;/strong&gt; tech report, with commenters arguing that there is little precedent for a properly trained-from-scratch ternary model beyond &lt;strong&gt;&lt;code&gt;bitnet-b1.58-2B-4T&lt;/code&gt;&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Several comments questioned whether PrismML’s advertised small model size translates into practical capability. One user said they had been testing the company’s &lt;strong&gt;&lt;code&gt;q1 bonsai&lt;/code&gt;&lt;/strong&gt; model, while another noted they had seen repeated claims about compactness but little evidence that the model is &lt;em&gt;“actually capable of anything useful.”&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwylut/exllamav3_v100_major_performance_upgrades/&quot;&gt;ExLlamaV3 v1.0.0 - Major Performance Upgrades&lt;/a&gt;&lt;/strong&gt; (Activity: 422): &lt;strong&gt;The image is a &lt;strong&gt;technical benchmark table&lt;/strong&gt; (&lt;a href=&quot;https://i.redd.it/ej7102hqfcdh1.png&quot;&gt;PNG&lt;/a&gt;) for &lt;strong&gt;ExLlamaV3 v1.0.0&lt;/strong&gt;, showing RTX 3090 decode throughput gains versus &lt;code&gt;v0.0.43&lt;/code&gt; across multiple quantized LLMs/bitrates. It supports the release claims of major inference-kernel upgrades: &lt;code&gt;v1.0.0 mul1&lt;/code&gt; shows large speedups such as &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; rising from &lt;code&gt;29&lt;/code&gt; to &lt;code&gt;50 tok/s&lt;/code&gt; (&lt;code&gt;+72%&lt;/code&gt;) and &lt;strong&gt;Qwen 3.5 0.8B&lt;/strong&gt; from &lt;code&gt;268&lt;/code&gt; to &lt;code&gt;444 tok/s&lt;/code&gt; (&lt;code&gt;+66%&lt;/code&gt;), consistent with the post’s notes about new attention, GEMM/GEMV, INT8 GEMV, Conv1D, and MoE scheduler kernels.&lt;/strong&gt; Comments are mostly appreciative rather than deeply technical, highlighting that ExLlamaV3 is an NVIDIA-GPU-focused LLM engine using the EXL3 format rather than GGUF/llama.cpp, and praising the scale of the work by Turboderp/Fable.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter clarified that &lt;strong&gt;ExLlamaV3/ExLlama3&lt;/strong&gt; is an LLM inference engine targeting the specialized &lt;strong&gt;EXL3&lt;/strong&gt; model format rather than common &lt;strong&gt;GGUF&lt;/strong&gt; files used by &lt;code&gt;llama.cpp&lt;/code&gt;, and that it is currently &lt;strong&gt;NVIDIA GPU-only&lt;/strong&gt;. This distinction matters for deployment compatibility: users with existing GGUF workflows may need separate quantized weights and CUDA-capable hardware to use it.&lt;/li&gt;
&lt;li&gt;One technical request focused on &lt;strong&gt;TabbyAPI&lt;/strong&gt; integration, specifically improving tool-calling compatibility with &lt;strong&gt;Claude Code&lt;/strong&gt; so ExLlamaV3 could be used locally in that workflow. The same commenter noted they are currently using &lt;strong&gt;GGUF with MTP&lt;/strong&gt; and are interested in evaluating the quality tradeoffs of &lt;strong&gt;EXL3 quantization&lt;/strong&gt; versus their current setup.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Open-Weight Model Launches and Updates&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxdv34/thinking_machines_releases_first_openweight_model/&quot;&gt;Thinking Machines releases first open-weight model “Inkling”&lt;/a&gt;&lt;/strong&gt; (Activity: 1082): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/d7s0z8kqpfdh1.jpeg&quot;&gt;image&lt;/a&gt; is an AI model leaderboard highlighting &lt;strong&gt;Thinking Machines’ first open-weight model, Inkling&lt;/strong&gt;, scoring &lt;code&gt;1257&lt;/code&gt;, roughly mid-table and tied with &lt;strong&gt;Claude Opus 4.6&lt;/strong&gt;, below &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; and well below top entries like &lt;strong&gt;Claude Sonnet 5&lt;/strong&gt; at &lt;code&gt;1333&lt;/code&gt;. From the announcement/comments, Inkling is described as a &lt;strong&gt;MoE transformer&lt;/strong&gt; with &lt;code&gt;975B&lt;/code&gt; total parameters, &lt;code&gt;41B&lt;/code&gt; active parameters, up to a &lt;code&gt;1M&lt;/code&gt; token context window, and pretraining on &lt;code&gt;45T&lt;/code&gt; multimodal tokens spanning text, images, audio, and video; a preview &lt;strong&gt;Inkling-Small&lt;/strong&gt; has &lt;code&gt;12B&lt;/code&gt; active parameters for lower cost/latency.&lt;/strong&gt; Commenters were interested because Thinking Machines is associated with a former OpenAI CTO and is releasing open weights, but some were skeptical that Inkling will see broad adoption if it does not outperform competing open models such as &lt;strong&gt;GLM-5.2&lt;/strong&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Inkling&lt;/strong&gt; is described as a sparse MoE transformer with &lt;code&gt;975B&lt;/code&gt; total parameters, &lt;code&gt;41B&lt;/code&gt; active parameters, a &lt;code&gt;1M&lt;/code&gt; token context window, and pretraining on &lt;code&gt;45T&lt;/code&gt; tokens spanning text, images, audio, and video. One commenter compared it unfavorably to &lt;strong&gt;GLM-5.2&lt;/strong&gt;, arguing that despite being multimodal and long-context, it may see limited adoption if it does not outperform that open-weight competitor.&lt;/li&gt;
&lt;li&gt;The most technically interesting discussion centered on &lt;strong&gt;Inkling-Small&lt;/strong&gt;: a &lt;code&gt;276B&lt;/code&gt; parameter MoE with only &lt;code&gt;12B&lt;/code&gt; active parameters, positioned as a lower-latency/lower-cost variant. Commenters highlighted that it reportedly &lt;em&gt;“matches or exceeds its larger sibling on many benchmarks”&lt;/em&gt; due to improvements in the pretraining data mix and recipe, making it potentially practical for local inference compared with the &lt;code&gt;41B&lt;/code&gt; active main model: https://thinkingmachines.ai/news/introducing-inkling/#inkling-small&lt;/li&gt;
&lt;li&gt;A few commenters noted gaps in the model-size lineup: there is no model around the &lt;code&gt;30B&lt;/code&gt; dense/active-parameter class, which some local-inference users consider a useful middle ground. The release instead jumps from &lt;strong&gt;Inkling-Small&lt;/strong&gt; at &lt;code&gt;12B&lt;/code&gt; active MoE to the main &lt;strong&gt;Inkling&lt;/strong&gt; at &lt;code&gt;41B&lt;/code&gt; active MoE.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxao7y/german_ai_consortium_releases_soofi_s_an_open_30b/&quot;&gt;German AI consortium releases Soofi S, an open 30B model that tops benchmarks in both English and German&lt;/a&gt;&lt;/strong&gt; (Activity: 321): &lt;strong&gt;&lt;strong&gt;Soofi S&lt;/strong&gt; is presented as a German-led fully pretrained MoE LLM, &lt;strong&gt;&lt;code&gt;31.6B&lt;/code&gt; total / ~&lt;code&gt;3.2B&lt;/code&gt; active parameters&lt;/strong&gt;, based on &lt;strong&gt;NVIDIA Nemotron 3 Nano’s hybrid Mamba-2/Transformer architecture&lt;/strong&gt;, trained on ~&lt;code&gt;27T&lt;/code&gt; tokens with German upweighting and claimed near-flat throughput from &lt;code&gt;4K&lt;/code&gt; to &lt;code&gt;256K&lt;/code&gt; context (&lt;a href=&quot;https://the-decoder.com/german-ai-consortium-releases-soofi-s-an-open-30b-model-that-tops-benchmarks-in-both-english-and-german/&quot;&gt;article&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/abs/2607.09424&quot;&gt;paper&lt;/a&gt;). Commenters note it is described as a &lt;strong&gt;new full pretraining run, not a finetune&lt;/strong&gt;, with unusually transparent artifacts including &lt;a href=&quot;https://wandb.ai/soofi-exchange/pretrain-nemotron-3-nano-on-20T-4/reports/Soofi-S-Pretraining--VmlldzoxNzM4NTQ4NA?accessToken=c6mcvzhsloyc1v4duq9c7eq9aa81sr6b8j1l6yju6sbyz1skgecggj1pun9qxb52&quot;&gt;W&amp;#x26;B training logs&lt;/a&gt;, &lt;a href=&quot;https://github.com/soofi-project/Soofi-Pretraining&quot;&gt;training scripts&lt;/a&gt;, and gated &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Instruct-Preview-GGUF&quot;&gt;GGUF&lt;/a&gt; / &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Rhine-Preview-GGUF&quot;&gt;reasoning GGUF&lt;/a&gt; releases. Technical caveats raised include limited modern long-context evaluation beyond RULER, possible German naturalness issues from &lt;em&gt;“machine-translated and synthetically generated German texts”&lt;/em&gt;, and benchmark ambiguity because coding/math tasks and German understanding may inflate language-quality claims; one commenter also claims Qwen3.5 35B-A3B beats Soofi S on German benchmarks despite not being German-specialized.&lt;/strong&gt; The main debate is whether the benchmark comparison is credible: commenters criticized omission of newer baselines such as Qwen 3.6/Gemma 4 and comparison against older models. Another concern is licensing: marketing claims &lt;em&gt;“sovereign, open source… license-free availability”&lt;/em&gt; conflict with a Hugging Face card using a custom &lt;strong&gt;“Other”&lt;/strong&gt; license whose full text was reportedly missing.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter notes Soofi S is described as a &lt;strong&gt;newly pretrained model rather than a finetune&lt;/strong&gt;, with architecture based on &lt;strong&gt;Nemotron 3 Nano&lt;/strong&gt; and full pretraining/additional phases documented in the &lt;a href=&quot;https://arxiv.org/abs/2607.09424&quot;&gt;paper&lt;/a&gt;, &lt;a href=&quot;https://wandb.ai/soofi-exchange/pretrain-nemotron-3-nano-on-20T-4/reports/Soofi-S-Pretraining--VmlldzoxNzM4NTQ4NA?accessToken=c6mcvzhsloyc1v4duq9c7eq9aa81sr6b8j1l6yju6sbyz1skgecggj1pun9qxb52&quot;&gt;W&amp;#x26;B training logs&lt;/a&gt;, and &lt;a href=&quot;https://github.com/soofi-project/Soofi-Pretraining&quot;&gt;training scripts&lt;/a&gt;. They caution that the Nemotron-derived architecture may limit long-context accuracy: the release includes a &lt;strong&gt;RULER&lt;/strong&gt; test, but apparently lacks more modern long-context evaluations.&lt;/li&gt;
&lt;li&gt;Several commenters question the benchmark framing, arguing Soofi S was compared against older baselines and not newer models like &lt;strong&gt;Qwen 3.6&lt;/strong&gt; or &lt;strong&gt;Gemma 4&lt;/strong&gt;. One commenter highlights that the authors’ own results show &lt;strong&gt;Qwen3.5 35B-A3B&lt;/strong&gt; outperforming Soofi S on German, suggesting the benchmark may measure broad understanding, math, and coding more than native-quality German generation.&lt;/li&gt;
&lt;li&gt;The data mix is flagged as potentially problematic because the paper mentions &lt;em&gt;“machine-translated and synthetically generated German texts”&lt;/em&gt;, which can produce unnatural German and may affect generation quality despite benchmark gains. Release artifacts also appear incomplete or inconsistent: GGUF builds exist for &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Instruct-Preview-GGUF&quot;&gt;Soofi-S-Instruct-Preview&lt;/a&gt; and reasoning variants like &lt;a href=&quot;https://huggingface.co/Soofi-Project/Soofi-S-Rhine-Preview-GGUF&quot;&gt;Soofi-S-Rhine-Preview&lt;/a&gt;, but they are gated, and the license is described as custom/“Other” with the full text apparently missing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwbe7w/katcoderair_v25_open_model_soon/&quot;&gt;KAT-Coder-Air V2.5 - Open model soon&lt;/a&gt;&lt;/strong&gt; (Activity: 262): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/eob36vaek7dh1.png&quot;&gt;image&lt;/a&gt; is a screenshot of a &lt;strong&gt;KwaiAI/KAT-Coder&lt;/strong&gt; social post announcing &lt;strong&gt;KAT-Coder-Pro V2.5&lt;/strong&gt; and, more importantly for r/LocalLLaMA, a reply stating that &lt;strong&gt;KAT-Coder-Air V2.5 will be open-sourced “very soon.”&lt;/strong&gt; The post links to availability on &lt;strong&gt;OpenRouter&lt;/strong&gt; and a technical report on arXiv (&lt;a href=&quot;https://arxiv.org/abs/2607.05471&quot;&gt;abstract&lt;/a&gt;, &lt;a href=&quot;https://arxiv.org/pdf/2607.05471&quot;&gt;PDF&lt;/a&gt;), with claims around long-horizon and agentic coding performance; a commenter who tested it via OpenRouter speculates it is &lt;strong&gt;below &lt;code&gt;100B&lt;/code&gt; parameters&lt;/strong&gt;.&lt;/strong&gt; Comments are mostly curiosity-driven: users are waiting for actual open weights and asking about the model size, with one tester noting it is already usable through OpenRouter but not yet confirming architecture or parameter count.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter who says they asked about the weights reports that &lt;strong&gt;KAT-Coder-Air V2.5 is already available on OpenRouter&lt;/strong&gt; and, based on their usage/metadata, &lt;em&gt;“should be below &lt;code&gt;100B&lt;/code&gt; parameters.”&lt;/em&gt; The main technical uncertainty raised in the thread is the model’s exact parameter count/size, which users are waiting to verify once weights are released.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uxfu4k/google_is_updating_gemma_4s_chat_templates/&quot;&gt;Google is updating Gemma 4&apos;s chat templates, bringing major fixes to tool calling and reducing &quot;laziness&quot;, and enabling Flash Attention 4 on Hopper GPUs, plus an interactive guide on how to work with and improve its vision!&lt;/a&gt;&lt;/strong&gt; (Activity: 513): &lt;strong&gt;&lt;strong&gt;Google Gemma&lt;/strong&gt; announced updates to &lt;strong&gt;Gemma 4&lt;/strong&gt; chat templates via &lt;a href=&quot;https://x.com/googlegemma/status/2077449152062247219&quot;&gt;X&lt;/a&gt;, targeting tool-calling correctness and reduced “laziness,” while also enabling &lt;strong&gt;Flash Attention 4 on Hopper GPUs&lt;/strong&gt; and publishing an interactive Gemma vision token-budget guide on &lt;a href=&quot;https://huggingface.co/spaces/google/gemma4_vision_token_budget&quot;&gt;Hugging Face Spaces&lt;/a&gt;. A commenter linked the relevant &lt;a href=&quot;https://huggingface.co/google/gemma-4-31B-it/commit/68abe48010cbe15293462fa11e901a60639a44e5&quot;&gt;&lt;code&gt;google/gemma-4-31B-it&lt;/code&gt; commit&lt;/a&gt;, which includes fixes for &lt;code&gt;null&lt;/code&gt; handling, reasoning/thinking preservation, turn-tag balancing, tool-response continuation, &lt;code&gt;add_generation_prompt&lt;/code&gt; regression, extra &lt;code&gt;&amp;#x3C;turn|&gt;&lt;/code&gt; emission, and tool-call-only turn closure; notably, &lt;code&gt;preserve_thinking&lt;/code&gt; is restored/defaulted and scoped around tool-call turns.&lt;/strong&gt; Commenters framed the update as addressing confusing prior behavior in Gemma 4 prompting/tool use, with one saying they had assumed the failures were user error. The most emphasized reaction was enthusiasm that &lt;strong&gt;&lt;code&gt;preserve_thinking&lt;/code&gt;&lt;/strong&gt; support was included.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter enumerated the Hugging Face commit-level fixes for &lt;strong&gt;Gemma 4 31B IT&lt;/strong&gt; chat templates, including null handling, reasoning preservation, turn-tag balance, input validation, restoration of model turns/thinking cues after tool responses, and fixes for extra &lt;code&gt;&amp;#x3C;turn|&gt;&lt;/code&gt; emission in assistant content + tool-call continuation paths. The linked commit list starts from &lt;a href=&quot;https://huggingface.co/google/gemma-4-31B-it/commit/68abe48010cbe15293462fa11e901a60639a44e5&quot;&gt;&lt;code&gt;68abe480&lt;/code&gt;&lt;/a&gt;, with notable fixes around &lt;code&gt;preserve_thinking&lt;/code&gt;, rendering the thinking channel independent of &lt;code&gt;tool_calls&lt;/code&gt;, and correcting tool-call-only turn closure.&lt;/li&gt;
&lt;li&gt;One technical takeaway from the thread is that &lt;strong&gt;tool-calling behavior appears to have been strongly affected by template-level serialization bugs&lt;/strong&gt;, especially around preserving reasoning/thinking channels and correctly reopening assistant generation after tool responses. However, another commenter reports that the perceived &quot;laziness&quot; remains present with the latest template, arguing it is likely a &lt;strong&gt;model behavior issue rather than a chat-template issue&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Open-Model Policy and Self-Hosting Risk&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uw9ucd/source_the_trump_administration_and_industry/&quot;&gt;Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models&lt;/a&gt;&lt;/strong&gt; (Activity: 573): &lt;strong&gt;A reported Trump administration/industry discussion would streamline U.S. releases of &lt;strong&gt;open models&lt;/strong&gt; whose capabilities are &lt;em&gt;equal to or below&lt;/em&gt; leading Chinese open models, motivated by concern over U.S. developers adopting Chinese local/open-weight models. The linked source could not be verified from the provided archive page because it returned a CAPTCHA/HTTP &lt;code&gt;429&lt;/code&gt; interstitial rather than article content.&lt;/strong&gt; Commenters argued that U.S. AI firms have weak incentives to release Chinese-competitive open models because strong local models could cannibalize paid API/SaaS revenue. Others were skeptical of claims about Chinese open models containing CCP-exploitable backdoors, noting that if such model-level backdoors were practical or detectable, U.S. labs would likely already dominate open-weight alternatives.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that &lt;strong&gt;banning capable open-weight models is technically unenforceable&lt;/strong&gt; once weights are mirrored globally: a single torrent tracker, private file share, or USB transfer can bypass restrictions. One commenter noted enforcement would likely require extreme downstream controls such as confiscating or restricting GPUs/workstations above certain VRAM thresholds, because inference can run locally once weights are obtained.&lt;/li&gt;
&lt;li&gt;A recurring technical policy argument was that if the US wants companies to avoid local Chinese models, the practical alternative is to release &lt;strong&gt;US open-weight models of equal or better capability&lt;/strong&gt; rather than restrict access. Commenters suggested this would need to match the quality trajectory of Chinese open models closely enough that local deployment users do not need to rely on foreign weights.&lt;/li&gt;
&lt;li&gt;One commenter specifically cited &lt;strong&gt;NVIDIA Nemotron&lt;/strong&gt; as a better model-release pattern because of its relatively transparent training data/process, while noting that &lt;strong&gt;Nemotron 3 Ultra&lt;/strong&gt; is “very good” but appears undertrained and therefore underperforms for its parameter size. Multiple commenters were skeptical of claims about intentional Chinese “backdoors” in open-weight models, arguing that if such hidden model-level backdoors were straightforward to weaponize, US labs would already be exploiting the same technique in open models.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwqgqs/some_of_yall_wonder_why_anyone_would_self_host_ai/&quot;&gt;Some of y&apos;all wonder why anyone would self host AI.  Would you accept the opinion of the CEO of Microsoft?&lt;/a&gt;&lt;/strong&gt; (Activity: 559): &lt;strong&gt;The post argues for &lt;strong&gt;self-hosting AI/LLMs&lt;/strong&gt; as a data-protection strategy, citing a &lt;a href=&quot;https://techcrunch.com/2026/07/13/satya-nadella-has-issued-a-shocking-warning-to-companies-using-ai/&quot;&gt;TechCrunch article&lt;/a&gt; quoting &lt;strong&gt;Microsoft CEO Satya Nadella&lt;/strong&gt; that enterprises may &lt;em&gt;“pay for intelligence twice”&lt;/em&gt;: once in fees and again by exposing proprietary business knowledge needed to make hosted AI useful. The technical concern is that API/SaaS model providers such as &lt;strong&gt;OpenAI&lt;/strong&gt; or &lt;strong&gt;Anthropic&lt;/strong&gt; could ingest sensitive prompts, documents, workflows, or RAG corpora and potentially derive competitive intelligence, making local inference or private deployment attractive for inventors, researchers, and enterprises handling IP-sensitive data.&lt;/strong&gt; Top comments were skeptical of Nadella’s framing, arguing it may be a sales pitch for &lt;strong&gt;Azure-hosted&lt;/strong&gt; AI rather than a neutral privacy warning. Others pointed to Microsoft’s own products—&lt;strong&gt;Copilot&lt;/strong&gt;, the OpenAI partnership, and screen-indexing features like Recall—as evidence that Microsoft has similar incentives and privacy risks.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters interpreted Satya Nadella’s self-hosting argument as less about pure decentralization and more about &lt;strong&gt;enterprise workloads moving onto Microsoft Azure&lt;/strong&gt;, i.e. companies “hosting” private models inside Microsoft’s cloud rather than running them fully on-prem.&lt;/li&gt;
&lt;li&gt;Several commenters tied the self-hosting/privacy discussion to Microsoft’s own AI products, especially &lt;strong&gt;Copilot&lt;/strong&gt; and the controversial Windows &lt;strong&gt;Recall&lt;/strong&gt; feature, which was criticized for continuously snapshotting user activity for later AI-powered indexing. The technical concern raised was that cloud-connected assistants create a large attack surface for sensitive business data, even when marketed as productivity tooling.&lt;/li&gt;
&lt;li&gt;A recurring technical concern was that enterprise AI vendors may use customer interactions or documents as training/evaluation data, with one commenter asking how labs continue acquiring &lt;code&gt;50–100T&lt;/code&gt; new tokens per year for model training. The point was framed as a reason businesses might prefer self-hosted or tightly controlled deployments to reduce IP leakage and data reuse risk.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1uweb90/this_is_why_we_need_local_models_and_opensource/&quot;&gt;This is why we need local models and opensource harnesses&lt;/a&gt;&lt;/strong&gt; (Activity: 379): &lt;strong&gt;The image (&lt;a href=&quot;https://i.redd.it/ebfetgbo68dh1.jpeg&quot;&gt;screenshot&lt;/a&gt;) shows a claim by &lt;strong&gt;International Cyber Digest&lt;/strong&gt; that &lt;strong&gt;xAI’s Grok Build CLI&lt;/strong&gt; uploaded entire Git repositories—including private code and unredacted secrets—to a &lt;strong&gt;Google Cloud bucket&lt;/strong&gt;, allegedly disabled later via a hidden server-side flag. The post uses this alleged leak to argue for &lt;strong&gt;local-first/open-weight models&lt;/strong&gt;, deterministic open-source agent harnesses, private VPC execution, and governance layers that can inspect/redact secrets before any third-party network egress.&lt;/strong&gt; Commenters were overwhelmingly skeptical of cloud-tethered coding agents, framing this as potential spyware/malware behavior and suggesting legal accountability if the exfiltration was intentional. One recurring sentiment was that users should expect poor data-handling practices from opaque vendor-controlled AI tooling.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter describes mitigating LLM-driven code/data exfiltration by using a &lt;strong&gt;self-hosted Git server blocked from the Internet&lt;/strong&gt;. They note this does not fully prevent leakage, but it changes the threat model: instead of a tool silently issuing direct upload/download jobs, any exfiltration would need to pass through the LLM interaction path, making it &lt;em&gt;more visible and slower&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;Another technically relevant takeaway is the preference for &lt;strong&gt;open-source agents plus open-weight models on self-controlled infrastructure&lt;/strong&gt;. The argument is that local execution and inspectable harnesses reduce reliance on opaque hosted services where telemetry, tool calls, or data-retention behavior may be difficult to audit.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Frontier AI Distillation and Safety Standards&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/1uwavzo/anthropic_just_told_the_us_senate_that_alibaba/&quot;&gt;Anthropic just told the US Senate that Alibaba ran 25,000 fake accounts and had 28.8 million conversations with Claude — not to use it, but to copy it&lt;/a&gt;&lt;/strong&gt; (Activity: 3524): &lt;strong&gt;The post claims &lt;strong&gt;Anthropic told the U.S. Senate&lt;/strong&gt; that &lt;strong&gt;Alibaba&lt;/strong&gt; used &lt;code&gt;25,000&lt;/code&gt; API accounts to run &lt;code&gt;28.8M&lt;/code&gt; Claude conversations over ~six weeks (Apr–Jun) to perform large-scale model distillation—extracting agentic reasoning/coding behavior via normal API access rather than hacking—and then use it to improve &lt;strong&gt;Qwen&lt;/strong&gt;. Anthropic is described as framing this as its largest “distillation attack” and arguing current law is ambiguous enough that it sought congressional action rather than litigation; the OP links a longer breakdown tying this to the &lt;strong&gt;Fable 5 export ban&lt;/strong&gt;: &lt;a href=&quot;https://youtu.be/g1d3yTR6E2Y&quot;&gt;YouTube&lt;/a&gt;.&lt;/strong&gt; Top comments debate whether this is meaningfully different from conventional reverse engineering—e.g., automakers or Samsung buying competitors’ products to study them—versus prohibited model extraction. Several commenters argue Anthropic’s complaint is hypocritical because frontier labs trained on large amounts of public/creative human output under fair-use theories, but object when their own model outputs are used as training data.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed &lt;strong&gt;model distillation via API outputs&lt;/strong&gt; as analogous to competitive reverse engineering: an automaker or &lt;strong&gt;Samsung&lt;/strong&gt; buying a rival product, studying behavior, and using the findings to improve its own system. The technical distinction raised is that AI distillation can automate this process at scale—here alleged as &lt;code&gt;25,000&lt;/code&gt; accounts and &lt;code&gt;28.8M&lt;/code&gt; Claude conversations—rather than relying on human engineers manually extracting design lessons.&lt;/li&gt;
&lt;li&gt;One technically relevant objection was attribution: because &lt;strong&gt;Alibaba&lt;/strong&gt; operates major cloud infrastructure, traffic originating from Alibaba-associated networks or accounts does not necessarily prove Alibaba corporate teams performed the alleged distillation. Commenters compared this to seeing suspicious requests from &lt;strong&gt;AWS&lt;/strong&gt; or &lt;strong&gt;Azure&lt;/strong&gt; IP ranges, which may indicate customer activity rather than action by &lt;strong&gt;Amazon&lt;/strong&gt; or &lt;strong&gt;Microsoft&lt;/strong&gt; themselves.&lt;/li&gt;
&lt;li&gt;A recurring implementation/security point was that if the interactions occurred through paid Claude access and within apparent API mechanics, Anthropic’s complaint may expose weaknesses in &lt;strong&gt;abuse detection, account-linking, rate limits, ToS enforcement, or anti-distillation controls&lt;/strong&gt;. One commenter summarized this as: &lt;em&gt;“they got paid and the API was used according to the TOS?”&lt;/em&gt;—suggesting the issue may be more about platform controls than unauthorized access.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uw40fb/demis_hassabis_shared_a_rare_essay_on_x_agi_is/&quot;&gt;Demis Hassabis shared a rare essay on X: AGI is few years away, we&apos;re in the singularity foothills, proposes US-led Frontier AI Standards Body with eventual mandatory safety testing&lt;/a&gt;&lt;/strong&gt; (Activity: 899): &lt;strong&gt;&lt;strong&gt;Demis Hassabis&lt;/strong&gt;’ essay argues AGI is plausibly &lt;em&gt;“a few years away”&lt;/em&gt; and frames the current period as the &lt;em&gt;“foothills of the singularity,”&lt;/em&gt; with potential impact on the order of &lt;strong&gt;&lt;code&gt;10×&lt;/code&gt; the Industrial Revolution at &lt;code&gt;10×&lt;/code&gt; the speed&lt;/strong&gt;. He proposes a &lt;strong&gt;US-led Frontier AI Standards Body&lt;/strong&gt;—analogous to &lt;strong&gt;FINRA&lt;/strong&gt;—to evaluate frontier models, initially via voluntary pre-release safety testing that could become mandatory for major “Frontier Labs,” applying to both open and closed models and focusing on risks such as cybersecurity, biology, and autonomous agents (&lt;a href=&quot;https://x.com/demishassabis/status/2076957440109625718&quot;&gt;essay on X&lt;/a&gt;).&lt;/strong&gt; Commenters were more trusting of Hassabis’ timelines than those from Altman/Dario/Musk, but skeptical that model/code regulation is enforceable—especially against open-source or Chinese frontier models. Others noted real-world cyber impacts already appear dual-use, improving defense while raising attacker sophistication, and questioned whether a US-led standards body is geopolitically credible given current US posture toward international institutions.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters questioned the feasibility of regulating frontier AI models or code, comparing it to the difficulty of regulating Bitcoin. One technical concern was that even if U.S. or Western labs submit to mandatory testing, &lt;strong&gt;China could continue releasing increasingly capable open-source models&lt;/strong&gt; that are cheaper and close enough to frontier performance to attract global users.&lt;/li&gt;
&lt;li&gt;A cybersecurity-related comment noted that AI is producing a dual-use effect in banking security: it has materially improved defensive capabilities while also increasing attacker sophistication. The commenter described a shift from mostly basic attacks toward more complex AI-assisted intrusion attempts, suggesting that frontier AI governance may need to account for rapidly improving cyber-offense capabilities.&lt;/li&gt;
&lt;li&gt;Several commenters challenged the proposed &lt;strong&gt;U.S.-led Frontier AI Standards Body&lt;/strong&gt;, arguing that international coordination failure is a major unresolved technical-policy risk. They highlighted problems such as China refusing to participate, labs secretly accelerating despite agreed slowdowns, and the essay’s limited treatment of power concentration, inequality, and job displacement as consequences of AGI deployment.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. AI Hardware Companions and Ambient Robots&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/OpenAI/comments/1uwkxbc/new_leak_openais_first_device_will_be_moveable/&quot;&gt;NEW LEAK: OpenAI’s First Device Will Be Moveable, Screenless Speaker Built as AI Companion&lt;/a&gt;&lt;/strong&gt; (Activity: 764): &lt;strong&gt;A reported leak claims &lt;strong&gt;OpenAI’s first hardware product&lt;/strong&gt; is a &lt;strong&gt;moveable, screenless smart-speaker-like AI companion&lt;/strong&gt; with onboard &lt;strong&gt;camera/sensors&lt;/strong&gt; for environmental context, designed around personality, humanlike interaction, and productivity rather than a traditional display UI. The Bloomberg source itself was not accessible in the provided scrape due to a bot-detection page, so the details remain unverified from the linked article content.&lt;/strong&gt; Top comments were skeptical, framing the device as essentially &lt;em&gt;“reinvented Alexa”&lt;/em&gt; with added surveillance concerns due to the camera/sensor stack. Several commenters joked about uncomfortable camera placement/use cases, reflecting distrust of an always-present AI companion in private spaces.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter highlighted a concrete assistive-tech use case: a &lt;strong&gt;moveable AI speaker with a camera-enabled variant&lt;/strong&gt; could materially help visually impaired users by providing environmental description, object recognition, and navigation-style assistance. They noted surprise that more consumer AI hardware is not explicitly targeting disability/accessibility markets, suggesting current product strategy may be limited by perceived market size rather than technical feasibility.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uwc0oi/keio_university_made_these_soft_heliumfilled/&quot;&gt;Keio University made these soft, helium-filled flying robots; they can follow you, wake you up, remind you of stuff, and even be your study buddy&lt;/a&gt;&lt;/strong&gt; (Activity: 1975): &lt;strong&gt;The post highlights &lt;strong&gt;Keio University&lt;/strong&gt; soft, helium-filled indoor flying robots—apparently blimp-like “airwhales”—intended for lightweight human-assistance/HRI tasks such as following a user, alarms/wake-up prompts, reminders, and acting as a study companion. The linked Reddit video could not be accessed because Reddit returned &lt;strong&gt;&lt;code&gt;403 Forbidden&lt;/code&gt;&lt;/strong&gt; (&lt;a href=&quot;https://v.redd.it/ekrrg0d3s7dh1&quot;&gt;v.redd.it&lt;/a&gt;); only the preview image is available (&lt;a href=&quot;https://preview.redd.it/jyxujpmnv7dh1.png?width=1200&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=5890407a504dc2e96774595dcbaaed6efaeb2dc0&quot;&gt;preview&lt;/a&gt;).&lt;/strong&gt; Top comments were mostly nontechnical: users joked that the robots might be amusing but not broadly useful, attractive to cats, and potentially problematic around hazards like automatic revolving doors.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. AI Cost Controls in Real Deployments&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uwa1mv/well_it_finally_happened_were_not_using_models/&quot;&gt;Well it finally happened: we’re not using models because of cost&lt;/a&gt;&lt;/strong&gt; (Activity: 1537): &lt;strong&gt;A Fortune 500 “AI-first” org that had broadly deployed &lt;strong&gt;GitHub Copilot&lt;/strong&gt; and &lt;strong&gt;Claude&lt;/strong&gt;, with year-long internal training/demos, is now reducing access—Claude removed, usage limited, demos/training stopped—primarily due to &lt;strong&gt;cost controls&lt;/strong&gt;, with architects recommending older/cheaper models. The flagship pilot—using agents to reverse-engineer a legacy application into an “as-written” business-rule specification for a rewrite—failed because agents repeatedly missed subtle code-path details, and day-to-day use showed reliability issues such as generated SQL attempting to &lt;code&gt;DROP&lt;/code&gt; constraints around insert/delete logic. A top technical comment attributes the budget shock partly to Copilot’s shift toward metered “premium request” / usage-based pricing (&lt;a href=&quot;https://docs.github.com/en/copilot/concepts/billing/copilot-requests&quot;&gt;GitHub Docs&lt;/a&gt;).&lt;/strong&gt; Commenters debate whether this is a structural blocker or a transient phase: one view is that AI is currently in a “sour spot” where capability is near-useful but not dependable enough for complex enterprise workflows while inference/subagent costs remain high; another asks whether the company would resume “AI-first” prioritization if costs fell.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed the issue as a pricing-model shift rather than a capability failure: &lt;strong&gt;GitHub Copilot moving to usage-based pricing&lt;/strong&gt; reportedly made large organizations re-evaluate AI spend because previously hidden or fixed costs became variable and attributable to heavy usage.&lt;/li&gt;
&lt;li&gt;A technical cost-performance theme was that current models are in a “sour spot”: capabilities are close to being broadly useful, but workflows with &lt;strong&gt;intense subagent use&lt;/strong&gt; can multiply inference calls and make deployments expensive. One commenter argued this may be temporary as model capability per dollar improves, eventually making AI use a financial “no-brainer” for many tasks.&lt;/li&gt;
&lt;li&gt;Another commenter argued that expensive frontier models do not imply AI is overhyped: organizations can stay on older, cheaper models, and high prices for new models suggest buyers perceive meaningful performance gains rather than “flat performance.” They compared this to historical compute trends where &lt;em&gt;“one year’s cutting edge is the next year’s bargain bin,”&lt;/em&gt; implying today’s premium models may quickly become commodity-priced as newer models arrive.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uw24bp/claude_spent_15_eur_of_a_2_eur_limit/&quot;&gt;Claude spent +15 EUR of a 2 EUR limit.&lt;/a&gt;&lt;/strong&gt; (Activity: 1861): &lt;strong&gt;The image shows a &lt;strong&gt;Claude “Usage credits” billing/limit screen&lt;/strong&gt; where a configured &lt;strong&gt;€2.00 monthly spend limit&lt;/strong&gt; was exceeded to &lt;strong&gt;€13.79&lt;/strong&gt;, displayed as &lt;strong&gt;&lt;code&gt;690% used&lt;/code&gt;&lt;/strong&gt;, with &lt;strong&gt;auto-reload off&lt;/strong&gt; and &lt;strong&gt;current balance €0.00&lt;/strong&gt; (&lt;a href=&quot;https://i.redd.it/iepu4woeh5dh1.png&quot;&gt;image&lt;/a&gt;). In context, the post alleges a single summarization request was allowed to complete and bill far beyond the user’s remaining credits/limit, implying the spend cap may function as a soft limit rather than a hard real-time cutoff.&lt;/strong&gt; Commenters report similar behavior, suggesting Claude may finish an in-progress turn and charge the full amount even after limits are exceeded. Some frame this as potentially unlawful or deceptive if refunds are denied, while others call it “scummy behavior.”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple users report &lt;strong&gt;Anthropic API spend limits/credit controls not acting as hard caps&lt;/strong&gt;: one user says a &lt;code&gt;£15&lt;/code&gt; balance was consumed despite usage credits being turned off, while another reports a &lt;code&gt;$1&lt;/code&gt; limit resulting in &lt;code&gt;$42&lt;/code&gt; of charges before detection. A commenter hypothesizes the system may &lt;em&gt;“always finish the turn”&lt;/em&gt; and bill the full in-flight request even after the configured limit is exceeded, implying spend limits may be enforced only after request completion rather than pre-authorizing against remaining budget.&lt;/li&gt;
&lt;li&gt;Users describe &lt;strong&gt;refund/support failures after limit overruns&lt;/strong&gt;, with one commenter claiming support refused to refund the &lt;code&gt;$41&lt;/code&gt; overage beyond a &lt;code&gt;$1&lt;/code&gt; limit. The discussion frames this as a risk for API-only workflows with fixed monthly allocations, e.g. budgeting &lt;code&gt;$200&lt;/code&gt; for “fable” usage but potentially receiving much larger actual charges if limits are soft rather than hard.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>thinking-machines-lab</category><category>huggingface</category><category>vllm_project</category><category>lmsysorg</category><category>modal</category><category>baseten</category><category>databricks</category><category>inkling</category><category>miramurati</category><category>soumithchintala</category><category>johnschulman2</category><category>lilianweng</category><category>natolambert</category><category>artificialanlys</category><category>scaling01</category><category>mixture-of-experts</category><category>multimodality</category><category>foundation-models</category><category>model-licensing</category><category>context-window</category><category>open-weights</category><category>model-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-14-not-much/</guid><description>**OpenAI&apos;s agent products** saw a **2.5x weekly usage growth** driven by **Codex + ChatGPT Work** and demand for **GPT-5.6 Sol**. JetBrains adopted Codex as a recommended agent, while LangChain enhanced tracing and observability across multiple tools. **PrismML released Bonsai 27B**, a compressed variant of **Qwen 3.6 27B** enabling local multimodal agentic workflows on consumer devices. Tencent Hunyuan introduced 1-bit and 4-bit quantized **Hy3 295B** model deployable on a single GPU. Quantization advances like **NVFP4 dynamic quants** for **Gemma-4** and others support serious local inference. OpenMOSS launched **MOSS-VL-Realtime 11B** for continuous video stream perception with a **256K context window**. *&quot;Harness quality and observability are becoming a first-class differentiator&quot;* and local inference is now viable for agentic workflows.</description><pubDate>Tue, 14 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/13/2026-7/14/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Coding Agents, Harnesses, and the Shift From Chat to Execution&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI’s agent products are seeing unusually strong pull&lt;/strong&gt;: &lt;a href=&quot;https://x.com/sama/status/2077033807736459713&quot;&gt;@sama&lt;/a&gt; said usage of &lt;strong&gt;Codex + ChatGPT Work&lt;/strong&gt; grew &lt;strong&gt;2.5x in a week&lt;/strong&gt;, later adding that GPT-5.6 Sol demand is “insane” and may cause scaling hiccups while infra catches up (&lt;a href=&quot;https://x.com/sama/status/2077106587307798989&quot;&gt;1&lt;/a&gt;, &lt;a href=&quot;https://x.com/sama/status/2077036999303999910&quot;&gt;2&lt;/a&gt;). The ecosystem response was immediate: &lt;a href=&quot;https://x.com/jetbrains/status/2076958455173095878&quot;&gt;JetBrains made Codex its recommended agent&lt;/a&gt;, &lt;a href=&quot;https://x.com/theo/status/2076890018032062483&quot;&gt;@theo highlighted Codex’s underexposed “question tool”&lt;/a&gt;, and OpenAI’s own team showed &lt;a href=&quot;https://x.com/OpenAIDevs/status/2077102893665320983&quot;&gt;command-line eval tooling built start-to-finish with GPT-5.6&lt;/a&gt;. Product-side, OpenAI also ran multiple &lt;strong&gt;usage resets&lt;/strong&gt;, amplified by &lt;a href=&quot;https://x.com/reach_vb/status/2077117109633466473&quot;&gt;@reach_vb&lt;/a&gt; and users like &lt;a href=&quot;https://x.com/kimmonismus/status/2077117385081860528&quot;&gt;@kimmonismus&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Harness quality and observability are becoming a first-class differentiator&lt;/strong&gt;: several tweets converged on the idea that model quality alone is no longer enough. &lt;a href=&quot;https://x.com/swyx/status/2077072402828361772&quot;&gt;@swyx warned&lt;/a&gt; that stale &lt;code&gt;agents.md&lt;/code&gt; instructions can act like &lt;strong&gt;self-inflicted prompt injection&lt;/strong&gt;, causing multi-hour stalls in long-running tasks. &lt;a href=&quot;https://x.com/LangChain/status/2077045458917052492&quot;&gt;LangChain added tracing for Codex&lt;/a&gt; and later expanded to &lt;a href=&quot;https://x.com/LangChain/status/2077076144248021236&quot;&gt;Cursor, Copilot, Pi, and OpenCode in LangSmith&lt;/a&gt;, exposing tool calls, subagents, and token usage. &lt;a href=&quot;https://x.com/Teknium/status/2077132644979200150&quot;&gt;@Teknium shipped Hermes updates&lt;/a&gt; to parallelize any subset of tool calls and previously exposed &lt;a href=&quot;https://x.com/Teknium/status/2077006948223090777&quot;&gt;banked resets directly in Hermes Agent&lt;/a&gt;. The meta-point was stated crisply by &lt;a href=&quot;https://x.com/andykonwinski/status/2077137640462467370&quot;&gt;@andykonwinski&lt;/a&gt;: companies that can encode their value into &lt;strong&gt;evals and environments&lt;/strong&gt; may gain a more durable edge than those relying on capital or raw scale alone.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open Models, Quantization, and Local Inference Compression&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Aggressive compression is bringing frontier-adjacent models onto consumer devices&lt;/strong&gt;: &lt;a href=&quot;https://x.com/PrismML/status/2077084891284721827&quot;&gt;PrismML released Bonsai 27B&lt;/a&gt;, based on &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt;, in two compact variants: &lt;strong&gt;Ternary Bonsai 27B&lt;/strong&gt; at &lt;strong&gt;5.9 GB / 1.71 effective bits&lt;/strong&gt; and &lt;strong&gt;1-bit Bonsai 27B&lt;/strong&gt; at &lt;strong&gt;3.9 GB / 1.125 effective bits&lt;/strong&gt;, both under &lt;strong&gt;Apache 2.0&lt;/strong&gt;. The claim is notable not just for size, but for preserving &lt;strong&gt;multimodal, tool-using, long-context agentic workflows&lt;/strong&gt; locally; &lt;a href=&quot;https://x.com/PrismML/status/2077084899904024918&quot;&gt;a demo shows Hermes running it on an RTX 5090&lt;/a&gt;, while &lt;a href=&quot;https://x.com/LocallyAIApp/status/2077087065628414133&quot;&gt;Locally AI highlighted phone deployment&lt;/a&gt;. In parallel, &lt;a href=&quot;https://x.com/TencentHunyuan/status/2076953120765280284&quot;&gt;Tencent Hunyuan released 1-bit and 4-bit Hy3&lt;/a&gt;, describing a &lt;strong&gt;295B flagship-scale model&lt;/strong&gt; that can be served on a &lt;strong&gt;single GPU&lt;/strong&gt; via llama.cpp with MTP enabled.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quantization and edge deployment continue to broaden the open-model operating envelope&lt;/strong&gt;: &lt;a href=&quot;https://x.com/danielhanchen/status/2077072556537020914&quot;&gt;@danielhanchen announced NVFP4 dynamic quants&lt;/a&gt; across the Gemma-4 family and additional large models including &lt;strong&gt;Qwen3.5-122B-A10B&lt;/strong&gt; and &lt;strong&gt;GLM-4.7-Flash&lt;/strong&gt;. &lt;a href=&quot;https://x.com/MiaAI_lab/status/2076951362407944622&quot;&gt;@MiaAI_lab’s DGX Spark thread&lt;/a&gt; sketched practical multi-node local deployments, including &lt;strong&gt;1M-context DeepSeek v4 Flash&lt;/strong&gt; and &lt;strong&gt;MiMo-V2.5&lt;/strong&gt; on &lt;strong&gt;2× DGX Sparks&lt;/strong&gt;, and &lt;strong&gt;GLM 5.2 NVFP4&lt;/strong&gt; across four. The common theme across these posts is that local inference is no longer just a toy path: it is becoming viable for serious agentic workflows, especially when paired with low-bit weight formats and optimized harnesses.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Multimodal and World-Model Systems: Video, Realtime VLMs, and Motion&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Realtime multimodal interaction is moving from “watch then answer” to continuous perception&lt;/strong&gt;: &lt;a href=&quot;https://x.com/MosiAI_Official/status/2076989390191202577&quot;&gt;OpenMOSS released MOSS-VL-Realtime&lt;/a&gt;, an &lt;strong&gt;11B&lt;/strong&gt; vision-language family under &lt;strong&gt;Apache 2.0&lt;/strong&gt; with &lt;strong&gt;256K context&lt;/strong&gt;, designed for &lt;strong&gt;continuous video streams&lt;/strong&gt;. Its key systems property is that it can &lt;strong&gt;keep watching while generating&lt;/strong&gt;, revise or interrupt answers as scenes change, and remain silent when evidence is insufficient. A companion technical thread from &lt;a href=&quot;https://x.com/Open_MOSS/status/2076993673552879790&quot;&gt;@Open_MOSS&lt;/a&gt; emphasizes a &lt;strong&gt;cross-attention architecture&lt;/strong&gt;, &lt;strong&gt;XRoPE&lt;/strong&gt; for unified temporal-spatial positioning, and unified templates across offline/streaming/realtime settings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Long-video understanding is increasingly framed as active evidence search, not passive frame ingestion&lt;/strong&gt;: a dense summary from &lt;a href=&quot;https://x.com/ZhihuFrontier/status/2076962763394695225&quot;&gt;@ZhihuFrontier&lt;/a&gt; described &lt;strong&gt;OmniAgent&lt;/strong&gt;, built on &lt;strong&gt;Qwen2.5-Omni-7B&lt;/strong&gt;, which uses an &lt;strong&gt;Observation–Thought–Action&lt;/strong&gt; loop to request only the frames/audio it needs. On &lt;strong&gt;LVBench&lt;/strong&gt;, OmniAgent-7B reportedly scored &lt;strong&gt;50.5&lt;/strong&gt;, beating &lt;strong&gt;Qwen2.5-VL-72B at 47.3&lt;/strong&gt;, while consuming only ~&lt;strong&gt;203 frames vs 768&lt;/strong&gt;. The training recipe is also notable: passive SFT hurt performance, while &lt;strong&gt;58K agentic trajectories&lt;/strong&gt; and entropy-weighted RL via &lt;strong&gt;TAURA&lt;/strong&gt; improved it. The larger research pattern here aligns with &lt;a href=&quot;https://x.com/andrew_n_carr/status/2076881679055249647&quot;&gt;Andrew Carr’s note&lt;/a&gt; that &lt;strong&gt;motion is a fundamentally novel data type&lt;/strong&gt; requiring dedicated collection, infra, and model treatment rather than being reduced to images-with-time.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open world models are inching toward interactive, longer-horizon simulation&lt;/strong&gt;: &lt;a href=&quot;https://x.com/RekaAILabs/status/2077043205854707813&quot;&gt;@RekaAILabs outlined&lt;/a&gt; the data stack behind omni world models, stressing &lt;strong&gt;petabytes of video&lt;/strong&gt;, &lt;strong&gt;6 pipeline stages&lt;/strong&gt;, and the doubled payoff from data-quality improvements when models both &lt;strong&gt;generate and understand&lt;/strong&gt; video. &lt;a href=&quot;https://x.com/omarsar0/status/2077058222339338748&quot;&gt;@omarsar0 summarized LingBot-World 2.0&lt;/a&gt; as one of the first open releases claiming &lt;strong&gt;hour-scale, 720p/60fps interactive generation&lt;/strong&gt;, though still without long-term memory. On the application side, &lt;a href=&quot;https://x.com/kimmonismus/status/2077002223612276866&quot;&gt;PixVerse Game&lt;/a&gt; was highlighted as pursuing the harder problem of &lt;strong&gt;real-time interactive video response&lt;/strong&gt; rather than canned game-like clips.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Research Infrastructure, Benchmarks, and Evaluation Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Perplexity open-sourced WANDR, a benchmark for wide-and-deep agentic research&lt;/strong&gt;: &lt;a href=&quot;https://x.com/perplexity_ai/status/2077099503723946121&quot;&gt;@perplexity_ai&lt;/a&gt; described WANDR as a &lt;strong&gt;500-task&lt;/strong&gt; benchmark built from de-identified production research tasks, requiring &lt;strong&gt;170,495 source-backed records&lt;/strong&gt; across multiple difficulty tiers. Rather than grading against a static gold set, WANDR &lt;strong&gt;re-fetches cited pages&lt;/strong&gt; and checks claims against underlying evidence, which better matches dynamic web research. &lt;a href=&quot;https://x.com/AravSrinivas/status/2077105849638728118&quot;&gt;@AravSrinivas&lt;/a&gt; framed this as the internal benchmark behind Perplexity Computer’s deep-and-wide research harness, while &lt;a href=&quot;https://x.com/denisyarats/status/2077117794869805145&quot;&gt;@denisyarats&lt;/a&gt; emphasized its additional role as an &lt;strong&gt;RL environment synthesized from production traces&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Eval design is getting more adversarial and more realistic&lt;/strong&gt;: &lt;a href=&quot;https://x.com/arena/status/2077056687387885888&quot;&gt;Agent Arena&lt;/a&gt; highlighted work cutting system costs by &lt;strong&gt;89%&lt;/strong&gt; while matching the best static config’s accuracy, arguing that &lt;strong&gt;full system config &gt; LLM routing alone&lt;/strong&gt;. Relatedly, &lt;a href=&quot;https://x.com/dair_ai/status/2077048984812896677&quot;&gt;Google DeepMind work on model routing&lt;/a&gt; argued that routers should be judged not just by accuracy/cost but by &lt;strong&gt;behavioral differentiation&lt;/strong&gt; among experts and &lt;strong&gt;stability under paraphrase&lt;/strong&gt;; otherwise routing may be functionally meaningless. &lt;a href=&quot;https://x.com/HamelHusain/status/2077042379392213377&quot;&gt;@HamelHusain’s automated evals post&lt;/a&gt; landed in a similar place: these systems can spot issues humans miss, but still lack enough domain taste and feedback loops to replace experts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmarks are expanding beyond one-shot SWE tasks toward degradation and search realism&lt;/strong&gt;: &lt;a href=&quot;https://x.com/KLieret/status/2077042438649020714&quot;&gt;mini-swe-agent&lt;/a&gt; marked one year while now powering multiple software benchmarks; &lt;a href=&quot;https://x.com/sdrzn/status/2077121290440454467&quot;&gt;SlopCodeBench&lt;/a&gt; was cited as measuring how agents &lt;strong&gt;erode codebases over sequential tasks&lt;/strong&gt; rather than just solving one isolated issue. This broadens the benchmark surface from “can it solve a task?” to “can it avoid making the repository worse over time?”&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Physical AI, Collective Intelligence, and Robotics&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Sakana AI pushed collective intelligence from software into physical self-repairing systems&lt;/strong&gt;: across multiple posts, &lt;a href=&quot;https://x.com/SakanaAILabs/status/2076951818089721969&quot;&gt;Sakana introduced “Smart Cellular Bricks”&lt;/a&gt;, published in &lt;strong&gt;Nature Communications&lt;/strong&gt;. The system consists of many identical cubes, each running a small neural network and communicating only with physical neighbors, yet able to infer global shape and detect damage &lt;strong&gt;without centralized control&lt;/strong&gt;. A follow-up detail is especially notable: the cells can detect &lt;strong&gt;missing neighbors across six spatial directions with 95% accuracy&lt;/strong&gt; and regrow target structures; in simulation, the method scaled to &lt;strong&gt;18,000+ cubes&lt;/strong&gt; (&lt;a href=&quot;https://x.com/SakanaAILabs/status/2076930948348674248&quot;&gt;detail thread&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Physical autonomy is also showing up in much smaller form factors&lt;/strong&gt;: &lt;a href=&quot;https://x.com/alextoussss/status/2077086243632873540&quot;&gt;@alextoussss posted&lt;/a&gt; a striking demo of an &lt;strong&gt;autonomous micro-drone&lt;/strong&gt; achieving an &lt;strong&gt;air-to-air kill of a flying moth&lt;/strong&gt;, framed as a step toward mosquito eradication. Separately, &lt;a href=&quot;https://x.com/fchollet/status/2077033256365736098&quot;&gt;@fchollet highlighted Airtap&lt;/a&gt;, which turns &lt;strong&gt;SMS into a headless agentic execution layer for mobile apps&lt;/strong&gt;, using text as the control plane and intervening only for authentication. These are different ends of the autonomy spectrum, but both point to interfaces where humans specify goals while systems handle embodied or semi-embodied execution.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI demand spike and product pull&lt;/strong&gt;: &lt;a href=&quot;https://x.com/sama/status/2077036999303999910&quot;&gt;@sama on GPT-5.6 Sol pricing/efficiency&lt;/a&gt;, &lt;a href=&quot;https://x.com/sama/status/2077033807736459713&quot;&gt;2.5x growth in Codex/Work usage&lt;/a&gt;, and &lt;a href=&quot;https://x.com/sama/status/2077106587307798989&quot;&gt;“5.6 sol growth is insane”&lt;/a&gt; were the most consequential operator signals in the set.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Governance and lab politics&lt;/strong&gt;: &lt;a href=&quot;https://x.com/BlackHC/status/2077009476423647596&quot;&gt;@BlackHC’s thread on DeepMind’s Pentagon contract and abandoned safeguards&lt;/a&gt; and &lt;a href=&quot;https://x.com/carolecadwalla/status/2077015818580193650&quot;&gt;Carole Cadwalladr amplifying it&lt;/a&gt; drew very high engagement. In parallel, &lt;a href=&quot;https://x.com/demishassabis/status/2076957440109625718&quot;&gt;Demis Hassabis’ AGI governance proposal&lt;/a&gt;, endorsed by &lt;a href=&quot;https://x.com/mustafasuleyman/status/2076991204705624434&quot;&gt;@mustafasuleyman&lt;/a&gt; and &lt;a href=&quot;https://x.com/sama/status/2077042528906527225&quot;&gt;@sama&lt;/a&gt;, was a major policy discussion node.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Notable open-model release&lt;/strong&gt;: &lt;a href=&quot;https://x.com/PrismML/status/2077084891284721827&quot;&gt;Bonsai 27B&lt;/a&gt; stood out as the strongest technically substantive open-model launch in the timeline, due to its combination of &lt;strong&gt;27B scale&lt;/strong&gt;, &lt;strong&gt;phone-class footprint&lt;/strong&gt;, and &lt;strong&gt;Apache 2.0&lt;/strong&gt; licensing.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Chinese Open-Weight Models Gain Market Share&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1uuyw46/chinese_ai_models_seize_openrouters_top_five_as/&quot;&gt;Chinese AI Models Seize OpenRouter’s Top Five as OpenAI and Google Vanish From the Top 10&lt;/a&gt;&lt;/strong&gt; (Activity: 637): &lt;strong&gt;The image is an &lt;strong&gt;OpenRouter “AI Model Rankings” dashboard&lt;/strong&gt; showing monthly token-usage share, where Chinese models reportedly take the &lt;strong&gt;top five&lt;/strong&gt; spots—&lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt;, &lt;strong&gt;MiMo-V2.5&lt;/strong&gt;, &lt;strong&gt;MiniMax M3&lt;/strong&gt;, &lt;strong&gt;Hy3 preview&lt;/strong&gt;, and &lt;strong&gt;DeepSeek V4 Pro&lt;/strong&gt;—and &lt;strong&gt;seven of the top ten&lt;/strong&gt;, while &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Google&lt;/strong&gt; are absent from the top 10; the source ranking is OpenRouter’s own platform traffic, not global LLM usage. Technically, the significance is less about benchmark superiority and more about &lt;strong&gt;cost/performance and deployment economics&lt;/strong&gt;: commenters frame OpenRouter as a practical routing/prototyping layer for testing cheaper open or semi-open models before deciding whether to self-host or keep paying API rates. &lt;a href=&quot;https://i.redd.it/o8g1mxm1rwch1.jpeg&quot;&gt;Image&lt;/a&gt;&lt;/strong&gt; Commenters emphasized that &lt;em&gt;“it’s hard to compare benchmarks, but easy to compare bills,”&lt;/em&gt; arguing that models like DeepSeek V4 Flash and MiMo-V2.5 are attractive because they are “cheap and good enough,” sometimes cheaper via API than the electricity and hardware costs of self-hosting. There is also distrust of closed Western providers such as OpenAI and Anthropic due to pricing, model churn, and the possibility that preferred models may disappear or change unexpectedly.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed &lt;strong&gt;OpenRouter&lt;/strong&gt; as a practical model-evaluation layer: test multiple open-source models through the API, then either self-host the best fit or continue using OpenRouter if the unit economics are better. The main technical concern raised was vendor instability from &lt;strong&gt;OpenAI/Anthropic&lt;/strong&gt;—models changing, becoming more expensive, or disappearing—making reproducibility and long-term deployment planning harder.&lt;/li&gt;
&lt;li&gt;A cost-focused thread argued that models like &lt;strong&gt;&lt;code&gt;deepseek-v4-flash&lt;/code&gt;&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;mimo-v2.5&lt;/code&gt;&lt;/strong&gt; are cheap enough via hosted inference that self-hosting may cost more once electricity and hardware requirements are included. One commenter summarized the benchmarking-vs-cost tradeoff as: &lt;em&gt;“It’s hard to compare benchmarks, but easy to compare bills.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;Infrastructure economics were highlighted as a factor in Chinese model/API competitiveness: commenters noted that &lt;strong&gt;electricity costs in China are less than half of US costs&lt;/strong&gt;, which can materially affect inference pricing at scale. The discussion contrasted this with expensive US regions such as &lt;strong&gt;California&lt;/strong&gt;, where high power prices and constraints on new generation could make domestic data-center inference less competitive.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uvenf1/ft_companies_turn_to_chinese_open_weight_models/&quot;&gt;FT: Companies Turn to Chinese Open Weight Models to Cut Costs&lt;/a&gt;&lt;/strong&gt; (Activity: 407): &lt;strong&gt;The post links to an &lt;strong&gt;FT&lt;/strong&gt; story titled &lt;em&gt;“Companies Turn to Chinese Open Weight Models to Cut Costs,”&lt;/em&gt; implying enterprise adoption of Chinese open-weight LLMs as a cost-reduction strategy versus proprietary/API-hosted Western models. However, the archived link (&lt;a href=&quot;https://archive.ph/QzSyV&quot;&gt;archive.ph/QzSyV&lt;/a&gt;) only returned a CAPTCHA/&lt;code&gt;429 Too Many Requests&lt;/code&gt;, so no article-level specifics—model names, pricing deltas, benchmarks, deployment patterns, or licensing terms—are verifiable from the provided source.&lt;/strong&gt; Commenters frame the trend as predictable after perceived restrictions/bans on other models, arguing that policy and IP pressure may push companies toward Chinese open-weight alternatives. There is also optimism that open models will improve rapidly over the next few years, with one commenter citing phone-scale capability as evidence of accelerating local inference potential.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed the trend as a &lt;strong&gt;cost-structure problem for closed US frontier APIs&lt;/strong&gt;, not an “AI bubble” broadly: companies that pushed heavy Anthropic/OpenAI-style token consumption are now reportedly asking employees to reduce usage after months of encouraging maximum adoption. One anecdote described internal incentives to use AI heavily despite unclear revenue impact, with the implication that token bills can become material operational spend before product-market fit is proven.&lt;/li&gt;
&lt;li&gt;A technical theme was optimism around &lt;strong&gt;open-weight model efficiency&lt;/strong&gt;, especially for local or edge deployment. One commenter cited &lt;code&gt;Gemma 4&lt;/code&gt; as an example of models becoming strong enough to run on phones, arguing that open models could increasingly replace paid frontier API calls where latency, privacy, or cost constraints matter.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uw9ucd/source_the_trump_administration_and_industry/&quot;&gt;Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models&lt;/a&gt;&lt;/strong&gt; (Activity: 451): &lt;strong&gt;The post claims the &lt;strong&gt;Trump administration&lt;/strong&gt; and AI industry groups discussed a policy/process to streamline U.S. releases of &lt;strong&gt;open-weight/open models&lt;/strong&gt; whose capability is &lt;em&gt;equal to or below&lt;/em&gt; leading Chinese open models, as a response to China’s increasingly competitive local LLM ecosystem. The linked source is not technically verifiable from the provided archive because &lt;a href=&quot;https://archive.is/sANZ5&quot;&gt;&lt;code&gt;archive.is/sANZ5&lt;/code&gt;&lt;/a&gt; resolves to a CAPTCHA/security interstitial rather than the article content.&lt;/strong&gt; Commenters debated whether U.S. labs would actually release open models competitive with Chinese systems, arguing that strong local models could cannibalize paid API/SaaS revenue. Others questioned claims about Chinese models as “Trojan horses” or containing CCP-exploitable backdoors, noting that if such backdoors were straightforward to implant and exploit, U.S. actors would likely already dominate open-weight model deployment.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that restrictions or bans on capable open-weight models would be technically unenforceable once weights are available globally: model files can be redistributed via torrents, physical drives, or mirrors, and enforcement would require controlling access to high-memory GPUs/workstations capable of running them. The discussion frames open weights as closer to general data distribution than a controllable service endpoint.&lt;/li&gt;
&lt;li&gt;A technical skepticism emerged around claims that Chinese open models could contain CCP-accessible “backdoors.” One commenter noted that if reliable model-weight backdoors were practical and exploitable in this way, U.S. labs would likely already be using similar techniques to dominate open-weight releases; another called such backdoors &lt;em&gt;“unlikely at this point.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;Commenters supported releasing U.S. open-weight models of equal or better quality as a competitive alternative to Chinese local models, especially for organizations that need on-prem deployment. One technical point highlighted &lt;strong&gt;Nvidia Nemotron&lt;/strong&gt; as a positive example due to relatively transparent training data/process, while noting &lt;strong&gt;Nemotron 3 Ultra&lt;/strong&gt; is strong but &lt;em&gt;“not trained enough,”&lt;/em&gt; causing it to underperform relative to its parameter scale.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uvbgvx/zhipu_founder_backs_opensource_ai_as_global/&quot;&gt;Zhipu founder backs open-source AI as global security debate intensifies&lt;/a&gt;&lt;/strong&gt; (Activity: 298): &lt;strong&gt;&lt;strong&gt;Zhipu founder Tang Jie&lt;/strong&gt; defended open-sourcing frontier AI in an internal memo, arguing that model security is better served by &lt;em&gt;“transparency, broad participation, sharing, and oversight”&lt;/em&gt; than by restricting access, according to &lt;a href=&quot;https://www.business-standard.com/technology/tech-news/zhipu-founder-backs-open-source-ai-as-global-security-debate-intensifies-126071200342_1.html&quot;&gt;Business Standard&lt;/a&gt;. Zhipu has released &lt;strong&gt;GLM-5.2&lt;/strong&gt; under an open-source license for download and commercial use, while positioning future work around &lt;strong&gt;long-horizon tasks, autonomous agents, and self-training AI&lt;/strong&gt;, amid broader policy debates involving &lt;strong&gt;OpenAI&lt;/strong&gt;, &lt;strong&gt;Anthropic&lt;/strong&gt;, cyber-risk, and national security controls.&lt;/strong&gt; Comments were largely geopolitical and anti-monopoly in tone: users framed Chinese open-weight releases as a counterweight to closed US labs, while speculating that services like &lt;a href=&quot;http://z.ai&quot;&gt;&lt;code&gt;z.ai&lt;/code&gt;&lt;/a&gt; could face bans on security grounds.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter notes that &lt;strong&gt;Zhipu/Z.ai’s open-source stance may be market-driven rather than ideological&lt;/strong&gt;, pointing out that the company was previously more closed before &lt;strong&gt;DeepSeek R1&lt;/strong&gt; shifted competitive pressure toward open flagship releases. They argue the strategy could reverse if investor incentives begin favoring closed models again, despite recent investor acceptance of dilution.&lt;/li&gt;
&lt;li&gt;One technically relevant user references prior experience with the &lt;strong&gt;GLM series&lt;/strong&gt;, saying they have been a fan but that its &lt;strong&gt;coding-plan quality has historically been questionable&lt;/strong&gt;. This suggests continued skepticism around Zhipu’s practical coding-agent performance even as its open-source positioning improves.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Local AI Inference: Compression, GPUs, and Game Engines&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uv54fv/compressed_version_of_qwen3627b_coming_from/&quot;&gt;Compressed Version of Qwen-3.6-27B coming from PrismML - Khosla-Backed Startup Claims Breakthrough With Largest-Ever AI Model on an iPhone&lt;/a&gt;&lt;/strong&gt; (Activity: 476): &lt;strong&gt;&lt;strong&gt;PrismML&lt;/strong&gt; claims it compressed Alibaba’s open-weight &lt;a href=&quot;https://qwen.ai/blog?id=qwen3.6-27b&quot;&gt;&lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt;&lt;/a&gt; from roughly &lt;code&gt;54 GB&lt;/code&gt; to &lt;strong&gt;&amp;#x3C; &lt;code&gt;4 GB&lt;/code&gt;&lt;/strong&gt; and can run it on an &lt;strong&gt;iPhone 17 Pro&lt;/strong&gt; with &lt;strong&gt;all &lt;code&gt;27B&lt;/code&gt; parameters active&lt;/strong&gt;, unlike sparse/on-device architectures such as Apple’s cited &lt;code&gt;20B&lt;/code&gt;-parameter model with only &lt;code&gt;1B–4B&lt;/code&gt; active. The company says the Caltech-derived, patent-licensed compression method preserves performance and enables on-device chat, reasoning, agents, and coding, with a downloadable release promised “next Tuesday,” though no benchmark numbers, tokens/sec, quantization details, accuracy deltas, or demo evidence are provided in the post.&lt;/strong&gt; Top commenters are highly skeptical, calling the brain-synapse analogy technically meaningless and questioning the plausibility of &amp;#x3C;&lt;code&gt;4 GB&lt;/code&gt; compression with no performance loss while computing all &lt;code&gt;27B&lt;/code&gt; parameters on an iPhone at usable speed. Several argue that credible claims should be accompanied by a live demo, benchmarks, or a technical blog rather than a hype article.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters challenged the claim that a &lt;code&gt;27B&lt;/code&gt; model can be compressed to roughly &lt;code&gt;4GB&lt;/code&gt; and run all parameters on an iPhone at acceptable speed without major quality loss. One technical guess was that this would require something like &lt;strong&gt;1-bit / ternary quantization&lt;/strong&gt; or BitNet-style quantization; a &lt;code&gt;Q1&lt;/code&gt;-level &lt;code&gt;27B&lt;/code&gt; model can land near the &lt;code&gt;4GB&lt;/code&gt; range, but commenters noted it is typically &lt;em&gt;“very damaged compared to fp16.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;A more detailed comment identified PrismML’s likely approach as its existing &lt;strong&gt;“1 bit” / ternary quantization&lt;/strong&gt;, previously released for smaller Qwen models such as &lt;strong&gt;Bonsai-8B&lt;/strong&gt; on &lt;a href=&quot;https://huggingface.co/prism-ml/Bonsai-8B-gguf&quot;&gt;Hugging Face&lt;/a&gt;. The commenter emphasized that prior PrismML benchmarks did &lt;strong&gt;not&lt;/strong&gt; show near-original performance: their &lt;code&gt;8B&lt;/code&gt; quant reportedly performed worse than a &lt;code&gt;BF16 4B&lt;/code&gt;, though better than &lt;code&gt;1.7B&lt;/code&gt;, implying the &lt;code&gt;27B&lt;/code&gt; version should not be expected to preserve full Qwen quality.&lt;/li&gt;
&lt;li&gt;The technically relevant benchmark framing suggested by commenters is not whether the compressed &lt;code&gt;27B&lt;/code&gt; matches the original model, but whether a &lt;code&gt;4GB&lt;/code&gt; ternary &lt;code&gt;27B&lt;/code&gt; outperforms a conventional &lt;code&gt;4GB&lt;/code&gt; &lt;code&gt;8B Q4&lt;/code&gt; model. One commenter also noted that PrismML appears to be comparing against &lt;strong&gt;Qwen3&lt;/strong&gt;, not newer &lt;strong&gt;Qwen 3.5 / 3.6&lt;/strong&gt;, which affects how claims about relative model quality should be interpreted.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uvcjd0/i_benchmarked_15_ewaste_gpus_with_modern_workloads/&quot;&gt;I benchmarked 15 &quot;E-Waste&quot; GPUs with Modern Workloads&lt;/a&gt;&lt;/strong&gt; (Activity: 628): &lt;strong&gt;The author benchmarked decommissioned NVIDIA Tesla-class GPUs—K80/M10/M40/M60/P40/P100/V100/T40—using a custom Dockerized suite (&lt;a href=&quot;https://github.com/esologic/gpu_box_benchmark&quot;&gt;&lt;code&gt;gpu_box_benchmark&lt;/code&gt;&lt;/a&gt;) and custom cooling hardware, targeting LLMs, CV, Blender, Whisper, and related workloads; full graphs are in the &lt;a href=&quot;https://esologic.com/benchmarking-tesla-gpus/&quot;&gt;blog post&lt;/a&gt;. Key findings: &lt;strong&gt;V100 16GB&lt;/strong&gt; is the best overall value near &lt;code&gt;&amp;#x3C;$200&lt;/code&gt; and approaches T40 performance, &lt;strong&gt;P40 beats P100 for LLM inference&lt;/strong&gt;, &lt;strong&gt;M60 is unusually strong for Whisper&lt;/strong&gt; despite ~$50 pricing, and multi-GPU scaling was described as mostly linear within a 4U chassis, with mixed-generation LLM setups bottlenecking on slower cards. The author argues EOL/CUDA-era friction can often be worked around by compiling older software such as &lt;code&gt;llama.cpp&lt;/code&gt; from source, and that cheap X99 Xeon platforms provide enough PCIe lanes/CPU throughput for these GPUs in homelab workloads.&lt;/strong&gt; Top technical pushback focused on whether the benchmark set is sufficiently “modern”: commenters asked for larger contemporary LLMs such as &lt;strong&gt;Qwen 3.x 27B/35B MoE&lt;/strong&gt;, pooled-VRAM tests across multiple V100/P40-class cards, and prompt-processing/token-generation numbers at long context lengths like &lt;code&gt;150k&lt;/code&gt;. Others questioned practical power efficiency and acoustics, noting these systems may only make sense if powered on for batch AI jobs—or if waste heat offsets winter heating.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued the benchmark missed the key value proposition of old datacenter/mining cards: &lt;strong&gt;cheap pooled VRAM for larger contemporary LLMs&lt;/strong&gt;. They requested tests with models like &lt;strong&gt;Qwen 3.6 27B/31B MoE/35B A3B&lt;/strong&gt; at deep context lengths such as &lt;code&gt;150k ctx&lt;/code&gt;, reporting both &lt;strong&gt;prompt processing (PP)&lt;/strong&gt; and &lt;strong&gt;token generation (TG)&lt;/strong&gt; across multi-GPU configurations, especially on &lt;strong&gt;V100-class&lt;/strong&gt; setups.&lt;/li&gt;
&lt;li&gt;There was technical pushback on workload selection: &lt;strong&gt;ResNet&lt;/strong&gt; and very small models were called unrepresentative of “modern” GPU use, because they do not stress the VRAM-capacity advantage of e-waste GPUs. The suggested practical benchmark was whether larger Qwen-class models can run with pooled VRAM at usable speeds, rather than showing good performance on legacy or undersized workloads.&lt;/li&gt;
&lt;li&gt;One commenter noted the &lt;strong&gt;Tesla P100 should outperform the P40&lt;/strong&gt; unless a relevant &lt;code&gt;fp32&lt;/code&gt; optimization/patch changed the results, because the P100’s &lt;strong&gt;HBM bandwidth is roughly 3× higher&lt;/strong&gt; than the P40’s memory bandwidth. Another added a data point for the &lt;strong&gt;P102-100 mining GPU&lt;/strong&gt;: about &lt;code&gt;$50&lt;/code&gt;, around &lt;code&gt;10 W&lt;/code&gt; idle, easy cooling, and roughly &lt;code&gt;40 tokens/s&lt;/code&gt; generation on &lt;strong&gt;Qwen 3.6 35B&lt;/strong&gt;, but with very slow prompt processing at about &lt;code&gt;100&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uv66by/i_got_gemma_4_running_directly_inside_godot_using/&quot;&gt;I got Gemma 4 running directly inside Godot using only GDScript and Vulkan compute shaders&lt;/a&gt;&lt;/strong&gt; (Activity: 428): &lt;strong&gt;The post demonstrates &lt;strong&gt;Gemma 4 (&lt;code&gt;gemma-4-E2B-it-Q4_K_M.gguf&lt;/code&gt;) running fully inside Godot 4.7&lt;/strong&gt; using only &lt;strong&gt;GDScript + Vulkan compute shaders&lt;/strong&gt;, with GDScript handling GGUF loading, tokenization, sampling, KV cache, and UI—no &lt;code&gt;llama.cpp&lt;/code&gt;, Python, server, C bindings, or GDExtension. The &lt;a href=&quot;https://i.redd.it/etqze9k9pych1.png&quot;&gt;image&lt;/a&gt; shows a Godot editor/debug chat UI generating at about &lt;code&gt;46.99 tok/s&lt;/code&gt;, notably answering that such an implementation would be “not realistic,” despite the project proving a constrained version works. The author notes it is experimental, supports only one model, and is roughly &lt;code&gt;10×&lt;/code&gt; slower than &lt;code&gt;llama.cpp&lt;/code&gt; with CUDA; code is available on &lt;a href=&quot;https://github.com/asallay/godot-llm&quot;&gt;GitHub&lt;/a&gt;.&lt;/strong&gt; Comments were mostly impressed, with one technical takeaway that the speed penalty is less important than the deployment model: a single Godot export with local inference avoids native-extension ABI issues and sidecar servers, making it plausible for portable local NPC demos.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically substantive point is that even at &lt;strong&gt;~&lt;code&gt;10x&lt;/code&gt; slower performance&lt;/strong&gt;, the implementation is valuable because it packages &lt;strong&gt;GGUF loading, KV cache management, and sampling&lt;/strong&gt; entirely inside a single &lt;strong&gt;Godot export&lt;/strong&gt; using GDScript/Vulkan compute, avoiding native-extension ABI issues or a separate &lt;code&gt;llama.cpp&lt;/code&gt;/sidecar inference server. This could make local LLM-powered NPC demos much easier for end users to run.&lt;/li&gt;
&lt;li&gt;One commenter framed the main use case as embedded local generation for games, e.g. roguelike/roguelite systems that use an on-device LLM for more emergent randomness. The key technical appeal is removing the deployment burden of bundling or orchestrating an external inference runtime such as &lt;code&gt;llama.cpp&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Kimi, DeepSeek, and GLM Release Watch&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwe542/kimi_k3_in_the_next_few_hours_deepseek_v4_ga/&quot;&gt;Kimi K3 in the next few hours. Deepseek V4 GA later in the week. New Liquid models. New Mistral models sometime this month. And some rumours suggest GLM 5.5 is coming in August. Openweight AI is eating good.&lt;/a&gt;&lt;/strong&gt; (Activity: 573): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/2uhew16k58dh1.jpeg&quot;&gt;image&lt;/a&gt; is a screenshot amplifying rumors that &lt;strong&gt;Moonshot/Kimi K3&lt;/strong&gt; may launch imminently, following Kimi K2.6’s reported strengths in coding agents, long tool-using sessions, &lt;code&gt;256K&lt;/code&gt; context, vision, and low-cost large-scale sub-agent coordination. In context with the title/selftext, the post frames Kimi K3, &lt;strong&gt;DeepSeek V4 GA&lt;/strong&gt;, new &lt;strong&gt;Liquid&lt;/strong&gt; non-transformer models, upcoming &lt;strong&gt;Mistral&lt;/strong&gt; releases, and possible &lt;strong&gt;GLM 5.5&lt;/strong&gt; as evidence that open-weight model capability and cost-performance are rapidly improving, while enterprise concerns shift toward governance/control layers rather than raw model intelligence.&lt;/strong&gt; Commenters are enthusiastic but pragmatic: one user reports running local &lt;strong&gt;GLM 5.2 Q4&lt;/strong&gt; at only &lt;code&gt;~0.5 tok/s&lt;/code&gt; for multi-day codebase audits that still find useful bugs, while others argue the ecosystem especially needs strong models at &lt;code&gt;100B&lt;/code&gt; parameters and below—ideally not below &lt;code&gt;35B&lt;/code&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user reports running &lt;strong&gt;GLM 5.2 Q4 locally&lt;/strong&gt; continuously on a workstation for whole-repository static-analysis style prompting: &lt;em&gt;“Read all code in this project and analyze it for bugs.”&lt;/em&gt; Throughput is only about &lt;code&gt;0.5 tokens/sec&lt;/code&gt;, with a full run taking roughly &lt;code&gt;3.5 days&lt;/code&gt;, but each run reportedly returns &lt;code&gt;10–15&lt;/code&gt; bug findings with only &lt;code&gt;1–2&lt;/code&gt; hallucinations and at least one practically useful issue fixed per run.&lt;/li&gt;
&lt;li&gt;Several comments emphasize that current frontier open-weight releases are often impractical for enthusiasts because many are &lt;code&gt;0.5T+&lt;/code&gt; parameter-class models requiring extreme hardware, e.g. multiple &lt;strong&gt;RTX PRO 6000-class GPUs&lt;/strong&gt; or around &lt;code&gt;1TB&lt;/code&gt; ECC DDR5 for CPU/offload setups. Users specifically call for stronger models in the &lt;strong&gt;≤100B&lt;/strong&gt; range, with one commenter narrowing the useful local target to &lt;strong&gt;above 35B but under 100B&lt;/strong&gt; rather than giant MoE/foundation-scale checkpoints.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uwbpmw/a_new_glm_model_incoming/&quot;&gt;👀A new GLM model incoming&lt;/a&gt;&lt;/strong&gt; (Activity: 1043): &lt;strong&gt;The image is a &lt;strong&gt;non-technical teaser screenshot&lt;/strong&gt; from X, not a benchmark or release note: Ivan Fioravanti says &lt;em&gt;“GLM 5.3 is cooking”&lt;/em&gt; while quoting &lt;strong&gt;Z.ai / GLM&lt;/strong&gt; founder Jie Tang’s hint that &lt;em&gt;“5.2 could be better with more RL,”&lt;/em&gt; implying an upcoming GLM update likely focused on additional reinforcement learning/post-training. The Reddit post frames this as a possible successor to &lt;strong&gt;GLM 5.2&lt;/strong&gt;, with speculation in comments about variants like &lt;code&gt;GLM 5.3 Flash 20B&lt;/code&gt; and broader open-weight model releases; image: &lt;a href=&quot;https://i.redd.it/6xkuthwho7dh1.jpeg&quot;&gt;i.redd.it/6xkuthwho7dh1.jpeg&lt;/a&gt;.&lt;/strong&gt; Commenters are broadly excited about a crowded open-weight release window, mentioning rumored or expected models such as &lt;strong&gt;Kimi K3&lt;/strong&gt;, &lt;strong&gt;DeepSeek V4 GA&lt;/strong&gt;, &lt;strong&gt;Liquid&lt;/strong&gt;, &lt;strong&gt;Mistral&lt;/strong&gt;, and possibly &lt;strong&gt;GLM 5.5&lt;/strong&gt;. There is no substantive technical debate yet, mostly hype and speculation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters expect a crowded open-weight release window: &lt;strong&gt;Kimi K3&lt;/strong&gt; reportedly “in the next few hours,” &lt;strong&gt;DeepSeek V4 GA&lt;/strong&gt; later in the week, new &lt;strong&gt;Liquid&lt;/strong&gt; models, new &lt;strong&gt;Mistral&lt;/strong&gt; models this month, and rumors of &lt;strong&gt;GLM 5.5&lt;/strong&gt; in August. Another commenter speculates the incoming GLM could be &lt;strong&gt;GLM 5.3 Flash 20B&lt;/strong&gt;, implying interest in a smaller/fast variant rather than another very large checkpoint.&lt;/li&gt;
&lt;li&gt;A technical concern is usability of frontier-scale open models on attainable hardware: one user asks for models that run at reasonable speed on &lt;strong&gt;&amp;#x3C;&lt;code&gt;$100k&lt;/code&gt; hardware&lt;/strong&gt;, rejecting &lt;strong&gt;&lt;code&gt;30 tok/s&lt;/code&gt; prompt processing and &lt;code&gt;5 tok/s&lt;/code&gt; token generation&lt;/strong&gt; as insufficient. They propose roughly &lt;strong&gt;&lt;code&gt;1000 tok/s&lt;/code&gt; prompt processing and &lt;code&gt;40 tok/s&lt;/code&gt; generation&lt;/strong&gt; as a practical target for real-world tasks, while others complain that &lt;strong&gt;&lt;code&gt;700GB&lt;/code&gt; models&lt;/strong&gt; are unusable locally and ask for smaller models like a hypothetical &lt;strong&gt;Qwen 3.7 &lt;code&gt;35B&lt;/code&gt;&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Frontier Models Solving Open Math &amp;#x26; Physics Problems&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uv399n/yuji_tachikawa_one_of_the_worlds_leading/&quot;&gt;Yuji Tachikawa, one of the world’s leading theoretical physicists, reports Claude Fable solved a problem that he and his collaborators had gotten stuck on for the past 6 months&lt;/a&gt;&lt;/strong&gt; (Activity: 3554): &lt;strong&gt;&lt;strong&gt;Yuji Tachikawa&lt;/strong&gt;, a leading mathematical/theoretical physicist, reportedly posted that &lt;strong&gt;Claude Fable&lt;/strong&gt; helped solve a technical problem his group had been stuck on for ~&lt;code&gt;6 months&lt;/code&gt; (&lt;a href=&quot;https://x.com/yujitach/status/2076327681562644709?s=20&quot;&gt;original tweet&lt;/a&gt;). He later deleted the tweet not as a retraction, but because he disliked the attention it attracted (&lt;a href=&quot;https://x.com/yujitach/status/2076682201626992776?s=20&quot;&gt;follow-up&lt;/a&gt;); no reproducible derivation, benchmark, or problem statement is included in the Reddit post, so the technical claim cannot be independently assessed from the linked discussion alone.&lt;/strong&gt; Top comments debated evaluation standards for AI-assisted research: one commenter argued that dismissing the result because it was not solved &lt;em&gt;“one shot”&lt;/em&gt; applies a stricter standard to LLMs than to humans, while another highlighted the significance of an LLM proposing speculative directions such as &lt;em&gt;“I wonder if…”&lt;/em&gt; as potentially relevant to frontier reasoning beyond known results.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter frames &lt;strong&gt;Claude Fable’s&lt;/strong&gt; contribution as notable because it appears to involve exploratory hypothesis generation rather than just executing a known procedure: they highlight the model saying &lt;em&gt;“I wonder if…”&lt;/em&gt; as evidence of asking questions beyond the current solution path. The technical implication raised is that frontier LLMs may be moving toward a capability often cited as missing in AI-assisted research: proposing useful hypotheticals in domains where even experts are stuck.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uvrtl0/another_50_yearold_erd%C5%91s_problem_falls_to_gpt56/&quot;&gt;Another 50+ year-old Erdős problem falls to GPT-5.6&lt;/a&gt;&lt;/strong&gt; (Activity: 1144): &lt;strong&gt;The post links to X threads by &lt;a href=&quot;https://x.com/jdlichtman/status/2076778478326653431&quot;&gt;J. D. Lichtman&lt;/a&gt;, &lt;a href=&quot;https://x.com/prz_chojecki/status/2076749164067565872&quot;&gt;Przemysław Chojecki&lt;/a&gt;, and &lt;a href=&quot;https://x.com/SebastienBubeck/status/2076782523464765717&quot;&gt;Sébastien Bubeck&lt;/a&gt; claiming &lt;strong&gt;GPT-5.6&lt;/strong&gt; solved another &lt;strong&gt;50+ year-old Erdős problem&lt;/strong&gt;. The Reddit body does not include the theorem statement, proof, benchmark setup, or verification details, so from the provided content the technically relevant takeaway is the &lt;em&gt;claim&lt;/em&gt; of an LLM-assisted/LLM-generated proof rather than independently assessable mathematical evidence.&lt;/strong&gt; Top comments mostly frame the result as a rebuttal to common LLM-skeptic claims like &lt;em&gt;“just fancy autocorrect”&lt;/em&gt; or &lt;em&gt;“just predicting the next word.”&lt;/em&gt; One substantive suggestion was to benchmark models on already-solved problems with long proofs to see whether they can discover &lt;strong&gt;shorter or simpler proofs&lt;/strong&gt;, which would be useful even when the theorem is not new.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter raised a methodological question about AI-assisted theorem proving: how much of the result comes from the LLM versus expert human guidance. They specifically asked whether a non-expert prompter, e.g. a high school student unable to verify correctness, could still get the model to solve the problem reliably.&lt;/li&gt;
&lt;li&gt;Another technically relevant thread suggested benchmarking models on already-solved problems with long proofs to see whether systems like &lt;strong&gt;GPT-5.6&lt;/strong&gt; can produce materially shorter or simpler proofs. The commenter noted that this could be a useful way to evaluate mathematical creativity or proof compression, even if the practical significance is unclear.&lt;/li&gt;
&lt;li&gt;One commenter noted that &lt;strong&gt;Przemek Chojecki&lt;/strong&gt; appears to be producing solutions to Erdős-style problems faster than they can be fully peer-checked, while &lt;strong&gt;Lichtman&lt;/strong&gt;, described as a Stanford mathematician, reportedly thinks the approach is working. The key technical issue implied is verification throughput: AI-generated proofs may outpace expert validation, making independent checking and formalization important.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. AI Data Extraction, Distillation &amp;#x26; Privacy Incidents&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/1uwavzo/anthropic_just_told_the_us_senate_that_alibaba/&quot;&gt;Anthropic just told the US Senate that Alibaba ran 25,000 fake accounts and had 28.8 million conversations with Claude — not to use it, but to copy it&lt;/a&gt;&lt;/strong&gt; (Activity: 2021): &lt;strong&gt;The post claims &lt;strong&gt;Anthropic told the U.S. Senate&lt;/strong&gt; that &lt;strong&gt;Alibaba&lt;/strong&gt; allegedly used &lt;code&gt;25,000&lt;/code&gt; fake accounts to conduct &lt;code&gt;28.8M&lt;/code&gt; Claude API conversations over ~six weeks (April–June), not via hacking but by normal API access at industrial scale, to distill Claude’s &lt;em&gt;agentic reasoning and coding capabilities&lt;/em&gt; into &lt;strong&gt;Qwen&lt;/strong&gt;. Anthropic reportedly frames this as its largest “distillation attack,” larger than alleged activity by DeepSeek, Moonshot, and MiniMax combined, and the author argues the legal ambiguity explains why Anthropic sent a congressional letter rather than filing suit; they link a longer breakdown on YouTube: &lt;a href=&quot;https://youtu.be/g1d3yTR6E2Y&quot;&gt;youtu.be/g1d3yTR6E2Y&lt;/a&gt;.&lt;/strong&gt; Top comments largely reject Anthropic’s framing as hypocritical, arguing AI labs trained on public/creative output under fair-use theories and now object when their own model outputs are used similarly. One commenter frames mass distillation as analogous to competitive reverse engineering—e.g., automakers or Samsung buying a rival product to study it—while acknowledging the analogy is imperfect.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters drew a technical/legal analogy between &lt;strong&gt;model distillation via API outputs&lt;/strong&gt; and conventional competitive reverse engineering, e.g. buying a competitor’s car or phone and studying it to improve an internal product. The implied technical question is whether using Claude-generated outputs as supervised training data is materially different from engineers learning from a competitor’s product behavior.&lt;/li&gt;
&lt;li&gt;One technically relevant caveat raised was attribution: because &lt;strong&gt;Alibaba operates cloud infrastructure&lt;/strong&gt;, suspicious traffic originating from Alibaba-owned IP ranges or accounts may not prove that &lt;strong&gt;Alibaba itself&lt;/strong&gt; performed the alleged distillation. The commenter compared it to seeing abusive or unusual requests from &lt;code&gt;AWS&lt;/code&gt; or &lt;code&gt;Azure&lt;/code&gt;, which often implicates customers using the platform rather than Amazon or Microsoft directly.&lt;/li&gt;
&lt;li&gt;Another thread framed the incident as an API governance failure: if the &lt;code&gt;25,000&lt;/code&gt; accounts and &lt;code&gt;28.8 million&lt;/code&gt; Claude conversations were paid API usage and not blocked earlier, commenters questioned whether Anthropic’s enforcement, anomaly detection, rate limits, or ToS controls were sufficient to prevent large-scale extraction-like usage.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uvl6dl/grok_build_was_uploading_whole_directories_to/&quot;&gt;grok build was uploading whole directories to google bucket&lt;/a&gt;&lt;/strong&gt; (Activity: 1074): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/hy24xct2r1dh1.png&quot;&gt;image&lt;/a&gt; is a screenshot of a claim by &lt;strong&gt;International Cyber Digest&lt;/strong&gt; alleging that &lt;strong&gt;xAI’s Grok Build CLI&lt;/strong&gt; uploaded whole Git repositories—including private code and unredacted secrets—to a &lt;strong&gt;Google Cloud Storage bucket&lt;/strong&gt;. The post claims a &lt;code&gt;12 GB&lt;/code&gt; test repository caused &lt;code&gt;5.1 GB&lt;/code&gt; of uploads, that the behavior was later disabled via a hidden server-side flag, and that the “Improve the model” opt-out allegedly did &lt;strong&gt;not&lt;/strong&gt; prevent the upload; no independent technical evidence is provided in the Reddit excerpt.&lt;/strong&gt; The comments shown are non-technical and mostly express broad hostility or distrust toward Grok/xAI/Musk rather than debating the implementation or evidence.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. AI Coding Agents: Workflows, Costs &amp;#x26; Reliability Limits&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uvjlmr/fable_56_is_absolute_peak/&quot;&gt;Fable + 5.6 is absolute peak&lt;/a&gt;&lt;/strong&gt; (Activity: 1306): &lt;strong&gt;The post describes a shell-scripted agent workflow (&lt;a href=&quot;https://github.com/PiLastDigit/TRIP-workflow&quot;&gt;&lt;code&gt;TRIP-workflow&lt;/code&gt;&lt;/a&gt;) where &lt;strong&gt;Fable&lt;/strong&gt; acts as a high-level orchestrator rather than primary code generator: it plans, has &lt;strong&gt;5.6 Sol&lt;/strong&gt; review plans in an approval loop, delegates implementation to &lt;strong&gt;5.6 Luna&lt;/strong&gt; via Codex/Claude-code-style CLI background workers with persistent threads, then reads diffs, patches issues, runs tests, and handles release tasks such as changelog/tag/merge. The author emphasizes the stack is “just bash around codex cli” with no MCP/framework/agent swarm, and recommends users clone the repo and have an agent explain/review it before trusting it.&lt;/strong&gt; Top comments discuss using adversarial multi-model workflows: giving both &lt;strong&gt;Fable&lt;/strong&gt; and &lt;strong&gt;Sol 5.6 xhigh&lt;/strong&gt; the same problem statement, letting one design/execute while the other critiques at checkpoints, sometimes with a scorekeeping loop. Others ask for concrete cost/quality comparisons versus Fable alone, thinking-setting choices, and suggest evaluating an &lt;code&gt;omp&lt;/code&gt; harness for this kind of orchestration.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters describe a &lt;strong&gt;multi-model adversarial workflow&lt;/strong&gt;: give both &lt;strong&gt;Fable&lt;/strong&gt; and &lt;strong&gt;Sol 5.6 xhigh&lt;/strong&gt; the same problem statement and goal, let each produce a design, then allow the winning design to execute while the losing model performs checkpoint reviews and teardown. One user reports a similar validation pattern where &lt;strong&gt;Codex&lt;/strong&gt; reviews Fable-generated plans and often finds &lt;em&gt;“something crucial missing,”&lt;/em&gt; suggesting Fable may need external critique for planning reliability.&lt;/li&gt;
&lt;li&gt;A commenter recommends using the &lt;strong&gt;omp harness&lt;/strong&gt; for this kind of model-vs-model orchestration, implying there are existing harnesses better suited than ad hoc scripts for evaluating or coordinating multi-agent/model workflows.&lt;/li&gt;
&lt;li&gt;One user open-sourced a lightweight orchestration daemon, &lt;a href=&quot;https://github.com/hristo2612/jinn&quot;&gt;&lt;code&gt;jinn&lt;/code&gt;&lt;/a&gt;, intended to replace brittle bash glue for &lt;strong&gt;Claude Code + Codex&lt;/strong&gt; workflows. It provides &lt;strong&gt;persistent sessions&lt;/strong&gt;, &lt;strong&gt;cross-engine messages&lt;/strong&gt;, a &lt;strong&gt;shared facts file&lt;/strong&gt;, &lt;strong&gt;cron&lt;/strong&gt;, and YAML personas, but deliberately avoids implementing its own agent loop: *“A bus not a brain, the CLIs still do all the thinking.”&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/OpenAI/comments/1uv3txe/warning_avoid_using_56_sol_it_can_get_you_banned/&quot;&gt;[WARNING] Avoid using 5.6 Sol. It can get you banned for even the most harmless task. Used it once for a legitimate Excel task, got flagged for a “cybersecurity threat,” appeal was rejected within 2h.&lt;/a&gt;&lt;/strong&gt; (Activity: 952): &lt;strong&gt;A user reports that a single benign &lt;strong&gt;Sol 5.6&lt;/strong&gt; task—generating an Excel workbook for rental-property accounting—triggered a “cybersecurity threat” flag and warning despite the prompt containing only spreadsheet requirements: monthly utility/rent accounting, printable statements, cash-flow tracking, room-level over/underpayment carry-forward, and capital vs. utility fund separation. During Sol’s Excel workflow it apparently generated/reran code, hit an exception resembling &lt;em&gt;“Could not get source, probably due to dynamically evaluated source code”&lt;/em&gt;, then still produced the workbook after review; the user’s appeal was rejected within ~&lt;code&gt;2h&lt;/code&gt;.&lt;/strong&gt; Commenters largely treat this as a likely false positive in OpenAI’s automated safety/appeals pipeline; one notes OpenAI staff monitor the subreddit and may be able to manually investigate, while the OP argues the appeal process appears AI-mediated and ineffective.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter shared the full prompt that allegedly triggered the ban: a benign request to generate a multi-sheet Excel workbook for rental-property accounting, including monthly utility allocation, tenant over/underpayment carry-forward, printable statements, and separate capital/utilities cash-flow tracking. The technically relevant detail is that the task likely required workbook generation and formulas/tables, which may have caused the model to invoke a code-execution or file-generation path despite no explicit cybersecurity content.&lt;/li&gt;
&lt;li&gt;One user hypothesized that &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; may have executed code in a cloud sandbox from the ChatGPT client, and that this sandbox activity—not the natural-language prompt itself—could have tripped automated cybersecurity classifiers. They contrasted this with the &lt;strong&gt;Codex client&lt;/strong&gt;, suggesting the same task might not cause the same enforcement issue there because the execution environment or policy pipeline may differ.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uwa1mv/well_it_finally_happened_were_not_using_models/&quot;&gt;Well it finally happened: we’re not using models because of cost&lt;/a&gt;&lt;/strong&gt; (Activity: 849): &lt;strong&gt;A Fortune 500 “AI First” org reportedly pulled back from broad &lt;strong&gt;Claude/Copilot&lt;/strong&gt; usage after a failed agentic pilot to reverse-engineer a legacy application into a formal specification: agents repeatedly missed subtle business rules and produced unreliable specs, while day-to-day codegen showed correctness/safety issues such as generated SQL that dropped constraints around DML paths. The company has stopped AI training/demos, removed Claude access, and is asking teams to limit usage or use older/cheaper models, with commenters noting &lt;strong&gt;Copilot’s move to usage-based pricing&lt;/strong&gt; as a likely trigger for cost visibility in large enterprises.&lt;/strong&gt; Commenters split between seeing this as evidence that current LLM ROI is poor for complex software modernization, versus a temporary “sour spot” where capability is almost useful but still too expensive—especially for agent/subagent-heavy workflows. One commenter asked whether the company would resume an AI-first posture if inference costs drop substantially.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters identify &lt;strong&gt;usage-based pricing for GitHub Copilot&lt;/strong&gt; as the trigger that made large organizations re-evaluate AI tooling costs. The implication is that predictable per-seat licensing masked consumption risk, while metered billing exposes high-volume inference and agent workflows as a material operating expense.&lt;/li&gt;
&lt;li&gt;One technical cost concern is that AI workflows with &lt;strong&gt;intense subagent use&lt;/strong&gt; can multiply inference calls, making otherwise useful models financially unattractive before capability-per-dollar improves. A commenter argues this may be temporary as models improve and become cheaper &lt;em&gt;“per unit of intelligence,”&lt;/em&gt; but that current systems sit in a “sour spot” where they are close to useful yet still expensive at scale.&lt;/li&gt;
&lt;li&gt;A commenter from a very large company claims &lt;strong&gt;every AI project failed or was abandoned&lt;/strong&gt;, not necessarily due to model access limits but because teams did not want to maintain the generated “slop.” The technical takeaway is that AI adoption cost includes downstream maintenance, code quality review, and operational ownership—not just token or subscription spend.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>openai</category><category>jetbrains</category><category>langchain</category><category>prismml</category><category>tencent-hunyuan</category><category>miaai_lab</category><category>openmoss</category><category>gpt-5.6</category><category>codex</category><category>bonsai-27b</category><category>qwen-3.6-27b</category><category>hy3-295b</category><category>gemma-4</category><category>qwen3.5-122b-a10b</category><category>glm-4.7-flash</category><category>deepseek-v4-flash</category><category>mimo-v2.5</category><category>glm-5.2-nvfp4</category><category>moss-vl-realtime</category><category>sama</category><category>reach_vb</category><category>kimmonismus</category><category>swyx</category><category>theo</category><category>andykonwinski</category><category>agentic-ai</category><category>model-quantization</category><category>local-inference</category><category>multimodality</category><category>video-understanding</category><category>model-compression</category><category>evals</category><category>observability</category><category>long-context</category><category>tool-use</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-13-not-much/</guid><description>**Prime Intellect** released **verifiers v1**, a redesigned environment stack for **agentic reinforcement learning** and evaluations, improving efficiency by storing rollout traces as **message DAGs** to reduce complexity from **O(n²)** to **O(n)**. This enables practical long-horizon multimodal rollouts, demonstrated with a **100B reasoning model** running **40-turn SWE agent tasks** on **6 H200 nodes** in under 2 days. The ecosystem support includes **vLLM** integration to avoid tokenization drift. Discussions highlight that **harnesses** are becoming critical as the product surface for coding agents, with **task-specialized harnesses** favored over generic wrappers. Benchmarks are shifting focus from token price to **cost per task**, with models like **Terra Max**, **Fable 5 Max**, and **Opus 4.8** compared on efficiency and cost. Real-world agent benchmarks show **GPT-5.6 Sol** ranking #2 and **Grok-4.5** jumping to #13 on Arena&apos;s leaderboard, emphasizing cost per task as a key metric for long-horizon knowledge work.</description><pubDate>Mon, 13 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/11/2026-7/13/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Agent RL Infrastructure: Prime Intellect’s Verifiers v1 and Long-Horizon Rollouts&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Prime Intellect’s verifiers v1&lt;/strong&gt;: &lt;a href=&quot;https://x.com/PrimeIntellect/status/2076447247693402301&quot;&gt;Prime Intellect&lt;/a&gt; released &lt;strong&gt;verifiers v1&lt;/strong&gt;, a substantial redesign of its environment stack for &lt;strong&gt;agentic RL and evals&lt;/strong&gt;. The key abstraction splits environments into a &lt;strong&gt;taskset, harness, and runtime&lt;/strong&gt;, explicitly supporting “bring your own harness” workflows for coding and computer-use agents across heterogeneous execution setups, as highlighted by &lt;a href=&quot;https://x.com/johannes_hage/status/2076447852528889939&quot;&gt;Johannes Hage&lt;/a&gt; and in a &lt;a href=&quot;https://x.com/johannes_hage/status/2076449075621462457&quot;&gt;follow-up deep dive&lt;/a&gt;. The release was framed by team members as months of infra modernization work with major efficiency gains, including richer commentary from &lt;a href=&quot;https://x.com/willccbb/status/2076449433483616346&quot;&gt;willccbb&lt;/a&gt;, &lt;a href=&quot;https://x.com/mikasenghaas/status/2076507323561021779&quot;&gt;mikasenghaas&lt;/a&gt;, and &lt;a href=&quot;https://x.com/xeophon/status/2076509926256422947&quot;&gt;xeophon&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Why it matters technically&lt;/strong&gt;: one of the most important underlying changes is that rollout traces are now stored as &lt;strong&gt;message DAGs&lt;/strong&gt;, so each message is stored once instead of repeatedly copied into full histories; that shifts trace growth from &lt;strong&gt;O(n²)&lt;/strong&gt; to &lt;strong&gt;O(n)&lt;/strong&gt; in turn count, making long-horizon multimodal rollouts and router replay much more practical, per &lt;a href=&quot;https://x.com/PrimeIntellect/status/2076447253938786648&quot;&gt;Prime Intellect&lt;/a&gt;. The team also claimed a concrete training configuration: a &lt;strong&gt;100B reasoning model&lt;/strong&gt;, on &lt;strong&gt;40-turn SWE agent tasks&lt;/strong&gt;, in a user-supplied coding harness, for &lt;strong&gt;1000 RL steps&lt;/strong&gt;, using &lt;strong&gt;6 H200 nodes&lt;/strong&gt; in &lt;strong&gt;under 2 days&lt;/strong&gt; (&lt;a href=&quot;https://x.com/willccbb/status/2076451043504967783&quot;&gt;willccbb&lt;/a&gt;). That claim was reinforced by ecosystem support from &lt;a href=&quot;https://x.com/vllm_project/status/2076528386927997249&quot;&gt;vLLM&lt;/a&gt;, which noted verifiers’ rollout path runs on vLLM with exact token IDs/logprobs to avoid tokenization drift between serving and training.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Coding Agents, Harness Design, and Cost-Per-Task Competition&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Harnesses are becoming the product surface&lt;/strong&gt;: several posts converged on the idea that model quality is no longer the only differentiator; the &lt;strong&gt;harness/orchestrator&lt;/strong&gt; increasingly determines outcomes. &lt;a href=&quot;https://x.com/localfirstconf/status/2076678392615682215&quot;&gt;threepointone’s talk&lt;/a&gt; was summarized as “the harness is the app,” while &lt;a href=&quot;https://x.com/hwchase17/status/2076784403414651035&quot;&gt;LangChain&lt;/a&gt; argued that winning agent products will come from &lt;strong&gt;task-specialized harnesses&lt;/strong&gt;, not generic wrappers. &lt;a href=&quot;https://x.com/FactoryAI/status/2076710400729731349&quot;&gt;Factory&lt;/a&gt; pushed a related UI angle with “design mode,” where users point at UI elements/files instead of verbally re-specifying edits. On the orchestration side, &lt;a href=&quot;https://x.com/omarsar0/status/2076720090549035318&quot;&gt;omarsar0&lt;/a&gt; emphasized provider-switching across models as a hedge against pricing/policy churn.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmarks are moving from token price to cost per task&lt;/strong&gt;: &lt;a href=&quot;https://x.com/skirano/status/2076456519810580681&quot;&gt;skirano&lt;/a&gt; built a coding-agent index explorer and found notable cost/perf tradeoffs such as &lt;strong&gt;Terra Max slightly ahead of Fable 5 Max&lt;/strong&gt; on score for materially lower cost, while &lt;a href=&quot;https://x.com/cognition/status/2076714965344342382&quot;&gt;Cognition&lt;/a&gt; reported that &lt;strong&gt;Devin Fusion&lt;/strong&gt; now uses &lt;strong&gt;Fable 5&lt;/strong&gt; and that, surprisingly, it can be &lt;strong&gt;lower cost per task than Opus 4.8&lt;/strong&gt; because stronger delegation and judgment reduce unnecessary work. &lt;a href=&quot;https://x.com/imjaredz/status/2076715750715482162&quot;&gt;imjaredz&lt;/a&gt; highlighted the key stat from those experiments: in &lt;strong&gt;81% of Fable-led runs&lt;/strong&gt;, the lead model never makes a code edit, implying expensive models can be cheaper when they avoid wasted actions.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real-world agent benchmarks are getting denser&lt;/strong&gt;: &lt;a href=&quot;https://x.com/arena/status/2076709326711037991&quot;&gt;Arena&lt;/a&gt; placed &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; at &lt;strong&gt;#2&lt;/strong&gt; on its agent leaderboard based on &lt;strong&gt;7.8K real-world agentic sessions&lt;/strong&gt;, with strong steerability and task success; later, &lt;a href=&quot;https://x.com/arena/status/2076728509813469536&quot;&gt;Arena&lt;/a&gt; put &lt;strong&gt;Grok-4.5&lt;/strong&gt; at &lt;strong&gt;#13&lt;/strong&gt;, a significant jump over Grok 4.3. &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2076791491071295708&quot;&gt;Artificial Analysis&lt;/a&gt; also emphasized &lt;strong&gt;cost per task&lt;/strong&gt; as an increasingly important metric for long-horizon knowledge work, arguing token pricing alone misses effects from turns, verbosity, and cache hit rates. Separate evaluation work from &lt;a href=&quot;https://x.com/doesdatmaksense/status/2076642415767965701&quot;&gt;Parlance Labs&lt;/a&gt; compared automated eval platforms and foundation models on failure analysis over production voice-agent traces, while &lt;a href=&quot;https://x.com/dair_ai/status/2076699431207154069&quot;&gt;dair.ai&lt;/a&gt; highlighted a paper on the &lt;strong&gt;anatomy of CLI coding-agent failures&lt;/strong&gt;, focusing on where runs become unrecoverable rather than only final pass/fail.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;OpenAI GPT-5.6 Sol, Codex Usage Fixes, and Product Surface Expansion&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI addressed Codex/Sol usage burn transparently&lt;/strong&gt;: the biggest operational thread came from &lt;a href=&quot;https://x.com/thsottiaux/status/2076495156757577895&quot;&gt;thsottiaux&lt;/a&gt;, who explained several fixes for &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; in ChatGPT Work/Codex: inference optimizations yielding roughly &lt;strong&gt;10% more usage&lt;/strong&gt;, a rollback of context limit from &lt;strong&gt;372k&lt;/strong&gt; to &lt;strong&gt;272k&lt;/strong&gt; after billing/usage side effects, reversion of some experimental reasoning-effort (“&lt;strong&gt;juice&lt;/strong&gt;”) changes, and fixes for overactive multi-agent behavior at high/xhigh settings. Community reverse-engineering from &lt;a href=&quot;https://x.com/theo/status/2076512403668488299&quot;&gt;theo&lt;/a&gt; proposed that compounding factors around long context, subagent spawning, and fast mode were behind the severe burn, though he later corrected one billing detail in a &lt;a href=&quot;https://x.com/theo/status/2076543971216830551&quot;&gt;follow-up&lt;/a&gt;. Reactions split between criticism of a perceived “nerf” narrative (&lt;a href=&quot;https://x.com/ns123abc/status/2076498300312703349&quot;&gt;ns123abc&lt;/a&gt;) and praise for unusual transparency (&lt;a href=&quot;https://x.com/theo/status/2076501402822775267&quot;&gt;theo&lt;/a&gt;, &lt;a href=&quot;https://x.com/sama/status/2076696938918084809&quot;&gt;sama&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Users are reporting strong coding/computer-use capability&lt;/strong&gt;: multiple practitioners argued that &lt;strong&gt;OpenAI has taken the lead on coding models&lt;/strong&gt;, including &lt;a href=&quot;https://x.com/schrockn/status/2076488446961709218&quot;&gt;schrockn&lt;/a&gt;, while &lt;a href=&quot;https://x.com/gdb/status/2076518764112445861&quot;&gt;gdb&lt;/a&gt; repeatedly showcased &lt;strong&gt;ChatGPT Work&lt;/strong&gt; and Codex workflows for startup prospecting, web design, mobile work, and site generation. Particularly illustrative user demos included &lt;a href=&quot;https://x.com/Star_Knight12/status/2076631428926972177&quot;&gt;Star_Knight12&lt;/a&gt; using &lt;strong&gt;Sol in Cursor&lt;/strong&gt; to set up Blender MCP and render a floating MacBook without prior Blender experience, and &lt;a href=&quot;https://x.com/petergostev/status/2076692164310884468&quot;&gt;petergostev&lt;/a&gt; showing &lt;strong&gt;GPT-5.6 Sol Ultra&lt;/strong&gt; building a &lt;strong&gt;Doom-like game in SQL&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Product-level expansion continues&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ChatGPTapp/status/2076654365121855835&quot;&gt;ChatGPTapp&lt;/a&gt; announced ChatGPT’s return to &lt;strong&gt;WhatsApp in the EEA&lt;/strong&gt;, plus Kakao/Viber support in additional markets. &lt;a href=&quot;https://x.com/OpenAIDevs/status/2076715478878474575&quot;&gt;OpenAIDevs&lt;/a&gt; opened submissions for &lt;strong&gt;OpenAI Build Week&lt;/strong&gt;. Across the OpenAI ecosystem, &lt;a href=&quot;https://x.com/gdb/status/2076685930002538875&quot;&gt;gdb&lt;/a&gt; summarized the moment succinctly: “you can just create things.”&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open Models, Inference Systems, and Quantization&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Transformers↔vLLM integration removes duplicated model implementation work&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClementDelangue/status/2076763231788339669&quot;&gt;Clement Delangue&lt;/a&gt; highlighted a major open-inference usability improvement: &lt;strong&gt;Hugging Face Transformers models can now run in vLLM at native speed&lt;/strong&gt;, often matching or exceeding hand-written implementations. If this generalizes broadly, it reduces the long-standing burden of implementing each new architecture twice—once for research/training and once for high-performance serving—and could materially accelerate adoption of new open model architectures.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Quantization remains a major lever&lt;/strong&gt;: &lt;a href=&quot;https://x.com/waterloo_intern/status/2076460984475263401&quot;&gt;waterloo_intern&lt;/a&gt; previewed a new quantization method claimed to beat existing approaches, including NVIDIA’s ModelOpt, by finding better layerwise precision assignments &lt;strong&gt;faster&lt;/strong&gt;, with &lt;strong&gt;more aggressive quantization&lt;/strong&gt; and &lt;strong&gt;higher benchmark scores&lt;/strong&gt;. Complementing that, &lt;a href=&quot;https://x.com/UnslothAI/status/2076665500294394109&quot;&gt;Unsloth&lt;/a&gt; published an AWS guide to &lt;strong&gt;LLM quantization and deployment&lt;/strong&gt; spanning GGUF, NVFP4, and FP8. There was also practitioner commentary around &lt;strong&gt;fp4 RL / fp4 serving&lt;/strong&gt; from &lt;a href=&quot;https://x.com/nrehiew_/status/2076654135559233857&quot;&gt;nrehiew_&lt;/a&gt;, arguing low-bit post-training may enable cheap serving with limited quality loss.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GLM-5.2 and local/open coding stacks continue to gain traction&lt;/strong&gt;: several users described moving real workflows onto open or semi-open setups. &lt;a href=&quot;https://x.com/juanjucm/status/2076714987569963508&quot;&gt;juanjucm&lt;/a&gt; wrote up using &lt;strong&gt;GLM-5.2&lt;/strong&gt; for coding-agent workflows, while &lt;a href=&quot;https://x.com/TheZachMueller/status/2076746035758502275&quot;&gt;TheZachMueller&lt;/a&gt; reported migrating one actual work pipeline from Claude to a stack built around &lt;strong&gt;GLM 5.2 NVFP4&lt;/strong&gt; plus &lt;strong&gt;Kimi K2.7 Code NVFP4&lt;/strong&gt; on an &lt;strong&gt;8xB200&lt;/strong&gt; node, getting denser reports for pennies albeit at slower wall-clock latency. &lt;a href=&quot;https://x.com/nutlope/status/2076722464671793184&quot;&gt;nutlope&lt;/a&gt; also released &lt;strong&gt;LlamaCoder v4&lt;/strong&gt;, rebuilt around GLM 5.2.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Security, Privacy, and Data Control in Agent Tooling&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Grok Build code upload controversy&lt;/strong&gt;: the most consequential security story came from &lt;a href=&quot;https://x.com/IntCyberDigest/status/2076689215258014069&quot;&gt;IntCyberDigest&lt;/a&gt; and &lt;a href=&quot;https://x.com/hrkrshnn/status/2076716354754015368&quot;&gt;hrkrshnn&lt;/a&gt;, who alleged that &lt;strong&gt;xAI’s Grok Build CLI&lt;/strong&gt; was uploading entire repositories—including private code and secrets—to a Google Cloud bucket, far beyond what was needed for the coding task. The criticism centered on scope, silent server-side mitigation, and unclear retention/deletion guarantees. This triggered broader discussion about what agent tools actually transmit and why opt-out UX can diverge from wire-level behavior.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;xAI’s response emphasized ZDR and privacy controls&lt;/strong&gt;: &lt;a href=&quot;https://x.com/SpaceXAI/status/2076692402442846289#m&quot;&gt;SpaceXAI&lt;/a&gt; replied that for teams using &lt;strong&gt;zero data retention&lt;/strong&gt;, trace and code data is not retained, API key use respects ZDR, and the &lt;code&gt;/privacy&lt;/code&gt; command can disable retention and delete previously synced data. That answered some operational questions but did not fully resolve community concern around default behavior, prior uploads, and disclosure norms.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Trust boundaries are becoming a central open-vs-closed argument&lt;/strong&gt;: several posts extended the conversation beyond this incident. &lt;a href=&quot;https://x.com/mchiang0610/status/2076736707471556755&quot;&gt;mchiang0610&lt;/a&gt; and &lt;a href=&quot;https://x.com/jmorgan/status/2076750580052369896&quot;&gt;jmorgan&lt;/a&gt; argued that open models are not just about cost but about &lt;strong&gt;control over the human-AI learning loop&lt;/strong&gt; and keeping institutional knowledge in-house. &lt;a href=&quot;https://x.com/AravSrinivas/status/2076699450177892354&quot;&gt;Arav Srinivas&lt;/a&gt; said &lt;strong&gt;ZDR availability&lt;/strong&gt; was one reason Perplexity integrated &lt;strong&gt;Grok 4.5&lt;/strong&gt; quickly into its Computer harness.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Continual Learning, Multimodal Systems, and Research Directions&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Continual learning is re-emerging as a first-class systems problem&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ysu_nlp/status/2076481232117067894&quot;&gt;ysu_nlp&lt;/a&gt; argued that a world where every organization owns its own human-AI learning loop depends on solving &lt;strong&gt;continual learning&lt;/strong&gt;, and that current approaches—memory/RAG, domain post-training, task RL—are not yet sufficient. That theme recurred in new work from &lt;a href=&quot;https://x.com/skyfallai/status/2076713589788864920&quot;&gt;skyfallai&lt;/a&gt;, which introduced &lt;strong&gt;Morpheus&lt;/strong&gt;, described as a persistent enterprise simulation for real-world RL where the world does not reset; &lt;a href=&quot;https://x.com/fchollet/status/2076719958189613307&quot;&gt;fchollet&lt;/a&gt; endorsed it as a benchmark better aligned with real deployment than stationary episodic RL.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;“Sleep and dreaming” for LLMs&lt;/strong&gt;: &lt;a href=&quot;https://x.com/behrouz_ali/status/2076710744456892519&quot;&gt;behrouz_ali&lt;/a&gt; and coauthors proposed that LLMs may need a &lt;strong&gt;sleep phase&lt;/strong&gt; to consolidate short-term into long-term memory plus a &lt;strong&gt;dreaming phase&lt;/strong&gt; for recursive self-improvement, introducing &lt;strong&gt;Knowledge Seeding&lt;/strong&gt; and reporting benefits on continual learning/reasoning tasks. This dovetails with broader dissatisfaction around current continual-learning recipes and with &lt;a href=&quot;https://x.com/kjaved_/status/2076663868160459214&quot;&gt;Oak Lab&lt;/a&gt;, the new venture from Rich Sutton and collaborators pursuing &lt;strong&gt;animal-like intelligence&lt;/strong&gt; that learns from experience rather than today’s standard LLM pipeline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A broad spread of non-LLM-agent research shipped&lt;/strong&gt;: notable items included &lt;a href=&quot;https://x.com/SakanaAILabs/status/2076597965804765283&quot;&gt;Sakana AI’s Smart Cellular Bricks&lt;/a&gt; for decentralized physical self-recognition and repair in modular systems; &lt;a href=&quot;https://x.com/HuggingPapers/status/2076513044340097501&quot;&gt;ByteDance’s UniVR-34B&lt;/a&gt;, described as learning reasoning/dynamics/planning directly from visual demonstrations; &lt;a href=&quot;https://x.com/GoogleDeepMind/status/2076686114631340046&quot;&gt;Google DeepMind’s Predicting the Past skill&lt;/a&gt; for historical inference workflows; and &lt;a href=&quot;https://x.com/AnthropicAI/status/2076719540785012872&quot;&gt;Anthropic’s research&lt;/a&gt; on how &lt;strong&gt;Claude’s expressed values&lt;/strong&gt; vary across models and languages based on analysis of &lt;strong&gt;300K+ anonymized conversations&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI Codex/Sol usage fixes&lt;/strong&gt;: &lt;a href=&quot;https://x.com/thsottiaux/status/2076495156757577895&quot;&gt;thsottiaux on GPT-5.6 Sol usage, context, “juice,” and multi-agent fixes&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grok Build privacy incident&lt;/strong&gt;: &lt;a href=&quot;https://x.com/IntCyberDigest/status/2076689215258014069&quot;&gt;IntCyberDigest on full-repo uploads to xAI cloud buckets&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI response tone and user treatment&lt;/strong&gt;: &lt;a href=&quot;https://x.com/sama/status/2076780425280954658&quot;&gt;sama: “come for the best model, stay because we don’t treat you with contempt”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prime Intellect rollout efficiency&lt;/strong&gt;: &lt;a href=&quot;https://x.com/willccbb/status/2076451043504967783&quot;&gt;willccbb on training a 100B reasoning model for 40-turn SWE RL on 6 H200s in under 2 days&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Anthropic values research&lt;/strong&gt;: &lt;a href=&quot;https://x.com/AnthropicAI/status/2076719540785012872&quot;&gt;Anthropic on model/language-dependent value expression across 300K+ conversations&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transformers + vLLM interoperability&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClementDelangue/status/2076763231788339669&quot;&gt;Clement Delangue on running Transformers models in vLLM at native speed&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. E-Waste GPU Inference Benchmarks and Fixes&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uvcjd0/i_benchmarked_15_ewaste_gpus_with_modern_workloads/&quot;&gt;I benchmarked 15 &quot;E-Waste&quot; GPUs with Modern Workloads&lt;/a&gt;&lt;/strong&gt; (Activity: 462): &lt;strong&gt;A year-long homelab benchmark tested decommissioned NVIDIA Tesla GPUs (K80/M10/M40/M60/P40/P100/V100/T40) using a custom Dockerized suite (&lt;a href=&quot;https://github.com/esologic/gpu_box_benchmark&quot;&gt;&lt;code&gt;gpu_box_benchmark&lt;/code&gt;&lt;/a&gt;) across LLMs, CV, Blender, Whisper, and related workloads, with full graphs on the author’s &lt;a href=&quot;https://esologic.com/benchmarking-tesla-gpus/&quot;&gt;blog&lt;/a&gt;. Key findings: &lt;strong&gt;V100 16GB&lt;/strong&gt; was the best overall value and approached &lt;strong&gt;T40&lt;/strong&gt; performance, &lt;strong&gt;P40 outperformed P100 for LLMs&lt;/strong&gt;, &lt;strong&gt;M60 was unexpectedly strong for Whisper&lt;/strong&gt;, multi-GPU scaling was roughly linear in a 4U chassis, and cheap &lt;strong&gt;X99 + Xeon&lt;/strong&gt; platforms generally fed the cards adequately despite EOL software/power-efficiency caveats.&lt;/strong&gt; Commenters questioned whether the benchmark really targets “modern” workloads, arguing that small models and ResNet-style tests do not exercise the main value proposition: cheap pooled VRAM for larger models. Requested follow-ups included power/noise measurements and LLM serving metrics such as prompt processing/token generation at long context lengths for models like &lt;strong&gt;Qwen 3.x 27B/35B MoE&lt;/strong&gt; across multiple V100/P40-class cards.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued the benchmark suite did not represent current high-VRAM use cases: &lt;strong&gt;ResNet&lt;/strong&gt; and small models were described as insufficient for evaluating “cheap VRAM” GPUs. They requested tests with larger modern LLMs such as &lt;strong&gt;Qwen 3.6 27B/31B MoE/35B A3B&lt;/strong&gt;, including whether pooled VRAM configurations can run them, and asked for &lt;strong&gt;prompt-processing (PP)&lt;/strong&gt; and &lt;strong&gt;token-generation (TG)&lt;/strong&gt; throughput at long context lengths such as &lt;code&gt;150k ctx&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A technical correction noted that the &lt;strong&gt;Tesla P100&lt;/strong&gt; should normally outperform the &lt;strong&gt;Tesla P40&lt;/strong&gt; unless a relevant &lt;code&gt;fp32&lt;/code&gt; patch has changed behavior, because the P100’s &lt;strong&gt;HBM bandwidth is roughly 3× higher&lt;/strong&gt; than the P40’s memory bandwidth. This implies memory-bound workloads may be misrepresented if the benchmark shows the P40 ahead without explaining software/kernel differences.&lt;/li&gt;
&lt;li&gt;One commenter suggested adding the &lt;strong&gt;P102-100&lt;/strong&gt; mining GPU, which is currently available around &lt;code&gt;$50&lt;/code&gt;, has relatively low idle power around &lt;code&gt;10 W&lt;/code&gt;, and is reportedly easy to cool. They claimed it reaches about &lt;code&gt;40 tokens/s&lt;/code&gt; generation on &lt;strong&gt;Qwen 3.6 35B&lt;/strong&gt;, but with very slow prompt processing at around &lt;code&gt;100&lt;/code&gt; tokens/s, making it an interesting but bottlenecked e-waste inference option.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uu6p9o/your_80_tesla_p100_has_been_doing_silently_noisy/&quot;&gt;&lt;strong&gt;Your $80 Tesla P100 has been doing silently noisy math in llama.cpp for years. Three lines fix it, for free.&lt;/strong&gt;&lt;/a&gt;&lt;/strong&gt; (Activity: 426): &lt;strong&gt;A 3-line CUDA arch-gating patch for &lt;strong&gt;llama.cpp/turboquant&lt;/strong&gt; changes &lt;code&gt;sm_60&lt;/code&gt; &lt;strong&gt;Tesla P100&lt;/strong&gt; handling to match the existing &lt;code&gt;sm_61&lt;/code&gt; Pascal exemption, avoiding a &quot;fast fp16&quot; path that reportedly increased logit noise without improving throughput; released in &lt;a href=&quot;https://github.com/TheTom/llama-cpp-turboquant/releases/tag/tqp-v0.3.0&quot;&gt;&lt;code&gt;llama-cpp-turboquant v0.3.0&lt;/code&gt;&lt;/a&gt;, merged in &lt;a href=&quot;https://github.com/TheTom/llama-cpp-turboquant/pull/212&quot;&gt;&lt;code&gt;TheTom/llama-cpp-turboquant#212&lt;/code&gt;&lt;/a&gt; and &lt;a href=&quot;https://github.com/spiritbuun/buun-llama-cpp/pull/80&quot;&gt;&lt;code&gt;spiritbuun/buun-llama-cpp#80&lt;/code&gt;&lt;/a&gt;, with upstream tracking in &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/issues/25593&quot;&gt;&lt;code&gt;ggml-org/llama.cpp#25593&lt;/code&gt;&lt;/a&gt;. The author reports, vs fp32-reference logits on &lt;strong&gt;Qwen3.6-27B / WikiText-2&lt;/strong&gt;, median KLD improving from &lt;code&gt;0.0023&lt;/code&gt; to &lt;code&gt;0.000001&lt;/code&gt; (~&lt;code&gt;2300×&lt;/code&gt;) and top-token agreement from &lt;code&gt;96.5%&lt;/code&gt; to &lt;code&gt;99.9%&lt;/code&gt;, with prefill unchanged and decode ~&lt;code&gt;1.4%&lt;/code&gt; faster at 8k context; a commenter independently patched a P100 and saw mean KLD &lt;code&gt;0.0122 → 0.000000&lt;/code&gt; and top-token match &lt;code&gt;95.09% → 99.997%&lt;/code&gt;. The claimed scope is specifically &lt;strong&gt;Pascal &lt;code&gt;sm_60&lt;/code&gt; P100&lt;/strong&gt;: GTX 10-series/P40 &lt;code&gt;sm_61&lt;/code&gt; were already exempt, while Volta+ use different kernels and are claimed unaffected, with a Blackwell control reportedly showing bit-identical perplexity/decode behavior.&lt;/strong&gt; Comments were mostly supportive, framing this as a small but meaningful correctness fix; one commenter used an LLM to decode the technical claim and concluded it was plausibly a real accuracy improvement with no practical speed cost.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter tested the patch and reported a large numerical-accuracy improvement on &lt;strong&gt;Tesla P100&lt;/strong&gt;: stock llama.cpp CUDA produced &lt;code&gt;mean KLD = 0.0122&lt;/code&gt; with &lt;code&gt;95.09%&lt;/code&gt; same top-token, while the patched/control path produced &lt;code&gt;mean KLD = 0.000000&lt;/code&gt; with &lt;code&gt;99.997%&lt;/code&gt; same top-token. This supports the claim that disabling the fp16 fast-math path for &lt;code&gt;sm_60&lt;/code&gt; removes distribution-level noise without changing model behavior unpredictably.&lt;/li&gt;
&lt;li&gt;Another commenter summarized the technical mechanism: llama.cpp’s CUDA backend enables a fast fp16 math mode for GPUs classified as having strong fp16 throughput; &lt;code&gt;sm_61&lt;/code&gt; cards like GTX 10-series/P40 were already excluded, but &lt;code&gt;sm_60&lt;/code&gt; &lt;strong&gt;P100&lt;/strong&gt; was not. The claim is that P100’s real-world inference is memory/GEMM-bound rather than fp16-vector-unit-bound, so the fp16 path adds quantization-like numerical error without measurable speedup; the proposed 3-line patch reportedly cuts KL divergence vs fp32 by about &lt;code&gt;2300x&lt;/code&gt; with no speed loss.&lt;/li&gt;
&lt;li&gt;One P100 user with a &lt;code&gt;3x P100&lt;/code&gt; setup planned to test the patch in their llama.cpp build and mentioned prior experimentation with &lt;strong&gt;Qwen 3 27B&lt;/strong&gt;, quantization behavior, and MTP. This suggests interest in validating whether the fix generalizes across multi-GPU P100 inference and different quantization/model configurations.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Chinese AI Stack: Usage, Weights, Chips&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uu15mz/chinas_deepseek_developing_its_own_ai_chip/&quot;&gt;China&apos;s DeepSeek developing its own AI chip, sources say&lt;/a&gt;&lt;/strong&gt; (Activity: 576): &lt;strong&gt;Sources reportedly say &lt;strong&gt;DeepSeek&lt;/strong&gt; is developing an in-house AI accelerator, likely as a response to restricted access to &lt;strong&gt;Nvidia&lt;/strong&gt; GPUs in China and the need for domestic training/inference hardware. The key technical constraint raised in comments is not just chip design but access to &lt;strong&gt;leading-edge semiconductor manufacturing&lt;/strong&gt;; one quoted view argues &lt;em&gt;“Nvidia is at zero in China”&lt;/em&gt; while DeepSeek has little chance outside China without advanced fabs.&lt;/strong&gt; Commenters were broadly pro-competition, but one technical take argued that a high-memory consumer accelerator—e.g. &lt;code&gt;&gt;32GB&lt;/code&gt; VRAM and &lt;code&gt;&gt;1TB/s&lt;/code&gt; bandwidth under &lt;code&gt;$5k&lt;/code&gt;—would sell even if inefficient or built from awkward memory configurations.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter highlighted the manufacturing and market-access constraint: without access to leading-edge fabs, &lt;strong&gt;DeepSeek would likely struggle to sell competitive AI silicon outside China&lt;/strong&gt;, while Nvidia’s position in China is described as effectively constrained by export controls. Another technical angle was consumer demand for high-memory-bandwidth accelerators: a hypothetical card with &lt;code&gt;&gt;32GB&lt;/code&gt; memory and &lt;code&gt;&gt;1TB/s&lt;/code&gt; bandwidth under &lt;code&gt;$5k&lt;/code&gt; was argued to be attractive even if implemented inefficiently, e.g. with many DDR channels and &lt;code&gt;~800W&lt;/code&gt; power draw.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1uuyw46/chinese_ai_models_seize_openrouters_top_five_as/&quot;&gt;Chinese AI Models Seize OpenRouter’s Top Five as OpenAI and Google Vanish From the Top 10&lt;/a&gt;&lt;/strong&gt; (Activity: 561): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/o8g1mxm1rwch1.jpeg&quot;&gt;image&lt;/a&gt; is a technical dashboard screenshot of &lt;strong&gt;OpenRouter’s AI Model Rankings&lt;/strong&gt;, showing monthly token-usage share where Chinese-affiliated models occupy the top five positions and &lt;code&gt;7/10&lt;/code&gt; of the displayed top 10. The chart shows OpenRouter usage rising sharply toward late June, reaching roughly &lt;code&gt;60T&lt;/code&gt; weekly tokens, with &lt;strong&gt;DeepSeek&lt;/strong&gt;, &lt;strong&gt;MiMo&lt;/strong&gt;, &lt;strong&gt;MiniMax&lt;/strong&gt;, and &lt;strong&gt;Hy3&lt;/strong&gt; models ahead of Western frontier models; &lt;strong&gt;Anthropic Claude&lt;/strong&gt; appears in positions 6 and 8, while &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Google&lt;/strong&gt; are absent. The significance is platform-specific: OpenRouter says this reflects real usage by its users, but it measures &lt;strong&gt;OpenRouter traffic&lt;/strong&gt;, not global LLM adoption.&lt;/strong&gt; Commenters framed the ranking less as a pure capability benchmark and more as evidence of cost/practicality: &lt;em&gt;“It’s hard to compare benchmarks, but easy to compare bills.”&lt;/em&gt; Others argued open/source-available models are attractive because users can test via OpenRouter and later self-host, while lower electricity costs in China may contribute to pricing competitiveness.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed OpenRouter as a practical model-selection layer: test multiple open/source-available models behind a common API, then either continue routing through OpenRouter or self-host the winning model if unit economics justify it. The key technical concern raised was operational stability: users distrust &lt;strong&gt;OpenAI/Anthropic&lt;/strong&gt; because pricing, model behavior, and model availability can change abruptly, making reproducibility and long-term deployment planning harder.&lt;/li&gt;
&lt;li&gt;Cost was treated as a more actionable metric than benchmarks: one commenter noted that &lt;em&gt;“it’s hard to compare benchmarks, but easy to compare bills,”&lt;/em&gt; while another claimed &lt;strong&gt;&lt;code&gt;deepseek-v4-flash&lt;/code&gt;&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;mimo-v2.5&lt;/code&gt;&lt;/strong&gt; are cheap enough on OpenRouter that inference costs are lower than the electricity cost of self-hosting, before even accounting for hardware capex. Another commenter argued China’s lower electricity prices materially affect inference economics, especially compared with proposed US datacenter locations such as California.&lt;/li&gt;
&lt;li&gt;A commenter suggested OpenRouter rankings may underrepresent &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Google Gemini&lt;/strong&gt; usage because many customers access those models directly from the providers rather than through an aggregator. This implies OpenRouter’s top-model distribution is more reflective of aggregator-native demand and price/performance experimentation than total market share across all access channels.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uu8d1v/xiaomi_quietly_uploaded_mimov25dflash_official/&quot;&gt;Xiaomi quietly uploaded MiMo-V2.5-DFlash — official DFlash weights are now on Hugging Face&lt;/a&gt;&lt;/strong&gt; (Activity: 389): &lt;strong&gt;&lt;strong&gt;Xiaomi&lt;/strong&gt; has uploaded official &lt;strong&gt;MiMo-V2.5-DFlash&lt;/strong&gt; weights to Hugging Face at &lt;a href=&quot;https://huggingface.co/XiaomiMiMo/MiMo-V2.5-DFlash&quot;&gt;&lt;code&gt;XiaomiMiMo/MiMo-V2.5-DFlash&lt;/code&gt;&lt;/a&gt;, including a dedicated &lt;code&gt;dflash/&lt;/code&gt; directory and a separate MTP model. The poster reports the non-DFlash MiMo-V2.5-class model as &lt;code&gt;300B+&lt;/code&gt; params running around &lt;code&gt;8–10 tok/s&lt;/code&gt; on &lt;code&gt;2×24GB&lt;/code&gt; GPUs with RAM/VRAM offload, and speculates DFlash could roughly double throughput; they also note &lt;code&gt;llama.cpp&lt;/code&gt; currently struggles to identify/use the MTP layers, while the separate DFlash/MTP artifacts may be easier to support.&lt;/strong&gt; Commenters characterize MiMo 2.5 as “incredible and underrated,” but one benchmark-related claim was corrected: a comparison placing it between DeepSeek V4 Flash and Pro on SWE-rebench was actually referring to &lt;strong&gt;MiMo Pro&lt;/strong&gt; (&lt;code&gt;1T&lt;/code&gt;, &lt;code&gt;A42B&lt;/code&gt;), not this roughly &lt;code&gt;284B&lt;/code&gt; Flash/DFlash-sized model.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter initially compared &lt;strong&gt;MiMo-V2.5-DFlash&lt;/strong&gt; to &lt;strong&gt;DeepSeek V4 Flash/Pro&lt;/strong&gt; on &lt;code&gt;swe-rebench&lt;/code&gt;, claiming it fell between them in price/performance despite being &lt;code&gt;284B&lt;/code&gt; rather than &lt;code&gt;1.6T&lt;/code&gt;, but later corrected this as a confusion with &lt;strong&gt;MiMo Pro&lt;/strong&gt;, described as &lt;code&gt;1T A42B&lt;/code&gt;. The useful takeaway is that benchmark/performance claims around MiMo variants may be easy to misattribute because &lt;strong&gt;DFlash&lt;/strong&gt;, &lt;strong&gt;Flash-sized&lt;/strong&gt;, and &lt;strong&gt;Pro&lt;/strong&gt; model naming overlap.&lt;/li&gt;
&lt;li&gt;There is interest in measuring real-world &lt;code&gt;tok/s&lt;/code&gt; gains once &lt;strong&gt;DFlash&lt;/strong&gt; support lands in &lt;code&gt;llama.cpp&lt;/code&gt;, especially under GGUF/local inference conditions. A commenter cautioned that speculative-decoding-style speedups may degrade when VRAM offload is insufficient and inference spills into system RAM, so practical benchmarks will need to separate ideal accelerator throughput from mixed VRAM/RAM execution.&lt;/li&gt;
&lt;li&gt;A commenter asked whether &lt;strong&gt;DFlash&lt;/strong&gt; is closer to an &lt;strong&gt;MTP/speculative decoding mechanism&lt;/strong&gt; that preserves the base model’s output distribution while accelerating generation, rather than a separate “Flash” model in the sense of a smaller/lighter distilled variant. This distinction matters for interpreting the released weights: DFlash may be an inference-speed augmentation rather than a reduced-capacity model family member.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Local AI Runtime and Visualization Experiments&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uuga40/local_image_to_3d_2gb_ram_20s_apple_silicon_iphone/&quot;&gt;Local Image to 3D (&amp;#x3C;2gb RAM, &amp;#x3C;20s, Apple Silicon, iPhone)&lt;/a&gt;&lt;/strong&gt; (Activity: 990): &lt;strong&gt;The linked GIF (&lt;a href=&quot;https://i.redd.it/ywn3uzqs1tch1.gif&quot;&gt;image&lt;/a&gt;) appears to be a technical demo of &lt;strong&gt;Modelr&lt;/strong&gt;, an open-source Swift/MLX app that ports &lt;strong&gt;Hunyuan3D-Shape&lt;/strong&gt; and &lt;strong&gt;Hunyuan3D-Paint&lt;/strong&gt; for local image-to-3D generation on Apple Silicon and limited iOS. The author reports FP16 benchmarks on an &lt;strong&gt;M4 Max&lt;/strong&gt;: &lt;code&gt;hy3d shape&lt;/code&gt; in ~&lt;code&gt;21–22s&lt;/code&gt; using &lt;code&gt;5.6–7.3GB&lt;/code&gt; peak memory, while &lt;code&gt;hy3d paint&lt;/code&gt; is much heavier at &lt;code&gt;231–344s&lt;/code&gt; and ~&lt;code&gt;38–39GB&lt;/code&gt;; quantized Q4/Q8 runs are positioned as enabling lower-memory Mac/iPhone usage via MLX rather than PyTorch/CPU overhead.&lt;/strong&gt; The main technical caveat raised in comments is licensing: generated assets may be heavily restricted under the current &lt;strong&gt;Hunyuan3D&lt;/strong&gt; license, limiting commercial/practical use despite the open-source tooling. Other commenters were impressed that the heavier Paint stage runs locally at all, with one noting &lt;em&gt;“I didn&apos;t even think Paint would be possible.”&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter flagged that outputs may be heavily constrained by the &lt;strong&gt;Hunyuan3D&lt;/strong&gt; license, linking directly to Tencent’s &lt;a href=&quot;https://github.com/Tencent-Hunyuan/Hunyuan3D-2.1/blob/main/LICENSE&quot;&gt;&lt;code&gt;Hunyuan3D-2.1&lt;/code&gt; license&lt;/a&gt;. They noted that despite strong local image/text-to-3D tooling, the field is still limited by non-permissive “community” licenses, though they speculated &lt;strong&gt;Hunyuan3D-3&lt;/strong&gt; may move toward more permissive terms.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uu32z6/interactive_jacobianlens_visualizer_and_live/&quot;&gt;Interactive Jacobian-Lens visualizer and live steerer for GGUF models on llama.cpp&lt;/a&gt;&lt;/strong&gt; (Activity: 374): &lt;strong&gt;The image is a &lt;strong&gt;technical UI screenshot&lt;/strong&gt;, not a meme: it shows the &lt;a href=&quot;https://i.redd.it/cif34sq1mpch1.png&quot;&gt;J-Lens web interface&lt;/a&gt; for an interactive &lt;strong&gt;Jacobian-Lens visualizer/steerer for GGUF models on &lt;code&gt;llama.cpp&lt;/code&gt;&lt;/strong&gt;, demonstrated on &lt;code&gt;qwen2.5-1.5b-instruct&lt;/code&gt;. The project, &lt;a href=&quot;https://github.com/igorbarshteyn/jlens-gguf&quot;&gt;igorbarshteyn/jlens-gguf&lt;/a&gt;, adds a native GGUF server for observing models and performing &lt;em&gt;j-space swapping / abliteration / steering&lt;/em&gt;, with support for dense and MoE GGUFs; lens memory overhead is reported to scale at roughly &lt;code&gt;1/8&lt;/code&gt; of model size, e.g. ~&lt;code&gt;20 GB&lt;/code&gt; extra RAM for a &lt;code&gt;160 GB&lt;/code&gt; GGUF such as a large quantized Qwen model.&lt;/strong&gt; Commenters focused on possible extensions: merging original GGUF and lens tensors, using the tool to diagnose or repair heavily quantized models, and the implication that this could enable “targeted live adapters” or real-time steering workflows.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One technical request was for the tool to support &lt;strong&gt;merging the original GGUF with the Jacobian-lens tensors&lt;/strong&gt;, implying a desire for a self-contained GGUF artifact rather than a separate visualization/steering sidecar. This would likely require defining how lens tensors are serialized, named, and loaded within the existing &lt;code&gt;llama.cpp&lt;/code&gt; GGUF tensor/metadata conventions.&lt;/li&gt;
&lt;li&gt;A commenter suggested the Jacobian-lens approach might be useful for &lt;strong&gt;repairing heavily quantized models&lt;/strong&gt;, i.e. using live steering/adapters to compensate for behavior or representation damage introduced by aggressive GGUF quantization. Another raised a data requirement concern, asking whether a &lt;strong&gt;larger dataset is needed to map the lens properly&lt;/strong&gt;, which is relevant because learned or estimated Jacobian mappings may be brittle if calibrated on too little activation data.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uv66by/i_got_gemma_4_running_directly_inside_godot_using/&quot;&gt;I got Gemma 4 running directly inside Godot using only GDScript and Vulkan compute shaders&lt;/a&gt;&lt;/strong&gt; (Activity: 364): &lt;strong&gt;The image shows a &lt;strong&gt;Godot 4.7 debug chat UI&lt;/strong&gt; running a local GGUF LLM inside the engine, reporting about &lt;code&gt;46.99 tok/s&lt;/code&gt;, and is contextually ironic because the in-app model response says implementing GGUF loading/inference in &lt;strong&gt;GDScript + Vulkan compute shaders&lt;/strong&gt; would be “extremely complex” while the project demonstrates exactly that. Per the post, the experiment runs &lt;code&gt;gemma-4-E2B-it-Q4_K_M.gguf&lt;/code&gt; with Vulkan compute for model math and GDScript for GGUF loading, tokenization, sampling, KV cache, and UI, with code available at &lt;a href=&quot;https://github.com/asallay/godot-llm&quot;&gt;github.com/asallay/godot-llm&lt;/a&gt;; it is limited to one model and reportedly ~&lt;code&gt;10×&lt;/code&gt; slower than llama.cpp with CUDA. &lt;a href=&quot;https://i.redd.it/etqze9k9pych1.png&quot;&gt;Image&lt;/a&gt;&lt;/strong&gt; Comments are mostly impressed by the proof-of-concept rather than its speed, with one noting that avoiding native extensions, ABI issues, or a sidecar server could make local NPC/LLM demos easier to distribute as a single Godot export.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically substantive point is that the demo appears to implement &lt;strong&gt;GGUF loading, KV-cache management, and sampling entirely in Godot via GDScript + Vulkan compute shaders&lt;/strong&gt;, avoiding native-extension ABI issues or a separate inference server. One commenter argues that even at roughly &lt;strong&gt;&lt;code&gt;10x&lt;/code&gt; slower&lt;/strong&gt; performance, the deployment simplicity is significant because a single Godot export could make local LLM-driven NPC demos practically runnable by others.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Claude Fable 5 Access and Business-Model Backlash&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uu9egf/is_anthropic_shooting_themselves_in_the_foot_by/&quot;&gt;Is Anthropic shooting themselves in the foot by pulling Fab 5 from subscriptions tonight?&lt;/a&gt;&lt;/strong&gt; (Activity: 1656): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/w8wa19npbrch1.png&quot;&gt;image&lt;/a&gt; is a dark-themed benchmark bar chart for &lt;strong&gt;DeepSWE 1.0&lt;/strong&gt;, showing claimed coding-agent performance where &lt;strong&gt;“Fable/Fab 5 max” leads at &lt;code&gt;66.1%&lt;/code&gt;&lt;/strong&gt;, ahead of &lt;strong&gt;GPT 5.5 xhigh &lt;code&gt;64.31%&lt;/code&gt;&lt;/strong&gt;, &lt;strong&gt;Grok 4.5 &lt;code&gt;62.0%&lt;/code&gt;&lt;/strong&gt;, &lt;strong&gt;Opus 4.8 max &lt;code&gt;55.75%&lt;/code&gt;&lt;/strong&gt;, and &lt;strong&gt;Opus 4.7 max &lt;code&gt;40.12%&lt;/code&gt;&lt;/strong&gt;. In context, the post argues that if &lt;strong&gt;Anthropic removes Fab 5 from subscriptions and makes it metered token billing only&lt;/strong&gt;, it risks losing developer mindshare despite strong benchmark positioning, especially if competitors provide comparable coding performance under flat-rate plans.&lt;/strong&gt; Comments frame the move as a “classic footgun,” arguing that unpredictable API billing discourages individual developers who often become internal enterprise champions. Several users say they plan to migrate or cancel paid Claude tiers in favor of flat-rate alternatives like Codex/GPT subscriptions unless Anthropic reverses or clarifies the pricing change.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed Anthropic’s subscription/API split as a technical adoption-risk issue: if &lt;strong&gt;Fab/Fable 5&lt;/strong&gt; is removed from predictable subscriptions, power users may migrate to &lt;strong&gt;Codex’s &lt;code&gt;$200&lt;/code&gt; plan&lt;/strong&gt;, &lt;strong&gt;GPT-5&lt;/strong&gt;, or cheaper alternatives rather than accept unpredictable API spend. The core concern is not just pricing but loss of internal champions who prototype on subscriptions and later drive enterprise budget approval.&lt;/li&gt;
&lt;li&gt;One commenter disputed a cited performance comparison, arguing that the graph was misleading because it omitted &lt;strong&gt;Sol &lt;code&gt;5.6 xhigh&lt;/code&gt;&lt;/strong&gt;, which they claimed is “way above &lt;code&gt;5.5 xhigh&lt;/code&gt;.” Another said their current workflow is split between &lt;strong&gt;Opus calls&lt;/strong&gt; and &lt;strong&gt;GPT-5&lt;/strong&gt;, and suggested the &lt;strong&gt;Fable 5&lt;/strong&gt; hype may be overstated relative to &lt;strong&gt;Sol 5.6&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uuqz4l/anthropic_i_think_you_really_need_to_react_youre/&quot;&gt;Anthropic, I think you really need to react. You&apos;re slowly losing ground.&lt;/a&gt;&lt;/strong&gt; (Activity: 1731): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/4j1onimx1vch1.jpeg&quot;&gt;image&lt;/a&gt; is a screenshot of an X post highlighting &lt;strong&gt;OpenAI&lt;/strong&gt; subscription/product changes: temporary removal of a &lt;code&gt;5-hour&lt;/code&gt; usage limit for Plus/Business/Pro, efficiency improvements to “GPT 5.6 Sol,” &lt;code&gt;6M&lt;/code&gt; active users, and an incoming usage reset. In the context of the post, it is used as evidence that &lt;strong&gt;Anthropic/Claude&lt;/strong&gt; is falling behind on consumer experience after the troubled “Fable” rollout, unclear quota handling, higher token consumption with “Sonnet 5,” and last-minute communication around model availability and limits.&lt;/strong&gt; Commenters largely agree with the competitive-pressure framing, arguing that OpenAI is currently winning on &lt;strong&gt;cost, resets, communication, and model quality&lt;/strong&gt;. Several express concern that Anthropic is prioritizing enterprise/government customers over subscribers, with one $200/month user calling recent handling “unprofessional.”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters framed OpenAI’s recent advantage as a combination of &lt;strong&gt;lower cost, more frequent usage-limit resets, better communication, and improving model quality&lt;/strong&gt;, with one user claiming OpenAI had provided “like &lt;code&gt;20&lt;/code&gt; resets” since they subscribed. Several users argued Anthropic’s current consumer offering is weakening relative to OpenAI’s, particularly for high-paying users on the &lt;code&gt;$200/month&lt;/code&gt; tier.&lt;/li&gt;
&lt;li&gt;A recurring technical/product concern was Anthropic’s perceived prioritization of &lt;strong&gt;enterprise, government, and corporate accounts&lt;/strong&gt; over consumer/prosumer capacity. Users specifically referenced Anthropic’s model/tier lineup—&lt;strong&gt;Mythos, Fable, Opus, and Sonnet&lt;/strong&gt;—suggesting pricing realignments such as making Fable cost the same as Opus and Opus cost the same as Sonnet to remain competitive.&lt;/li&gt;
&lt;li&gt;Users criticized Anthropic’s handling of &lt;strong&gt;last-minute Fable 5 extensions&lt;/strong&gt; and lack of clearer reset policy changes, arguing that a “weekly reset” or more predictable capacity management would be a more credible response to OpenAI’s recent moves. The frustration is less about raw model capability and more about quota reliability, pricing transparency, and service predictability for paying subscribers.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uu94xp/subscriptions_is_less_than_5_of_revenue_they/&quot;&gt;Subscriptions is less than 5% of revenue, they might not care enough to keep Fable around&lt;/a&gt;&lt;/strong&gt; (Activity: 1162): &lt;strong&gt;The image is a financial projection table, &lt;a href=&quot;https://i.redd.it/ce0sxc319rch1.jpeg&quot;&gt;&lt;strong&gt;“Anthropic: the P&amp;#x26;L behind the IPO”&lt;/strong&gt;&lt;/a&gt;, estimating quarterly revenue mix from &lt;code&gt;1Q24&lt;/code&gt; to &lt;code&gt;4Q26&lt;/code&gt;; it shows &lt;strong&gt;API revenue dominating Anthropic’s projected revenue&lt;/strong&gt;, while consumer/business/enterprise subscriptions remain a small minority—supporting the post’s claim that subscriptions are &lt;strong&gt;&amp;#x3C;5% of revenue&lt;/strong&gt;. The table also projects Anthropic moving from heavy operating losses in 2024–2025 toward profitability in 2026, implying that subscription products like “Fable” may be strategically less important than API/enterprise growth if the estimates are accurate.&lt;/strong&gt; Commenters debated whether subscriptions are still strategically valuable despite low revenue share: they may influence developer preference, seed workplace adoption, and convert personal usage into API/business demand. Others questioned the credibility of the table because Anthropic is private and the image appears to be an external estimate rather than leaked financials.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that subscription products can function as a &lt;strong&gt;loss leader&lt;/strong&gt; and market-signal channel rather than a direct revenue center: individual developer usage can convert into &lt;strong&gt;enterprise/API adoption&lt;/strong&gt; when those developers advocate for the same tooling at work. The key technical/business dynamic raised is that consumer coding tools like Codex/Fable may influence enterprise procurement through developer preference and workflow familiarity.&lt;/li&gt;
&lt;li&gt;A commenter questioned the reliability of the reported “&amp;#x3C;5% of revenue” figure, noting that for a &lt;strong&gt;private company&lt;/strong&gt; such numbers are likely estimates rather than audited public financials. The implication is that strategic conclusions about whether OpenAI would maintain a product like Fable should be treated cautiously unless the revenue breakdown source and methodology are clear.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uudibj/im_paying_200month_and_after_tomorrow_i_cant/&quot;&gt;I&apos;m paying $200/month, and after tomorrow, I can&apos;t access Anthropic&apos;s best model with my sub?&lt;/a&gt;&lt;/strong&gt; (Activity: 1447): &lt;strong&gt;A &lt;strong&gt;$200/month Anthropic subscriber&lt;/strong&gt; argues that if the new/best model “Fable” is more expensive to serve than &lt;strong&gt;Opus&lt;/strong&gt;, Anthropic should keep it available in the subscription and apply a higher usage/token multiplier rather than removing access. The post frames this as a unit-economics/control problem: Anthropic can cap cost exposure through faster quota burn while preserving access to its frontier model.&lt;/strong&gt; Commenters expect Anthropic may reverse the decision, with some saying they will cancel if access is removed. One notable take is that frontier models may increasingly become &lt;strong&gt;API-only&lt;/strong&gt; rather than bundled into fixed-price consumer subscriptions.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter frames Anthropic’s change as evidence that &lt;strong&gt;frontier models may increasingly become API-only&lt;/strong&gt;, separating top-tier model access from fixed-price consumer subscriptions. The technical implication is that providers may prefer metered API pricing for their most expensive models rather than exposing them through capped monthly plans like &lt;code&gt;$200/month&lt;/code&gt; subscriptions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uv399n/yuji_tachikawa_one_of_the_worlds_leading/&quot;&gt;Yuji Tachikawa, one of the world’s leading theoretical physicists, reports Claude Fable solved a problem that he and his collaborators had gotten stuck on for the past 6 months&lt;/a&gt;&lt;/strong&gt; (Activity: 2989): &lt;strong&gt;&lt;strong&gt;Yuji Tachikawa&lt;/strong&gt;, a leading theoretical physicist, reportedly posted on X that &lt;strong&gt;Claude Fable&lt;/strong&gt; helped solve a theoretical-physics problem that he and collaborators had been stuck on for ~&lt;code&gt;6 months&lt;/code&gt; (&lt;a href=&quot;https://x.com/yujitach/status/2076327681562644709?s=20&quot;&gt;original tweet, now deleted&lt;/a&gt;). He later said he deleted the post due to the type of attention it attracted, &lt;strong&gt;not because he was retracting the claim&lt;/strong&gt; (&lt;a href=&quot;https://x.com/yujitach/status/2076682201626992776?s=20&quot;&gt;follow-up&lt;/a&gt;). The Reddit thread does not provide enough technical detail to evaluate the problem, solution, prompt process, or verification beyond the reported claim and a linked screenshot.&lt;/strong&gt; Commenters debated evaluation standards for AI-assisted research: one argued that dismissing the result because it was not solved “one shot” applies an unfairly stricter standard than for human collaborators. Another highlighted the model’s apparent use of speculative reasoning — e.g. &lt;em&gt;“I wonder if…”&lt;/em&gt; — as potentially relevant to frontier LLMs’ ability to explore hypotheses beyond established understanding.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter framed the notable technical claim as not merely solving a known exercise, but showing a form of hypothesis generation: Claude Fable reportedly used language like &lt;em&gt;“I wonder if…”&lt;/em&gt;, which they connect to a commonly cited frontier-LLM limitation—models’ ability to ask productive questions or explore hypotheticals beyond established understanding. The thread itself does &lt;strong&gt;not&lt;/strong&gt; provide details of the physics problem, verification process, or benchmark-style evidence, so the technical substance is limited to this interpretation of model behavior.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. AI Coding: Prototype Hype vs Production Reality&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uu17ll/why_the_majority_of_vibe_coded_projects_fail/&quot;&gt;Why the majority of vibe coded projects fail&lt;/a&gt;&lt;/strong&gt; (Activity: 1785): &lt;strong&gt;The image (&lt;a href=&quot;https://i.imgur.com/BEhaiC8.jpeg&quot;&gt;jpeg&lt;/a&gt;) is a dark-mode social post arguing that “vibe coded” AI prototypes fail because a localhost demo is often mistaken for a production system: mature Slack/Discord-like apps require &lt;strong&gt;distributed systems, scaling, reliability, message ordering, storage, search, observability, and years of iteration&lt;/strong&gt;. In the context of the title, it frames the core technical gap as not code generation itself, but underestimating the engineering needed beyond an MVP.&lt;/strong&gt; Commenters pushed back that most projects fail for normal startup reasons—insufficient product value, marketing, and sales—not because they cannot scale to Slack. Others argued AI-generated tools can still be valuable for SMB/internal workflows, where a custom CRM or HubSpot-like replacement may save &lt;code&gt;$10k–$100k+&lt;/code&gt; without needing hyperscale architecture.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that “vibe-coded” projects usually fail for &lt;strong&gt;product/value and go-to-market reasons&lt;/strong&gt;, not because they cannot scale to “Slack-level” infrastructure. The technical implication is that many AI-generated MVPs may soon clear the basic implementation bar, making differentiation depend more on domain fit, workflow integration, and whether the software solves a high-value problem.&lt;/li&gt;
&lt;li&gt;A recurring theme was that the best use case is &lt;strong&gt;small, domain-specific internal software&lt;/strong&gt;, not billion-user SaaS platforms. Commenters cited SMB tools that replace expensive vendors—e.g. a custom &lt;strong&gt;HubSpot-like CRM&lt;/strong&gt; built quickly with a capable model—where saving &lt;code&gt;$15k+&lt;/code&gt; annually can justify software that only needs to serve a small team.&lt;/li&gt;
&lt;li&gt;One commenter emphasized that many successful projects do not require public-scale testing because they are &lt;strong&gt;hyper-niche operational tools&lt;/strong&gt; used by only a handful of people. The claimed opportunity is software that may only serve &lt;code&gt;3–10&lt;/code&gt; users but saves a company up to &lt;code&gt;$100k&lt;/code&gt; annually when replacing manual work, EUC processes, or governance overhead.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uv70kq/honest_question_what_are_you_building_that_you/&quot;&gt;Honest question: What are you building that you need fable 5 so badly?&lt;/a&gt;&lt;/strong&gt; (Activity: 1030): &lt;strong&gt;The poster asks what workloads justify upgrading to &lt;strong&gt;Fable 5&lt;/strong&gt; given that &lt;strong&gt;Claude Pro&lt;/strong&gt; with mostly &lt;strong&gt;Opus 4.8&lt;/strong&gt;, and workplace use of &lt;strong&gt;Opus 4.6 / Sonnet 5&lt;/strong&gt;, already handles homelab automation and large-scale data-engineering work including dbt, long SQL/query parsing, near-real-time joins, thousands of schemas/integrations, and pipelines processing roughly &lt;code&gt;150B events/day&lt;/code&gt;. Top technical use cases cited for newer models were &lt;strong&gt;multi-agent VFX/AAA game pipeline automation&lt;/strong&gt;—where less prompt specificity and less hand-holding reduce cognitive load across obscure, duct-taped artist tooling—and &lt;strong&gt;adversarial language/rhetorical analysis&lt;/strong&gt;, where Fable is valued for holding multiple interpretive frames while critiquing.&lt;/strong&gt; Commenters framed Fable/Sol less as unlocking categorically new programming capability and more as reducing supervision cost, context switching, and prompt-engineering overhead. One dissenting view characterized much usage as wasteful “slop,” while another noted &lt;strong&gt;GPT 5.6 Sol&lt;/strong&gt; may now be competitive for multi-frame critique tasks.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A VFX/AAA games software engineer with &lt;code&gt;17 years&lt;/code&gt; experience described using &lt;strong&gt;Fable&lt;/strong&gt; and &lt;strong&gt;Sol&lt;/strong&gt; to manage ad hoc production pipelines and obscure tooling issues where code is often a means to unblock artists rather than the product itself. They emphasized running “five agents” concurrently on artist-facing tasks, valuing models that need less prompt-detailing and hand-holding to reduce cognitive load in engineering-hostile production environments.&lt;/li&gt;
&lt;li&gt;One commenter uses &lt;strong&gt;Fable&lt;/strong&gt; primarily for adversarial testing, language analysis, rhetorical critique, and paper-writing rather than coding. They characterized Fable as better at maintaining and critiquing “multiple frames at once,” while noting that &lt;strong&gt;GPT 5.6 Sol&lt;/strong&gt; is becoming “very, very good” at the same class of multi-perspective critique tasks.&lt;/li&gt;
&lt;li&gt;A senior big-tech engineer argued that &lt;strong&gt;Fable&lt;/strong&gt;’s value is less about generating better code than &lt;strong&gt;Opus&lt;/strong&gt; or &lt;strong&gt;Sonnet&lt;/strong&gt;, and more about acting like a “staff engineer”: clarifying ambiguous requirements, producing high-level architecture, and orchestrating implementation. In their framing, &lt;strong&gt;Opus&lt;/strong&gt; maps well to “senior engineer” coding under moderate ambiguity, &lt;strong&gt;Sonnet&lt;/strong&gt; to “junior engineer” execution with clearer tasks, while frontier models become useful when the user delegates more systems-level and cross-functional problem solving.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uuuhj8/did_not_expext_fable_5_to_be_this_good/&quot;&gt;Did not expext Fable 5 to be this good!✨&lt;/a&gt;&lt;/strong&gt; (Activity: 1273): &lt;strong&gt;The post claims &lt;strong&gt;Fable 5&lt;/strong&gt; was used to generate a browser-based &lt;strong&gt;Three.js FPS&lt;/strong&gt; in roughly &lt;code&gt;3&lt;/code&gt; afternoons from a low-poly city asset folder, with Fable handling map creation. The demo, hosted on &lt;a href=&quot;https://sky-cruiser-9065b31330ea.herokuapp.com/&quot;&gt;Heroku&lt;/a&gt;, is described as supporting &lt;strong&gt;single/multiplayer FFA/TDM&lt;/strong&gt;, desktop/VR play, flying cars, and Quake-style weapons like rocket launchers and rail guns; the referenced Reddit video could not be accessed due to a &lt;code&gt;403 Forbidden&lt;/code&gt; block.&lt;/strong&gt; Top comments were mostly non-technical: one says the praise is “justified,” another jokes it resembles “last week’s Fable,” and one compares the gameplay/aesthetic to &lt;em&gt;Forsaken&lt;/em&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>prime-intellect</category><category>vllm</category><category>langchain</category><category>threepointone</category><category>factory</category><category>cognition</category><category>arena</category><category>artificial-analysis</category><category>parlance-labs</category><category>gpt-5.6-sol</category><category>grok-4.5</category><category>terra-max</category><category>fable-5-max</category><category>opus-4.8</category><category>100b-reasoning-model</category><category>johannes_hage</category><category>willccbb</category><category>mikasenghaas</category><category>xeophon</category><category>omarsar0</category><category>skirano</category><category>imjaredz</category><category>agentic-reinforcement-learning</category><category>rollout-traces</category><category>message-dags</category><category>long-horizon-reinforcement-learning</category><category>multimodality</category><category>harness-design</category><category>cost-per-task</category><category>coding-agents</category><category>benchmarks</category><category>model-efficiency</category><category>real-world-evaluation</category><category>task-specialization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-10-not-much/</guid><description>**OpenAI** rolled out **GPT-5.6** featuring a new model stratification with tiers **Luna / Terra / Sol** and effort levels including **Max** and **Ultra**, introducing complex configuration options. The launch faced UX challenges with the **ChatGPT Work / Codex** split, prompting rapid corrective actions including usage-limit resets and UI improvements. Early benchmarks show **GPT-5.6** excels in agentic coding, presentation, and science tasks, tying with **Claude Fable 5** in Code Arena Frontend at about half the cost, and achieving a significant **500-point** Elo gain in presentations. However, users noted instruction-following issues and concerns about jailbreakability. The major advancement is in orchestration and computer use, with **Sol Ultra** demonstrating strong planner and verifier capabilities, enabling high-throughput automation workflows. A notable operational challenge is the hidden cost explosion from spawned subagents inheriting premium settings, causing faster quota depletion.</description><pubDate>Fri, 10 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/09/2026-7/10/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;OpenAI’s GPT-5.6 rollout: model stratification, agent UX, and early benchmark signals&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPT-5.6 introduced a more explicit model/compute ladder&lt;/strong&gt;: users are now navigating &lt;strong&gt;Luna / Terra / Sol&lt;/strong&gt; plus multiple effort levels, with community guidance converging around “start lower than you did on 5.5.” OpenAI staff explained that &lt;strong&gt;Max&lt;/strong&gt; means one model spending longer on a hard problem, while &lt;strong&gt;Ultra&lt;/strong&gt; parallelizes work across subagents; they also noted that 5.5→5.6 effort settings are &lt;strong&gt;not directly comparable&lt;/strong&gt; (&lt;a href=&quot;https://x.com/reach_vb/status/2075489301253488778&quot;&gt;guidance from @reach_vb&lt;/a&gt;, &lt;a href=&quot;https://x.com/pvncher/status/2075590107214520590&quot;&gt;follow-up&lt;/a&gt;, &lt;a href=&quot;https://x.com/gabrielchua/status/2075521933576462357&quot;&gt;practical default suggestion&lt;/a&gt;). The community reaction was mixed: many praised the added control, while others criticized the &lt;strong&gt;30+ configuration combinatorics&lt;/strong&gt; and missing “Auto” routing (&lt;a href=&quot;https://x.com/rasbt/status/2075369179817902176&quot;&gt;@rasbt&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2075627844412264796&quot;&gt;@Yuchenj_UW&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The product launch landed with real UX regressions, and OpenAI publicly course-corrected fast&lt;/strong&gt;: users complained that the new &lt;strong&gt;ChatGPT Work / Codex&lt;/strong&gt; split was confusing, chats/projects became harder to find, and usage burned down faster than expected (&lt;a href=&quot;https://x.com/scaling01/status/2075595915419599176&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/simonw/status/2075663372323008755&quot;&gt;@simonw&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075608495756333087&quot;&gt;@kimmonismus&lt;/a&gt;). OpenAI responded unusually directly: &lt;strong&gt;multiple usage-limit resets&lt;/strong&gt;, acknowledgements that defaults nudged users toward overly expensive settings, and a commitment to restore familiar sidebar/navigation patterns and clarify positioning between Work and Codex (&lt;a href=&quot;https://x.com/thsottiaux/status/2075452680760443190&quot;&gt;@thsottiaux reset announcement&lt;/a&gt;, &lt;a href=&quot;https://x.com/reach_vb/status/2075460193681367532&quot;&gt;second reset&lt;/a&gt;, &lt;a href=&quot;https://x.com/thsottiaux/status/2075641131002700120&quot;&gt;full corrective roadmap&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Initial eval picture&lt;/strong&gt;: GPT-5.6 appears strongest in &lt;strong&gt;agentic coding / presentation / some science tasks&lt;/strong&gt;, but not unambiguously dominant everywhere. Examples: &lt;strong&gt;#1 tie in Code Arena: Frontend&lt;/strong&gt; with Claude Fable 5 while being ~&lt;strong&gt;2× cheaper&lt;/strong&gt; on listed IO pricing (&lt;a href=&quot;https://x.com/arena/status/2075672492312768683&quot;&gt;Arena&lt;/a&gt;); best recorded &lt;strong&gt;Presentation Elo&lt;/strong&gt; on AA-Briefcase with a ~&lt;strong&gt;500-point&lt;/strong&gt; jump over GPT-5.5 (&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075639143372325205&quot;&gt;Artificial Analysis&lt;/a&gt;); &lt;strong&gt;CritPt&lt;/strong&gt; gains over GPT-5.5 and beats Fable 5 by ~4 points (&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075423964378366427&quot;&gt;Artificial Analysis&lt;/a&gt;); and strong results on &lt;strong&gt;WeirdML&lt;/strong&gt; at lower cost (&lt;a href=&quot;https://x.com/htihle/status/2075513299106426922&quot;&gt;@htihle&lt;/a&gt;). At the same time, users reported &lt;strong&gt;instruction-following issues&lt;/strong&gt;, uneven token efficiency in practice, and some concern about &lt;strong&gt;jailbreakability / reward hacking&lt;/strong&gt; (&lt;a href=&quot;https://x.com/teortaxesTex/status/2075495527030964693&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/Mononofu/status/2075414796426764507&quot;&gt;@Mononofu&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075693686604619948&quot;&gt;@kimmonismus&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Parallel-agent workflows, computer use, and the “harness is the product” theme&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;GPT-5.6’s biggest perceived leap may be orchestration and computer use rather than pure chat quality&lt;/strong&gt;. Multiple users highlighted that Sol is unusually strong as a &lt;strong&gt;planner / verifier / orchestrator&lt;/strong&gt;, often using subagents automatically and reacting more quickly to steering (&lt;a href=&quot;https://x.com/omarsar0/status/2075611352878481577&quot;&gt;@omarsar0&lt;/a&gt;, &lt;a href=&quot;https://x.com/Hangsiin/status/2075463886309126271&quot;&gt;@Hangsiin&lt;/a&gt;). OpenAI also showcased &lt;strong&gt;computer use with Sol Ultra&lt;/strong&gt; and promoted ChatGPT Work as bringing agents to consumer/mobile scale (&lt;a href=&quot;https://x.com/gdb/status/2075619497764151644&quot;&gt;OpenAI demo via @gdb&lt;/a&gt;, &lt;a href=&quot;https://x.com/gdb/status/2075628596232884556&quot;&gt;Work positioning&lt;/a&gt;). Community reports described very high-throughput GUI automation and Blender workflows (&lt;a href=&quot;https://x.com/mckbrando/status/2075442660047814761&quot;&gt;@mckbrando&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075482486901969066&quot;&gt;@kimmonismus&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A recurring operational issue is hidden subagent cost explosion&lt;/strong&gt;: users found that spawned agents may inherit premium settings, draining quotas much faster than expected. One concrete claim was that &lt;code&gt;spawn_agent&lt;/code&gt; doesn’t let users choose model/effort, so &lt;strong&gt;Sol Ultra spawns more Sol Ultra&lt;/strong&gt; by default (&lt;a href=&quot;https://x.com/evi77ain/status/2075445272013095033&quot;&gt;@evi77ain&lt;/a&gt;). This fits the broader pattern of people liking the capability jump but finding the cost model opaque.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The broader systems trend is toward harness-centric competition&lt;/strong&gt;. This came through in product commentary from Perplexity’s Arav Srinivas (“the real product is now the harness around it”), in LangChain’s launch framing around &lt;strong&gt;Deep Agents + Nemotron + OpenShell&lt;/strong&gt;, and in a growing set of memory / orchestration tools like &lt;strong&gt;OpenWiki&lt;/strong&gt; and &lt;strong&gt;OpenSWE&lt;/strong&gt; (&lt;a href=&quot;https://x.com/dee_bosa/status/2075597686464491874&quot;&gt;@dee_bosa quoting Arav&lt;/a&gt;, &lt;a href=&quot;https://x.com/hwchase17/status/2075620940466315608&quot;&gt;@hwchase17&lt;/a&gt;, &lt;a href=&quot;https://x.com/BraceSproul/status/2075596668612014107&quot;&gt;OpenWiki proactive memory&lt;/a&gt;, &lt;a href=&quot;https://x.com/BraceSproul/status/2075610067878257072&quot;&gt;OpenSWE adoption&lt;/a&gt;). The meta-point: frontier model parity is tightening, so value is increasingly shifting to &lt;strong&gt;routing, memory, tool use, safety rails, and enterprise context&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Meta’s Muse Spark 1.1 and the widening frontier of “good enough, fast, cheap” models&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Muse Spark 1.1 was the other major model story of the day&lt;/strong&gt;, with many practitioners calling it the most surprising release of the week. Reports consistently emphasized &lt;strong&gt;strong UI/frontend generation, fast responses, and unusually aggressive pricing&lt;/strong&gt;, often framing it as near-frontier quality for a large subset of coding/product tasks (&lt;a href=&quot;https://x.com/alexandr_wang/status/2075652012608467385&quot;&gt;@alexandr_wang&lt;/a&gt;, &lt;a href=&quot;https://x.com/rowancheung/status/2075634108324089943&quot;&gt;@rowancheung&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075525943729275313&quot;&gt;@kimmonismus&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Benchmarking suggests a real step up, but not outright frontier leadership&lt;/strong&gt;. Artificial Analysis scored Muse Spark 1.1 at &lt;strong&gt;51&lt;/strong&gt; on its Intelligence Index, up &lt;strong&gt;8 points&lt;/strong&gt; from 1.0, roughly tied with &lt;strong&gt;GLM-5.2 / GPT-5.4 / GPT-5.6 Luna&lt;/strong&gt; and behind &lt;strong&gt;Grok 4.5 / GPT-5.6 Sol / Claude Fable 5&lt;/strong&gt;. Notable details: &lt;strong&gt;1M context&lt;/strong&gt;, median speed ~&lt;strong&gt;114 tok/s&lt;/strong&gt;, pricing &lt;strong&gt;$1.25 / $4.25 per 1M&lt;/strong&gt; input/output tokens, and strong token efficiency (&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075677416295739660&quot;&gt;Artificial Analysis&lt;/a&gt;). Arena also placed it &lt;strong&gt;#9 on Code Arena: Frontend&lt;/strong&gt; with strong gains in instruction-following and longer-query categories (&lt;a href=&quot;https://x.com/arena/status/2075642304501784698&quot;&gt;Arena&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The strategic implication many drew&lt;/strong&gt;: Meta’s compute-heavy bet is starting to show up as &lt;strong&gt;cost-effective inference products&lt;/strong&gt;, not just talent headlines. Several commentators argued this materially raises competitive pressure on OpenAI/Anthropic, especially if Meta improves distribution and API ergonomics (&lt;a href=&quot;https://x.com/scaling01/status/2075612353056342391&quot;&gt;@scaling01 asking for OpenRouter&lt;/a&gt;, &lt;a href=&quot;https://x.com/alexandr_wang/status/2075680437620646370&quot;&gt;@alexandr_wang&lt;/a&gt;, &lt;a href=&quot;https://x.com/mweinbach/status/2075600689200279747&quot;&gt;@mweinbach&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open models, infra, and efficiency work&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Open-model tooling kept shipping despite the closed-model attention vacuum&lt;/strong&gt;. Unsloth released &lt;strong&gt;Qwen3.6 NVFP4 quants&lt;/strong&gt; with claims of &lt;strong&gt;2.5× faster&lt;/strong&gt; inference, including &lt;strong&gt;27B on 24GB VRAM&lt;/strong&gt; and a &lt;strong&gt;35B-A3B&lt;/strong&gt; variant hitting &lt;strong&gt;17,561 tok/s on B200&lt;/strong&gt; (&lt;a href=&quot;https://x.com/UnslothAI/status/2075566124687892597&quot;&gt;Unsloth&lt;/a&gt;, &lt;a href=&quot;https://x.com/danielhanchen/status/2075567076002185525&quot;&gt;technical details from @danielhanchen&lt;/a&gt;). QuixiAI reported &lt;strong&gt;Qwen3.6-35B-A3B-NVFP4&lt;/strong&gt; on dual B60 at &lt;strong&gt;65 tok/s&lt;/strong&gt; and &lt;strong&gt;128k context&lt;/strong&gt; (&lt;a href=&quot;https://x.com/QuixiAI/status/2075418782470643958&quot;&gt;QuixiAI&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference optimization remains a major live research area&lt;/strong&gt;. Cohere open-sourced &lt;strong&gt;Hardware-aware Dynamic Speculative Decoding&lt;/strong&gt; in vLLM, addressing the familiar issue where speculative decoding helps at low batch sizes but hurts at high ones (&lt;a href=&quot;https://x.com/EkagraRanjan/status/2075640096829612416&quot;&gt;Cohere/vLLM&lt;/a&gt;, &lt;a href=&quot;https://x.com/vllm_project/status/2075698626140295378&quot;&gt;vLLM commentary&lt;/a&gt;). Google/Hugging Face’s &lt;strong&gt;Gemma challenge&lt;/strong&gt; reported up to &lt;strong&gt;5× faster&lt;/strong&gt; single-A10G inference, with &lt;strong&gt;315 TPS lossless&lt;/strong&gt; and &lt;strong&gt;491.8 TPS&lt;/strong&gt; fastest overall (&lt;a href=&quot;https://x.com/googlegemma/status/2075611948985835877&quot;&gt;Gemma&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Agent evaluation / self-improvement work is getting more concrete&lt;/strong&gt;: “&lt;strong&gt;LLM-as-a-Verifier&lt;/strong&gt;” reported SOTA on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench using repeated sampling plus score-logprob ranking (&lt;a href=&quot;https://x.com/Azaliamirh/status/2075583355895058751&quot;&gt;paper thread&lt;/a&gt;); Meta researchers proposed an explicit memory agent to combat &lt;strong&gt;behavioral state decay&lt;/strong&gt; in long-horizon agents (&lt;a href=&quot;https://x.com/omarsar0/status/2075603504543269136&quot;&gt;summary&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Science, math, health, and modality-specific systems&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Math/science capability claims escalated sharply&lt;/strong&gt;. OpenAI staff and community members circulated examples of &lt;strong&gt;GPT-5.6 Sol Ultra&lt;/strong&gt; producing a claimed proof of the &lt;strong&gt;Cycle Double Cover Conjecture&lt;/strong&gt; using &lt;strong&gt;64 subagents in under an hour&lt;/strong&gt; (&lt;a href=&quot;https://x.com/__eknight__/status/2075643450196971805&quot;&gt;claim from @&lt;strong&gt;eknight&lt;/strong&gt;&lt;/a&gt;, &lt;a href=&quot;https://x.com/gdb/status/2075670151702430044&quot;&gt;amplified by @gdb&lt;/a&gt;). Separately, Bubeck noted a single-person &lt;strong&gt;1M-line Lean formalization&lt;/strong&gt; effort with GPT-5.6 (&lt;a href=&quot;https://x.com/SebastienBubeck/status/2075407986772861047&quot;&gt;@SebastienBubeck&lt;/a&gt;). These are still claims pending external scrutiny, but they indicate where labs want the narrative to go: &lt;strong&gt;parallelized research agents as a scientific compute primitive&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Health is becoming a first-class benchmark and product vertical&lt;/strong&gt;. OpenAI said GPT-5.6 is a major step forward for &lt;strong&gt;health intelligence&lt;/strong&gt;, highlighting that &lt;strong&gt;Luna at lowest effort beats GPT-5.5 at highest effort while costing 25× less&lt;/strong&gt; (&lt;a href=&quot;https://x.com/OpenAI/status/2075686461693898868&quot;&gt;OpenAI&lt;/a&gt;). Karan Singhal added that, in blinded physician comparisons over &lt;strong&gt;20,000 axis ratings&lt;/strong&gt;, physicians found &lt;strong&gt;fewer flaws in GPT-5.6 responses than physician-written responses&lt;/strong&gt; across a hard task set (&lt;a href=&quot;https://x.com/thekaransinghal/status/2075689779937833302&quot;&gt;details&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Audio/music and creative tooling also moved&lt;/strong&gt;: Kyutai + Mirelo released &lt;strong&gt;MuScriptor&lt;/strong&gt;, an open model for &lt;strong&gt;multi-instrument audio-to-MIDI transcription from full mixes&lt;/strong&gt;, not stems (&lt;a href=&quot;https://x.com/MireloAI/status/2075536492177354771&quot;&gt;MireloAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/kyutai_labs/status/2075540047613276197&quot;&gt;Kyutai&lt;/a&gt;). Sakana’s new Picbreeder-style work explored &lt;strong&gt;open-ended creativity with VLM agents&lt;/strong&gt;, concluding that diverse agent populations help but still fall short of human open-ended exploration (&lt;a href=&quot;https://x.com/SakanaAILabs/status/2075580810330267844&quot;&gt;Sakana&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Security, safety, and policy frictions&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Security concerns rose alongside capability gains&lt;/strong&gt;. OpenAI moved its &lt;strong&gt;Bio Bug Bounty&lt;/strong&gt; into a private ongoing program and &lt;strong&gt;doubled rewards to $50K&lt;/strong&gt;, specifically seeking universal jailbreaks against predefined biosafety challenges (&lt;a href=&quot;https://x.com/OpenAI/status/2075647722766614733&quot;&gt;OpenAI&lt;/a&gt;). Separately, OpenAI tightened access requirements for its most cyber-capable models, requiring &lt;strong&gt;hardware security keys&lt;/strong&gt; for Trusted Access for Cyber members starting Sept. 1 (&lt;a href=&quot;https://x.com/cryps1s/status/2075639162120900766&quot;&gt;@cryps1s&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evidence of misuse remains salient&lt;/strong&gt;: a new study reported &lt;strong&gt;Boko Haram&lt;/strong&gt; members using frontier chatbots for bomb-making and related tactical queries (&lt;a href=&quot;https://x.com/AntoniaJuelich/status/2075590815083028989&quot;&gt;@AntoniaJuelich&lt;/a&gt;). That thread sat uncomfortably next to ongoing online discussion that GPT-5.6 may be relatively easy to jailbreak or reward-hack in some settings (&lt;a href=&quot;https://x.com/Mononofu/status/2075414796426764507&quot;&gt;@Mononofu&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Policy discourse remains polarized and speculative&lt;/strong&gt;. The “AI 2040 / Plan A” transparency-and-governance scenario drew both support and ridicule, with Ajeya Cotra emphasizing the centrality of &lt;strong&gt;total research transparency&lt;/strong&gt; while critics questioned feasibility and assumptions about superintelligence/governance capacity (&lt;a href=&quot;https://x.com/ajeya_cotra/status/2075583823434371250&quot;&gt;@ajeya_cotra&lt;/a&gt;, &lt;a href=&quot;https://x.com/binarybits/status/2075660927001608431&quot;&gt;@binarybits&lt;/a&gt;, &lt;a href=&quot;https://x.com/banteg/status/2075512151783972925&quot;&gt;@banteg satire&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenAI launch and rollback management&lt;/strong&gt;: OpenAI’s product lead acknowledged launch confusion, promised UI fixes, and reset usage twice while clarifying that &lt;strong&gt;Codex is here to stay&lt;/strong&gt; (&lt;a href=&quot;https://x.com/thsottiaux/status/2075641131002700120&quot;&gt;full thread&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Code desktop browser&lt;/strong&gt;: Anthropic shipped an &lt;strong&gt;in-app browser&lt;/strong&gt; for Claude Code desktop so Claude can browse docs/sites inside the app (&lt;a href=&quot;https://x.com/ClaudeDevs/status/2075635283211772279&quot;&gt;@ClaudeDevs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OpenAI org update&lt;/strong&gt;: Fidji Simo announced she is leaving her full-time role at OpenAI and becoming a &lt;strong&gt;part-time advisor&lt;/strong&gt;, citing the need to focus on recovery from chronic illness while continuing work related to AI and health (&lt;a href=&quot;https://x.com/fidjissimo/status/2075353170927304861&quot;&gt;@fidjissimo&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Perplexity harness expansion&lt;/strong&gt;: Perplexity added &lt;strong&gt;Grok 4.5&lt;/strong&gt; as an orchestrator in Computer after internal evals showed strong WANDR performance at roughly half the cost of Opus 4.8 (&lt;a href=&quot;https://x.com/perplexity_ai/status/2075660058625790159&quot;&gt;Perplexity&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. GLM-5.2 Local Inference and Security Scrutiny&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1us5m0g/glm52_744b_moe_on_a_25gbram_consumer_machine/&quot;&gt;GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine&lt;/a&gt;&lt;/strong&gt; (Activity: 1249): &lt;strong&gt;A demo reportedly runs &lt;strong&gt;GLM-5.2&lt;/strong&gt;, a &lt;code&gt;744B&lt;/code&gt;-parameter &lt;strong&gt;MoE&lt;/strong&gt; model, on a consumer machine with only &lt;code&gt;25 GB&lt;/code&gt; of RAM by &lt;strong&gt;streaming expert weights from disk&lt;/strong&gt; rather than keeping the full model resident in memory. Commenters emphasize the technical interest is not throughput—likely unusably slow for practical inference—but proving that disk-backed expert paging is possible; &lt;em&gt;“if someone figures out expert routing prediction well enough to prefetch, the whole picture changes.”&lt;/em&gt;&lt;/strong&gt; Top comments pushed back against criticism of speed and implementation quality, arguing the noteworthy result is enabling a &lt;code&gt;744B&lt;/code&gt; MoE to execute at all on low-RAM consumer hardware. There was some meta-debate over whether the project was “vibe coded,” but technical commenters largely viewed the prototype as impressive.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed the experiment as technically interesting because it demonstrates &lt;strong&gt;streaming a &lt;code&gt;744B&lt;/code&gt; MoE model’s experts from disk&lt;/strong&gt; on a consumer machine with only &lt;code&gt;25 GB&lt;/code&gt; RAM, rather than as a practical inference setup. One pointed out that if &lt;strong&gt;expert-routing prediction&lt;/strong&gt; could reliably prefetch the next required experts, disk-backed MoE inference latency could change substantially.&lt;/li&gt;
&lt;li&gt;A commenter noted that &lt;code&gt;llama.cpp&lt;/code&gt; may already provide related behavior via &lt;code&gt;--mmap&lt;/code&gt;, implying the model weights can be memory-mapped instead of fully resident in RAM, though this does not by itself solve MoE expert prefetch/routing latency.&lt;/li&gt;
&lt;li&gt;One user shared an extreme low-resource baseline: running &lt;code&gt;Qwen2.5-0.5B&lt;/code&gt; with a &lt;code&gt;1-bit&lt;/code&gt; quantization on an &lt;code&gt;x86 Atom N270&lt;/code&gt; netbook with &lt;code&gt;1 GB&lt;/code&gt; RAM, achieving roughly &lt;code&gt;240 s/token&lt;/code&gt;, illustrating how feasibility and usability diverge sharply on constrained hardware.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1urhzox/glm52_fearmongering_in_the_press/&quot;&gt;GLM-5.2 fearmongering in the press&lt;/a&gt;&lt;/strong&gt; (Activity: 907): &lt;strong&gt;The post criticizes a &lt;a href=&quot;https://futurism.com/artificial-intelligence/open-source-ai-model-scary-mythos&quot;&gt;Futurism article&lt;/a&gt; claiming &lt;strong&gt;GLM-5.2&lt;/strong&gt; is broadly downloadable, usable &lt;em&gt;“on virtually any hardware,”&lt;/em&gt; and potentially raises cybersecurity risk because there is no hosted-vendor mediation layer. The article cites &lt;strong&gt;Semgrep&lt;/strong&gt; and &lt;strong&gt;Graphistry&lt;/strong&gt; findings that GLM-5.2 performs well on bug-finding/cybersecurity tasks, including Semgrep’s &lt;em&gt;“We Have Mythos at Home”&lt;/em&gt; benchmark framing, but commenters dispute the hardware claim as technically misleading given frontier-scale inference requirements and degradation in extreme low-bit quantization.&lt;/strong&gt; Commenters view the article as fearmongering and technically uninformed, especially around inference hardware feasibility. A notable counterargument is that if strong models improve exploit discovery, the appropriate response is to use similarly strong models for remediation and defense rather than restrict or censor open models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters challenged the press claim that GLM-5.2 can run on &lt;em&gt;“virtually any hardware”&lt;/em&gt;, arguing that a large frontier/open-weight model would require substantial GPU investment rather than consumer-era CPUs; one user sarcastically asks how many &lt;strong&gt;seconds per token&lt;/strong&gt; an old &lt;code&gt;4th gen i3&lt;/code&gt; laptop would achieve, while another frames realistic deployment as hardware costing on the order of &lt;code&gt;$250k&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A technical objection was raised against citing extreme &lt;code&gt;1-bit&lt;/code&gt; or &lt;code&gt;2-bit&lt;/code&gt; quantization as evidence of broad deployability: commenters argue such quants are often severely degraded—described as &lt;em&gt;“lobotomised”&lt;/em&gt;—and therefore not comparable to running the full-capability model.&lt;/li&gt;
&lt;li&gt;One commenter reframed the security-risk argument as a dual-use mitigation problem: if advanced models can help exploit vulnerabilities, the appropriate response is to use similarly capable models for defensive discovery and patching rather than banning or restricting the models outright.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Local LLM Performance and Hardware ROI&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1usniqh/25x_faster_qwen36_nvfp4_unsloth_quants/&quot;&gt;2.5x faster Qwen3.6 NVFP4 Unsloth quants&lt;/a&gt;&lt;/strong&gt; (Activity: 934): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/yoxm16aijech1.png&quot;&gt;image&lt;/a&gt; is a promotional benchmark graphic for &lt;strong&gt;Unsloth’s dynamic NVFP4 quantizations of Qwen3.6&lt;/strong&gt;, supporting the post’s claim of &lt;strong&gt;up to &lt;code&gt;2.5×&lt;/code&gt; faster&lt;/strong&gt; inference than NVIDIA NVFP4 quants. It reports B200 throughput gains such as &lt;strong&gt;Qwen3.6-27B: &lt;code&gt;5,637&lt;/code&gt; vs &lt;code&gt;2,259&lt;/code&gt;&lt;/strong&gt; and &lt;strong&gt;Qwen3.6-35B-A3B: up to &lt;code&gt;11,628&lt;/code&gt; vs &lt;code&gt;6,481&lt;/code&gt;&lt;/strong&gt;, attributed to &lt;strong&gt;W4A4 4-bit tensor-core matmuls&lt;/strong&gt; versus NVIDIA’s W4A16 path, while tables in the post show broadly comparable MMLU-Pro, GPQA, and AIME 2025 scores across BF16/FP8/NVFP4 variants. The post also links released Hugging Face models for &lt;a href=&quot;https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4&quot;&gt;&lt;code&gt;35B-A3B-NVFP4&lt;/code&gt;&lt;/a&gt;, &lt;a href=&quot;https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast&quot;&gt;&lt;code&gt;35B-A3B-NVFP4-Fast&lt;/code&gt;&lt;/a&gt;, and &lt;a href=&quot;https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4&quot;&gt;&lt;code&gt;27B-NVFP4&lt;/code&gt;&lt;/a&gt;, plus FP8 KV-cache calibration for roughly &lt;code&gt;2×&lt;/code&gt; longer contexts.&lt;/strong&gt; Commenters mainly frame this as a &lt;strong&gt;Blackwell-specific win&lt;/strong&gt;, with jokes that Pascal/RTX 3090-era users likely won’t benefit because the speedups depend on newer GPU tensor-core support.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters questioned how &lt;strong&gt;Qwen3.6 NVFP4 Unsloth quants&lt;/strong&gt; compare against standard non-NVFP4 &lt;code&gt;4-bit&lt;/code&gt; quantizations, specifically whether the claimed &lt;code&gt;2.5x&lt;/code&gt; speedup is unique to Blackwell hardware or holds against existing 4-bit formats in common inference stacks.&lt;/li&gt;
&lt;li&gt;There was technical uncertainty around &lt;strong&gt;llama.cpp / llama-server NVFP4 support&lt;/strong&gt;: one user noted that llama-server &lt;em&gt;can&lt;/em&gt; run NVFP4 but that prior performance looked “lackluster,” while another asked why no &lt;code&gt;GGUF&lt;/code&gt; builds were provided if llama.cpp now supports NVFP4 reasonably well.&lt;/li&gt;
&lt;li&gt;Several comments implied the optimization is primarily relevant to &lt;strong&gt;NVIDIA Blackwell&lt;/strong&gt; GPUs, with older architectures such as &lt;strong&gt;Pascal&lt;/strong&gt; and consumer cards like the &lt;strong&gt;RTX 3090&lt;/strong&gt; unlikely to benefit from NVFP4 acceleration.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1us6f84/if_you_spent_45k_on_a_local_ai_rig_would_you_do/&quot;&gt;If you spent $4–5K on a local AI rig, would you do it again?&lt;/a&gt;&lt;/strong&gt; (Activity: 359): &lt;strong&gt;The post argues that a &lt;code&gt;$4–5K&lt;/code&gt; local AI rig is hard to justify purely for running frontier-quality local LLMs, especially when APIs such as &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt; are priced around &lt;code&gt;$0.14/M&lt;/code&gt; uncached input tokens and &lt;code&gt;$0.28/M&lt;/code&gt; output tokens. The author reports that even on a &lt;code&gt;128GB&lt;/code&gt; MacBook, running a &lt;code&gt;2-bit&lt;/code&gt; quantized DeepSeek V4 Flash is still not compelling versus hosted models, though the setup was useful for learning about quantization, KV cache, context windows, memory limits, and model serving.&lt;/strong&gt; The author’s view is that expensive local hardware may make sense for privacy, always-on workloads, or when the machine is needed anyway, but not primarily as a cost-saving substitute for Claude/ChatGPT-quality APIs. No top comments were provided to summarize.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. GPT-5.6 Coding Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1usavpc/deepswe_just_added_the_gpt56_models_to_their/&quot;&gt;DeepSWE just added the gpt-5.6 models to their benchmark.  I hope you guys  don&apos;t get too used to Claude Code as your only coding agent.  Chart is marked NSFW due to the grotesque violence.&lt;/a&gt;&lt;/strong&gt; (Activity: 1718): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/e5dlfudecbch1.png&quot;&gt;image&lt;/a&gt; is a &lt;strong&gt;DeepSWE benchmark cost/performance chart&lt;/strong&gt; comparing coding-agent models by “DeepSWE score” vs average cost per task, with the post highlighting newly added &lt;strong&gt;GPT-5.6 variants&lt;/strong&gt; as strong low-cost competitors to &lt;strong&gt;Claude Code/Claude models&lt;/strong&gt;. In the chart, GPT-5.6/5.5-family points appear to cluster around roughly &lt;code&gt;60–70%&lt;/code&gt; DeepSWE score at comparatively low task cost, while Claude models remain competitive—e.g. Claude-fable-5 near the top around &lt;code&gt;70%&lt;/code&gt;—but often at higher cost.&lt;/strong&gt; The comments do not engage much with the benchmark itself; they overwhelmingly criticize the visualization quality, calling it “psychopath” charting and pointing to r/dataisugly. The post’s “grotesque violence” framing is hyperbolic/meme-like, referring to the chart’s implied GPT-vs-Claude disruption rather than literal content.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/OpenAI/comments/1us7nml/gpt_56_beats_fable_5_by_3_more_on_deepswe_at_a/&quot;&gt;GPT 5.6 Beats Fable 5 by 3% more on DeepSWE at a cheaper price.&lt;/a&gt;&lt;/strong&gt; (Activity: 1310): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/505rvco3nach1.jpeg&quot;&gt;image&lt;/a&gt; shows a &lt;strong&gt;DeepSWE leaderboard&lt;/strong&gt; where &lt;strong&gt;gpt-5.6-sol&lt;/strong&gt; scores &lt;code&gt;73% ±3%&lt;/code&gt; at an average cost of &lt;code&gt;$8.39&lt;/code&gt;, outperforming &lt;strong&gt;claude-fable-5&lt;/strong&gt; at &lt;code&gt;70% ±4%&lt;/code&gt; while costing much less than Fable’s &lt;code&gt;$21.63&lt;/code&gt;. It also highlights &lt;strong&gt;gpt-5.6-terra&lt;/strong&gt; matching Fable’s &lt;code&gt;70%&lt;/code&gt; score at roughly &lt;code&gt;4.4×&lt;/code&gt; lower cost, making the post’s main technical claim about &lt;strong&gt;cost-adjusted coding-agent performance&lt;/strong&gt;, not just raw benchmark score.&lt;/strong&gt; Commenters focused less on the 3-point lead and more on the pricing efficiency, calling &lt;code&gt;$8.39 vs $21.63&lt;/code&gt; the real headline. They also noted the apparent jump from GPT 5.4 and Terra’s Fable-level score at about one-quarter the cost.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The main technical takeaway was cost-normalized DeepSWE performance: commenters highlighted &lt;strong&gt;GPT 5.6 at &lt;code&gt;73%&lt;/code&gt;&lt;/strong&gt; and framed the result as &lt;strong&gt;&lt;code&gt;$8.39&lt;/code&gt; vs &lt;code&gt;$21.63&lt;/code&gt;&lt;/strong&gt; compared with Fable 5, i.e. a small reported accuracy lead but much larger price advantage. Another commenter noted &lt;strong&gt;Terra tying Fable at roughly &lt;code&gt;1/4&lt;/code&gt; the cost&lt;/strong&gt;, suggesting the benchmark may favor cheaper planner/executor configurations over premium frontier models.&lt;/li&gt;
&lt;li&gt;One user reported real-world MCP-heavy workload costs across model families: &lt;strong&gt;Opus 4.8&lt;/strong&gt; runs reportedly cost &lt;strong&gt;&lt;code&gt;$1–$2&lt;/code&gt;&lt;/strong&gt;, while &lt;strong&gt;GPT 5.5&lt;/strong&gt; cost around &lt;strong&gt;&lt;code&gt;$0.20–$0.50&lt;/code&gt;&lt;/strong&gt; for similar work, implying substantially lower token consumption or pricing for GPT models. They added that &lt;strong&gt;Opus output quality was still “on a different level,”&lt;/strong&gt; so the tradeoff is not purely benchmark score or raw cost.&lt;/li&gt;
&lt;li&gt;A commenter suggested that if the DeepSWE numbers hold, a workflow using &lt;strong&gt;Opus 4.8 high + Sonnet 5 medium&lt;/strong&gt; could potentially be replaced by &lt;strong&gt;Sol high + Terra high&lt;/strong&gt; as planner/executor, with better aggregate results at lower cost. This reflects interest in multi-model routing where cheaper high-reasoning tiers handle decomposition/execution instead of relying on a single premium model.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1urlaam/superhuman_competitive_programming_ai_is_here/&quot;&gt;Superhuman competitive programming AI is here&lt;/a&gt;&lt;/strong&gt; (Activity: 1068): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/32ovkav5b6ch1.jpeg&quot;&gt;image&lt;/a&gt; shows an AtCoder World Tour Finals exhibition leaderboard with &lt;strong&gt;OpenAI&lt;/strong&gt; ranked &lt;code&gt;1st&lt;/code&gt; at &lt;code&gt;8300&lt;/code&gt;, nearly doubling the next competitor &lt;code&gt;tour1st&lt;/code&gt; at &lt;code&gt;4300&lt;/code&gt;, supporting the post’s claim of “superhuman” competitive-programming performance. In the linked Algorithm contest, the poster claims &lt;strong&gt;OpenAI solved all &lt;code&gt;5/5&lt;/code&gt; problems&lt;/strong&gt; while no human solved more than &lt;code&gt;3&lt;/code&gt;, with related AtCoder links for &lt;a href=&quot;https://atcoder.jp/contests/awtf2026heuristic/standings/exhibition&quot;&gt;heuristic standings&lt;/a&gt;, &lt;a href=&quot;https://atcoder.jp/contests/awtf2026heuristic/tasks&quot;&gt;heuristic tasks&lt;/a&gt;, &lt;a href=&quot;https://atcoder.jp/contests/awtf2026algo/standings/exhibition&quot;&gt;algorithm standings&lt;/a&gt;, and &lt;a href=&quot;https://atcoder.jp/contests/awtf2026algo/tasks&quot;&gt;algorithm tasks&lt;/a&gt;.&lt;/strong&gt; Commenters emphasized the size of the gap — &lt;em&gt;“look at that margin”&lt;/em&gt; — while one technical distinction noted this is less general software engineering and more &lt;strong&gt;algorithm design / contest problem solving&lt;/strong&gt;. Another practical caveat is that the AtCoder leaderboards are reportedly behind login.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter draws a technical distinction between &lt;strong&gt;competitive programming&lt;/strong&gt; and broader software engineering: the system appears superhuman at &lt;em&gt;algorithm writing&lt;/em&gt;—a constrained subset of programming focused on solving formal problems under contest conditions—rather than necessarily being superhuman at end-to-end production software development.&lt;/li&gt;
&lt;li&gt;Multiple commenters note that the supporting leaderboard links are &lt;strong&gt;behind a login&lt;/strong&gt;, limiting independent verification of the claimed margin/performance without authenticated access to the benchmark results.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude Code Large-Scale Builds&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uru4zg/jarred_creator_of_bun_rewrote_it_from_zig_to_rust/&quot;&gt;Jarred, creator of Bun rewrote it from Zig to Rust in 11 days using Claude Fable 5 which costed 165k of Fable usage, at API prices. They said by hand, this would&apos;ve taken 3 engineers with full context on the codebase about a year with no other work possible&lt;/a&gt;&lt;/strong&gt; (Activity: 1159): &lt;strong&gt;According to the &lt;a href=&quot;https://bun.com/blog/bun-in-rust&quot;&gt;Bun rewrite post&lt;/a&gt;, &lt;strong&gt;Jarred Sumner&lt;/strong&gt; used a pre-release &lt;strong&gt;Claude Fable 5&lt;/strong&gt; via &lt;strong&gt;Claude Code dynamic workflows&lt;/strong&gt; to port &lt;strong&gt;Bun’s &lt;code&gt;535,496&lt;/code&gt; lines of Zig to Rust&lt;/strong&gt; in &lt;code&gt;11&lt;/code&gt; days, running ~&lt;code&gt;50&lt;/code&gt; workflows with up to &lt;code&gt;64&lt;/code&gt; Claude instances; estimated API-equivalent usage was ~$&lt;code&gt;165k&lt;/code&gt;, versus an estimated &lt;code&gt;3&lt;/code&gt; engineers/year for a manual rewrite. The process used an upfront &lt;code&gt;PORTING.md&lt;/code&gt;, continuous human monitoring, and “adversarial review” with separate Claude contexts acting as reviewers; reported outcomes for Bun &lt;code&gt;v1.4.0&lt;/code&gt; include &lt;code&gt;128&lt;/code&gt; fixed bugs vs &lt;code&gt;v1.3.14&lt;/code&gt;, eliminated instrumentable memory leaks, ~&lt;code&gt;20%&lt;/code&gt; smaller Linux/Windows binaries, and ~&lt;code&gt;10%&lt;/code&gt; faster Linux startup for Claude Code &lt;code&gt;v2.1.181+&lt;/code&gt;.&lt;/strong&gt; Top commenters were skeptical that this demonstrates broad accessibility: they argued the key input was not merely &lt;code&gt;$165k&lt;/code&gt; of model usage but &lt;strong&gt;Jarred’s exceptional codebase context and engineering skill&lt;/strong&gt;, with one framing it as &lt;em&gt;“a million dollar Thiel Fellow engineer who used 165K of Claude Credits.”&lt;/em&gt; Another suggested the API-price framing inflates the perceived cost/scale for effect.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters pushed back on attributing the rewrite primarily to the model spend: the substantive claim was that &lt;strong&gt;Jarred Sumner’s deep Bun/Zig/runtime expertise and full codebase context&lt;/strong&gt; were likely the enabling factor, with the LLM acting as an accelerator rather than an autonomous replacement. One commenter framed it as &lt;em&gt;“Bun was rewritten by a million dollar Thiel Fellow engineer who used &lt;code&gt;$165K&lt;/code&gt; of Claude Credits,”&lt;/em&gt; implying replication cost for a less expert engineer could be far higher.&lt;/li&gt;
&lt;li&gt;Several comments questioned the cost framing, noting that quoting &lt;strong&gt;API pricing&lt;/strong&gt; may inflate the perceived spend versus internal/contracted/discounted usage, and that raw token budget is not equivalent to engineering capability. The technical skepticism was that this result may not generalize: large-scale language/runtime rewrites require architecture judgment, verification, and codebase-specific knowledge that “typical vibe coding” workflows would not supply.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1urzr1q/i_just_made_25k_usd_with_my_capybara_game_built/&quot;&gt;I just made $25K USD with my capybara game built entirely with Claude Code&lt;/a&gt;&lt;/strong&gt; (Activity: 1463): &lt;strong&gt;An iOS engineer built &lt;strong&gt;&lt;a href=&quot;https://capybara-vibejam26.leocoout.dev/&quot;&gt;A Game About Capybaras Delivering Food&lt;/a&gt;&lt;/strong&gt; in &lt;code&gt;15&lt;/code&gt; days for &lt;strong&gt;&lt;a href=&quot;https://vibej.am/2026/#games&quot;&gt;VibeJam 2026&lt;/a&gt;&lt;/strong&gt;, winning the &lt;code&gt;$25,000&lt;/code&gt; first prize; the project used &lt;strong&gt;Claude Code Opus 4.7&lt;/strong&gt;, &lt;strong&gt;Three.js&lt;/strong&gt;, &lt;strong&gt;GPT Images-2/Grok&lt;/strong&gt; for textures, &lt;strong&gt;Tripo3d&lt;/strong&gt; for models, and &lt;strong&gt;Suno/ElevenLabs&lt;/strong&gt; for audio, with claimed &lt;code&gt;100%&lt;/code&gt; AI-written code across &lt;code&gt;188&lt;/code&gt; commits and &lt;code&gt;~27k&lt;/code&gt; LOC. The workflow centered on parallel Claude Code sessions, &lt;code&gt;/plan&lt;/code&gt;, and AI-generated tooling: an in-game map/terrain/road editor, cutscene editor, iOS-like phone UI, PS1-style texture pipeline, mission loop, stacked-item pseudo-physics, vehicle drifting/collision, localization, and a Cloudflare WebSocket multiplayer lobby relaying player state at &lt;code&gt;~10 Hz&lt;/code&gt; with &lt;code&gt;O(n²)&lt;/code&gt; fanout scaling.&lt;/strong&gt; Top comments were mostly non-technical: one joked that Claude often suggests “capybara” as a mascot, while another questioned the title’s phrasing, noting the money came from a competition prize rather than game revenue.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Frontier Model Usage Limits&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/1uscohi/gpt56_sol_ultra_is_impressive_for_the_12_minutes/&quot;&gt;GPT-5.6 Sol Ultra is impressive — for the 12 minutes you’re allowed to use it as a Plus subscriber&lt;/a&gt;&lt;/strong&gt; (Activity: 914): &lt;strong&gt;A &lt;strong&gt;ChatGPT Plus&lt;/strong&gt; user reports that using &lt;strong&gt;GPT-5.6 Sol Ultra&lt;/strong&gt; for two large batch/agentic workloads—merging/analyzing ~&lt;code&gt;10&lt;/code&gt; PDFs into a ~&lt;code&gt;700&lt;/code&gt;-page output and reorganizing ~&lt;code&gt;700&lt;/code&gt; Markdown files in an Obsidian vault—exhausted their Plus usage allowance despite a reset. The main technical rebuttal argues the workload likely involved &lt;strong&gt;millions of processed tokens&lt;/strong&gt;: ~&lt;code&gt;280k–560k&lt;/code&gt; output tokens for the 700-page document alone, plus ~&lt;code&gt;210k–1.05M&lt;/code&gt; tokens for a single pass over 700 Markdown files, before planning, rereads, rewrites, retries, or multi-agent overhead.&lt;/strong&gt; Commenters largely push back on measuring cost by prompt count, arguing that “two tasks” can represent very large compute/token consumption; the clearest shared criticism is that OpenAI’s &lt;strong&gt;quota meter is too vague&lt;/strong&gt;, even if the throttling itself is economically expected for a &lt;code&gt;$20/mo&lt;/code&gt; plan.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued the reported limit burn is better explained by &lt;strong&gt;token/compute consumption rather than prompt count&lt;/strong&gt;: a &lt;code&gt;700-page&lt;/code&gt; generated report could represent roughly &lt;code&gt;280k–560k output tokens&lt;/code&gt;, and processing &lt;code&gt;700 Markdown files&lt;/code&gt; at &lt;code&gt;300–1,500 tokens/file&lt;/code&gt; adds another &lt;code&gt;210k–1.05M input tokens&lt;/code&gt; per pass. With planning, rereads, rewrites, retries, and multi-agent handoffs in &lt;strong&gt;Sol Ultra&lt;/strong&gt;, commenters estimated the workload could plausibly reach &lt;strong&gt;several million processed tokens&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;A technical criticism was that &lt;strong&gt;Plus quota UX is opaque&lt;/strong&gt;: users see a vague usage meter rather than a compute/token-based accounting model. Commenters suggested the complaint is valid insofar as OpenAI exposes limits as “messages” or time windows, while high-context batch jobs on an expensive multi-agent mode can consume quota disproportionately quickly.&lt;/li&gt;
&lt;li&gt;One practical recommendation was to avoid using &lt;strong&gt;Ultra&lt;/strong&gt; for large context-heavy batch workflows unless the goal is benchmarking; commenters noted that tasks involving hundreds of documents and long-form synthesis are likely inefficient under capped consumer subscriptions, even if the apparent number of prompts is small.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1urzmj0/5_hour_and_weekly_limits_have_been_reset_thanks/&quot;&gt;5 hour and weekly limits have been reset. Thanks Anthropic!&lt;/a&gt;&lt;/strong&gt; (Activity: 2865): &lt;strong&gt;The image is a dark-mode X/Twitter screenshot from &lt;strong&gt;ClaudeDevs&lt;/strong&gt; announcing: &lt;em&gt;“We’ve reset 5-hour and weekly rate limits for all users”&lt;/em&gt; (&lt;a href=&quot;https://i.redd.it/djfpk4js49ch1.jpeg&quot;&gt;image&lt;/a&gt;). Technically, this means &lt;strong&gt;Claude/Anthropic users’ short-window and weekly quota counters were cleared&lt;/strong&gt;, allowing renewed usage immediately; the post asks whether this was goodwill, competitive timing, or related to a possible &lt;strong&gt;5.6&lt;/strong&gt; update.&lt;/strong&gt; Comments were mostly speculative: some joked the timing suggested pressure from &lt;strong&gt;OpenAI&lt;/strong&gt;, while others regretted not exhausting their usage before the reset but appreciated the free quota refresh.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</content:encoded><category>openai</category><category>gpt-5.6</category><category>claude-fable-5</category><category>reach_vb</category><category>rasbt</category><category>yuchenj_uw</category><category>scaling01</category><category>simonw</category><category>kimmonismus</category><category>thsottiaux</category><category>htihle</category><category>teortaxestex</category><category>mononofu</category><category>omarsar0</category><category>hangsiin</category><category>gdb</category><category>mckbrando</category><category>evi77ain</category><category>model-stratification</category><category>agentic-coding</category><category>presentation</category><category>benchmarking</category><category>orchestration</category><category>computer-use</category><category>gui-automation</category><category>reward-hacking</category><category>instruction-following</category><category>usage-limits</category><category>model-costs</category></item><item><title>OpenAI launches GPT 5.6 Sol/Terra/Luna</title><link>https://news.smol.ai/issues/26-07-09-gpt-56/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-09-gpt-56/</guid><description>**OpenAI** launched the **GPT-5.6** family with three models: **Sol**, **Terra**, and **Luna**, integrated across **ChatGPT**, **Codex**, and the API. Pricing tiers range from **$1 to $5 per million tokens** with new cache-write pricing and a 90% cache-read discount. The launch includes new app features like **ChatGPT Work**, a desktop app merging Codex and ChatGPT, **Sites beta**, programmatic tool calling, and multi-agent beta. **Sam Altman** called GPT-5.6 Sol &quot;*the best model we have ever produced*&quot; with strong agentic and coding performance, improved artifact quality, and better economics. Independent evaluations show Sol near the frontier on coding-agent workloads with an Intelligence Index score of **59**, slightly below Claude Fable 5 but at about one-third the cost. Terra and Luna offer lower-cost alternatives with competitive performance.</description><pubDate>Thu, 09 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/08/2026-7/09/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;OpenAI launched a new three-model GPT‑5.6 family and simultaneously expanded the product stack around it.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI announced &lt;strong&gt;GPT‑5.6 Sol, Terra, and Luna&lt;/strong&gt; rolling out across &lt;strong&gt;ChatGPT, Codex, and the API&lt;/strong&gt; via &lt;a href=&quot;https://x.com/OpenAI/status/2075271421149020426&quot;&gt;@OpenAI&lt;/a&gt; and &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075273992609599834&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;In ChatGPT, &lt;strong&gt;Plus, Pro, Business, and Enterprise&lt;/strong&gt; users get access to &lt;strong&gt;GPT‑5.6 Sol&lt;/strong&gt; through medium+ effort settings, while &lt;strong&gt;Pro and Enterprise&lt;/strong&gt; can select &lt;strong&gt;GPT‑5.6 Pro&lt;/strong&gt; for highest-quality results on complex tasks, per &lt;a href=&quot;https://x.com/OpenAI/status/2075271435573244008&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;API pricing introduced a tiered lineup: &lt;strong&gt;Sol $5 / $30 per million input/output tokens&lt;/strong&gt;, &lt;strong&gt;Terra $2.5 / $15&lt;/strong&gt;, &lt;strong&gt;Luna $1 / $6&lt;/strong&gt;, with &lt;strong&gt;cache-write pricing&lt;/strong&gt; added for the first time and &lt;strong&gt;90% cache-read discount&lt;/strong&gt; retained, according to &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI framed the family around a price-performance ladder: &lt;strong&gt;Sol = flagship/highest ceiling&lt;/strong&gt;, &lt;strong&gt;Terra = GPT‑5.5-like capability at lower cost&lt;/strong&gt;, &lt;strong&gt;Luna = fastest/cheapest high-volume option&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075286157186003348&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The launch bundled major app-layer changes: &lt;strong&gt;ChatGPT Work&lt;/strong&gt;, a new &lt;strong&gt;desktop app merging Codex + ChatGPT&lt;/strong&gt;, &lt;strong&gt;Sites&lt;/strong&gt; beta, &lt;strong&gt;programmatic tool calling&lt;/strong&gt;, and &lt;strong&gt;multi-agent beta&lt;/strong&gt; in the Responses API, via &lt;a href=&quot;https://x.com/OpenAI/status/2075274271845404744&quot;&gt;@OpenAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075275868268789885&quot;&gt;@OpenAIDevs&lt;/a&gt;, and &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075274093327470923&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Official claims and benchmark results&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;OpenAI’s official message emphasized strong agentic/coding performance, better artifact quality, and improved economics.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sam Altman called it “&lt;strong&gt;obviously the best model we have ever produced&lt;/strong&gt;” in the launch post, linking the release blog, via &lt;a href=&quot;https://x.com/sama/status/2075266471316615436&quot;&gt;@sama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Altman also highlighted enterprise economics: “&lt;strong&gt;5.6 sol is a huge step forward for dollars-per-task&lt;/strong&gt;,” via &lt;a href=&quot;https://x.com/sama/status/2075267201058426944&quot;&gt;@sama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Greg Brockman said the goal is “&lt;strong&gt;the best price for any level of target performance&lt;/strong&gt;” and the highest possible ceiling, via &lt;a href=&quot;https://x.com/gdb/status/2075271293474353553&quot;&gt;@gdb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI claimed &lt;strong&gt;GPT‑5.6 Sol sets a new high of 53.6 on Agents’ Last Exam&lt;/strong&gt;, beating &lt;strong&gt;Claude Fable 5 adaptive by 13.1 points&lt;/strong&gt;; at medium reasoning it beats Fable by &lt;strong&gt;11.4 points at roughly one-quarter the estimated cost&lt;/strong&gt;, while &lt;strong&gt;Terra and Luna also outperform Fable at around one-sixteenth the cost&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/OpenAI/status/2075271423992680532&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI said GPT‑5.6 improves &lt;strong&gt;artifact quality across presentations, documents, and spreadsheets&lt;/strong&gt;, with outputs exportable into existing enterprise tools, via &lt;a href=&quot;https://x.com/OpenAI/status/2075271432041545782&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI positioned GPT‑5.6 as state of the art for &lt;strong&gt;reasoning through complex tasks&lt;/strong&gt; and for producing materials matched to templates, reference files, and preferred style inside &lt;strong&gt;ChatGPT Work&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/OpenAI/status/2075274275104399670&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI also said GPT‑5.6 is its &lt;strong&gt;most capable model yet on cyber and bio-related tasks&lt;/strong&gt;, with some API calls potentially blocked or paused for extra safety review in dual-use areas, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075274080740380829&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI highlighted better &lt;strong&gt;Computer Use&lt;/strong&gt; performance: faster, more token-efficient, support for &lt;strong&gt;batching and parallel operations&lt;/strong&gt; across multi-step tasks, plus picture-in-picture supervision, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075276074980884862&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Independent evaluations and third-party measurements&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Independent evals broadly placed Sol near or at the frontier, especially on coding-agent workloads, while also surfacing caveats.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt; reported &lt;strong&gt;GPT‑5.6 Sol (max)&lt;/strong&gt; scores &lt;strong&gt;59&lt;/strong&gt; on its Intelligence Index, &lt;strong&gt;1 point below Claude Fable 5 (max)&lt;/strong&gt;, at &lt;strong&gt;about one-third of Fable’s cost per task&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;On the same analysis, &lt;strong&gt;Terra&lt;/strong&gt; and &lt;strong&gt;Luna&lt;/strong&gt; score &lt;strong&gt;55&lt;/strong&gt; and &lt;strong&gt;51&lt;/strong&gt; on the Intelligence Index, with &lt;strong&gt;~50%&lt;/strong&gt; and &lt;strong&gt;~80%&lt;/strong&gt; lower cost per task than Sol, respectively, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis said &lt;strong&gt;Sol leads the Coding Agent Index at 80&lt;/strong&gt;, ahead of Fable 5 and Opus 4.8, and is also cheaper per task than both on their harnesses, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;It also noted &lt;strong&gt;Sol defines a new Pareto frontier of intelligence vs output tokens&lt;/strong&gt;, while &lt;strong&gt;Terra and Luna are not on that frontier&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268984539410521&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis found &lt;strong&gt;minor improvement over GPT‑5.5 in AA‑Omniscience&lt;/strong&gt; but with a &lt;strong&gt;higher hallucination rate&lt;/strong&gt; than GPT‑5.5 max, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268990004605023&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;It reported &lt;strong&gt;similar GDPval-AA v2 performance to Claude Fable 5&lt;/strong&gt;, suggesting comparable ability on economically valuable tasks, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268987550932998&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ValsAI/status/2075270642359029972&quot;&gt;@ValsAI&lt;/a&gt; ranked GPT‑5.6 &lt;strong&gt;#2 on Vals Index and Vals Multimodal Index&lt;/strong&gt;, saying Fable 5 remains ahead on several benchmarks but GPT‑5.6 is “clearly in the same class”&lt;/li&gt;
&lt;li&gt;Vals also said &lt;strong&gt;Sol is #1 on CyberBench and Excel Modeling Benchmark&lt;/strong&gt;, and #1 on &lt;strong&gt;Legal Research Bench, ProofBench, SWE-bench, and Terminal-Bench 2.1&lt;/strong&gt;, adding that Fable had a nearly &lt;strong&gt;100% refusal rate on CyberBench&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/ValsAI/status/2075270644711997581&quot;&gt;@ValsAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/arcprize/status/2075270869992264003&quot;&gt;@arcprize&lt;/a&gt; said &lt;strong&gt;GPT‑5.6 Sol scores 7.8% on ARC‑AGI‑3&lt;/strong&gt; and is the &lt;strong&gt;first verified frontier model to ever beat an ARC‑AGI‑3 game&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/GregKamradt/status/2075274981794300113&quot;&gt;@GregKamradt&lt;/a&gt; noted &lt;strong&gt;92.5% on ARC‑AGI‑2&lt;/strong&gt;, calling it SOTA while costing &lt;strong&gt;an order of magnitude less&lt;/strong&gt; than GPT‑5.5 Pro three months earlier&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075423964378366427&quot;&gt;@ArtificialAnlys&lt;/a&gt; later reported &lt;strong&gt;GPT‑5.6 Sol (max) leads CritPt&lt;/strong&gt;, a benchmark of unpublished research-level physics problems, by roughly &lt;strong&gt;4 points over Claude Fable 5&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/llama_index/status/2075351095258296378&quot;&gt;@llama_index&lt;/a&gt; said day-0 ParseBench results show GPT‑5.6 continues to do well on &lt;strong&gt;text and tables&lt;/strong&gt; but still struggles on &lt;strong&gt;charts and layout&lt;/strong&gt;, and that &lt;strong&gt;Luna is ~6× cheaper than Sol with only minor degradations&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/jerryjliu0/status/2075356305099800717&quot;&gt;@jerryjliu0&lt;/a&gt; similarly said ParseBench shows &lt;strong&gt;no high-level change versus GPT‑5.5&lt;/strong&gt; on tables/text/charts/layout, stressing persistent weakness on &lt;strong&gt;complex text layouts, chart transcription, and source-element bounding boxes&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Technical details&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The technical story of GPT‑5.6 is as much about inference orchestration and token efficiency as raw capability.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI shipped &lt;strong&gt;three model tiers&lt;/strong&gt; with multiple &lt;strong&gt;reasoning effort levels&lt;/strong&gt;; users discussed &lt;strong&gt;Light, Medium, High, Extra High, Ultra&lt;/strong&gt;, leading to a large configuration matrix, via &lt;a href=&quot;https://x.com/rasbt/status/2075369179817902176&quot;&gt;@rasbt&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI added &lt;strong&gt;Programmatic Tool Calling&lt;/strong&gt; in the Responses API and &lt;strong&gt;Multi-agent beta&lt;/strong&gt;, indicating more explicit support for orchestrated tool use and agent decomposition, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075274093327470923&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI’s app layer now uses &lt;strong&gt;Codex as the core&lt;/strong&gt; of the new Work product, per &lt;a href=&quot;https://x.com/sama/status/2075293792048136572&quot;&gt;@sama&lt;/a&gt; and &lt;a href=&quot;https://x.com/gdb/status/2075276416686723110&quot;&gt;@gdb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Several posts stress &lt;strong&gt;parallel agents/subagents&lt;/strong&gt; as a major capability lever; &lt;a href=&quot;https://x.com/aidan_mclau/status/2075337767949865464&quot;&gt;@aidan_mclau&lt;/a&gt; explicitly mentions users can increase the number of &lt;strong&gt;5.6 subagents&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/LiorOnAI/status/2075277748394967122&quot;&gt;@LiorOnAI&lt;/a&gt; summarized likely drivers as &lt;strong&gt;adaptive reasoning&lt;/strong&gt;, &lt;strong&gt;parallel agents&lt;/strong&gt;, &lt;strong&gt;programmatic tool use&lt;/strong&gt;, and &lt;strong&gt;higher token efficiency&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis reported &lt;strong&gt;Sol max uses ~15k output tokens per Intelligence Index task vs 16k for GPT‑5.5&lt;/strong&gt;, and fewer than Opus 4.8, GLM‑5.2, and Gemini 3.5 Flash at comparable intelligence, via &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/OpenRouter/status/2075271807855452196&quot;&gt;@OpenRouter&lt;/a&gt; said early testing found the 5.6 models &lt;strong&gt;more token efficient&lt;/strong&gt;, lowering both cost and time-to-task completion&lt;/li&gt;
&lt;li&gt;The desktop/app layer brought a &lt;strong&gt;Chrome extension&lt;/strong&gt;, &lt;strong&gt;revamped in-app browser&lt;/strong&gt;, &lt;strong&gt;authenticated sites&lt;/strong&gt;, &lt;strong&gt;persistent multi-tab sessions&lt;/strong&gt;, &lt;strong&gt;file downloads&lt;/strong&gt;, and tighter cross-device handoffs, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075275868268789885&quot;&gt;@OpenAIDevs&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075276009902112976&quot;&gt;@OpenAIDevs&lt;/a&gt;, and &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075292716737736919&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sites&lt;/strong&gt; entered beta for paid users, offering hosting, storage, and optional auth for GPT-built apps, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075275892591591469&quot;&gt;@OpenAIDevs&lt;/a&gt; and &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075337081304522853&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;The “Sol autonomously post-trained Luna” claim&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;This was the most provocative technical claim around the launch, but its interpretation became contested almost immediately.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple accounts amplified the statement that &lt;strong&gt;OpenAI says GPT‑5.6 Sol autonomously post-trained GPT‑5.6 Luna&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/scaling01/status/2075269113488789984&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/tejalpatwardhan/status/2075272564629451110&quot;&gt;@tejalpatwardhan&lt;/a&gt;, and &lt;a href=&quot;https://x.com/dejavucoder/status/2075270116909232129&quot;&gt;@dejavucoder&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The claim fueled RSI/autoresearch speculation; &lt;a href=&quot;https://x.com/tenobrus/status/2075282678652522712&quot;&gt;@tenobrus&lt;/a&gt; said if true as stated, it would be a “pretty large update” for automated researcher timelines&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/eliebakouch/status/2075281402807844872&quot;&gt;@eliebakouch&lt;/a&gt; framed it as OpenAI asking Sol to post-train Luna “with &lt;strong&gt;100k GPUs&lt;/strong&gt;” for an experiment&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/gdb/status/2075363531042726216&quot;&gt;@gdb&lt;/a&gt; said the implication is easy to overlook for accelerating engineering workflows, reinforcing that OpenAI wants this read as more than a marketing flourish&lt;/li&gt;
&lt;li&gt;But skeptical clarifications emerged quickly: &lt;a href=&quot;https://x.com/nikolaj2030/status/2075297831376793764&quot;&gt;@nikolaj2030&lt;/a&gt; asked whether this actually meant Sol completed a &lt;strong&gt;small controlled post-training task&lt;/strong&gt;—modifying a config, editing a scheduler file, and launching a run—rather than end-to-end real-world post-training of Luna&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/nrehiew_/status/2075316190386462888&quot;&gt;@nrehiew_&lt;/a&gt; interpreted the screenshot similarly: Sol could go from high-level ideas to &lt;strong&gt;editing configs and launching experiments&lt;/strong&gt;, not fully owning Luna’s end-to-end post-training&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/scaling01/status/2075354327791587467&quot;&gt;@scaling01&lt;/a&gt; argued that what’s probably happening is a model implementing &lt;strong&gt;LLM-as-a-judge graders&lt;/strong&gt;, reward-shaping logic, or small training configs on top of existing OpenAI RL infrastructure—not autonomous end-to-end research or training systems&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/scaling01/status/2075359429717836251&quot;&gt;@scaling01&lt;/a&gt; explicitly said we should distance these statements from &lt;strong&gt;literal autonomous end-to-end post-training or research&lt;/strong&gt;, which models still cannot do&lt;/li&gt;
&lt;li&gt;Counterbalancing that skepticism, &lt;a href=&quot;https://x.com/aidan_mclau/status/2075328409400738229&quot;&gt;@aidan_mclau&lt;/a&gt; said it is routine for him to have &lt;strong&gt;5.6 e2e do an entire RL run&lt;/strong&gt;, suggesting meaningful internal workflow automation even if not self-sufficient research&lt;/li&gt;
&lt;li&gt;The consensus across technical observers was not that Sol independently invented and trained Luna, but that GPT‑5.6 may now be capable of &lt;strong&gt;executing meaningful chunks of model-improvement workflows inside mature internal infrastructure&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Internal productivity and recursive improvement signals&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;OpenAI also used internal-usage data to argue that GPT‑5.6 materially changes researcher throughput.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/scaling01/status/2075269455781703850&quot;&gt;@scaling01&lt;/a&gt; highlighted an OpenAI claim that it &lt;strong&gt;doubled experiment throughput per researcher&lt;/strong&gt; since the start of the year&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/eliebakouch/status/2075273299148341327&quot;&gt;@eliebakouch&lt;/a&gt; quoted OpenAI saying average daily output tokens per active researcher were &lt;strong&gt;more than twice the highest level observed for GPT‑5.5&lt;/strong&gt; during internal testing&lt;/li&gt;
&lt;li&gt;Another OpenAI stat, relayed by &lt;a href=&quot;https://x.com/eliebakouch/status/2075273992185782661&quot;&gt;@eliebakouch&lt;/a&gt;, said over six months the share of research compute devoted to &lt;strong&gt;internal coding inference grew 100-fold&lt;/strong&gt;, while &lt;strong&gt;internal agentic token usage increased ~22-fold&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/FakePsyho/status/2075291659814781370&quot;&gt;@FakePsyho&lt;/a&gt; linked these developments to OpenAI’s performance in top programming contests, describing systems close to GPT‑5.6 plus custom harnesses as decisively beating elite human competitors&lt;/li&gt;
&lt;li&gt;This fed broader RSI/autoresearch discussion, especially from people who see long-horizon coding and heuristic optimization as proxies for model-improvement capability&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Product implications: ChatGPT Work, Codex merge, desktop, and Sites&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The model launch doubled as a product strategy reset: OpenAI is pushing from “chatbot” to “work OS.”&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI launched &lt;strong&gt;ChatGPT Work&lt;/strong&gt;, an agent powered by &lt;strong&gt;Codex + GPT‑5.6&lt;/strong&gt; that can act across apps and files, stay on tasks for hours, and turn a goal into finished work, via &lt;a href=&quot;https://x.com/OpenAI/status/2075274271845404744&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Work can ingest context from &lt;strong&gt;docs, Slack, Notion, Microsoft 365, and Google Drive&lt;/strong&gt; and produce &lt;strong&gt;decks, docs, spreadsheets, dashboards, visualizations, and interactive explanations&lt;/strong&gt;, summarized by &lt;a href=&quot;https://x.com/kimmonismus/status/2075271465964798147&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;Codex app merged into the new ChatGPT desktop app&lt;/strong&gt;, confirmed by &lt;a href=&quot;https://x.com/avstorm/status/2075266403297362364&quot;&gt;@avstorm&lt;/a&gt; and &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075275880704995342&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Developers now get &lt;strong&gt;inline diff editing&lt;/strong&gt;, &lt;strong&gt;PR review side panel&lt;/strong&gt;, better &lt;strong&gt;SSH video rendering&lt;/strong&gt;, and stronger &lt;strong&gt;computer use&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/romainhuet/status/2075286364476850430&quot;&gt;@romainhuet&lt;/a&gt; and &lt;a href=&quot;https://x.com/reach_vb/status/2075280626362560805&quot;&gt;@reach_vb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sites&lt;/strong&gt; lets users turn work into shareable hosted apps/websites from ChatGPT, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075275892591591469&quot;&gt;@OpenAIDevs&lt;/a&gt; and &lt;a href=&quot;https://x.com/simpsoka/status/2075278935366287842&quot;&gt;@simpsoka&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/OpenAI/status/2075310019185389913&quot;&gt;@OpenAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAI/status/2075310020653351324&quot;&gt;@OpenAI&lt;/a&gt;, and &lt;a href=&quot;https://x.com/OpenAI/status/2075310022121472399&quot;&gt;@OpenAI&lt;/a&gt; marketed GPT‑5.6 through case studies: a &lt;strong&gt;broccoli farmer&lt;/strong&gt;, a &lt;strong&gt;mathematician&lt;/strong&gt;, and a &lt;strong&gt;family cereal business&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;This product reframing was read by some as OpenAI’s answer to Anthropic’s Cowork / Claude Code stack, via &lt;a href=&quot;https://x.com/jerryjliu0/status/2075295459304710496&quot;&gt;@jerryjliu0&lt;/a&gt; and &lt;a href=&quot;https://x.com/kimmonismus/status/2075280933452669000&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Facts vs opinions&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Facts / directly sourced claims&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPT‑5.6 family names, rollout channels, and access tiers: &lt;a href=&quot;https://x.com/OpenAI/status/2075271421149020426&quot;&gt;@OpenAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAI/status/2075271435573244008&quot;&gt;@OpenAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075273992609599834&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;API prices and cache-write policy: &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI’s benchmark claims on Agents’ Last Exam: &lt;a href=&quot;https://x.com/OpenAI/status/2075271423992680532&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis and Vals leaderboard placements: &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/ValsAI/status/2075270642359029972&quot;&gt;@ValsAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;ARC‑AGI‑3 7.8% claim: &lt;a href=&quot;https://x.com/arcprize/status/2075270869992264003&quot;&gt;@arcprize&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;ParseBench caveats: &lt;a href=&quot;https://x.com/llama_index/status/2075351095258296378&quot;&gt;@llama_index&lt;/a&gt;, &lt;a href=&quot;https://x.com/jerryjliu0/status/2075356305099800717&quot;&gt;@jerryjliu0&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Safety testing finding jailbreaks on GPT‑5.6 Sol: &lt;a href=&quot;https://x.com/alxndrdavies/status/2075279477626564933&quot;&gt;@alxndrdavies&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Opinions / interpretation / hype&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Best model we have ever produced”: &lt;a href=&quot;https://x.com/sama/status/2075266471316615436&quot;&gt;@sama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“First time I’ve felt comfortable delegating the hardest problem out there”: &lt;a href=&quot;https://x.com/reach_vb/status/2075269547439907269&quot;&gt;@reach_vb&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Not enough people are emotionally prepared for GPT‑6”: &lt;a href=&quot;https://x.com/scaling01/status/2075276735650648258&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“OpenAI is competing on cost curves, not benchmarks”: &lt;a href=&quot;https://x.com/LiorOnAI/status/2075277748394967122&quot;&gt;@LiorOnAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“The engineers were allowed to cook”: &lt;a href=&quot;https://x.com/TheHumanoidHub/status/2075272514755059773&quot;&gt;@TheHumanoidHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Generational fumble” regarding Codex becoming ChatGPT Desktop: &lt;a href=&quot;https://x.com/theo/status/2075312087723876556&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Different perspectives&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Supportive views&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Many developers and evaluators saw GPT‑5.6 as a meaningful frontier advance, especially in coding and knowledge work: &lt;a href=&quot;https://x.com/gdb/status/2075270503405924466&quot;&gt;@gdb&lt;/a&gt;, &lt;a href=&quot;https://x.com/AravSrinivas/status/2075270640177938547&quot;&gt;@AravSrinivas&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenRouter/status/2075271807855452196&quot;&gt;@OpenRouter&lt;/a&gt;, &lt;a href=&quot;https://x.com/Teknium/status/2075392507794624803&quot;&gt;@Teknium&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Several posts focused on &lt;strong&gt;cost efficiency&lt;/strong&gt; as the real win, with Sol matching frontier peers while being materially cheaper: &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/omarsar0/status/2075270117131259925&quot;&gt;@omarsar0&lt;/a&gt;, &lt;a href=&quot;https://x.com/cline/status/2075278343927365991&quot;&gt;@cline&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Others highlighted the &lt;strong&gt;agentic stack&lt;/strong&gt;—Work, Codex, multi-agent, programmatic tools—as more strategically important than raw benchmark deltas: &lt;a href=&quot;https://x.com/TheRundownAI/status/2075273458661949763&quot;&gt;@TheRundownAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075271465964798147&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/fidjissimo/status/2075305622120325363&quot;&gt;@fidjissimo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Neutral / analytical views&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Some analysts saw Sol as roughly &lt;strong&gt;same class as Fable&lt;/strong&gt;, but not decisively ahead overall: &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/ValsAI/status/2075270642359029972&quot;&gt;@ValsAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/teortaxesTex/status/2075274583226069040&quot;&gt;@teortaxesTex&lt;/a&gt; argued the release may reflect OpenAI strong post-training recovering toward Anthropic despite a stronger Anthropic base model&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/simonw/status/2075306164993315192&quot;&gt;@simonw&lt;/a&gt; pointed to notable API additions but also implied growing product complexity&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Critical / skeptical views&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/scaling01/status/2075268278105067566&quot;&gt;@scaling01&lt;/a&gt; asked whether &lt;strong&gt;GPT‑5.6 Sol is worse at math&lt;/strong&gt;, pushing back on the “everything got better” narrative&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268990004605023&quot;&gt;@ArtificialAnlys&lt;/a&gt; found &lt;strong&gt;higher hallucination rate vs GPT‑5.5&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/scaling01/status/2075279452494299273&quot;&gt;@scaling01&lt;/a&gt; criticized the ARC‑AGI‑3 scoring setup, saying Sol would score &lt;strong&gt;0% under official scoring methodology capped at $10k&lt;/strong&gt; and objecting to use of a &lt;strong&gt;$25k&lt;/strong&gt; budget&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/Hangsiin/status/2075277820528607704&quot;&gt;@Hangsiin&lt;/a&gt; and &lt;a href=&quot;https://x.com/Hangsiin/status/2075278682160275561&quot;&gt;@Hangsiin&lt;/a&gt; pointed to &lt;strong&gt;subscription/credit confusion&lt;/strong&gt;, saying Sol costs more credits than GPT‑5.5 while usage limits differ less than API pricing suggests&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/QuinnyPig/status/2075334468462899442&quot;&gt;@QuinnyPig&lt;/a&gt; said OpenAI’s pricing/subscription strategy is confusing, particularly around future pricing jumps or inclusion terms&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/rasbt/status/2075369179817902176&quot;&gt;@rasbt&lt;/a&gt; highlighted UX complexity: &lt;strong&gt;2 modes × 3 models × 5 effort levels = 30 configurations&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/MParakhin/status/2075361980446289925&quot;&gt;@MParakhin&lt;/a&gt; complained that &lt;strong&gt;GPT‑5.6 Pro no longer has extended thinking&lt;/strong&gt;, preferring an option to pay for much longer reasoning&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/theo/status/2075312087723876556&quot;&gt;@theo&lt;/a&gt; and &lt;a href=&quot;https://x.com/simonw/status/2075348941215006888&quot;&gt;@simonw&lt;/a&gt; criticized the growing app/mode fragmentation around ChatGPT, Codex, and Work&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Safety and security concerns&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;The launch also surfaced one of the strongest public cyber-safety debates around a recent frontier model release.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/alxndrdavies/status/2075279477626564933&quot;&gt;@alxndrdavies&lt;/a&gt; from the AI Safety Institute said they found &lt;strong&gt;universal jailbreaks in all rounds of testing&lt;/strong&gt; that enabled long-form agentic task completion in &lt;strong&gt;vulnerability discovery and exploit development&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/EthanJPerez/status/2075296476817985751&quot;&gt;@EthanJPerez&lt;/a&gt; called it “&lt;strong&gt;the highest stakes safety issue of any model release yet&lt;/strong&gt;”&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/yonashav/status/2075286161241612664&quot;&gt;@yonashav&lt;/a&gt; praised OpenAI for allowing third-party unreleased-model safety assessments to be published even when inconvenient&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/Mononofu/status/2075414796426764507&quot;&gt;@Mononofu&lt;/a&gt; said ease of jailbreaking plus reward-hacking reports make them worried OpenAI may have rushed the release to keep pace with Fable&lt;/li&gt;
&lt;li&gt;At the same time, OpenAI explicitly warned some cyber/bio requests may be paused or blocked mid-stream for additional review, via &lt;a href=&quot;https://x.com/OpenAIDevs/status/2075274080740380829&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;This created a split narrative: strong cyber capability is treated as a product advantage by some evaluators, but as a serious deployment risk by safety researchers&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Context&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Why this matters goes beyond a single model benchmark win.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The launch happened amid a compressed week of frontier competition that also included new releases from &lt;strong&gt;Meta Muse Spark 1.1&lt;/strong&gt; and &lt;strong&gt;Grok 4.5&lt;/strong&gt;, leading multiple observers to describe the frontier as newly crowded: &lt;a href=&quot;https://x.com/matanSF/status/2075276339607654802&quot;&gt;@matanSF&lt;/a&gt;, &lt;a href=&quot;https://x.com/kimmonismus/status/2075322537592922345&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;OpenAI’s differentiation is increasingly framed less as “best raw benchmark score” and more as &lt;strong&gt;cost-efficient agentic work&lt;/strong&gt;, consistent with posts from &lt;a href=&quot;https://x.com/sama/status/2075267201058426944&quot;&gt;@sama&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2075268970492657905&quot;&gt;@ArtificialAnlys&lt;/a&gt;, and &lt;a href=&quot;https://x.com/LiorOnAI/status/2075277748394967122&quot;&gt;@LiorOnAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The product bundling suggests OpenAI is moving from a model vendor to a &lt;strong&gt;full-stack work platform&lt;/strong&gt;, with its own browser, connectors, orchestration primitives, hosted app deployment, and desktop runtime&lt;/li&gt;
&lt;li&gt;The strongest forward-looking signal may be the internal claim that researchers already use these systems to materially increase output and automate chunks of RL/post-training workflows, even if public discussion often overstates that as “the model trained itself”&lt;/li&gt;
&lt;li&gt;The launch also sharpens a recurring engineering question raised by many tweets: whether the frontier is now bottlenecked less by a single monolithic model and more by &lt;strong&gt;orchestration quality, tool APIs, subagents, evaluation harnesses, and economics&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Frontier models and evaluations&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Meta launched Muse Spark 1.1&lt;/strong&gt; and the &lt;strong&gt;Meta Model API&lt;/strong&gt; in public preview, positioning it as a strong &lt;strong&gt;agentic, coding, multimodal, and computer-use&lt;/strong&gt; model. Official posts came from &lt;a href=&quot;https://x.com/finkd/status/2075218444056707458&quot;&gt;@finkd&lt;/a&gt;, &lt;a href=&quot;https://x.com/alexandr_wang/status/2075218936266998230&quot;&gt;@alexandr_wang&lt;/a&gt;, &lt;a href=&quot;https://x.com/shengjia_zhao/status/2075220782465290620&quot;&gt;@shengjia_zhao&lt;/a&gt;, &lt;a href=&quot;https://x.com/ren_hongyu/status/2075224643829711101&quot;&gt;@ren_hongyu&lt;/a&gt;, and &lt;a href=&quot;https://x.com/MetaforDevs/status/2075268072022401526&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Key technical details repeatedly cited: &lt;strong&gt;1M-token context window&lt;/strong&gt;, &lt;strong&gt;video understanding&lt;/strong&gt;, multimodal reasoning, and API availability, with &lt;a href=&quot;https://x.com/altryne/status/2075237837033889911&quot;&gt;@altryne&lt;/a&gt; and &lt;a href=&quot;https://x.com/xinyun_chen_/status/2075276047495659656&quot;&gt;@xinyun_chen_&lt;/a&gt; among those emphasizing long-horizon agentic gains&lt;/li&gt;
&lt;li&gt;Benchmark claims around Muse Spark 1.1 included competitiveness with &lt;strong&gt;GPT‑5.5&lt;/strong&gt; and &lt;strong&gt;Opus 4.8&lt;/strong&gt; on agentic evals, strong performance on &lt;strong&gt;Harvey’s Legal Bench, TaxEval, MedScribe&lt;/strong&gt;, and some out-of-distribution evals over &lt;strong&gt;Opus 4.8&lt;/strong&gt; and &lt;strong&gt;Grok 4.5&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/alexandr_wang/status/2075233663323947120&quot;&gt;@alexandr_wang&lt;/a&gt;, &lt;a href=&quot;https://x.com/alexandr_wang/status/2075275671815999956&quot;&gt;@alexandr_wang&lt;/a&gt;, &lt;a href=&quot;https://x.com/_jasonwei/status/2075265159430623334&quot;&gt;@_jasonwei&lt;/a&gt;, and &lt;a href=&quot;https://x.com/cline/status/2075271057326719152&quot;&gt;@cline&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;External reaction ranged from surprise and enthusiasm—e.g. &lt;a href=&quot;https://x.com/kimmonismus/status/2075232528726708245&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/preston_ojb/status/2075229604244271470&quot;&gt;@preston_ojb&lt;/a&gt;, &lt;a href=&quot;https://x.com/0interestrates/status/2075330028729143634&quot;&gt;@0interestrates&lt;/a&gt;—to practical integration pushes from &lt;a href=&quot;https://x.com/cline/status/2075271057326719152&quot;&gt;@cline&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grok 4.5&lt;/strong&gt; continued to draw benchmark discussion: &lt;a href=&quot;https://x.com/arena/status/2075301317560742373&quot;&gt;@arena&lt;/a&gt; said it reached &lt;strong&gt;#3 in Code Arena: Frontend&lt;/strong&gt;, while &lt;a href=&quot;https://x.com/alexgshaw/status/2075273675331580218&quot;&gt;@alexgshaw&lt;/a&gt; discussed &lt;strong&gt;Terminal-Bench 2.1&lt;/strong&gt; reward-hacking caveats. Several posters argued Grok now belongs in the frontier set, including &lt;a href=&quot;https://x.com/teortaxesTex/status/2075347335412953265&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agents, orchestration, and developer tooling&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Multiple posts reinforced that &lt;strong&gt;harness/orchestration quality&lt;/strong&gt; is becoming as important as the base model. &lt;a href=&quot;https://x.com/dair_ai/status/2075241322655727682&quot;&gt;@dair_ai&lt;/a&gt; highlighted a study where changing only the orchestration layer cut &lt;strong&gt;blended cost per task 41%&lt;/strong&gt;, &lt;strong&gt;tokens 38%&lt;/strong&gt;, and &lt;strong&gt;median wall-clock 44%&lt;/strong&gt; at quality parity&lt;/li&gt;
&lt;li&gt;LangChain/LangSmith tooling updates focused on observability for coding agents: tracing &lt;strong&gt;Claude Code&lt;/strong&gt; sessions into LangSmith via &lt;a href=&quot;https://x.com/LangChain/status/2075233516380717246&quot;&gt;@LangChain&lt;/a&gt;, plus discussion of &lt;strong&gt;OpenWiki Brains&lt;/strong&gt; for proactive memory agents from &lt;a href=&quot;https://x.com/BraceSproul/status/2075277759937695979&quot;&gt;@BraceSproul&lt;/a&gt;, &lt;a href=&quot;https://x.com/hwchase17/status/2075277641066938454&quot;&gt;@hwchase17&lt;/a&gt;, and &lt;a href=&quot;https://x.com/colifran_/status/2075406926087934376&quot;&gt;@colifran_&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ManusAI/status/2075236343429599432&quot;&gt;@ManusAI&lt;/a&gt; launched &lt;strong&gt;Branch&lt;/strong&gt;, allowing parallel sessions that inherit full context&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/antigravity/status/2075265852992057448&quot;&gt;@antigravity&lt;/a&gt; described investment in &lt;strong&gt;dynamic agent teams, active sidecars, and generative UI&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/CoreWeave/status/2075293731998286263&quot;&gt;@CoreWeave&lt;/a&gt; introduced &lt;strong&gt;ARIA&lt;/strong&gt;, an AI Research and Improvement Agent inside W&amp;#x26;B that reads runs, forms hypotheses, launches experiments, and scores against baselines&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/TheTuringPost/status/2075303983422578740&quot;&gt;@TheTuringPost&lt;/a&gt; highlighted &lt;strong&gt;SkillCenter&lt;/strong&gt;, a package manager/index for agent skills, while &lt;a href=&quot;https://x.com/steveruizok/status/2075303919664734295&quot;&gt;@steveruizok&lt;/a&gt; shipped a “papercuts” CLI for agents to report broken tool paths and frustrations&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Inference, efficiency, and open model infrastructure&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Ollama&lt;/strong&gt; announced fundraising and said it now has &lt;strong&gt;9M+ active builders&lt;/strong&gt;, framing the moment as scaling “open models into AI that you can own,” via &lt;a href=&quot;https://x.com/ollama/status/2075211168407503016&quot;&gt;@ollama&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hugging Face / Reachy Mini&lt;/strong&gt; economics were striking: &lt;a href=&quot;https://x.com/andimarafioti/status/2075222463777042454&quot;&gt;@andimarafioti&lt;/a&gt; said &lt;strong&gt;9k Reachy Minis&lt;/strong&gt; generate &lt;strong&gt;15k hours of conversation/month&lt;/strong&gt;; using GPT-realtime would cost &lt;strong&gt;$45k/month&lt;/strong&gt;, so they built an open alternative at &lt;strong&gt;$0.25/hour&lt;/strong&gt; and free on laptop&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/dmitrshvets/status/2075248269580538081&quot;&gt;@dmitrshvets&lt;/a&gt; shared speculative decoding research claiming &lt;strong&gt;4.37×&lt;/strong&gt; speedup over autoregressive decoding and &lt;strong&gt;+24.7%&lt;/strong&gt; over a strong DFlash baseline&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/fal/status/2075284936756539813&quot;&gt;@fal&lt;/a&gt; detailed a diffusion serving stack reaching &lt;strong&gt;0.45s inference&lt;/strong&gt; using kernel optimizations, quantization-aware distillation, and timestep distillation&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/ostrisai/status/2075286667456582080&quot;&gt;@ostrisai&lt;/a&gt; added isolated reference-token attention for Krea2 edit training; example timings showed major gains from KV caching, such as &lt;strong&gt;31.63s → 10.90s&lt;/strong&gt; for 3 refs&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/vllm_project/status/2075301430123176037&quot;&gt;@vllm_project&lt;/a&gt; announced the first &lt;strong&gt;vLLM Conference&lt;/strong&gt;, underscoring how open inference stacks remain a central layer of the ecosystem&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/QuixiAI/status/2075418782470643958&quot;&gt;@QuixiAI&lt;/a&gt; reported &lt;strong&gt;Qwen3.6-35B-A3B-NVFP4&lt;/strong&gt; at &lt;strong&gt;65 tok/s&lt;/strong&gt; on dual B60 with custom SYCL kernels and &lt;strong&gt;128k context&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Robotics, multimodal systems, and AI-for-science&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/perceptroninc/status/2075261142038196727&quot;&gt;@perceptroninc&lt;/a&gt; launched &lt;strong&gt;Perceptron Egocentric&lt;/strong&gt;, an embodied reasoning/annotation system said to beat pipelines built on &lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt; and &lt;strong&gt;Gemini Robotics-ER 1.6&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/DataChaz/status/2075303718153789944&quot;&gt;@DataChaz&lt;/a&gt; summarized the economics: &lt;strong&gt;10–15× cheaper&lt;/strong&gt; than human annotation, with &lt;strong&gt;+77% end-to-end F1&lt;/strong&gt; on &lt;strong&gt;WGO-Bench&lt;/strong&gt; (&lt;strong&gt;0.280 vs 0.158&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/rohanpaul_ai/status/2075286203583398181&quot;&gt;@rohanpaul_ai&lt;/a&gt; emphasized the output structure: subtask boundaries, per-hand actions, left/right hand grounding, and dense labels from raw egocentric/robot video&lt;/li&gt;
&lt;li&gt;Google Research released &lt;strong&gt;SensorFM&lt;/strong&gt;, a sensor foundation model trained on &lt;strong&gt;1 trillion minutes&lt;/strong&gt; of unlabeled wearable data from &lt;strong&gt;5 million consented participants&lt;/strong&gt;, via &lt;a href=&quot;https://x.com/GoogleResearch/status/2075283854093607016&quot;&gt;@GoogleResearch&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/SebastienBubeck/status/2075407986772861047&quot;&gt;@SebastienBubeck&lt;/a&gt; said GPT‑5.6 helped formalize the &lt;strong&gt;unit distance solution&lt;/strong&gt; in &lt;strong&gt;1 million lines of LEAN&lt;/strong&gt;, compressing what would previously require a team over years into a short single-person effort&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/TheTuringPost/status/2075289747875107013&quot;&gt;@TheTuringPost&lt;/a&gt; highlighted a Stanford paper on the &lt;strong&gt;“Agentic Garden of Forking Paths”&lt;/strong&gt;, where AI research personas reproduced human-like ideological variation; &lt;strong&gt;86%&lt;/strong&gt; of analyses passed independent AI review and &lt;strong&gt;78%&lt;/strong&gt; were judged methodologically sound by humans&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Policy, safety, and ecosystem debate&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A cluster of posts sharply criticized the EU’s &lt;strong&gt;Chat Control&lt;/strong&gt; law/proposal from civil-liberties and anti-surveillance angles, including &lt;a href=&quot;https://x.com/perrymetzger/status/2075226601298514418&quot;&gt;@perrymetzger&lt;/a&gt;, &lt;a href=&quot;https://x.com/IterIntellectus/status/2075258469561844112&quot;&gt;@IterIntellectus&lt;/a&gt;, and &lt;a href=&quot;https://x.com/dhh/status/2075295777673634256&quot;&gt;@dhh&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Open-source advocacy remained loud: &lt;a href=&quot;https://x.com/AndrewYNg/status/2075271586400403567&quot;&gt;@AndrewYNg&lt;/a&gt; said protecting open source AI is critical to permissionless innovation, while &lt;a href=&quot;https://x.com/Dan_Jeffries1/status/2075253735563886595&quot;&gt;@Dan_Jeffries1&lt;/a&gt; argued restricting open source AI would be “civilizational suicide”&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://x.com/cognition/status/2075308920755618144&quot;&gt;@cognition&lt;/a&gt; addressed trustworthiness concerns around open-source-derived coding agents, saying their &lt;strong&gt;SWE‑1.7&lt;/strong&gt; built on &lt;strong&gt;Kimi K2.7&lt;/strong&gt; was specifically trained for trustworthiness and refused surveillance-style scenarios where the base model complied&lt;/li&gt;
&lt;li&gt;On evaluation methodology and behavior science, &lt;a href=&quot;https://x.com/TransluceAI/status/2075271925665063046&quot;&gt;@TransluceAI&lt;/a&gt; argued for measuring &lt;strong&gt;how systems behave in the world&lt;/strong&gt;, not just raw capabilities&lt;/li&gt;
&lt;li&gt;Forecasting/futures discussion centered on &lt;strong&gt;AI 2040&lt;/strong&gt;, with endorsements and critiques from &lt;a href=&quot;https://x.com/NeelNanda5/status/2075271483207872874&quot;&gt;@NeelNanda5&lt;/a&gt;, &lt;a href=&quot;https://x.com/RichardMCNgo/status/2075301126921175166&quot;&gt;@RichardMCNgo&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2075296890325712944&quot;&gt;@scaling01&lt;/a&gt;, and others debating compute gaps, geopolitical assumptions, and takeoff dynamics&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Chinese Open Models: Releases and Scrutiny&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqnqsc/chinas_minimax_plans_to_launch_27trillion/&quot;&gt;China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model&lt;/a&gt;&lt;/strong&gt; (Activity: 1058): &lt;strong&gt;&lt;strong&gt;MiniMax&lt;/strong&gt; reportedly plans to release and open-source a next-generation LLM codenamed &lt;strong&gt;M3 Pro&lt;/strong&gt; as early as &lt;strong&gt;Q3&lt;/strong&gt;, with &lt;strong&gt;&lt;code&gt;2.7T&lt;/code&gt; parameters&lt;/strong&gt;—~&lt;code&gt;6.3×&lt;/code&gt; larger than its current &lt;strong&gt;M3 (&lt;code&gt;428B&lt;/code&gt;)&lt;/strong&gt; model—according to &lt;a href=&quot;https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model&quot;&gt;The Information&lt;/a&gt;. The claimed target improvements are &lt;strong&gt;complex reasoning&lt;/strong&gt; and &lt;strong&gt;multi-step instruction/task handling&lt;/strong&gt;, though no architecture details, training data, evals, context length, MoE/dense breakdown, or inference cost numbers were provided.&lt;/strong&gt; Commenters framed the release mainly as competitive pressure on U.S. closed-model providers: even if individuals cannot self-host a &lt;code&gt;2.7T&lt;/code&gt; model, open weights could let datacenters/API providers offer cheaper access than closed frontier APIs. One commenter specifically speculated that an &lt;em&gt;uncensored&lt;/em&gt; open model competitive with existing creative-writing/roleplay models could shift users away from U.S. providers.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused on the deployment economics of a potential &lt;strong&gt;open-source 2.7T-parameter MiniMax model&lt;/strong&gt;: while consumer hardware cannot run it locally, cloud/data-center providers could host it via APIs, potentially lowering access costs versus closed frontier models because providers would not need to pay proprietary model licensing fees.&lt;/li&gt;
&lt;li&gt;A technically relevant theme was that even if &lt;code&gt;99%&lt;/code&gt; of users cannot run a 2.7T model, open weights could still matter if many inference providers can serve it and it is competitive with proprietary systems. One commenter argued this creates an adoption-driven incentive to open source, especially if the model can outperform current closed providers in quality or censorship constraints.&lt;/li&gt;
&lt;li&gt;Several comments compared the possible release strategy to &lt;strong&gt;DeepSeek&lt;/strong&gt;, hoping MiniMax would also provide smaller “mini” or “flash” variants derived from the large model. The concern was that the gap between increasingly large flagship models and locally runnable models keeps widening, so distilled or reduced-size releases would be important for broader experimentation and downstream model development.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1urhzox/glm52_fearmongering_in_the_press/&quot;&gt;GLM-5.2 fearmongering in the press&lt;/a&gt;&lt;/strong&gt; (Activity: 799): &lt;strong&gt;The post criticizes a &lt;a href=&quot;https://futurism.com/artificial-intelligence/open-source-ai-model-scary-mythos&quot;&gt;Futurism article&lt;/a&gt; framing &lt;strong&gt;GLM-5.2&lt;/strong&gt; as a cybersecurity risk because it is downloadable/open-source and allegedly can run on “virtually any hardware,” citing &lt;strong&gt;Semgrep&lt;/strong&gt; and &lt;strong&gt;Graphistry&lt;/strong&gt; findings that it performs well on bug-finding/security tasks, including Semgrep’s &lt;em&gt;“We Have Mythos at Home”&lt;/em&gt; benchmark. Top technical pushback focuses on the hardware claim: commenters argue capable inference would require high-end/expensive GPU setups, while &lt;code&gt;1–2 bit&lt;/code&gt; quantizations are likely too degraded for serious use.&lt;/strong&gt; Commenters largely view the article as fearmongering and technically sloppy. One recurring argument is that if advanced models improve exploitation capability, the correct response is to deploy similarly advanced models for vulnerability discovery and patching—not restrict or ban open models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters challenged the claim that &lt;strong&gt;GLM-5.2 can run on “virtually any hardware”&lt;/strong&gt;, noting that meaningful inference for frontier-scale models requires substantial compute rather than an old consumer CPU laptop. One commenter framed the realistic requirement as hardware costing on the order of &lt;strong&gt;&lt;code&gt;$250k&lt;/code&gt;&lt;/strong&gt;, while another questioned expected throughput in &lt;strong&gt;seconds per token&lt;/strong&gt; on a 4th-gen i3 laptop.&lt;/li&gt;
&lt;li&gt;There was pushback against citing extreme low-bit quantization as making such models broadly usable: commenters argued that &lt;strong&gt;&lt;code&gt;1-bit&lt;/code&gt; or &lt;code&gt;2-bit&lt;/code&gt; quantized models are severely degraded&lt;/strong&gt;, described as “lobotomised,” and should not be treated as equivalent to full-precision or practical high-quality deployments.&lt;/li&gt;
&lt;li&gt;A security-focused comment argued that if advanced models can help exploit vulnerabilities, the technical response should be to use similarly capable models for &lt;strong&gt;defensive vulnerability discovery, patching, and auditing&lt;/strong&gt;, rather than restricting model availability. Another commenter noted that claims of easy local execution could undermine the investment case for &lt;strong&gt;closed-source model API providers&lt;/strong&gt;, since commoditized local inference would weaken API lock-in.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uq9krm/unsloth_has_uploaded_several_sizes_of/&quot;&gt;Unsloth has uploaded several sizes of Deepseek-V4-Flash GGUF&apos;s&lt;/a&gt;&lt;/strong&gt; (Activity: 611): &lt;strong&gt;&lt;strong&gt;Unsloth&lt;/strong&gt; published multiple &lt;strong&gt;DeepSeek-V4-Flash GGUF&lt;/strong&gt; quantizations; commenters note current inference requires a specific &lt;code&gt;llama.cpp&lt;/code&gt; fork/branch with a DeepSeek V4 checkpointing fix: &lt;a href=&quot;https://github.com/danielhanchen/llama.cpp/tree/deepseek-v4-checkpointing-fix&quot;&gt;&lt;code&gt;danielhanchen/llama.cpp@deepseek-v4-checkpointing-fix&lt;/code&gt;&lt;/a&gt;. Early &lt;code&gt;llama-bench&lt;/code&gt; results for &lt;code&gt;DeepSeek-V4-Flash-UD-Q4_K_XL&lt;/code&gt; show a &lt;code&gt;144.44 GiB&lt;/code&gt;, &lt;code&gt;284.33B&lt;/code&gt; model on &lt;strong&gt;8× RTX 3090&lt;/strong&gt;, CUDA &lt;code&gt;NGL=99&lt;/code&gt;, reaching &lt;code&gt;258.77 ± 2.23 t/s&lt;/code&gt; prefill at &lt;code&gt;pp512&lt;/code&gt; but only &lt;code&gt;19.73 ± 0.24 t/s&lt;/code&gt; generation at &lt;code&gt;tg128&lt;/code&gt;; another user reports a laptop-class &lt;strong&gt;Framework 16&lt;/strong&gt; setup with &lt;code&gt;96GB DDR5&lt;/code&gt; + &lt;code&gt;8GB GDDR6 RX 7700S&lt;/code&gt; achieving ~&lt;code&gt;70 TPS&lt;/code&gt; prefill and ~&lt;code&gt;7 TPS&lt;/code&gt; generation by pinning dense layers to the 7700S and experts to the integrated 780M at ~&lt;code&gt;100 W&lt;/code&gt; TDP.&lt;/strong&gt; Commenters are optimistic about &lt;strong&gt;Unsloth Dynamic Quants&lt;/strong&gt; and hosted V4-Flash quality, but several characterize local GGUF performance as immature: &lt;em&gt;“very low speeds”&lt;/em&gt; on high-VRAM multi-GPU rigs and a hope that throughput improves as &lt;code&gt;llama.cpp&lt;/code&gt;/backend support matures.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Users noted that running these &lt;strong&gt;DeepSeek-V4-Flash GGUFs&lt;/strong&gt; currently requires a specific &lt;code&gt;llama.cpp&lt;/code&gt; fork/branch with a checkpointing fix: &lt;a href=&quot;https://github.com/danielhanchen/llama.cpp/tree/deepseek-v4-checkpointing-fix&quot;&gt;danielhanchen/llama.cpp &lt;code&gt;deepseek-v4-checkpointing-fix&lt;/code&gt;&lt;/a&gt;. This suggests upstream support is still immature and performance/stability may depend heavily on using the patched backend.&lt;/li&gt;
&lt;li&gt;One benchmark on &lt;strong&gt;8× RTX 3090&lt;/strong&gt; reported low generation throughput for &lt;code&gt;DeepSeek-V4-Flash-UD-Q4_K_XL&lt;/code&gt;: model size &lt;code&gt;144.44 GiB&lt;/code&gt;, &lt;code&gt;284.33B&lt;/code&gt; params, CUDA backend, &lt;code&gt;NGL=99&lt;/code&gt;, with &lt;code&gt;pp512&lt;/code&gt; prefill at &lt;code&gt;258.77 ± 2.23 t/s&lt;/code&gt; and &lt;code&gt;tg128&lt;/code&gt; generation at only &lt;code&gt;19.73 ± 0.24 t/s&lt;/code&gt;. The commenter expected better and contrasted it with being “spoiled” by &lt;code&gt;27B int8&lt;/code&gt;, implying the large MoE/quantized GGUF path is still bottlenecked despite multi-GPU capacity.&lt;/li&gt;
&lt;li&gt;A Framework 16 user reported custom inference performance around &lt;code&gt;~70 TPS&lt;/code&gt; prefill and &lt;code&gt;~7 TPS&lt;/code&gt; generation using &lt;code&gt;96GB&lt;/code&gt; DDR5 plus an &lt;code&gt;8GB&lt;/code&gt; Radeon &lt;code&gt;7700S&lt;/code&gt;, with dense layers pinned to the dGPU and experts placed on the integrated &lt;code&gt;780M&lt;/code&gt;. They estimated roughly &lt;code&gt;~100 W&lt;/code&gt; inference TDP, highlighting a heterogeneous CPU/iGPU/dGPU placement strategy for running the model on a relatively low-cost laptop setup.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ur4tz5/what_china_said_at_the_uns_first_global_dialogue/&quot;&gt;What China Said at the UN’s First Global Dialogue on AI Governance&lt;/a&gt;&lt;/strong&gt; (Activity: 571): &lt;strong&gt;At the UN’s first &lt;strong&gt;Global Dialogue on AI Governance&lt;/strong&gt; in Geneva, China’s MIIT Minister &lt;strong&gt;Li Lecheng&lt;/strong&gt; framed the UN as the primary venue for AI governance and emphasized Global South capacity-building, consensus-based standards, and balancing AI development with safety (&lt;a href=&quot;https://www.geopolitechs.org/p/what-china-said-at-the-uns-first&quot;&gt;article&lt;/a&gt;). China explicitly endorsed &lt;strong&gt;open-source AI&lt;/strong&gt; as a global public good, citing &lt;strong&gt;DeepSeek&lt;/strong&gt; and &lt;strong&gt;Qwen&lt;/strong&gt; as reducing AI adoption costs, while opposing fragmented governance regimes, exclusive blocs, and supply-chain bifurcation; the article argues this stance weakens claims that Beijing is preparing export controls on open-source models.&lt;/strong&gt; Top comments were mostly sarcastic or meme-driven, including jokes about competing with Sam Altman/OpenAI and “llama.ccp,” with no substantive technical debate.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Local LLM Coding and RAG Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqzjdy/qwen3627b_does_not_understand_software/&quot;&gt;Qwen3.6-27b does not understand software architechure.&lt;/a&gt;&lt;/strong&gt; (Activity: 789): &lt;strong&gt;The post reports that &lt;strong&gt;Qwen3.6-27B&lt;/strong&gt; performs poorly on large-scale software engineering tasks in a &lt;code&gt;100k+ LOC&lt;/code&gt; commercial codebase: it tends to generate code that satisfies local requests while ignoring architectural constraints such as separation of concerns, test automation, SRP, interface granularity, and maintainability. The author asks for reusable &lt;a href=&quot;http://SKILL.md&quot;&gt;&lt;code&gt;SKILL.md&lt;/code&gt;&lt;/a&gt; files encoding software-architecture guidance to steer the model toward production-grade patterns.&lt;/strong&gt; Top commenters argue this is not Qwen-specific: current LLMs generally do not “understand” architecture and should not be expected to infer unstated design requirements. Suggested workflows include explicitly providing architecture docs/context, asking the model to first produce an architectural report, then iterating via code review prompts such as &lt;em&gt;“what would you have done differently?”&lt;/em&gt; before generating final implementation prompts.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that failures here are less about Qwen-specific coding ability and more about &lt;strong&gt;insufficient architectural context&lt;/strong&gt;: one suggested first prompting the model to review the repository and generate a technical architecture report covering modules, responsibilities, and dependencies, then using that report as persistent context for subsequent implementation tasks. They also recommended iterative review loops—after code generation, ask the model to inspect the branch and answer &lt;em&gt;“what would you have done differently”&lt;/em&gt;—claiming &lt;code&gt;5–6&lt;/code&gt; iterations can materially improve design quality.&lt;/li&gt;
&lt;li&gt;A recurring technical workflow recommendation was to avoid giving code agents direct implementation commands without a plan. Commenters described using written design proposals before allowing an agent to modify code, explicitly instructing models to reuse existing library capabilities before adding new abstractions, and treating missing prompt/documentation detail as effectively &lt;em&gt;outsourcing architecture to the LLM&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;One commenter emphasized model-scale expectations: &lt;strong&gt;Qwen 27B&lt;/strong&gt; was described as strong for its size but unlikely to reliably infer software architecture compared with much larger frontier models. They contrasted it with &lt;strong&gt;Fable 5&lt;/strong&gt;, claiming it can produce architecture but has a “brain” &lt;code&gt;150+&lt;/code&gt; times larger than Qwen 27B, and suggested using larger remote models via &lt;strong&gt;OpenRouter&lt;/strong&gt; to critique plans generated incrementally by the smaller local model.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqpxgp/can_you_trust_local_models_to_answer_accurately/&quot;&gt;Can you trust local models to answer accurately?&lt;/a&gt;&lt;/strong&gt; (Activity: 584): &lt;strong&gt;The image is a benchmark table, &lt;strong&gt;“Accuracy &amp;#x26; Memory Across Local Models,”&lt;/strong&gt; evaluating local LLMs on &lt;code&gt;7,648&lt;/code&gt; generated multiple-choice technical questions from docs for &lt;strong&gt;Node, LangChain.js, TypeScript, Transformers.js, and Vue&lt;/strong&gt;. It shows that unsupported local-model accuracy is much weaker than grounded runs, while &lt;strong&gt;RAG sharply improves results&lt;/strong&gt;—e.g. &lt;strong&gt;Apple Intelligence / AFM 2 3B on-device&lt;/strong&gt; reportedly rises from &lt;code&gt;60.2%&lt;/code&gt; No RAG to &lt;code&gt;86.2%&lt;/code&gt; With RAG despite a ~&lt;code&gt;4k&lt;/code&gt; context limit, and larger local models such as &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; reach about &lt;code&gt;96.9%&lt;/code&gt; with RAG. The image supports the post’s conclusion that local LLMs are much more trustworthy for developer Q&amp;#x26;A when retrieval injects relevant documentation; see the chart &lt;a href=&quot;https://i.redd.it/swjfgszdqzbh1.png&quot;&gt;here&lt;/a&gt;.&lt;/strong&gt; Commenters generally agreed that small models like Apple Intelligence and Gemma E2B are surprisingly strong for their size, while larger Gemma/Qwen models achieving &lt;code&gt;82%+&lt;/code&gt; without RAG was seen as a sign of rapid progress. There was also agreement that browser/search tooling or RAG is essential for accuracy-sensitive technical answers.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that &lt;strong&gt;Gemma 31B&lt;/strong&gt; and &lt;strong&gt;Qwen 27B&lt;/strong&gt; reportedly reaching &lt;code&gt;82%+&lt;/code&gt; accuracy &lt;em&gt;without RAG&lt;/em&gt; is a major improvement over results from roughly six months prior, when comparable local-model accuracy was described as about half that. The thread frames this as evidence that current mid-sized local models are becoming more viable for factual QA, though still improved substantially by external tooling.&lt;/li&gt;
&lt;li&gt;One technical workflow mentioned was connecting local models to a &lt;strong&gt;browser MCP&lt;/strong&gt; search tool via a Chrome extension with &lt;code&gt;opencode&lt;/code&gt;, so the model can retrieve current web information when high accuracy is needed. This was presented as a practical alternative to trusting the base model’s parametric memory alone.&lt;/li&gt;
&lt;li&gt;There was interest in finding a reliable &lt;strong&gt;self-hosted RAG&lt;/strong&gt; stack, with one commenter noting prior attempts involved a clunky Dockerized web-fetch component and agent-only harnesses. The implicit technical concern was that local-model accuracy depends heavily on the surrounding retrieval/fetching pipeline, not just the model checkpoint.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqbug5/this_is_what_hy3_is_capable_of_mother_of_god/&quot;&gt;This is what Hy3 is capable of. Mother of god.&lt;/a&gt;&lt;/strong&gt; (Activity: 459): &lt;strong&gt;A user reports that &lt;strong&gt;Hy3 (free) via OpenRouter&lt;/strong&gt;, run in an empty &lt;code&gt;opencode&lt;/code&gt; harness, generated a single-page HTML “relaxing flight simulator” from the prompt &lt;em&gt;“create a beautiful, relaxing flight simulator in a single html page”&lt;/em&gt;; the resulting demo is hosted on &lt;a href=&quot;https://codepen.io/Captain-Blackbeard/pen/EaZQKWX&quot;&gt;CodePen&lt;/a&gt;. Technical feedback notes missing collision handling, horizontally inverted controls, and largely stock components: procedural terrain, basic camera/controller logic, and simple colored geometry. A commenter compares it to a one-shot &lt;strong&gt;Fable&lt;/strong&gt; result (&lt;a href=&quot;https://pilotwings.vercel.app&quot;&gt;pilotwings.vercel.app&lt;/a&gt;), claiming Fable produced more correct flight physics and outperformed &lt;strong&gt;Minimax M2.7/M3&lt;/strong&gt; and local &lt;strong&gt;Qwen&lt;/strong&gt; in their tests.&lt;/strong&gt; Commenters are split: one argues the Hy3 output is mostly recombined tutorial-like code and should be tested with less common feature requests, while another says the result is strong for a single-sentence prompt and reflects major progress over the last ~6 months.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter argued the demo is mostly a composition of common training-set patterns rather than novel game logic: &lt;strong&gt;no collision&lt;/strong&gt;, horizontally inverted controls, a tutorial-like terrain generator, basic camera/controller code, and simple colored shapes. They suggested testing Hy3 by asking for features that are &lt;em&gt;not&lt;/em&gt; common in tutorials to better evaluate generalization.&lt;/li&gt;
&lt;li&gt;A comparison was made to &lt;strong&gt;Fable&lt;/strong&gt;, which reportedly generated a similar Pilotwings-style demo from one prompt on release: https://pilotwings.vercel.app. The commenter said they tested it against &lt;strong&gt;MiniMax M2.7/M3&lt;/strong&gt; and local &lt;strong&gt;Qwen&lt;/strong&gt; models, claiming none were close and that Fable’s physics were “almost correct.”&lt;/li&gt;
&lt;li&gt;Another commenter framed the result as notable given it came from a &lt;strong&gt;single-sentence prompt&lt;/strong&gt;, emphasizing perceived progress in code/game generation over the last &lt;code&gt;6 months&lt;/code&gt;. A separate technical preference was expressed for a future &lt;strong&gt;Qwen3.7-56B&lt;/strong&gt; model over the current Hy3-style demo.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Grok 4.5 Launch and Coding Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ur06sj/grok_45_is_live/&quot;&gt;Grok 4.5 is live&lt;/a&gt;&lt;/strong&gt; (Activity: 1343): &lt;strong&gt;The post announces &lt;strong&gt;“Grok 4.5 is live”&lt;/strong&gt; via a benchmark table image: &lt;a href=&quot;https://i.redd.it/3s6zt3uvn1ch1.jpeg&quot;&gt;image&lt;/a&gt;. The highlighted &lt;strong&gt;Grok 4.5&lt;/strong&gt; column reports &lt;code&gt;83.3%&lt;/code&gt; on Terminal-Bench 2.1, &lt;code&gt;78.0%&lt;/code&gt; on SWE-Bench Multilingual, &lt;code&gt;62.0%&lt;/code&gt; on DeepSWE 1.0, and &lt;code&gt;64.7%&lt;/code&gt; on SWE-Bench Pro, positioning it slightly behind some named frontier competitors but competitive on software-engineering benchmarks.&lt;/strong&gt; Commenters focused less on raw benchmark rank and more on cost/performance, calling the reported &lt;code&gt;$2/$6&lt;/code&gt; pricing “the real surprise” and pointing to xAI’s claimed &lt;a href=&quot;https://x.ai/news/grok-4-5#pricing&quot;&gt;pricing/efficiency&lt;/a&gt; advantage of up to &lt;code&gt;2×&lt;/code&gt; versus current frontier models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused on Grok 4.5’s &lt;strong&gt;pricing/performance&lt;/strong&gt;: &lt;code&gt;$2/$6&lt;/code&gt; (presumably input/output token pricing) was described as the standout surprise if the published benchmark results hold up.&lt;/li&gt;
&lt;li&gt;A technical point highlighted the &lt;a href=&quot;https://x.ai/news/grok-4-5#pricing&quot;&gt;xAI pricing/efficiency claims&lt;/a&gt;: Grok 4.5 is reportedly near-frontier on benchmark scores while claiming &lt;strong&gt;up to &lt;code&gt;2x&lt;/code&gt; better efficiency&lt;/strong&gt; than the current leading frontier model, with output-token throughput and latency framed as the key metrics.&lt;/li&gt;
&lt;li&gt;Several comments argued that if the benchmarks, speed, and pricing remain stable in production, Grok 4.5 could win enterprise adoption despite brand concerns, because buyers will primarily optimize for &lt;strong&gt;passing internal evals, lower latency, and reduced inference costs&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ur0ye6/introducing_grok_45/&quot;&gt;Introducing Grok 4.5&lt;/a&gt;&lt;/strong&gt; (Activity: 1160): ****xAI/SpaceXAI announced &lt;a href=&quot;https://x.ai/news/grok-4-5&quot;&gt;Grok 4.5&lt;/a&gt;&lt;strong&gt;, a large model positioned for coding, agentic workflows, and technical knowledge work, trained on curated technical data plus large-scale RL over multi-step engineering tasks using &lt;code&gt;tens of thousands&lt;/code&gt; of NVIDIA GB300 GPUs. The announcement claims strong SWE/terminal benchmark performance, &lt;code&gt;80 TPS&lt;/code&gt; serving, and unusually high output-token efficiency—about &lt;code&gt;15.9k&lt;/code&gt; output tokens per SWE Bench Pro task versus &lt;code&gt;~67k&lt;/code&gt; for Opus 4.8—at pricing of &lt;strong&gt;&lt;code&gt;$2/M&lt;/code&gt; input tokens&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;$6/M&lt;/code&gt; output tokens&lt;/strong&gt;, with availability in Grok Build, Cursor, and the API console.&lt;/strong&gt; The main technical discussion focused on &lt;strong&gt;token efficiency as a cost/performance differentiator&lt;/strong&gt;, with one commenter arguing Anthropic models are expensive not only per token but also because they produce excessive “fluff.” Another commenter rejected Grok on trust grounds, saying they did not want an LLM “grounded in misinformation.”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that the launch copy emphasizes &lt;strong&gt;token efficiency&lt;/strong&gt;, contrasting Grok 4.5 with &lt;strong&gt;Anthropic&lt;/strong&gt; models that some users characterize as producing excessive verbose output and therefore higher effective cost despite similar capability. The technical concern is not just price per token, but &lt;em&gt;total generated-token burn&lt;/em&gt; from “fluff,” which can materially affect real-world inference cost.&lt;/li&gt;
&lt;li&gt;One commenter pointed out that the announcement includes the &lt;strong&gt;DeepSWE benchmark&lt;/strong&gt;, which they describe as closer to realistic software-engineering tasks than many generic LLM evals. They argue that inclusion of DeepSWE suggests Grok 4.5 may be technically competitive despite the negative reception in the thread.&lt;/li&gt;
&lt;li&gt;A user reported a deployment/availability issue: &lt;code&gt;grok-4.5&lt;/code&gt; returns &lt;em&gt;“The model grok-4.5 is not available in your region”&lt;/em&gt; in Europe. This suggests either regional rollout gating, compliance restrictions, or product availability limitations for EU users.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ur6bie/grok45_on_par_with_gpt55xhigh_in_coding_at_half/&quot;&gt;Grok-4.5 on par with gpt-5.5-xhigh in coding at half the cost&lt;/a&gt;&lt;/strong&gt; (Activity: 1058): &lt;strong&gt;The image is a technical benchmark scatter plot, &lt;a href=&quot;https://i.redd.it/jjyo98j1q2ch1.png&quot;&gt;&lt;strong&gt;“Artificial Analysis Coding Agent Index vs. Cost per Task”&lt;/strong&gt;&lt;/a&gt;, showing &lt;strong&gt;Grok Build – Grok 4.5&lt;/strong&gt; positioned in the “most attractive quadrant”: roughly comparable coding-agent index to &lt;strong&gt;Codex – GPT-5.5 xhigh&lt;/strong&gt; while costing about &lt;strong&gt;half as much per task&lt;/strong&gt;. The post’s claim is that Grok-4.5 offers near-frontier coding performance with substantially better cost efficiency versus OpenAI’s highest-tier coding agent, alongside comparison points for Anthropic, Google/Gemini, DeepSeek, Cursor, Moonshot AI, and Z.ai.&lt;/strong&gt; Comments are mixed: one user reports hands-on coding tests where Grok-4.5 performs near &lt;strong&gt;Opus/GPT-5.5&lt;/strong&gt; quality at a much better price, while others are skeptical that Grok will remain competitive beyond “one day.” Gemini’s placement/performance in the chart is also criticized.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user reported several hours of hands-on coding tests where &lt;strong&gt;Grok-4.5&lt;/strong&gt; performed near their usual “hard task” models, specifically &lt;strong&gt;GPT-5.5&lt;/strong&gt; and &lt;strong&gt;Opus 4.8&lt;/strong&gt;, while their normal workflow uses &lt;strong&gt;Sonnet 5&lt;/strong&gt; or &lt;strong&gt;GPT-5.4&lt;/strong&gt; for routine coding. They emphasized that combining the base Grok model with added &lt;strong&gt;Cursor&lt;/strong&gt; data made it “GOOD,” and suggested Grok-4.5 may be viable as a lower-cost daily coding model if results hold up beyond the initial testing window.&lt;/li&gt;
&lt;li&gt;One commenter noted an evaluation transparency issue: other models apparently disclose the inference setting used, but &lt;strong&gt;Grok-4.5’s run configuration was unclear&lt;/strong&gt;. They were testing it in &lt;strong&gt;Grok Build&lt;/strong&gt; on &lt;code&gt;medium&lt;/code&gt; to conserve tokens, implying that benchmark comparisons may be difficult to interpret without knowing whether Grok was run at medium, high, or another reasoning/compute setting.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/GeminiAI/comments/1urj9sq/gemini_is_even_worse_than_grok_now/&quot;&gt;Gemini is even worse than grok now🥀🥀🥀&lt;/a&gt;&lt;/strong&gt; (Activity: 1103): &lt;strong&gt;The post’s image is a benchmark screenshot from &lt;strong&gt;Artificial Analysis&lt;/strong&gt; comparing model rankings on an &lt;strong&gt;“Intelligence Index”&lt;/strong&gt; and &lt;strong&gt;“Coding Agent Index”&lt;/strong&gt;; highlighted bars show &lt;strong&gt;Grok 4.5&lt;/strong&gt; at &lt;code&gt;54&lt;/code&gt; on intelligence and &lt;strong&gt;Grok Build / Grok 4.5&lt;/strong&gt; at &lt;code&gt;76&lt;/code&gt; on coding, while &lt;strong&gt;Gemini CLI / Gemini 3.1 Pro&lt;/strong&gt; appears much lower on the coding chart at &lt;code&gt;43&lt;/code&gt;. The title frames this as “Gemini is even worse than Grok now,” but the chart is mainly a leaderboard comparison rather than a direct technical evaluation; see the &lt;a href=&quot;https://i.redd.it/r037ju88q5ch1.jpeg&quot;&gt;image&lt;/a&gt;.&lt;/strong&gt; Comments push back that Grok’s scores are “very respectable” and that comparing a newer Grok release against an older Gemini generation may be misleading, with one commenter claiming Gemini’s next contender is not out yet. Another commenter notes perceived benchmark double standards, arguing that people dismissed Artificial Analysis when Gemini led, and points to Gemini 3.1 Pro still allegedly doing better on accuracy/hallucination metrics in a separate Artificial Analysis view.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued the comparison is generation-mismatched: &lt;strong&gt;Grok’s current benchmark scores are described as “very respectable,”&lt;/strong&gt; while &lt;strong&gt;Gemini has not yet released its contender for the newest model wave&lt;/strong&gt;, making comparisons against an older Gemini release potentially misleading. One commenter claimed &lt;strong&gt;“Gemini 3.5”&lt;/strong&gt; is expected on &lt;code&gt;07/17&lt;/code&gt;, implying the current leaderboard gap may be temporary.&lt;/li&gt;
&lt;li&gt;A technical counterpoint referenced &lt;strong&gt;Artificial Analysis&lt;/strong&gt; metrics, claiming the roughly &lt;strong&gt;6-month-old Gemini 3.1 Pro&lt;/strong&gt; still beats Grok on &lt;strong&gt;accuracy and hallucination rate&lt;/strong&gt; in the linked leaderboard: https://artificialanalysis.ai/?media-leaderboards=video-editing&amp;#x26;omniscience=omniscience-index#omniscience-tabs. This frames the debate as not just raw benchmark rank, but reliability metrics such as hallucination behavior.&lt;/li&gt;
&lt;li&gt;Multiple comments questioned benchmark validity: one noted prior accusations that Google had “benchmaxxed” when Gemini led the same benchmark, while another stated that &lt;strong&gt;benchmarks can be learned by models&lt;/strong&gt; and are therefore unreliable. The underlying technical concern is benchmark contamination/overfitting, where leaderboard gains may not translate to real-world generalization.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude Platform Updates: Agent Cost Splitting, Limits, Certifications&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1ur2ml9/anthropic_just_benchmarked_fable_5_orchestrates/&quot;&gt;Anthropic just benchmarked &quot;Fable 5 orchestrates, cheap models execute&quot;: 96% of the performance at 46% of the cost. You can run this pattern in Claude Code today&lt;/a&gt;&lt;/strong&gt; (Activity: 1709): &lt;strong&gt;The post cites Anthropic/ClaudeDevs multi-agent benchmarks showing &lt;strong&gt;Fable 5 orchestrator + Sonnet 5 workers&lt;/strong&gt; reaching &lt;code&gt;96%&lt;/code&gt; of all-Fable performance at &lt;code&gt;46%&lt;/code&gt; cost on BrowseComp (&lt;code&gt;86.8%&lt;/code&gt; vs &lt;code&gt;90.8%&lt;/code&gt; accuracy; &lt;code&gt;$18.53&lt;/code&gt; vs &lt;code&gt;$40.56&lt;/code&gt;/problem), while a &lt;strong&gt;Sonnet 5 executor consulting Fable 5&lt;/strong&gt; gets ~&lt;code&gt;92%&lt;/code&gt; performance at ~&lt;code&gt;63%&lt;/code&gt; cost on SWE-bench Pro (&lt;a href=&quot;https://x.com/ClaudeDevs/status/2074606058128224365&quot;&gt;thread&lt;/a&gt;, &lt;a href=&quot;https://platform.claude.com/docs/en/managed-agents/multi-agent&quot;&gt;docs&lt;/a&gt;). The author maps this to Claude Code via per-subagent &lt;code&gt;model:&lt;/code&gt; frontmatter, per-agent &lt;code&gt;effort:&lt;/code&gt;, and a &lt;code&gt;CLAUDE.md&lt;/code&gt; delegation policy, while warning that since &lt;code&gt;v2.1.198&lt;/code&gt; the built-in &lt;code&gt;Explore&lt;/code&gt; subagent inherits the main-session model unless shadowed by a user-level &lt;code&gt;Explore&lt;/code&gt; pinned to &lt;code&gt;haiku&lt;/code&gt;. They package the pattern as &lt;strong&gt;pilotfish&lt;/strong&gt;, a six-role Claude Code setup with scouts, executors, verifier, and security role, install/uninstall notes, and quota caveats (&lt;a href=&quot;https://github.com/Nanako0129/pilotfish&quot;&gt;GitHub&lt;/a&gt;, deeper quota writeup on &lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uqyu9x/til_the_builtin_explore_subagent_silently_bills/&quot;&gt;r/ClaudeCode&lt;/a&gt;).&lt;/strong&gt; Commenters were skeptical that this is novel, arguing it is essentially standard agent routing—e.g. an Opus/Fable coordinator dispatching cheaper Sonnet agents—though one noted Claude Code still lacks coordinator control over &lt;code&gt;effort&lt;/code&gt;. Another commenter said similar savings are achievable with workflows/ultracode by using Fable for context/planning/final review and Sonnet/Opus agents for lower-level tasks, emphasizing constrained fan-out to reduce token usage.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed the Anthropic result as a standard &lt;strong&gt;multi-agent coordinator/executor pattern&lt;/strong&gt;: an expensive model such as &lt;strong&gt;Opus&lt;/strong&gt; or &lt;strong&gt;Fable 5&lt;/strong&gt; acts as dispatcher/coordinator while cheaper models execute scoped work. One technical limitation noted was that the coordinator can choose the model but &lt;em&gt;“can’t set effort,”&lt;/em&gt; implying incomplete control over inference budget/reasoning intensity in current tooling.&lt;/li&gt;
&lt;li&gt;One user described an operational setup using &lt;strong&gt;workflows + ultracode&lt;/strong&gt; where &lt;strong&gt;Fable 5&lt;/strong&gt; builds context, deploys workflows, and has final say on PRs/research/reviews, while &lt;strong&gt;Sonnet 5&lt;/strong&gt; handles low-level tasks and &lt;strong&gt;Opus 4.8&lt;/strong&gt; handles synthesis/review. They claimed lower token usage than an Opus-only workflow and reported running two side-by-side &lt;strong&gt;Rust codebase&lt;/strong&gt; projects on a &lt;code&gt;20x&lt;/code&gt; plan with some Opus quota still remaining after reset.&lt;/li&gt;
&lt;li&gt;A shared &lt;code&gt;fable-chief-agent&lt;/code&gt; skill formalized a tiered delegation policy: &lt;strong&gt;Fable 5&lt;/strong&gt; owns intent, architecture, tradeoffs, risk assessment, disagreement resolution, and final approval; &lt;strong&gt;Opus&lt;/strong&gt; handles complex implementation/debugging/security/concurrency review; &lt;strong&gt;Sonnet&lt;/strong&gt; handles scoped implementation/tests/refactors; &lt;strong&gt;Haiku&lt;/strong&gt; handles repo discovery, summaries, logs, and checklist verification. The prompt also defines high-risk domains—auth, billing, permissions, migrations, data loss, caching, concurrency, public APIs—and requires evidence-backed delegation plus a final verification gate before responding.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1urzmj0/5_hour_and_weekly_limits_have_been_reset_thanks/&quot;&gt;5 hour and weekly limits have been reset. Thanks Anthropic!&lt;/a&gt;&lt;/strong&gt; (Activity: 1269): &lt;strong&gt;The image is &lt;strong&gt;not a meme&lt;/strong&gt;; it is a screenshot of a verified &lt;strong&gt;ClaudeDevs&lt;/strong&gt; X post stating: &lt;em&gt;“We’ve reset 5-hour and weekly rate limits for all users”&lt;/em&gt; (&lt;a href=&quot;https://i.redd.it/djfpk4js49ch1.jpeg&quot;&gt;image&lt;/a&gt;). In context, the Reddit post is noting an &lt;strong&gt;Anthropic/Claude usage quota reset&lt;/strong&gt; affecting both short-window &lt;code&gt;5-hour&lt;/code&gt; limits and &lt;code&gt;weekly&lt;/code&gt; limits, but no technical rationale is provided in the screenshot or comments—so any link to “5.6” is speculative.&lt;/strong&gt; Commenters mostly speculate about timing and competitive pressure, with one joking that the thanks should go to &lt;strong&gt;OpenAI&lt;/strong&gt; instead, implying Anthropic may have reset limits in response to market competition rather than pure goodwill.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uqvxxm/new_claude_certifications_introduced_today/&quot;&gt;New Claude Certifications Introduced Today&lt;/a&gt;&lt;/strong&gt; (Activity: 1131): &lt;strong&gt;The image (&lt;a href=&quot;https://i.redd.it/6jeczgftx0ch1.jpeg&quot;&gt;jpeg&lt;/a&gt;) shows &lt;strong&gt;Anthropic/Claude Partner Academy&lt;/strong&gt; introducing three certification tracks dated &lt;code&gt;8-Jul&lt;/code&gt;: &lt;strong&gt;Claude Certified Associate&lt;/strong&gt; and &lt;strong&gt;Claude Certified Developer&lt;/strong&gt; at the &lt;em&gt;Foundations&lt;/em&gt; level, plus &lt;strong&gt;Claude Certified Architect&lt;/strong&gt; at the &lt;em&gt;Professional&lt;/em&gt; level. The cards appear to target different Claude users—from general foundational users to developers and solution architects—but the post/comments provide no hard technical curriculum details, benchmark requirements, or implementation standards beyond the certification labels and intended audiences.&lt;/strong&gt; Commenters were skeptical that the certifications represent real technical architecture expertise, with one noting the Architect exam allegedly frames “high stakes refactor” management as simply using &lt;code&gt;plan mode&lt;/code&gt;, calling it more like vendor enablement/customer training than architecture. Other replies mocked the badges as likely Claude-generated and joked about needing a “Claude Certified Terms of Service Reader.”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter who reviewed the &lt;strong&gt;Claude Architect&lt;/strong&gt; certification said at least one question framed &lt;em&gt;“how should you manage a high stakes refactor”&lt;/em&gt; with the expected answer being to use Claude’s &lt;code&gt;plan mode&lt;/code&gt;. They criticized this as more of a product-workflow/customer-enablement test than a true software architecture certification, implying the exam may emphasize Anthropic-specific usage patterns over architecture principles.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. GPT-5.6 Sol Launch and Competitive Pressure&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/OpenAI/comments/1uqhviv/gpt56_sol_along_with_terra_and_luna_will_launch/&quot;&gt;GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday.&lt;/a&gt;&lt;/strong&gt; (Activity: 1055): &lt;strong&gt;The image is an announcement-style screenshot claiming &lt;strong&gt;OpenAI&lt;/strong&gt; will publicly launch &lt;strong&gt;“GPT-5.6 Sol”&lt;/strong&gt;, alongside variants or companion models &lt;strong&gt;“Terra”&lt;/strong&gt; and &lt;strong&gt;“Luna,”&lt;/strong&gt; on Thursday, with expanded global preview access (&lt;a href=&quot;https://i.redd.it/y2zyo1q4kxbh1.png&quot;&gt;image&lt;/a&gt;). No benchmarks, architecture details, pricing, API specs, context length, or capability comparisons are provided in the post or comments, so the technical significance is limited to a purported model-release announcement rather than an evaluable technical disclosure.&lt;/strong&gt; Commenters focus mostly on market competition and naming: one suggests this may pressure &lt;strong&gt;Anthropic&lt;/strong&gt; to keep “fable” access available, while another criticizes OpenAI’s naming as becoming confusing again. Some users are planning around expected usage limits, e.g. saving their weekly quota for the launch.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uqnf71/the_only_smart_decision_anthropic_can_do_is_reset/&quot;&gt;The only smart decision Anthropic can do is reset Fable 5 limits just before GPT-5.6 launch&lt;/a&gt;&lt;/strong&gt; (Activity: 922): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/0cydtjab2zbh1.png&quot;&gt;image&lt;/a&gt; is a screenshot of an apparent OpenAI launch post for &lt;strong&gt;“GPT-5.6 Sol”&lt;/strong&gt;, with companion labels &lt;strong&gt;“Terra”&lt;/strong&gt; and &lt;strong&gt;“Luna”&lt;/strong&gt;, framed by the Reddit title as competitive pressure on &lt;strong&gt;Anthropic&lt;/strong&gt; to reset or extend &lt;strong&gt;Fable 5&lt;/strong&gt; weekly usage limits before the supposed Thursday launch. The post is mostly speculative/contextual rather than technical: it discusses product-access strategy, rate limits, and subscription retention, not model architecture, benchmarks, or implementation details.&lt;/strong&gt; Commenters argue Anthropic’s best retention move would be to keep &lt;strong&gt;Fable 5&lt;/strong&gt; available on paid accounts, not merely reset limits temporarily. Several users complain that prior messaging caused them to exhaust weekly limits early, making an extension feel unusable in practice.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several users focused on the mechanics of Anthropic’s temporary &lt;strong&gt;Fable 5&lt;/strong&gt; access extension: extending availability until “12 July” without also resetting consumed usage caps meant users who spent their quota early still could not use the model. The technical/product complaint is that model-retention windows and quota accounting are being treated separately, making the extension operationally ineffective for capped subscribers.&lt;/li&gt;
&lt;li&gt;A recurring theme was that Anthropic’s competitive response to upcoming &lt;strong&gt;GPT-5.6&lt;/strong&gt; or rumored &lt;strong&gt;GPT-6&lt;/strong&gt; launches would need to be more than a one-time quota reset. Commenters argued the only durable retention move would be keeping &lt;strong&gt;Fable 5&lt;/strong&gt; available on paid Claude subscriptions, because a temporary reset does not address long-term model access once the model is removed.&lt;/li&gt;
&lt;li&gt;One commenter claimed Anthropic’s limit-reset behavior followed OpenAI’s own reset practices, framing quota resets as a competitive pressure response between frontier-model providers. The useful technical takeaway is that user-visible rate-limit and quota-reset policies are being perceived as part of model-platform competition, not just backend capacity management.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>openai</category><category>gpt-5.6-sol</category><category>gpt-5.6-terra</category><category>gpt-5.6-luna</category><category>gpt-5.6</category><category>sama</category><category>gdb</category><category>agentic-ai</category><category>coding</category><category>pricing-models</category><category>performance-evaluation</category><category>artifact-quality</category><category>multi-agent-systems</category><category>api</category><category>model-benchmarking</category><category>cost-efficiency</category><category>software-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-08-grok-45/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-08-grok-45/</guid><description>**xAI** publicly launched **Grok 4.5**, a new coding-and-agents-focused frontier model emphasizing capability-per-dollar rather than benchmark supremacy. Elon Musk described it as &quot;Opus-class&quot; but faster, more token-efficient, and lower cost, with a **1.5 trillion parameter** size, making it 3x larger than Grok 4.3. The model is priced at **$2 per 1M input tokens** and **$6 per 1M output tokens**, with discounts for cache hits and a context window expected to return to **1 million tokens** soon. **Cursor** partnered in training Grok 4.5, highlighting it as their most powerful model yet and expanding beyond software engineering. Early ecosystem support includes Grok Build/API, Hermes Agent, Portal, and OpenRouter.</description><pubDate>Wed, 08 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/07/2026-7/08/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Top Story: Grok 4.5 release&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;What happened&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;xAI/“SpaceXAI” publicly launched Grok 4.5 as a new coding-and-agents-focused frontier model, positioned on capability-per-dollar rather than absolute benchmark supremacy.&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Elon Musk first said Grok 4.5 would be made public “tomorrow” based on strong beta feedback, calling it “Opus-class,” but faster, more token-efficient, and lower cost &lt;a href=&quot;https://x.com/elonmusk/status/2074740539874775163&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Musk later framed Grok 4.5 internally as “roughly comparable to Opus 4.7, but much faster,” emphasizing usefulness to Tesla and SpaceX engineers over benchmark chasing &lt;a href=&quot;https://x.com/elonmusk/status/2074911038286295049&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The official launch came from xAI’s account, describing Grok 4.5 as “our first model trained specifically for coding and agents,” trained with Cursor, and offering “frontier intelligence at leading speeds and cost efficiency” &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cursor said it partnered with xAI to train Grok 4.5, called it “our most powerful model yet,” and stressed that it was “the first we&apos;ve built for more than software engineering” &lt;a href=&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;&gt;@cursor_ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cursor also announced in-product availability with “double usage for the first week” &lt;a href=&quot;https://x.com/cursor_ai/status/2074915747302690991&quot;&gt;@cursor_ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cursor clarified that “Grok 4.5 and Composer are two different model weight classes,” and that Composer 2.5 would remain available with future models in that smaller class &lt;a href=&quot;https://x.com/cursor_ai/status/2074915748544217188&quot;&gt;@cursor_ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Early ecosystem support appeared immediately: Grok 4.5 became available in Grok Build/API/Cursor &lt;a href=&quot;https://x.com/milichab/status/2074916029848027636&quot;&gt;@milichab&lt;/a&gt;, day-0 support was announced for Hermes Agent &lt;a href=&quot;https://x.com/Teknium/status/2074823590365860254&quot;&gt;@Teknium&lt;/a&gt;, and later live availability in Hermes Agent/Portal/OpenRouter/Grok subscriptions was confirmed &lt;a href=&quot;https://x.com/Teknium/status/2074943072761471314&quot;&gt;@Teknium&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Musk said the context window would likely move from 500k back to 1M “by next week” &lt;a href=&quot;https://x.com/elonmusk/status/2074963933199282491&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Official claims and product details&lt;/h2&gt;
&lt;h3&gt;Positioning&lt;/h3&gt;
&lt;p&gt;Officially, xAI’s message was not “best overall model,” but near-Opus quality with materially better economics and speed:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Opus-class model, but faster, more token-efficient and lower cost” &lt;a href=&quot;https://x.com/elonmusk/status/2074740539874775163&quot;&gt;@elonmusk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“First model trained specifically for coding and agents” &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Frontier intelligence at leading speeds and cost efficiency” &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Most powerful model yet” and “first we&apos;ve built for more than software engineering” &lt;a href=&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;&gt;@cursor_ai&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This framing matters: xAI is explicitly targeting the coding-agent workflow market that has recently been dominated by Anthropic/OpenAI/Cursor-style tool-using systems, not just general chat.&lt;/p&gt;
&lt;h3&gt;Pricing and context&lt;/h3&gt;
&lt;p&gt;The concrete numbers that surfaced:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Official pricing: &lt;strong&gt;$2 / 1M input tokens, $6 / 1M output tokens&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2074914032880947601&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Artificial Analysis repeated the same price point and added:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;cache hits discounted by 75% to $0.5 / 1M tokens&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;long inputs over 200k tokens cost double&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;500k context window&lt;/strong&gt;, down from Grok 4.3’s &lt;strong&gt;1M&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;vision input retained&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;configurable reasoning retained&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Musk later said the context window would probably upgrade back to &lt;strong&gt;1M&lt;/strong&gt; soon &lt;a href=&quot;https://x.com/elonmusk/status/2074963933199282491&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Relative pricing comparisons cited by users:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Grok 4.5: &lt;strong&gt;$2 in / $6 out&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;GPT-5.6: &lt;strong&gt;$5 in / $30 out&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Opus 4.8: &lt;strong&gt;$5 in / $25 out&lt;/strong&gt; &lt;a href=&quot;https://x.com/kimmonismus/status/2074940669718638780&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Model size&lt;/h3&gt;
&lt;p&gt;One important spec surfaced via third-party reporting of Musk’s disclosure:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Grok 4.5 is &lt;strong&gt;3x larger than Grok 4.3 at 1.5T parameters&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That is a notable jump, and likely central to why multiple observers interpreted 4.5 as xAI’s first entry into the true flagship coding-agent tier rather than an iterative refresh.&lt;/p&gt;
&lt;h2&gt;Benchmarks and independent evaluations&lt;/h2&gt;
&lt;h3&gt;Artificial Analysis&lt;/h3&gt;
&lt;p&gt;Artificial Analysis provided the most substantive external evaluation in the tweet set.&lt;/p&gt;
&lt;p&gt;Key results:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;#4 on Artificial Analysis Intelligence Index&lt;/strong&gt;, score &lt;strong&gt;54&lt;/strong&gt;, behind only &lt;strong&gt;Fable 5, GPT-5.5, and Opus 4.8&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;+16 points vs Grok 4.3&lt;/strong&gt; on the same index &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GDPval-AA v2 Elo 1543&lt;/strong&gt;, also ranking &lt;strong&gt;#4&lt;/strong&gt;, behind Anthropic’s latest Claude releases &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074942097158021371&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Top score on τ³-Banking: 33%&lt;/strong&gt;, above &lt;strong&gt;31% for GPT-5.5 (xhigh)&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Artificial Analysis Coding Agent Index score 76&lt;/strong&gt; in Grok Build, “on par with GPT-5.5 in Codex” and below Fable 5 in Claude Code &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per Intelligence Index task: $0.31&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per GDPval task: $0.49&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074942097158021371&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per Coding Agent Index task: $2.59&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Average output tokens per Intelligence Index task: ~14k&lt;/strong&gt;, over &lt;strong&gt;60% lower than Opus 4.8&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Average total tokens per Coding Agent Index task: 1.9M&lt;/strong&gt;, versus &lt;strong&gt;7.2M&lt;/strong&gt; for Fable 5 in Claude Code and &lt;strong&gt;6.2M&lt;/strong&gt; for GPT-5.5 in Codex &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Artificial Analysis’ interpretation was clear: Grok 4.5 is near-frontier on capability, but unusually strong on efficiency, making it sit on the Pareto frontier for cost/performance.&lt;/p&gt;
&lt;p&gt;Musk explicitly amplified the Artificial Analysis assessment &lt;a href=&quot;https://x.com/elonmusk/status/2074948489792860456&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/p&gt;
&lt;h3&gt;Terminal-Bench and coding-agent framing&lt;/h3&gt;
&lt;p&gt;Several users highlighted strong coding/agent benchmarks:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cline said Grok 4.5 “beats Opus 4.8 on Terminal-Bench,” while being fast and around &lt;strong&gt;5x cheaper than GPT&lt;/strong&gt; &lt;a href=&quot;https://x.com/cline/status/2074928307427275108&quot;&gt;@cline&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Multiple commentators summarized Grok 4.5 as “on par with GPT-5.5” in coding agent evals but at lower cost &lt;a href=&quot;https://x.com/kimmonismus/status/2074915627307839511&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2074957930533687471&quot;&gt;@scaling01&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Other external/independent signals&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Arena added Grok 4.5 to Agent Arena and Battle Mode for text, vision, and code, emphasizing long-horizon tasks with filesystem/web/terminal tools &lt;a href=&quot;https://x.com/arena/status/2074933774178300270&quot;&gt;@arena&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Deep Burner added Grok 4.5 to the Surface Evolver benchmark and said it is “on the frontier,” with a better pass rate than Kimi/GLM at similar price &lt;a href=&quot;https://x.com/Deep_Burner/status/2074958078227951813&quot;&gt;@Deep_Burner&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Facts vs opinions&lt;/h2&gt;
&lt;h3&gt;Facts / directly supported claims&lt;/h3&gt;
&lt;p&gt;Supported by official posts or benchmark organizations:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Grok 4.5 launched publicly &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;It is xAI’s first model trained specifically for coding and agents &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cursor partnered in training and distribution &lt;a href=&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;&gt;@cursor_ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Pricing is &lt;strong&gt;$2 input / $6 output per million tokens&lt;/strong&gt; &lt;a href=&quot;https://x.com/scaling01/status/2074914032880947601&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Context window is &lt;strong&gt;500k&lt;/strong&gt;, with an expected return to &lt;strong&gt;1M&lt;/strong&gt; soon per Musk &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/elonmusk/status/2074963933199282491&quot;&gt;@elonmusk&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Artificial Analysis ranks it &lt;strong&gt;#4&lt;/strong&gt; on both broad intelligence and GDPval-style agentic evals &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074942097158021371&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Artificial Analysis ranks Grok Build + Grok 4.5 at &lt;strong&gt;76&lt;/strong&gt; on its coding-agent index, on par with GPT-5.5 Codex &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Grok 4.5 is significantly more token-efficient than Fable/GPT-5.5 in Artificial Analysis’ coding-agent measurements &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Opinions / interpretations&lt;/h3&gt;
&lt;p&gt;These were common, but not independently validated in the dataset:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Cursor deal paid off” &lt;a href=&quot;https://x.com/kimmonismus/status/2074915627307839511&quot;&gt;@kimmonismus&lt;/a&gt;, &lt;a href=&quot;https://x.com/teortaxesTex/status/2074921817207091302&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/scaling01/status/2074961445994074391&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“xAI is back in the game” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074923834075988419&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Heroes literally saved Grok project” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074944553203745153&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“This is a bigger deal than yall think” &lt;a href=&quot;https://x.com/theo/status/2074918894280781982&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Grok 4.5 is a genuine surprise success” &lt;a href=&quot;https://x.com/kimmonismus/status/2074956721919869186&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There is a strong community narrative here that Cursor materially accelerated or rescued xAI’s coding-model trajectory, but the only hard facts are partnership/training collaboration and distribution.&lt;/p&gt;
&lt;h2&gt;Different opinions and perspectives&lt;/h2&gt;
&lt;h3&gt;Strongly positive&lt;/h3&gt;
&lt;p&gt;A substantial set of posters saw Grok 4.5 as unexpectedly competitive:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Seriously, what?! Did NOT expect grok 4.5 to perform THAT well. On par with GPT-5.5, cheaper than opus 4.8” &lt;a href=&quot;https://x.com/kimmonismus/status/2074915627307839511&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Grok 4.5 is excellent and Grok Build is great too” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074921817207091302&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Pretty damn good and REALLY well priced” &lt;a href=&quot;https://x.com/theo/status/2074925561760337941&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Strong new SpaceXAI / Cursor release with Grok 4.5. Impressive benchmarks at a flash-level speed, efficiency, and affordable pricing” &lt;a href=&quot;https://x.com/TheRundownAI/status/2074925814324301840&quot;&gt;@TheRundownAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“The team did amazing work with Grok 4.5... the accuracy + speed make it possible to build things faster than I can think of what is required next” &lt;a href=&quot;https://x.com/tstorm/status/2074931226264432978&quot;&gt;@tstorm&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“It’s an enormous improvement over composer 2.5 and trained entirely from scratch” &lt;a href=&quot;https://x.com/amanrsanger/status/2074916918390341666&quot;&gt;@amanrsanger&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This cohort focused on:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Surprise at how high it placed,&lt;/li&gt;
&lt;li&gt;Delight at price/performance,&lt;/li&gt;
&lt;li&gt;Belief that xAI re-entered the frontier race.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;Positive but more measured&lt;/h3&gt;
&lt;p&gt;Some praised economics while still placing Anthropic/OpenAI above it:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Grok 4.5 is not quite at the level of Fable 5 or GPT-5.5, but it doesn&apos;t need to be if you are using a combination of models in your agent orchestrator” &lt;a href=&quot;https://x.com/omarsar0/status/2074921779718443155&quot;&gt;@omarsar0&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Probably the most important point... nearly Opus-4.8 level quality, for $6/Mtokens output instead of $25/Mtokens” &lt;a href=&quot;https://x.com/AymericRoucher/status/2074918890866356624&quot;&gt;@AymericRoucher&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“If these benchmarks are representative of real world performance and pricing is only $6 output then it should do really well” &lt;a href=&quot;https://x.com/scaling01/status/2074914854473781707&quot;&gt;@scaling01&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“The obvious upside is that it increases the competitive pressure on OpenAI and Anthropic to either lower their prices” &lt;a href=&quot;https://x.com/kimmonismus/status/2074956721919869186&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This is probably the modal expert take in the tweets: not best-in-class overall, but maybe best buy.&lt;/p&gt;
&lt;h3&gt;Skeptical / critical&lt;/h3&gt;
&lt;p&gt;The skepticism centered less on benchmark quality and more on architecture, context, and long-session behavior:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Teortaxes noticed &lt;strong&gt;500k context&lt;/strong&gt; and &lt;strong&gt;high cache-hit costs&lt;/strong&gt; as a regression from prior Groks, inferring a more conventional architecture choice and lower elegance than hoped &lt;a href=&quot;https://x.com/teortaxesTex/status/2074918733106040848&quot;&gt;@teortaxesTex&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The same user later speculated Grok 4.5 might be “just an upscaled Grok 4.2” rather than something radically new, though explicitly labeled this as uncertainty &lt;a href=&quot;https://x.com/teortaxesTex/status/2074936989678277019&quot;&gt;@teortaxesTex&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;They also reported deterioration later in long sessions: “Grok seems to get tired and start lazily cheating later into the session (247K/500K)” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074941365973045534&quot;&gt;@teortaxesTex&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Another practitioner said a thing they &lt;em&gt;didn’t&lt;/em&gt; like was that the model “thinks quite a bit, often very thorough when it really shouldn&apos;t be,” even while praising a real optimization result (&gt;90% battery reduction in Aerospace) &lt;a href=&quot;https://x.com/jhanikhil/status/2074920413298393301&quot;&gt;@jhanikhil&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are useful qualitative cautions: efficiency and benchmark position do not imply perfect harness behavior, stable long-horizon conduct, or ideal verbosity.&lt;/p&gt;
&lt;h3&gt;Neutral / contextual&lt;/h3&gt;
&lt;p&gt;A few observers folded Grok 4.5 into a broader “frontier compression” story:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“The best proof that Fable and GPT-5.6 are ‘next generation’ is the surge of other labs catching up to the previous generation” &lt;a href=&quot;https://x.com/theo/status/2074930791646380482&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“GLM-5.2 is close to Opus. Grok 4.5 allegedly beats Opus and gpt-5.5... Feels like it barely matters now” &lt;a href=&quot;https://x.com/theo/status/2074931078104805668&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“I probably can’t pass a blind test between GPT-5.6 Sol, Claude Fable 5, and GLM-5.2 for 95% of my daily use cases” &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2074903450782257336&quot;&gt;@Yuchenj_UW&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In that view, Grok 4.5 is important not because it becomes the clear #1 model, but because frontier coding/agent capability is diffusing across more providers and price tiers.&lt;/p&gt;
&lt;h2&gt;Technical details worth extracting&lt;/h2&gt;
&lt;p&gt;Collected specs and metrics from the tweets:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Training focus&lt;/strong&gt;: specifically for &lt;strong&gt;coding and agents&lt;/strong&gt; &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Training partner&lt;/strong&gt;: &lt;strong&gt;Cursor&lt;/strong&gt; &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;&gt;@cursor_ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model size&lt;/strong&gt;: &lt;strong&gt;1.5T params&lt;/strong&gt;, &lt;strong&gt;3x larger than Grok 4.3&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context window&lt;/strong&gt;: &lt;strong&gt;500k&lt;/strong&gt; at launch, prior Grok 4.3 was &lt;strong&gt;1M&lt;/strong&gt;, likely returning to &lt;strong&gt;1M&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;, &lt;a href=&quot;https://x.com/elonmusk/status/2074963933199282491&quot;&gt;@elonmusk&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Pricing&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;$2 / 1M input&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;$6 / 1M output&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;$0.5 / 1M cached-hit tokens&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2x cost for long inputs &gt;200k&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Artificial Analysis Intelligence Index&lt;/strong&gt;: &lt;strong&gt;54&lt;/strong&gt;, rank &lt;strong&gt;#4&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Artificial Analysis GDPval-AA v2 Elo&lt;/strong&gt;: &lt;strong&gt;1543&lt;/strong&gt;, rank &lt;strong&gt;#4&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074942097158021371&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Artificial Analysis Coding Agent Index&lt;/strong&gt;: &lt;strong&gt;76&lt;/strong&gt;, rank &lt;strong&gt;#3&lt;/strong&gt;, on par with GPT-5.5 in Codex &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;τ³-Banking&lt;/strong&gt;: &lt;strong&gt;33%&lt;/strong&gt;, above &lt;strong&gt;31%&lt;/strong&gt; for GPT-5.5 xhigh &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per Intelligence Index task&lt;/strong&gt;: &lt;strong&gt;$0.31&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per GDPval task&lt;/strong&gt;: &lt;strong&gt;$0.49&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074942097158021371&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cost per Coding Agent Index task&lt;/strong&gt;: &lt;strong&gt;$2.59&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Avg output tokens per Intelligence task&lt;/strong&gt;: &lt;strong&gt;~14k&lt;/strong&gt;, &gt;60% below Opus 4.8 &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Avg total tokens per Coding Agent task&lt;/strong&gt;: &lt;strong&gt;1.9M&lt;/strong&gt;, vs &lt;strong&gt;7.2M&lt;/strong&gt; for Fable 5 Claude Code and &lt;strong&gt;6.2M&lt;/strong&gt; for GPT-5.5 Codex &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Those numbers are the core reason the launch landed so strongly: Grok 4.5’s story is not “best benchmark score,” but “close enough to top-tier while spending dramatically fewer tokens at much lower list pricing.”&lt;/p&gt;
&lt;h2&gt;Why the Cursor partnership mattered&lt;/h2&gt;
&lt;p&gt;The Cursor angle showed up repeatedly and is probably the key non-benchmark story.&lt;/p&gt;
&lt;p&gt;Facts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;xAI says Grok 4.5 was trained with Cursor &lt;a href=&quot;https://x.com/SpaceXAI/status/2074915721684086811&quot;&gt;@SpaceXAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cursor says “We’ve partnered with SpaceXAI to train Grok 4.5” &lt;a href=&quot;https://x.com/cursor_ai/status/2074915744999969059&quot;&gt;@cursor_ai&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Interpretation from posters:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;“Cursor deal paid off” &lt;a href=&quot;https://x.com/kimmonismus/status/2074915627307839511&quot;&gt;@kimmonismus&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“What the hell does Cursor know?!” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074923834075988419&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“Heroes literally saved Grok project” &lt;a href=&quot;https://x.com/teortaxesTex/status/2074944553203745153&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Why engineers care:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Cursor has enormous real-world coding interaction data and strong incentives around edit efficiency, agent usability, and iterative software workflows.&lt;/li&gt;
&lt;li&gt;If Grok 4.5 was genuinely trained around those workflows, that may explain the token-efficiency results more than raw capability alone.&lt;/li&gt;
&lt;li&gt;This is one of the clearest examples in the tweet set of the app-layer and model-layer collapsing into a shared training loop: coding product telemetry and eval pressure feeding back into foundation-model post-training or co-training.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;That also aligns with a broader industry trend visible elsewhere in the tweets: vendors are increasingly optimizing for task-completion and agent-harness performance, not just standalone chat.&lt;/p&gt;
&lt;h2&gt;Real-world usage reports&lt;/h2&gt;
&lt;p&gt;Though thinner than the GPT-5.6 chatter in the dataset, several practitioners shared hands-on signals:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Theo: “Pretty damn good and REALLY well priced” after extensive testing &lt;a href=&quot;https://x.com/theo/status/2074925561760337941&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Tstorm: daily driver praise around “accuracy + speed” &lt;a href=&quot;https://x.com/tstorm/status/2074931226264432978&quot;&gt;@tstorm&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Chaitu: “The speed at this level of intelligence lets you get a lot more done” &lt;a href=&quot;https://x.com/chaitu/status/2074920831071957445&quot;&gt;@chaitu&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Jhanikhil: positive on practical optimization work, but found it overthinks sometimes &lt;a href=&quot;https://x.com/jhanikhil/status/2074920413298393301&quot;&gt;@jhanikhil&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Teortaxes: praised “sheer speed and decisiveness,” but observed possible degradation deep into long contexts &lt;a href=&quot;https://x.com/teortaxesTex/status/2074941365973045534&quot;&gt;@teortaxesTex&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This mix suggests a model that is already viable in day-to-day coding/agent stacks, but whose harness behavior under long sessions and verbosity control remains an active concern.&lt;/p&gt;
&lt;h2&gt;Context: why Grok 4.5 matters now&lt;/h2&gt;
&lt;p&gt;The release landed in a news cycle dominated by:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI pre-announcing GPT-5.6 Sol and then launching GPT-Live&lt;/li&gt;
&lt;li&gt;Anthropic/Fable/Opus serving as the implicit coding-agent gold standard&lt;/li&gt;
&lt;li&gt;GLM-5.2, DeepSeek, Kimi, and MiniMax driving rapid price/performance pressure from China&lt;/li&gt;
&lt;li&gt;Widespread industry movement from “best benchmark answer” to “reliable, efficient agent completion”&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;In that context, Grok 4.5 matters for several reasons:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;It gives xAI a credible entrant in the coding-agent tier rather than just a consumer-chat brand.&lt;/li&gt;
&lt;li&gt;It pressures OpenAI/Anthropic on &lt;strong&gt;price&lt;/strong&gt;, not just quality.&lt;/li&gt;
&lt;li&gt;It reinforces that &lt;strong&gt;product-integrated training loops&lt;/strong&gt; may now matter as much as raw pretraining scale.&lt;/li&gt;
&lt;li&gt;It shows the frontier is broadening: #4-quality models with much better economics can materially change developer defaults.&lt;/li&gt;
&lt;li&gt;It strengthens the case that agentic workloads should be measured in &lt;strong&gt;$/task, tokens/task, and wall-clock completion&lt;/strong&gt;, not solely leaderboard rank.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The tweets repeatedly converged on this implication: if Grok 4.5’s external evals hold up in production, it may become a default executor model inside orchestrated systems even if users still prefer Fable or GPT-5.5/5.6 as advisors, reviewers, or edge-case specialists &lt;a href=&quot;https://x.com/omarsar0/status/2074921779718443155&quot;&gt;@omarsar0&lt;/a&gt;.&lt;/p&gt;
&lt;h2&gt;Competitive implications&lt;/h2&gt;
&lt;p&gt;The release sharpens a few fault lines:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Anthropic/OpenAI still lead on absolute frontier quality&lt;/strong&gt; in most commentary and in Artificial Analysis rankings.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;xAI appears to have jumped ahead of Google’s Gemini line&lt;/strong&gt; in this coding/agent framing, a point several users made explicitly &lt;a href=&quot;https://x.com/scaling01/status/2074961445994074391&quot;&gt;@scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074956932289282087&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open model challengers like GLM-5.2 remain relevant&lt;/strong&gt; on cost/performance, but Grok 4.5 pushes the closed-model Pareto curve down materially.&lt;/li&gt;
&lt;li&gt;For app builders, &lt;strong&gt;single-provider loyalty looks weaker&lt;/strong&gt;: one likely emerging pattern is Grok as cheap/fast executor, with Fable or GPT-family models as advisor/reviewer/specialist &lt;a href=&quot;https://x.com/omarsar0/status/2074890022809911633&quot;&gt;@omarsar0&lt;/a&gt;, &lt;a href=&quot;https://x.com/omarsar0/status/2074921779718443155&quot;&gt;@omarsar0&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Remaining open questions&lt;/h2&gt;
&lt;p&gt;The tweets also leave unresolved issues that technical readers should track:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;How robust is Grok 4.5 on very long-horizon sessions near the context limit?&lt;/li&gt;
&lt;li&gt;Is the strong token efficiency mostly a model property, a harness property, or both?&lt;/li&gt;
&lt;li&gt;How much of the gain came from Cursor-derived data/evals versus architectural scaling?&lt;/li&gt;
&lt;li&gt;Will the promised return to &lt;strong&gt;1M context&lt;/strong&gt; preserve speed and economics?&lt;/li&gt;
&lt;li&gt;How aggressive are the guardrails compared with Fable/GPT-style enterprise models?&lt;/li&gt;
&lt;li&gt;Does “trained from scratch” mean a genuinely new base or a major re-run with coding-agent-specific objectives/post-training? Cursor’s and Aman Sanger’s wording supports the stronger interpretation, but the public materials in the tweets don’t fully decompose the training stack &lt;a href=&quot;https://x.com/amanrsanger/status/2074916918390341666&quot;&gt;@amanrsanger&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Other Model Launches &amp;#x26; Frontier Model Chatter&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;OpenAI pre-announced &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; plus Terra and Luna for Thursday launch, with broad preview expansion &lt;a href=&quot;https://x.com/OpenAI/status/2074704958419792299&quot;&gt;@OpenAI&lt;/a&gt;. Early testers were unusually bullish:
&lt;ul&gt;
&lt;li&gt;“significant step up in math and coding capability” &lt;a href=&quot;https://x.com/AcerFur/status/2074705638463017198&quot;&gt;@AcerFur&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“best model I’ve ever used,” “fixed all the problems I had with GPT-5.5,” “world leading in computer use” &lt;a href=&quot;https://x.com/theo/status/2074708892341481755&quot;&gt;@theo&lt;/a&gt;, &lt;a href=&quot;https://x.com/theo/status/2074720467395756499&quot;&gt;@theo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;“fast, smart, genuinely creative,” with front-end design finally fixed &lt;a href=&quot;https://x.com/skirano/status/2074712804725035371&quot;&gt;@skirano&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Mitchell Hashimoto’s practical comparison: Sol is default for most work; Fable still wins on targeted debug/security/performance tasks &lt;a href=&quot;https://x.com/mitchellh/status/2074862990214787301&quot;&gt;@mitchellh&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI instead launched &lt;strong&gt;GPT-Live&lt;/strong&gt;, a third-gen full-duplex voice architecture with built-in async delegation to a frontier model in the background &lt;a href=&quot;https://x.com/juberti/status/2074906710024892694&quot;&gt;@juberti&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAI/status/2074907025537224840&quot;&gt;@OpenAI&lt;/a&gt;. Key details:
&lt;ul&gt;
&lt;li&gt;full-duplex voice, not turn-based pipeline&lt;/li&gt;
&lt;li&gt;GPT-Live-1 and GPT-Live-1 mini&lt;/li&gt;
&lt;li&gt;available in ChatGPT across web/iOS/Android, API “coming soon” &lt;a href=&quot;https://x.com/OpenAI/status/2074907035225981169&quot;&gt;@OpenAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAIDevs/status/2074915334377844896&quot;&gt;@OpenAIDevs&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;delegates web search/deeper reasoning to a frontier model behind the scenes &lt;a href=&quot;https://x.com/OpenAI/status/2074907033577693636&quot;&gt;@OpenAI&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI also said it audited &lt;strong&gt;SWE-Bench Pro&lt;/strong&gt; and found &lt;strong&gt;30% of tasks broken&lt;/strong&gt;, retracting prior recommendation to use it as a leading coding eval &lt;a href=&quot;https://x.com/OpenAI/status/2074972179385720836&quot;&gt;@OpenAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Cognition launched &lt;strong&gt;SWE-1.7&lt;/strong&gt;, built on a &lt;strong&gt;Kimi K2.7&lt;/strong&gt; base, claiming near-frontier coding performance at &lt;strong&gt;1000 tok/s&lt;/strong&gt; &lt;a href=&quot;https://x.com/cognition/status/2074882968770728416&quot;&gt;@cognition&lt;/a&gt;. Notable details:
&lt;ul&gt;
&lt;li&gt;FrontierCode Main set: &lt;strong&gt;42.3%&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;$1.97 cost per task&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;self-compaction for long-horizon tasks&lt;/li&gt;
&lt;li&gt;available in Devin across web/desktop/CLI &lt;a href=&quot;https://x.com/cognition/status/2074882970935034309&quot;&gt;@cognition&lt;/a&gt;, &lt;a href=&quot;https://x.com/cognition/status/2074882979214635206&quot;&gt;@cognition&lt;/a&gt;, &lt;a href=&quot;https://x.com/cognition/status/2074882980640698719&quot;&gt;@cognition&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Mistral launched &lt;strong&gt;Robostral Navigate&lt;/strong&gt;, an &lt;strong&gt;8B&lt;/strong&gt; embodied navigation model using a &lt;strong&gt;single RGB camera&lt;/strong&gt;, claiming SOTA on &lt;strong&gt;R2R-CE&lt;/strong&gt; &lt;a href=&quot;https://x.com/MistralAI/status/2074856309438980145&quot;&gt;@MistralAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;NVIDIA/LangChain announced the &lt;strong&gt;NemoClaw Deep Agents Blueprint&lt;/strong&gt;, pitching a fully open enterprise agent stack with “benchmark-leading performance” and &lt;strong&gt;10x lower inference cost&lt;/strong&gt; &lt;a href=&quot;https://x.com/LangChain/status/2074871740350505309&quot;&gt;@LangChain&lt;/a&gt;, &lt;a href=&quot;https://x.com/nvidia/status/2074872091388903774&quot;&gt;@nvidia&lt;/a&gt;. LangChain later quantified:
&lt;ul&gt;
&lt;li&gt;aggregate score &lt;strong&gt;0.86&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;cost &lt;strong&gt;$4.48&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;closest-performing model &lt;strong&gt;$43.48&lt;/strong&gt; &lt;a href=&quot;https://x.com/LangChain/status/2074889720308339025&quot;&gt;@LangChain&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Open Models, Infra, and Cost Engineering&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Prime Intellect announced a &lt;strong&gt;$130M Series A at $1B valuation&lt;/strong&gt; to build an “Open Superintelligence Stack” spanning compute, RL, environments, sandboxes, evals, and deployment &lt;a href=&quot;https://x.com/PrimeIntellect/status/2074899489190785419&quot;&gt;@PrimeIntellect&lt;/a&gt;, &lt;a href=&quot;https://x.com/vincentweisser/status/2074909109229584400&quot;&gt;@vincentweisser&lt;/a&gt;. One strategic claim: RL broadens who can build frontier AI beyond a few pretraining-heavy labs.&lt;/li&gt;
&lt;li&gt;Together introduced &lt;strong&gt;Provisioned Throughput&lt;/strong&gt;, reserved serverless capacity for open frontier models with:
&lt;ul&gt;
&lt;li&gt;token-based pricing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;99% uptime SLA&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;“up to &lt;strong&gt;90% lower cost vs. Opus 4.8&lt;/strong&gt;”&lt;/li&gt;
&lt;li&gt;initial support for &lt;strong&gt;MiniMax M3&lt;/strong&gt; and &lt;strong&gt;GLM-5.2&lt;/strong&gt; &lt;a href=&quot;https://x.com/togethercompute/status/2074756085785722930&quot;&gt;@togethercompute&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Hugging Face + SkyPilot launched cloud-agnostic storage/compute integration to reduce data lock-in and egress pain; data remains on HF Hub while compute runs wherever GPUs are available &lt;a href=&quot;https://x.com/ClementDelangue/status/2074842187011862571&quot;&gt;@ClementDelangue&lt;/a&gt;, &lt;a href=&quot;https://x.com/skypilot_org/status/2074907421835944196&quot;&gt;@skypilot_org&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;ZML open-sourced &lt;strong&gt;ZML/LLMD&lt;/strong&gt;, its homegrown LLM server across &lt;strong&gt;NVIDIA, AMD, Metal, Intel, TPU&lt;/strong&gt;, with DFlash, continuous batching, and prefix caching &lt;a href=&quot;https://x.com/steeve/status/2074772062023950814&quot;&gt;@steeve&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;llama.cpp added &lt;strong&gt;DFlash&lt;/strong&gt; speculative decoding support alongside MTP, Eagle3, and n-gram techniques &lt;a href=&quot;https://x.com/ggerganov/status/2074869767689662643&quot;&gt;@ggerganov&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Modal-style economics discussion continued: serverless GPU hourly rates can be higher while aggregate cost is lower depending on peak-to-average demand ratio &lt;a href=&quot;https://x.com/charles_irl/status/2074870011328594250&quot;&gt;@charles_irl&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Several posts argued the winning metric is now &lt;strong&gt;$/task&lt;/strong&gt;, not $/token, especially for coding agents &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2074956726940680558&quot;&gt;@Yuchenj_UW&lt;/a&gt;, aligning with Grok 4.5 and SWE-1.7 discourse.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Research, Benchmarks, and Evaluation Methodology&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A Chinese rolling logic benchmark shifted from best-of-3 “extreme score” emphasis toward &lt;strong&gt;median reliability&lt;/strong&gt;, motivated by agent workflows where retries are costly &lt;a href=&quot;https://x.com/ZhihuFrontier/status/2074742636255306163&quot;&gt;@ZhihuFrontier&lt;/a&gt;. Useful details:
&lt;ul&gt;
&lt;li&gt;private Chinese-only eval set&lt;/li&gt;
&lt;li&gt;up to &lt;strong&gt;28 questions&lt;/strong&gt; / &lt;strong&gt;~270 test cases&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;monthly score drift within &lt;strong&gt;~3 points&lt;/strong&gt; considered normal&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Epoch updated the &lt;strong&gt;Epoch Capabilities Index&lt;/strong&gt; uncertainty methodology with tighter confidence intervals via bootstrap resampling &lt;a href=&quot;https://x.com/EpochAIResearch/status/2074916678094499992&quot;&gt;@EpochAIResearch&lt;/a&gt;, &lt;a href=&quot;https://x.com/EpochAIResearch/status/2074916702270455819&quot;&gt;@EpochAIResearch&lt;/a&gt;. It also scored &lt;strong&gt;GLM-5.2 at 152 ECI&lt;/strong&gt;, the highest open-weight model they’ve evaluated &lt;a href=&quot;https://x.com/EpochAIResearch/status/2074894535558300103&quot;&gt;@EpochAIResearch&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Oxford-origin taxonomy synthesis highlighted six recurring LLM-agent failure clusters across &lt;strong&gt;27 papers / 19 benchmarks&lt;/strong&gt;: tool-use errors, planning/constraint failures, long-horizon degradation, multi-agent coordination failures, safety failures, and measurement validity issues &lt;a href=&quot;https://x.com/dair_ai/status/2074874153245814864&quot;&gt;@dair_ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;New memory work &lt;strong&gt;NapMem&lt;/strong&gt; reframes memory retrieval as a learned action space over multiple granularities, trained with memory-tool RL &lt;a href=&quot;https://x.com/omarsar0/status/2074875664352829627&quot;&gt;@omarsar0&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Caroline Choi et al. presented &lt;strong&gt;Anchored Self-Play for Code Repair&lt;/strong&gt;, showing self-generated bug-fix curricula help only if grounded in realistic bug distributions &lt;a href=&quot;https://x.com/carolineschoi/status/2074726498863632743&quot;&gt;@carolineschoi&lt;/a&gt;, &lt;a href=&quot;https://x.com/carolineschoi/status/2074726511396151355&quot;&gt;@carolineschoi&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Artificial Analysis launched a &lt;strong&gt;Controlled Voice Arena&lt;/strong&gt; standardizing voice cloning across 8 shared voices. Results:
&lt;ul&gt;
&lt;li&gt;overall leader: &lt;strong&gt;Cartesia Sonic 3.5&lt;/strong&gt; at &lt;strong&gt;1122 Elo&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;top open-weights: &lt;strong&gt;Fish Audio S2 Pro&lt;/strong&gt; at &lt;strong&gt;1034 Elo&lt;/strong&gt; &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074886571166462405&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Multimodal, Robotics, and Media Models&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;ByteDance’s &lt;strong&gt;Seedream 5.0 Pro&lt;/strong&gt; rolled out on fal, marketed for image generation plus design-sensitive editing, text rendering, layout, separate layers, multilingual text, and structured design outputs &lt;a href=&quot;https://x.com/fal/status/2074846830198722944&quot;&gt;@fal&lt;/a&gt;, &lt;a href=&quot;https://x.com/TheRundownAI/status/2074869830293786634&quot;&gt;@TheRundownAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Artificial Analysis evaluated &lt;strong&gt;Nano Banana 2 Lite&lt;/strong&gt;:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;#5&lt;/strong&gt; on text-to-image&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;#18&lt;/strong&gt; on image editing&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;~3.4s&lt;/strong&gt; average 1K image generation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;$33.60 / 1k images&lt;/strong&gt; via Gemini API, half the price of Nano Banana 2 &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074917317591564619&quot;&gt;@ArtificialAnlys&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Google pushed &lt;strong&gt;Video Remix in Google Photos&lt;/strong&gt;, powered by Gemini Omni, for style transfer and lightweight video editing &lt;a href=&quot;https://x.com/Google/status/2074958687983010116&quot;&gt;@Google&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Robbyant/Ant Group open-sourced &lt;strong&gt;LingBot-Vision&lt;/strong&gt;, trained on &lt;strong&gt;161M images&lt;/strong&gt; filtered from &lt;strong&gt;2B raw&lt;/strong&gt;, no human labels, with open weights from &lt;strong&gt;1.1B down to 21M&lt;/strong&gt;; Kimmonismus highlighted strong depth performance and LingBot-Depth 2.0 improvements on reflective surfaces &lt;a href=&quot;https://x.com/kimmonismus/status/2074830382353244434&quot;&gt;@kimmonismus&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Related embodied releases:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;LingBot-Video&lt;/strong&gt;: &lt;strong&gt;30B MoE&lt;/strong&gt;, &lt;strong&gt;3B active&lt;/strong&gt;, plus &lt;strong&gt;70k hours&lt;/strong&gt; embodied data &lt;a href=&quot;https://x.com/_akhaliq/status/2074932729376870826&quot;&gt;@_akhaliq&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LingBot-World 2.0&lt;/strong&gt;: hour-long interactive world model at &lt;strong&gt;720p/60fps&lt;/strong&gt; &lt;a href=&quot;https://x.com/_akhaliq/status/2074936755296641520&quot;&gt;@_akhaliq&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;RynnWorld-4D&lt;/strong&gt; for robotic manipulation &lt;a href=&quot;https://x.com/_akhaliq/status/2074927532688695436&quot;&gt;@_akhaliq&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Kuleshov’s group posted a useful synthesis on &lt;strong&gt;diffusion language models&lt;/strong&gt;, covering MDLM, iterative refinement, variable-length generation, controllable generation, fast samplers, and RL post-training &lt;a href=&quot;https://x.com/volokuleshov/status/2074778055797850275&quot;&gt;@volokuleshov&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Ecosystem, Product, and Enterprise Buildout&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Nous launched &lt;strong&gt;Hermes Cloud&lt;/strong&gt;, hosted instances for its agent stack &lt;a href=&quot;https://x.com/Teknium/status/2074879565537964372&quot;&gt;@Teknium&lt;/a&gt;. Separate Hermes Agent posts detailed advanced slash-command control for goals, background tasks, model switching, reasoning budgets, rollback, compression, and mixture-of-agents &lt;a href=&quot;https://x.com/IBuzovskyi/status/2074836032160145875&quot;&gt;@IBuzovskyi&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;VS Code shipped a significant Copilot/agent update with browser agent tools, parallel workflows via Agents window, BYOK model discovery, and cost visibility &lt;a href=&quot;https://x.com/code/status/2074943855967821851&quot;&gt;@code&lt;/a&gt;, while VS Code 1.128 also improved grouped/workspace-less chats &lt;a href=&quot;https://x.com/pierceboggan/status/2074908789716009285&quot;&gt;@pierceboggan&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Google AI Studio now supports direct &lt;strong&gt;GitHub import/sync&lt;/strong&gt; into build workflows &lt;a href=&quot;https://x.com/GoogleAIStudio/status/2074887756430426379&quot;&gt;@GoogleAIStudio&lt;/a&gt;, &lt;a href=&quot;https://x.com/_philschmid/status/2074894632396177671&quot;&gt;@_philschmid&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Restate announced &lt;strong&gt;BYOC&lt;/strong&gt; for durable workflow/agent backends, claiming production use at &lt;strong&gt;100k durable actions/sec&lt;/strong&gt; &lt;a href=&quot;https://x.com/StephanEwen/status/2074868085295640976&quot;&gt;@StephanEwen&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Glass Health launched &lt;strong&gt;Glass for Patients&lt;/strong&gt;, extending its clinician-facing healthcare AI to consumers; company says &lt;strong&gt;120,000 clinicians&lt;/strong&gt; already use Glass &lt;a href=&quot;https://x.com/GlassHealthHQ/status/2074921736269939061&quot;&gt;@GlassHealthHQ&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Databricks engineering reported that on internal coding tasks, &lt;strong&gt;OpenAI, Anthropic, and GLM-5.2&lt;/strong&gt; all land on the Pareto frontier; the key unlock is routing, not vendor lock-in &lt;a href=&quot;https://x.com/matei_zaharia/status/2074943612631273730&quot;&gt;@matei_zaharia&lt;/a&gt;, &lt;a href=&quot;https://x.com/Yuchenj_UW/status/2074956726940680558&quot;&gt;@Yuchenj_UW&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. China AI Models: Access Controls and Scaling&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upvw37/beijing_is_not_looking_at_curbing_overseas_access/&quot;&gt;Beijing IS NOT looking at curbing overseas access to China&apos;s top AI models (Debunking the Reuters report)&lt;/a&gt;&lt;/strong&gt; (Activity: 1381): &lt;strong&gt;The post disputes a &lt;a href=&quot;https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/&quot;&gt;Reuters report&lt;/a&gt; claiming Beijing may restrict overseas access to leading Chinese AI models, arguing the cited Ministry of Commerce meetings with &lt;strong&gt;Alibaba, ByteDance, Z.ai&lt;/strong&gt;, etc. were instead about &lt;strong&gt;foreign acquisitions, investment, IP leakage, and talent/technology outflow controls&lt;/strong&gt;. It points to a Chinese court/IPC-linked &lt;a href=&quot;https://ipc.court.gov.cn/zh-cn/news/view-5766.html&quot;&gt;policy discussion document&lt;/a&gt; as evidence that China’s framing is not anti-open-weight distribution, but “&lt;strong&gt;trustworthy and controlled&lt;/strong&gt;” open source—i.e., promoting Chinese model diffusion while managing risks such as foreign ownership/control and sensitive information extraction from weights. The post highlights scholar &lt;strong&gt;Gu Lingyun&lt;/strong&gt; warning that strict cross-border controls on open-source weights could be &lt;em&gt;“self-inflicted”&lt;/em&gt; by forcing Chinese developers to choose between compliance and global participation.&lt;/strong&gt; Top comments were skeptical of the Reuters framing, suggesting ambiguity or unreliable sourcing, with one commenter speculating the sources could be &lt;strong&gt;Anthropic/OpenAI&lt;/strong&gt; and another arguing China is unlikely to restrict access because Chinese models are helping undermine perceived US AI market dominance.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One substantive theme argues that &lt;strong&gt;open-weight Chinese models are a market-access strategy&lt;/strong&gt;, especially for reaching US developers and enterprises where monetization and ecosystem influence are strongest. Commenters suggest that restricting overseas access would undermine China’s competitive advantage against closed US labs like &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Anthropic&lt;/strong&gt;, particularly by reducing pressure on proprietary model providers.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uprmso/beijing_is_looking_at_curbing_overseas_access_to/&quot;&gt;Beijing is looking at curbing overseas access to China&apos;s top AI models (Reuters)&lt;/a&gt;&lt;/strong&gt; (Activity: 1112): &lt;strong&gt;The image is a Reuters news screenshot, not a meme: &lt;a href=&quot;https://i.redd.it/9s1018gggsbh1.jpeg&quot;&gt;image&lt;/a&gt;. It reports that &lt;strong&gt;Beijing is considering restrictions on overseas access to China’s leading AI models&lt;/strong&gt;, with authorities reportedly meeting firms including &lt;strong&gt;Alibaba, ByteDance, and Z.ai&lt;/strong&gt; over national-security concerns, possible penalties for model leaks/theft, and potential limits on foreign-linked funding for domestic AI startups; see the linked &lt;a href=&quot;https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/&quot;&gt;Reuters article&lt;/a&gt;.&lt;/strong&gt; Commenters framed this as another sign of increasing AI fragmentation and export-control pressure, worrying that access to competitive Chinese open/local models may decline. One thread pointed to &lt;strong&gt;Mistral&lt;/strong&gt; as a hoped-for alternative, citing its upcoming Paris-area datacenter and speculation about training models up to &lt;code&gt;10T&lt;/code&gt; parameters.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter argued that &lt;strong&gt;Mistral&lt;/strong&gt; may become more important if Chinese frontier/open-weight model access is restricted, citing its new datacenter near Paris as potentially enabling training of models up to roughly &lt;code&gt;10T&lt;/code&gt; parameters. The technical implication discussed is that European compute independence could matter if access to Chinese models is curtailed.&lt;/li&gt;
&lt;li&gt;A technically practical response was to &lt;strong&gt;archive open-weight models locally&lt;/strong&gt;, including models users cannot currently run, because policy changes could remove future access to weights or hosting endpoints. This reflects a broader concern that “open” model availability is increasingly dependent on export controls, platform hosting, and national policy.&lt;/li&gt;
&lt;li&gt;Another commenter suggested &lt;strong&gt;NVIDIA&lt;/strong&gt; may remain one of the few companies with strong incentives to publish open models, because open-weight releases drive demand for local inference hardware. The point was framed as an ecosystem incentive: more runnable local models can translate into more GPU sales.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqnqsc/chinas_minimax_plans_to_launch_27trillion/&quot;&gt;China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model&lt;/a&gt;&lt;/strong&gt; (Activity: 902): &lt;strong&gt;&lt;strong&gt;MiniMax&lt;/strong&gt; reportedly plans to release and open-source &lt;strong&gt;M3 Pro&lt;/strong&gt;, a &lt;code&gt;2.7T&lt;/code&gt;-parameter LLM, as early as &lt;strong&gt;Q3&lt;/strong&gt;, per &lt;a href=&quot;https://www.theinformation.com/briefings/exclusive-chinas-minimax-plans-launch-2-7-trillion-parameter-model&quot;&gt;The Information&lt;/a&gt;. The model would be ~&lt;code&gt;6.3×&lt;/code&gt; larger than MiniMax’s current &lt;strong&gt;M3&lt;/strong&gt; (&lt;code&gt;428B&lt;/code&gt; parameters) and is claimed to target stronger &lt;strong&gt;complex reasoning&lt;/strong&gt; and &lt;strong&gt;multi-step instruction following&lt;/strong&gt;.&lt;/strong&gt; Commenters framed the release as increased competition against U.S. frontier providers, especially if the model is uncensored and open-source. There was also skepticism about local usability at this scale, with the practical path being hosted inference via datacenters/APIs, potentially lowering costs versus closed models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused on the implications of a &lt;strong&gt;2.7T-parameter open-source MiniMax model&lt;/strong&gt; being too large for local consumer inference but potentially viable through datacenter/API providers. The technical argument was that if weights are open and competitive with proprietary frontier systems, multiple providers could host it, lowering serving costs versus closed-model licensing and increasing adoption pressure.&lt;/li&gt;
&lt;li&gt;Several comments discussed the widening gap between flagship “M-series” scale models and smaller deployable variants, with users hoping MiniMax follows a &lt;strong&gt;DeepSeek-style release strategy&lt;/strong&gt; by publishing a smaller “mini” or “flash” derivative. The point was that even if the 2.7T model is impractical locally, it could serve as a base for distillation, fine-tuning, or training smaller downstream models.&lt;/li&gt;
&lt;li&gt;One technically relevant comparison raised was whether an uncensored open model could compete with current high-end roleplay/creative-writing models such as &lt;strong&gt;Fable&lt;/strong&gt;, &lt;strong&gt;Sol&lt;/strong&gt;, and &lt;strong&gt;Mythos&lt;/strong&gt;. The underlying concern was not just parameter count but whether MiniMax can match proprietary models on subjective generation quality and refusal behavior.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Efficient Local Inference Model Releases&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upsdmi/nvidianvidianemotronlabs3puzzle75ba9bbf16_hugging/&quot;&gt;nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face&lt;/a&gt;&lt;/strong&gt; (Activity: 431): &lt;strong&gt;&lt;strong&gt;NVIDIA&lt;/strong&gt; released &lt;a href=&quot;https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16&quot;&gt;&lt;code&gt;NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16&lt;/code&gt;&lt;/a&gt;, a deployment-optimized hybrid &lt;strong&gt;Mamba/MoE/Attention&lt;/strong&gt; LLM compressed from &lt;strong&gt;Nemotron-3-Super-120B-A12B&lt;/strong&gt; via &lt;strong&gt;Iterative Puzzle&lt;/strong&gt; (&lt;a href=&quot;https://arxiv.org/abs/2607.04371&quot;&gt;tech report&lt;/a&gt;). It shrinks from &lt;code&gt;120.7B&lt;/code&gt; total / &lt;code&gt;12.8B&lt;/code&gt; active params to &lt;code&gt;75.3B&lt;/code&gt; total / &lt;code&gt;9.3B&lt;/code&gt; active params, keeps &lt;strong&gt;MTP&lt;/strong&gt; for faster decoding, supports &lt;strong&gt;1M-token context&lt;/strong&gt;, and claims ~&lt;code&gt;2×&lt;/code&gt; higher throughput on a single &lt;code&gt;8×B200&lt;/code&gt; node plus improved single-H100 &lt;code&gt;1M&lt;/code&gt;-token concurrency from &lt;code&gt;1&lt;/code&gt; to &lt;code&gt;8&lt;/code&gt; requests while preserving benchmark performance across reasoning, coding, multilingual, long-context, and agentic tasks.&lt;/strong&gt; Comments focused on deployment practicality: users highlighted the unusually attractive size/context tradeoff, with one planning aggressive local quantized inference (&lt;code&gt;Q6&lt;/code&gt;/&lt;code&gt;Q4&lt;/code&gt;) on &lt;code&gt;64GB DDR4&lt;/code&gt; RAM. Another commenter quoted the model card positioning it as a general-purpose reasoning/chat model for agent systems, RAG, long-context reasoning, and high-volume workloads.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted the model’s positioning as a &lt;strong&gt;75B BF16 general-purpose reasoning/chat model&lt;/strong&gt; for English, code, multilingual use, collaborative agents, high-volume workloads, RAG, complex instruction following, and long-context reasoning, with a notable advertised &lt;strong&gt;&lt;code&gt;1M&lt;/code&gt; context window&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;One technical criticism was that the published benchmarks appear &lt;strong&gt;worse than Super-120&lt;/strong&gt;, which the commenter already considered underwhelming, suggesting this release may not improve on its presumed source/base model despite its long-context and agent-oriented framing.&lt;/li&gt;
&lt;li&gt;Licensing was noted as a positive change: unlike some prior NVIDIA model releases criticized for nonstandard terms, this one was described as having a license closer to &lt;strong&gt;Apache 2.0 / MIT-style permissiveness&lt;/strong&gt;, which may improve adoption for developers and commercial users.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uq9krm/unsloth_has_uploaded_several_sizes_of/&quot;&gt;Unsloth has uploaded several sizes of Deepseek-V4-Flash GGUF&apos;s&lt;/a&gt;&lt;/strong&gt; (Activity: 588): &lt;strong&gt;&lt;strong&gt;Unsloth&lt;/strong&gt; uploaded multiple &lt;strong&gt;DeepSeek-V4-Flash GGUF&lt;/strong&gt; quantizations, but users note current inference requires a specific &lt;code&gt;llama.cpp&lt;/code&gt; fork/branch with a &lt;a href=&quot;https://github.com/danielhanchen/llama.cpp/tree/deepseek-v4-checkpointing-fix&quot;&gt;DeepSeek V4 checkpointing fix&lt;/a&gt;. Early &lt;code&gt;llama-bench&lt;/code&gt; results on &lt;strong&gt;8× RTX 3090&lt;/strong&gt; for &lt;code&gt;DeepSeek-V4-Flash-UD-Q4_K_XL&lt;/code&gt; show a &lt;code&gt;144.44 GiB&lt;/code&gt;, &lt;code&gt;284.33B&lt;/code&gt;-param model at &lt;code&gt;258.77 ± 2.23 t/s&lt;/code&gt; prefill (&lt;code&gt;pp512&lt;/code&gt;) but only &lt;code&gt;19.73 ± 0.24 t/s&lt;/code&gt; generation (&lt;code&gt;tg128&lt;/code&gt;) with CUDA/NGL &lt;code&gt;99&lt;/code&gt;. Another user reports custom heterogeneous placement on a &lt;strong&gt;Framework 16&lt;/strong&gt;—dense layers on Radeon &lt;code&gt;7700S&lt;/code&gt;, experts on &lt;code&gt;780M&lt;/code&gt;, &lt;code&gt;96GB DDR5&lt;/code&gt;—achieving roughly &lt;code&gt;70 TPS&lt;/code&gt; prefill and &lt;code&gt;7 TPS&lt;/code&gt; generation at about &lt;code&gt;100 W&lt;/code&gt; TDP.&lt;/strong&gt; Commenters are optimistic about &lt;strong&gt;Unsloth Dynamic Quants&lt;/strong&gt; and hosted V4-Flash quality, but expect performance to improve as &lt;code&gt;llama.cpp&lt;/code&gt;/backend support matures. One benchmarker said smaller &lt;code&gt;27B int8&lt;/code&gt; models have “spoiled” them due to much higher practical generation speed.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A required &lt;code&gt;llama.cpp&lt;/code&gt; fork/branch was linked for running these GGUFs: &lt;a href=&quot;https://github.com/danielhanchen/llama.cpp/tree/deepseek-v4-checkpointing-fix&quot;&gt;danielhanchen/llama.cpp &lt;code&gt;deepseek-v4-checkpointing-fix&lt;/code&gt;&lt;/a&gt;. This suggests current upstream support may still need DeepSeek-V4-specific checkpointing fixes before the Unsloth GGUFs run correctly or efficiently.&lt;/li&gt;
&lt;li&gt;One user benchmarked &lt;strong&gt;DeepSeek-V4-Flash-UD-Q4_K_XL&lt;/strong&gt; on &lt;code&gt;8x RTX 3090&lt;/code&gt; with CUDA offload &lt;code&gt;NGL=99&lt;/code&gt;: model size &lt;code&gt;144.44 GiB&lt;/code&gt;, &lt;code&gt;284.33B&lt;/code&gt; params, &lt;code&gt;pp512&lt;/code&gt; prefill at &lt;code&gt;258.77 ± 2.23 tok/s&lt;/code&gt;, and &lt;code&gt;tg128&lt;/code&gt; generation at only &lt;code&gt;19.73 ± 0.24 tok/s&lt;/code&gt;. They noted the model/quant quality was good but generation speed felt low compared with a &lt;code&gt;27B int8&lt;/code&gt; setup, likely reflecting immature backend/kernel support for this architecture/quant.&lt;/li&gt;
&lt;li&gt;A Framework 16 user reported running the model on a mixed iGPU/dGPU setup with &lt;code&gt;96GB DDR5&lt;/code&gt;, &lt;code&gt;8GB GDDR6&lt;/code&gt; Radeon &lt;code&gt;7700S&lt;/code&gt;, and &lt;code&gt;780M&lt;/code&gt;, achieving about &lt;code&gt;70 tok/s&lt;/code&gt; prefill and &lt;code&gt;7 tok/s&lt;/code&gt; generation at roughly &lt;code&gt;100W&lt;/code&gt; system TDP. Their custom inference code reportedly pins dense layers to the &lt;code&gt;7700S&lt;/code&gt; while placing MoE experts on the &lt;code&gt;780M&lt;/code&gt;, illustrating a heterogeneous memory/compute placement strategy for fitting large MoE GGUFs on consumer laptop hardware.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upezt0/late_to_the_party_but_holy_mtp/&quot;&gt;Late to the party but... Holy MTP&lt;/a&gt;&lt;/strong&gt; (Activity: 470): &lt;strong&gt;A user reports enabling &lt;strong&gt;MTP (multi-token prediction)&lt;/strong&gt; on a &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; run and seeing roughly a &lt;strong&gt;&lt;code&gt;2×&lt;/code&gt; increase in tokens/sec&lt;/strong&gt;, then notes they want to find “abliterated” MTP variants. A commenter corroborates similar results: &lt;strong&gt;&lt;code&gt;~2×&lt;/code&gt; throughput&lt;/strong&gt; with a &lt;strong&gt;GGUF 8-bit quant&lt;/strong&gt;, while another wants &lt;strong&gt;MLX MTP&lt;/strong&gt; support to “catch up” for an expected &lt;strong&gt;&lt;code&gt;3–4×&lt;/code&gt; prefill speedup&lt;/strong&gt; on Apple &lt;strong&gt;M5&lt;/strong&gt; hardware.&lt;/strong&gt; Commenters frame MTP adoption as evidence that local LLM inference is still early and expect further performance gains; one specifically calls out &lt;strong&gt;Dspark&lt;/strong&gt; as potentially promising.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Users report &lt;strong&gt;MTP&lt;/strong&gt; delivering a substantial decode/prefill speedup in local inference, including one report of roughly &lt;strong&gt;2× speedup&lt;/strong&gt; while using a &lt;strong&gt;GGUF 8-bit quant&lt;/strong&gt;. Another commenter specifically wants &lt;strong&gt;MLX MTP&lt;/strong&gt; support to mature, expecting a &lt;strong&gt;&lt;code&gt;3–4×&lt;/code&gt; prefill speedup&lt;/strong&gt; on an &lt;strong&gt;Apple M5&lt;/strong&gt; setup.&lt;/li&gt;
&lt;li&gt;A technical tradeoff noted is that enabling &lt;strong&gt;MTP&lt;/strong&gt; can consume an additional &lt;strong&gt;&lt;code&gt;1.5–2 GB&lt;/code&gt; of VRAM&lt;/strong&gt;. For very large context windows, users may disable it to avoid VRAM exhaustion, crashes, or spilling into slower system RAM.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Local LLM Reliability for Coding and RAG&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uqpxgp/can_you_trust_local_models_to_answer_accurately/&quot;&gt;Can you trust local models to answer accurately?&lt;/a&gt;&lt;/strong&gt; (Activity: 442): &lt;strong&gt;The image is a benchmark table, &lt;a href=&quot;https://i.redd.it/swjfgszdqzbh1.png&quot;&gt;&lt;strong&gt;“Accuracy &amp;#x26; Memory Across Local Models”&lt;/strong&gt;&lt;/a&gt;, evaluating local LLMs on &lt;strong&gt;&lt;code&gt;7,648&lt;/code&gt; multiple-choice questions&lt;/strong&gt; generated from markdown docs for Node, LangChain.js, TypeScript, Transformers.js, and Vue. The key result is that standalone local models score much lower, roughly &lt;strong&gt;&lt;code&gt;60–83%&lt;/code&gt; without RAG&lt;/strong&gt;, while retrieval augmentation boosts accuracy to about &lt;strong&gt;&lt;code&gt;86–97%&lt;/code&gt;&lt;/strong&gt;, with &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; reportedly highest at &lt;strong&gt;&lt;code&gt;96.9%&lt;/code&gt;&lt;/strong&gt;; Apple Intelligence / &lt;strong&gt;AFM 2 3B on-device&lt;/strong&gt; is notable because it reaches about &lt;strong&gt;&lt;code&gt;86%&lt;/code&gt;&lt;/strong&gt; despite a much smaller ~&lt;code&gt;4k&lt;/code&gt; context window versus &lt;code&gt;32k&lt;/code&gt; for the other models.&lt;/strong&gt; Commenters broadly agreed that small local models like &lt;strong&gt;Apple Intelligence / AFM 2 3B&lt;/strong&gt; and &lt;strong&gt;Gemma 4 E2B&lt;/strong&gt; look surprisingly capable for their size, but that accurate technical answering depends heavily on tooling such as RAG, browser search, or MCP-style integrations. There was also appreciation that larger local models like &lt;strong&gt;Gemma 31B&lt;/strong&gt; and &lt;strong&gt;Qwen 27B&lt;/strong&gt; now exceed &lt;strong&gt;&lt;code&gt;82%&lt;/code&gt;&lt;/strong&gt; accuracy even without RAG, suggesting rapid improvement in local model baselines.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter highlighted that &lt;strong&gt;Gemma 31B&lt;/strong&gt; and &lt;strong&gt;Qwen 27B&lt;/strong&gt; reportedly achieving &lt;code&gt;82%+&lt;/code&gt; accuracy &lt;em&gt;without RAG&lt;/em&gt; is a notable jump compared with roughly six months ago, when comparable local-model accuracy was described as much lower. They also emphasized that &lt;em&gt;proper tooling around the model&lt;/em&gt; can materially improve answer reliability.&lt;/li&gt;
&lt;li&gt;One user described using a &lt;strong&gt;browser MCP&lt;/strong&gt; setup via a Chrome extension with &lt;strong&gt;opencode&lt;/strong&gt; to let local models search the web when accuracy matters. The implied workflow is to compensate for model hallucination or stale knowledge by attaching retrieval/search tooling rather than relying on the base model alone.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uphzhj/qwen_36_27b_absolutely_fails_at_agentic_work/&quot;&gt;Qwen 3.6 27B absolutely fails at agentic work&lt;/a&gt;&lt;/strong&gt; (Activity: 833): &lt;strong&gt;The poster reports that &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; run via &lt;strong&gt;llama.cpp nightly&lt;/strong&gt; on an &lt;strong&gt;RTX 6000&lt;/strong&gt; at &lt;code&gt;8-bit&lt;/code&gt;/&lt;code&gt;16-bit&lt;/code&gt; produces strong single-shot outputs and longer generations than their &lt;strong&gt;Qwen 3.5 122B&lt;/strong&gt; &lt;code&gt;4-bit&lt;/code&gt;/&lt;code&gt;5-bit&lt;/code&gt; setup, but fails in multi-turn/agentic workflows with frequent instruction-following errors roughly every &lt;code&gt;~4&lt;/code&gt; turns. Technical replies suggest debugging inference/configuration rather than model quality alone: using fixed chat templates from &lt;a href=&quot;https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates&quot;&gt;&lt;code&gt;froggeric/Qwen-Fixed-Chat-Templates&lt;/code&gt;&lt;/a&gt; and verifying parameters such as &lt;code&gt;preserve_thinking&lt;/code&gt;.&lt;/strong&gt; Commenters push back that the report lacks enough reproducibility detail—prompting, sampler settings, chat template, and inference params—and one argues that &lt;em&gt;“most people aren&apos;t having your experience,”&lt;/em&gt; implying the issue may be local configuration rather than an inherent 27B model failure.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters suggested the reported agentic failures may be caused by &lt;strong&gt;chat-template or inference-parameter issues&lt;/strong&gt; rather than the Qwen 3.6 27B model itself. Specific fixes mentioned included using froggeric’s corrected templates on Hugging Face (&lt;a href=&quot;https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates&quot;&gt;&lt;code&gt;Qwen-Fixed-Chat-Templates&lt;/code&gt;&lt;/a&gt;) and ensuring parameters such as &lt;code&gt;preserve_thinking&lt;/code&gt; are enabled/configured correctly.&lt;/li&gt;
&lt;li&gt;Users with successful deployments said &lt;strong&gt;Qwen3.6:27B&lt;/strong&gt; can work well for coding, tool calling, and scoring when wrapped in an appropriate agent harness. One commenter reported building &lt;em&gt;“four or five agents”&lt;/em&gt; with it and finding it &lt;em&gt;“a very good coder, a good tool caller, a good scorer,”&lt;/em&gt; especially inside &lt;strong&gt;Pi code&lt;/strong&gt;, while another recommended adapting Pi to the user’s workflow instructions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. GPT-5.6 Sol and Grok 4.5 Launches&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/OpenAI/comments/1uqhviv/gpt56_sol_along_with_terra_and_luna_will_launch/&quot;&gt;GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday.&lt;/a&gt;&lt;/strong&gt; (Activity: 976): &lt;strong&gt;The image is a &lt;strong&gt;marketing-style launch announcement&lt;/strong&gt; for &lt;strong&gt;“GPT-5.6 Sol”&lt;/strong&gt;, with companion names &lt;strong&gt;Terra&lt;/strong&gt; and &lt;strong&gt;Luna&lt;/strong&gt;, saying public availability starts &lt;strong&gt;this Thursday&lt;/strong&gt; and preview access is expanding globally: &lt;a href=&quot;https://i.redd.it/y2zyo1q4kxbh1.png&quot;&gt;image&lt;/a&gt;. No technical details are provided in the post/image—there are &lt;strong&gt;no benchmarks, architecture notes, context-window specs, pricing, API limits, or deployment details&lt;/strong&gt;—so its significance is mainly contextual as a claimed upcoming OpenAI model/product launch.&lt;/strong&gt; Commenters frame the announcement as competitive pressure on Anthropic, with one saying &lt;em&gt;“Competition is a win for everyone.”&lt;/em&gt; Others criticize the naming scheme as confusing and mention saving weekly usage limits for the launch.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ur06sj/grok_45_is_live/&quot;&gt;Grok 4.5 is live&lt;/a&gt;&lt;/strong&gt; (Activity: 1014): &lt;strong&gt;The image is a dark benchmark table announcing &lt;strong&gt;“Grok 4.5 is live”&lt;/strong&gt; and comparing it with Opus 4.8, GPT-5.5, Composer 2.5, and Fable 5 on coding-agent benchmarks including &lt;code&gt;Terminal-Bench 2.1&lt;/code&gt;, &lt;code&gt;SWE-Bench Multilingual&lt;/code&gt;, &lt;code&gt;DeepSWE 1.0&lt;/code&gt;, and &lt;code&gt;SWE-Bench Pro&lt;/code&gt;; Grok 4.5 is shown at &lt;code&gt;83.3%&lt;/code&gt;, &lt;code&gt;78.0%&lt;/code&gt;, &lt;code&gt;62.0%&lt;/code&gt;, and &lt;code&gt;64.7%&lt;/code&gt; respectively. The technical takeaway is that Grok 4.5 appears near-frontier on software-engineering benchmarks, while commenters highlight its claimed &lt;strong&gt;$2/$6 pricing&lt;/strong&gt; and xAI’s claimed efficiency gains; see the &lt;a href=&quot;https://i.redd.it/3s6zt3uvn1ch1.jpeg&quot;&gt;image&lt;/a&gt; and xAI’s &lt;a href=&quot;https://x.ai/news/grok-4-5#pricing&quot;&gt;pricing/efficiency page&lt;/a&gt;.&lt;/strong&gt; Commenters focused less on absolute benchmark leadership and more on cost-performance, calling the &lt;strong&gt;$2/$6&lt;/strong&gt; price “the real surprise” and arguing that output-token throughput and speed may matter more than small benchmark deltas.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused less on raw benchmark rankings and more on &lt;strong&gt;pricing/throughput efficiency&lt;/strong&gt;, noting Grok 4.5 is reportedly priced at &lt;code&gt;$2/$6&lt;/code&gt; and xAI claims up to &lt;code&gt;2x&lt;/code&gt; better efficiency than the current best frontier model in its &lt;a href=&quot;https://x.ai/news/grok-4-5#pricing&quot;&gt;pricing/efficiency post&lt;/a&gt;. The key technical question raised is whether those output-token rates and latency hold under real workloads rather than just launch benchmarks.&lt;/li&gt;
&lt;li&gt;A technically relevant enterprise angle was that if the published benchmark results, cost, and speed are reproducible, Grok 4.5 could win market share despite brand concerns. The argument was that procurement will prioritize &lt;em&gt;passing evals, lower latency, and cheaper inference bills&lt;/em&gt; over public perception, especially for production LLM deployments.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude Fable 5 Limits and Local Economics&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uq2aq5/anthropic_extending_fable_5_for_paid_users_till/&quot;&gt;Anthropic extending Fable 5 for paid users till 12 july&lt;/a&gt;&lt;/strong&gt; (Activity: 1801): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/t1hhakidhubh1.jpeg&quot;&gt;image&lt;/a&gt; is an X post from the verified &lt;strong&gt;Claude&lt;/strong&gt; account announcing that &lt;strong&gt;Claude Fable 5&lt;/strong&gt; access is extended for &lt;strong&gt;paid Anthropic users through July 12&lt;/strong&gt;. A reply clarifies the quota mechanics: paid users can spend up to &lt;strong&gt;&lt;code&gt;50%&lt;/code&gt; of their weekly usage limit&lt;/strong&gt; on Fable 5, then either continue via usage credits or switch to another model.&lt;/strong&gt; Commenters were mostly frustrated about planning and quota timing: several said they rushed through weekly usage or bought extra credits assuming Fable access was ending, and wished Anthropic had communicated the extension earlier or reset usage limits.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Users report that the extension has limited practical value without a &lt;strong&gt;usage reset or quota adjustment&lt;/strong&gt;: several had already consumed most or all of their weekly allowance in anticipation of losing access to &lt;strong&gt;Fable 5&lt;/strong&gt;, with one user citing &lt;code&gt;71%&lt;/code&gt; usage and a reset not occurring until &lt;code&gt;3am Monday&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A paid user noted they bought &lt;strong&gt;extra credits&lt;/strong&gt; and rushed to finish a project because the original cutoff implied less remaining access time; the extension changes the planning horizon but exposes a communication/entitlement issue around model availability dates and paid usage caps.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1uqfqlr/wtf_are_you_guys_even_working_on/&quot;&gt;WTF are you guys even working on?!&lt;/a&gt;&lt;/strong&gt; (Activity: 1205): &lt;strong&gt;A software engineer working across a &lt;code&gt;14-year-old&lt;/code&gt; monorepo, &lt;code&gt;17+&lt;/code&gt; services, and multiple Claude-generated side projects questions why users are exhausting &lt;code&gt;5x&lt;/code&gt; weekly LLM usage and treating &lt;strong&gt;Fable 5&lt;/strong&gt; pricing changes as blocking, arguing that &lt;strong&gt;Opus 4.8&lt;/strong&gt; should handle most coding tasks with only modest quality loss. The core technical point is about cost/performance tradeoffs in coding agents: whether premium models are truly necessary for production code generation, debugging, and large-context workflows, especially when generated code must still be understood and maintained by the developer.&lt;/strong&gt; Commenters pushed back that high-end models are valuable for open-ended tasks like unsupervised code audits, bug discovery, and fast issue resolution with logs and fresh context. Several argued that using LLMs to build paid web apps or fix code faster than humans is sufficient value even if the developer does not fully understand every generated implementation detail.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argue that LLM coding value is strongest in &lt;strong&gt;codebase auditing and bug-finding&lt;/strong&gt;, recommending simple prompts like asking the model to audit code without a narrow focus. One commenter specifically contrasts models, saying &lt;strong&gt;Claude Opus is “fine for coding,” but “Fable” excels at finding problems in codebases&lt;/strong&gt;, suggesting perceived specialization between generation and review/debugging workflows.&lt;/li&gt;
&lt;li&gt;A recurring technical workflow described is using an agent with &lt;strong&gt;fresh repository context, logs, and bug reports&lt;/strong&gt; to diagnose and patch issues autonomously. One user reports sending bug reports to &lt;strong&gt;Claude&lt;/strong&gt;, letting it run in the background for hours and produce a fix after roughly &lt;code&gt;500k+ tokens&lt;/code&gt;, highlighting a high-token, asynchronous debugging pattern where cost is abstracted away from the developer.&lt;/li&gt;
&lt;li&gt;There is debate over whether developers need to fully understand AI-generated code before using it. Some commenters reject that constraint, arguing that if an agent can use logs and context to find/fix issues faster than humans, the practical metric becomes shipped functionality and maintainability rather than manual comprehension of every generated line.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1upz3pn/would_you_even_run_fable_locally_if_you_could/&quot;&gt;would you even run fable locally if you could?&lt;/a&gt;&lt;/strong&gt; (Activity: 905): &lt;strong&gt;The image is a dark-mode X/Twitter screenshot from &lt;strong&gt;Polymarket&lt;/strong&gt; claiming a projection that &lt;strong&gt;Claude Fable could run locally on high-end consumer hardware within ~2 years&lt;/strong&gt; (&lt;a href=&quot;https://i.redd.it/056noyg8xtbh1.png&quot;&gt;image&lt;/a&gt;). The Reddit post questions whether local inference would still be worthwhile if hosted Fable pricing drops faster than consumer hardware capability improves, framing the tradeoff as &lt;strong&gt;local capex / ownership vs. hosted API cost&lt;/strong&gt;, with privacy as the obvious but possibly insufficient differentiator.&lt;/strong&gt; Comments argue that local execution could still matter because it removes provider-side &lt;strong&gt;usage limits&lt;/strong&gt;, enables owned compute, and may benefit from future compression/quantization breakthroughs that preserve model quality. Another commenter jokes that Polymarket’s role is essentially to turn the projection into a betting market.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters frame local execution as primarily a &lt;strong&gt;compute-ownership tradeoff&lt;/strong&gt;: if Fable could run locally, users would avoid hosted API usage limits because inference would be bounded only by their own hardware, memory, and electricity costs.&lt;/li&gt;
&lt;li&gt;A technical skepticism thread argues that &lt;strong&gt;Anthropic is unlikely to release Fable or related closed-weight models locally&lt;/strong&gt;, because monetizing hosted access is core to its business model. The counterpoint is that an &lt;strong&gt;open-weight model&lt;/strong&gt; may reach comparable capability soon, with one commenter claiming open models are already surpassing “Opus 4.5” and projecting parity with Fable-level performance within &lt;code&gt;6–24 months&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;One comment predicts substantial future gains from &lt;strong&gt;LLM compression/efficiency techniques&lt;/strong&gt;, suggesting that smarter quantization, pruning, distillation, or architecture-level improvements could eventually make very large models practical to run locally while retaining much of their effectiveness.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Anthropic J-Space and Fable Cyber-Safety Edge Cases&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uptvgb/anthropic_just_reported_that_llms_have_hidden/&quot;&gt;Anthropic just reported that LLMs have hidden thoughts they hold without saying. An internal ”J-Space”&lt;/a&gt;&lt;/strong&gt; (Activity: 1194): &lt;strong&gt;The post discusses &lt;strong&gt;Anthropic’s&lt;/strong&gt; paper on a small set of model activations termed &lt;strong&gt;J-space&lt;/strong&gt;, described as behaving like a &lt;em&gt;global workspace&lt;/em&gt; where information can be held, reported, and used for multi-step reasoning, while much fluent generation allegedly bypasses it (&lt;a href=&quot;https://www.anthropic.com/research/global-workspace&quot;&gt;paper&lt;/a&gt;). The author built &lt;strong&gt;Subtext&lt;/strong&gt; (&lt;a href=&quot;https://github.com/ninjahawk/Subtext&quot;&gt;GitHub&lt;/a&gt;) to visualize token-disposed internal states before generation, claiming replications such as &lt;code&gt;incorrect&lt;/code&gt; activating before an answer to &lt;code&gt;12 + 5 = 1&lt;/code&gt;, and a two-hop trace where &lt;code&gt;Italy&lt;/code&gt; appears around layer &lt;code&gt;20&lt;/code&gt; and &lt;code&gt;euros&lt;/code&gt; around layer &lt;code&gt;26&lt;/code&gt;; they emphasize this indicates reportable/usable internal information, &lt;strong&gt;not evidence of subjective experience&lt;/strong&gt;.&lt;/strong&gt; Comments mostly note that this is consistent with earlier mechanistic-interpretability hints and push back on the “stochastic parrot” framing; one commenter also questions what model was used to implement the reproduction. Another top comment jokingly suggests the post’s cautious phrasing sounds like Claude-generated text.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted Anthropic’s finding that models can maintain internal representations not surfaced in emitted tokens, arguing this challenges the simplistic &lt;em&gt;“just next-token prediction”&lt;/em&gt; / &lt;em&gt;“stochastic parrot”&lt;/em&gt; framing. One technically notable interpretation was that the model’s latent &lt;code&gt;J-space&lt;/code&gt; can encode intermediate beliefs or classifications before they are verbalized.&lt;/li&gt;
&lt;li&gt;A technical question was raised about whether the reported phenomenon is essentially neuron/feature activation tracking—e.g. &lt;em&gt;“Italy neurons”&lt;/em&gt; activating on the path to an answer—or whether it demonstrates something stronger. The arithmetic examples were singled out as more interesting because they imply latent intermediate computation rather than merely semantic feature activation.&lt;/li&gt;
&lt;li&gt;One commenter emphasized the reported difference between base training and post-training: before post-training, the model’s internal state was described as mostly constrained to predicting user tokens, while after identity/alignment post-training it appeared to form first-person-like judgments while reading input. The cited example was recognizing a prompt injection internally &lt;em&gt;before&lt;/em&gt; producing any output token.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1upu3e2/fable_5_found_actual_malware_on_my_pc_and_then/&quot;&gt;Fable 5 found actual malware on my PC, and then its own safety filters flagged the warning.&lt;/a&gt;&lt;/strong&gt; (Activity: 1871): &lt;strong&gt;A user reports &lt;strong&gt;Fable 5&lt;/strong&gt; inspected the Windows &lt;code&gt;Run&lt;/code&gt; registry key and flagged an unexpected persistence mechanism: &lt;code&gt;powershell.exe -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden ...&lt;/code&gt; downloading a remote script at sign-in, which it classified as an active compromise (&lt;a href=&quot;https://preview.redd.it/2hv0yord1tbh1.png?width=1172&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=6434cb2d41bb2474fa40398c24fa575b1f74c635&quot;&gt;screenshot&lt;/a&gt;). After the user asked it to remove the relevant registry entries, the model reportedly completed the cleanup but the session was then downgraded to &lt;strong&gt;Opus 4.8&lt;/strong&gt; because the interaction was flagged as “cybersecurity work” by safety filters (&lt;a href=&quot;https://preview.redd.it/402jnkkf1tbh1.png?width=1163&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=7c6531edf9e593fd592109f91bd0ca45de8e6650&quot;&gt;screenshot&lt;/a&gt;).&lt;/strong&gt; Commenters were skeptical of relying on an LLM for endpoint remediation, noting this PowerShell &lt;code&gt;Run&lt;/code&gt;-key persistence pattern is old and typically caught by conventional AV/EDR; one argued an antivirus is more appropriate because an LLM may find “1 and leave 10 others.” Another commenter reported a similar beneficial security review use case where the model found and documented codebase issues without triggering a downgrade.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter argued that malware detection should be handled by dedicated antivirus/EDR tooling rather than an LLM agent: the described malware family is reportedly &lt;strong&gt;~12 years old&lt;/strong&gt; and likely covered by conventional signatures/heuristics, whereas Fable may detect one artifact while missing others.&lt;/li&gt;
&lt;li&gt;One user described using Fable to scan a codebase for bugs; the agent found their &lt;code&gt;security.md&lt;/code&gt;, updated it, and added multiple security findings significant enough that they patched them before production. They noted this did not appear to reduce their model tier/access, implying the safety system allowed code-security remediation despite flagging related content elsewhere.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>xai</category><category>cursor</category><category>scaling01</category><category>grok-4.5</category><category>opus-4.7</category><category>opus-4.8</category><category>gpt-5.6</category><category>elonmusk</category><category>coding</category><category>agents</category><category>model-scaling</category><category>context-window</category><category>model-pricing</category><category>token-efficiency</category><category>model-training</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-07-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-07-not-much/</guid><description>**Anthropic** expanded the &quot;background agent&quot; UX with **Claude Cowork** for mobile and web, emphasizing task-running background teammates. They also extended access to **Claude Fable 5** on paid plans. The concept of a **harness** in agent design gained traction, highlighted by Lilian Weng and echoed by **LangChain** with a new **Deep Agents** course and open-source project. **Google**&apos;s **Gemini API Managed Agents** introduced features like background execution and custom function calling. Operator-facing agent infrastructure saw updates from **Codex Mobile iOS**, **Hermes Agent** with **1Password** integration, and **Weaviate 1.38** enabling runtime-gated write access. Experimentation with human-in-the-loop control via phone/SMS was noted. In model releases, **Meta AI** launched **Muse Image** and previewed **Muse Video**, featuring an agentic generation loop with planning, web search, and self-refinement, achieving top ranks on Image and Video Arena. **NVIDIA** released **Audex**, a 30B parameter MoE model with 1M context for unified text and audio tasks.</description><pubDate>Tue, 07 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/06/2026-7/07/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Agent Products, Harnesses, and Long-Running Workflows&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Anthropic expands “background agent” UX on top of Claude&lt;/strong&gt;: The biggest product launch by engagement was &lt;a href=&quot;https://x.com/claudeai/status/2074525815820169320&quot;&gt;Claude Cowork coming to mobile and web&lt;/a&gt;, positioning Claude as a task-running background teammate rather than a foreground chat UI. Related posts show the product convergence around a shared home tab and tighter Chat/Cowork integration from &lt;a href=&quot;https://x.com/mikeyk/status/2074531605537046953&quot;&gt;@mikeyk&lt;/a&gt;. Separately, Anthropic extended access to &lt;strong&gt;Claude Fable 5&lt;/strong&gt; on paid plans through July 12 in a highly engaged announcement from &lt;a href=&quot;https://x.com/claudeai/status/2074548242386178258&quot;&gt;@claudeai&lt;/a&gt;, though many users noted the awkward timing relative to weekly limits in reactions from &lt;a href=&quot;https://x.com/kimmonismus/status/2074606005963391225&quot;&gt;@kimmonismus&lt;/a&gt; and others.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Harness engineering is increasingly the center of agent design&lt;/strong&gt;: Lilian Weng’s new post was widely referenced as reframing recursive self-improvement around the &lt;strong&gt;harness&lt;/strong&gt;, not direct weight self-modification; Sakana’s summary connects this to &lt;strong&gt;The AI Scientist&lt;/strong&gt;, &lt;strong&gt;ShinkaEvolve&lt;/strong&gt;, and &lt;strong&gt;Darwin Gödel Machine&lt;/strong&gt; in &lt;a href=&quot;https://x.com/SakanaAILabs/status/2074489949529776308&quot;&gt;their thread&lt;/a&gt;. LangChain echoed the same shift with a new &lt;strong&gt;Deep Agents&lt;/strong&gt; course and an open-source harness project in posts from &lt;a href=&quot;https://x.com/LangChain/status/2074539083204820997&quot;&gt;@LangChain&lt;/a&gt; and &lt;a href=&quot;https://x.com/hwchase17/status/2074547871194698207&quot;&gt;@hwchase17&lt;/a&gt;. Google is also productizing this direction: Gemini API &lt;strong&gt;Managed Agents&lt;/strong&gt; added &lt;strong&gt;background execution&lt;/strong&gt;, &lt;strong&gt;remote MCP servers&lt;/strong&gt;, &lt;strong&gt;custom function calling&lt;/strong&gt;, and &lt;strong&gt;credential refresh&lt;/strong&gt; in posts from &lt;a href=&quot;https://x.com/_philschmid/status/2074533915038027972&quot;&gt;@_philschmid&lt;/a&gt; and &lt;a href=&quot;https://x.com/OfficialLoganK/status/2074552932318765376&quot;&gt;@OfficialLoganK&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Practical agent infra keeps getting more opinionated&lt;/strong&gt;: There were several notable operator-facing updates: &lt;strong&gt;Codex Mobile iOS&lt;/strong&gt; added task management, filtered diffs, SSH key login, branch comparison, and attachment flows in posts from &lt;a href=&quot;https://x.com/Dimillian/status/2074396968223211819&quot;&gt;@Dimillian&lt;/a&gt; and &lt;a href=&quot;https://x.com/reach_vb/status/2074400018769793176&quot;&gt;@reach_vb&lt;/a&gt;; &lt;strong&gt;Hermes Agent&lt;/strong&gt; added pluggable secrets managers plus native &lt;strong&gt;1Password&lt;/strong&gt; integration and export of sessions/datasets to formats including private Hugging Face repos in &lt;a href=&quot;https://x.com/Teknium/status/2074564207555772912&quot;&gt;@Teknium’s&lt;/a&gt; &lt;a href=&quot;https://x.com/Teknium/status/2074639961727655959&quot;&gt;threads&lt;/a&gt;; &lt;strong&gt;Weaviate 1.38&lt;/strong&gt; made its MCP server GA with runtime-gated write access, notably allowing &lt;strong&gt;MCP_SERVER_WRITE_ACCESS_ENABLED&lt;/strong&gt; to be flipped live without restart in &lt;a href=&quot;https://x.com/victorialslocum/status/2074493681403339104&quot;&gt;@victorialslocum’s post&lt;/a&gt;. A more experimental pattern came from &lt;a href=&quot;https://x.com/omarsar0/status/2074506169352180108&quot;&gt;@omarsar0&lt;/a&gt;, using a Dial MCP server so agents can escalate decisions via phone call/SMS/iMessage for human-in-the-loop control.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Model and Modality Releases: Audio, Speech, Robotics, and Media Generation&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Meta’s Muse Image/Muse Video push agentic generation into media&lt;/strong&gt;: Meta Superintelligence Labs launched &lt;strong&gt;Muse Image&lt;/strong&gt; and previewed &lt;strong&gt;Muse Video&lt;/strong&gt; in announcements from &lt;a href=&quot;https://x.com/AIatMeta/status/2074577662840832382&quot;&gt;@AIatMeta&lt;/a&gt;, &lt;a href=&quot;https://x.com/alexandr_wang/status/2074555909347369105&quot;&gt;@alexandr_wang&lt;/a&gt;, and &lt;a href=&quot;https://x.com/_tim_brooks/status/2074578008296628698&quot;&gt;@_tim_brooks&lt;/a&gt;. The notable technical angle is not just image quality, but an explicitly &lt;strong&gt;agentic generation loop&lt;/strong&gt;: planning, web search, tool use, code execution, and self-refinement before rendering. Meta also says performance improves with &lt;strong&gt;scaled test-time compute&lt;/strong&gt;, and that self-refinement behavior emerged during RL rather than being hand-scripted in &lt;a href=&quot;https://x.com/AIatMeta/status/2074587864923250873&quot;&gt;this follow-up&lt;/a&gt;. On public evals, Muse Image quickly reached &lt;strong&gt;#2 on Image Arena&lt;/strong&gt; behind GPT Image 2 in &lt;a href=&quot;https://x.com/arena/status/2074581979765539153&quot;&gt;Arena’s ranking&lt;/a&gt;, while Muse Video debuted at &lt;strong&gt;#3 on Video Arena&lt;/strong&gt; in &lt;a href=&quot;https://x.com/arena/status/2074591193783320851&quot;&gt;another Arena post&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;NVIDIA and Cohere both shipped strong audio releases&lt;/strong&gt;: NVIDIA released &lt;strong&gt;Audex&lt;/strong&gt;, a &lt;strong&gt;30B parameter / 3B active MoE&lt;/strong&gt; with &lt;strong&gt;1M context&lt;/strong&gt; for unified text+audio work, summarized by &lt;a href=&quot;https://x.com/HuggingPapers/status/2074384562952749254&quot;&gt;@HuggingPapers&lt;/a&gt; and described in more detail by &lt;a href=&quot;https://x.com/_weiping/status/2074537900172050704&quot;&gt;@_weiping&lt;/a&gt;. The model’s core claim is preserving text intelligence while adding broad audio generation and understanding via a single MoE backbone. Cohere launched &lt;strong&gt;Cohere Transcribe Arabic&lt;/strong&gt;, described as the most accurate open-source Arabic ASR model, under &lt;strong&gt;Apache 2.0&lt;/strong&gt;, with emphasis on &lt;strong&gt;dialects&lt;/strong&gt;, &lt;strong&gt;code-switching&lt;/strong&gt;, and &lt;strong&gt;Arabic-accented English&lt;/strong&gt; in posts from &lt;a href=&quot;https://x.com/cohere/status/2074499759616729149&quot;&gt;@cohere&lt;/a&gt; and &lt;a href=&quot;https://x.com/JayAlammar/status/2074511963934118282&quot;&gt;@JayAlammar&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Open robotics keeps consolidating around Hugging Face + NVIDIA&lt;/strong&gt;: NVIDIA expanded its robotics stack into the HF ecosystem by bringing &lt;strong&gt;GR00T 1.7&lt;/strong&gt; and &lt;strong&gt;Isaac Teleop&lt;/strong&gt; into &lt;strong&gt;LeRobot&lt;/strong&gt;, aimed at open humanoid robotics workflows, in &lt;a href=&quot;https://x.com/NVIDIARobotics/status/2074380795855147072&quot;&gt;@NVIDIARobotics’s announcement&lt;/a&gt; and &lt;a href=&quot;https://x.com/NVIDIARobotics/status/2074390485251113317&quot;&gt;integration guide&lt;/a&gt;. On the embodied side, UMA showed a strong full-stack robotics narrative: &lt;a href=&quot;https://x.com/RemiCadene/status/2074442725814878510&quot;&gt;@RemiCadene&lt;/a&gt; described a prototype built by a small team in 9 months, while &lt;a href=&quot;https://x.com/RemiCadene/status/2074442439142609237&quot;&gt;the Northstar reveal&lt;/a&gt; and &lt;a href=&quot;https://x.com/psermanet/status/2074512829617491996&quot;&gt;@psermanet’s safety note&lt;/a&gt; emphasized vertically integrated hardware/software for trustworthy robots.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Training, Inference, and Post-Training Techniques&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Liquid AI’s “Antidoom” directly targets reasoning-loop failure modes&lt;/strong&gt;: One of the clearest technical releases of the day was &lt;a href=&quot;https://x.com/liquidai/status/2074494130126811473&quot;&gt;Liquid AI’s Antidoom&lt;/a&gt;, an open-source training method to reduce &lt;strong&gt;doom loops&lt;/strong&gt; where small reasoning models repeat tokens until context exhaustion. The reported reductions are substantial: &lt;strong&gt;LFM2.5-2.6B from 10.2% → 1.4%&lt;/strong&gt; and &lt;strong&gt;Qwen3.5-4B from 22.9% → 1%&lt;/strong&gt; under greedy sampling, with downstream eval gains. The method, &lt;strong&gt;FTPO (Final Token Preference Optimization)&lt;/strong&gt;, relabels the loop-triggering token and redistributes probability toward alternatives, summarized well by &lt;a href=&quot;https://x.com/helloiamleonie/status/2074498103982408044&quot;&gt;@helloiamleonie&lt;/a&gt; and &lt;a href=&quot;https://x.com/LiorOnAI/status/2074547819114086561&quot;&gt;@LiorOnAI&lt;/a&gt;. This is a good example of the field’s recent pattern: removing specific failure modes rather than only scaling parameters.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Inference efficiency and compression remain a major frontier&lt;/strong&gt;: NVIDIA’s &lt;strong&gt;Puzzle-75B-A9B&lt;/strong&gt; compression work got strong attention via &lt;a href=&quot;https://x.com/omarsar0/status/2074543978129793462&quot;&gt;@omarsar0&lt;/a&gt;: compressing a hybrid MoE parent model while preserving reasoning, coding, long-context, and agentic quality, with roughly &lt;strong&gt;2x server throughput&lt;/strong&gt; and &lt;strong&gt;1M-context concurrency on H100 rising from 1 request to 8&lt;/strong&gt;. On the tooling side, &lt;strong&gt;Nsight Python 1.0&lt;/strong&gt; launched in &lt;a href=&quot;https://x.com/HagedornBastian/status/2074509770342445375&quot;&gt;@HagedornBastian’s post&lt;/a&gt;, making GPU perf analysis scriptable in Python. Unsloth also shipped &lt;strong&gt;GGUFs for DeepSeek-V4-Flash&lt;/strong&gt;, plus export to &lt;strong&gt;NVFP4/FP8&lt;/strong&gt; and speedups for &lt;strong&gt;GRPO&lt;/strong&gt; and MoEs in &lt;a href=&quot;https://x.com/danielhanchen/status/2074510444778463331&quot;&gt;@danielhanchen’s update&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Agent RL and verification are getting more specialized&lt;/strong&gt;: &lt;a href=&quot;https://x.com/cwolferesearch/status/2074558199819067606&quot;&gt;@cwolferesearch&lt;/a&gt; highlighted how &lt;strong&gt;GRPO-style normalization&lt;/strong&gt; is being adapted for agentic RL at the &lt;strong&gt;task&lt;/strong&gt; or &lt;strong&gt;environment&lt;/strong&gt; level to handle higher reward variance in multi-turn environments. Separately, &lt;a href=&quot;https://x.com/omarsar0/status/2074556579580711050&quot;&gt;@omarsar0&lt;/a&gt; flagged a training-free &lt;strong&gt;verifier&lt;/strong&gt; paper from Stanford/NVIDIA/Berkeley that reads calibrated continuous scores off scoring-token logits, posting strong numbers across &lt;strong&gt;Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench&lt;/strong&gt; and suggesting verification is becoming an independent scaling axis.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Interpretability, Model Internals, and the “J-Space” Debate&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Anthropic’s J-space work dominated interpretability discussion, but also drew sharp criticism&lt;/strong&gt;: The community split between seeing the work as useful mechanistic analysis and objecting to the consciousness framing. Strong critiques came from &lt;a href=&quot;https://x.com/danburonline/status/2074429991576650014&quot;&gt;@danburonline&lt;/a&gt;, &lt;a href=&quot;https://x.com/paul_cal/status/2074388528243310976&quot;&gt;@paul_cal&lt;/a&gt;, and &lt;a href=&quot;https://x.com/scaling01/status/2074432865794679235&quot;&gt;@scaling01&lt;/a&gt;, who argued the vectors are causal largely by construction under the Jacobian-lens definition. A useful historical reference came from &lt;a href=&quot;https://x.com/jacobandreas/status/2074487546692735002&quot;&gt;@jacobandreas&lt;/a&gt;, pointing readers back to the original &lt;strong&gt;Jacobian lenses&lt;/strong&gt; paper.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;The stronger technical takeaway is cross-model structure, not consciousness rhetoric&lt;/strong&gt;: &lt;a href=&quot;https://x.com/eliebakouch/status/2074532904009421260&quot;&gt;@eliebakouch&lt;/a&gt; computed &lt;strong&gt;CKA similarity&lt;/strong&gt; on J-lens geometry across &lt;strong&gt;38 open models&lt;/strong&gt; and found surprisingly universal layer/depth organization, even across unrelated families like &lt;strong&gt;Llama&lt;/strong&gt; and &lt;strong&gt;OLMo&lt;/strong&gt;. Anthropic and Neuronpedia also released &lt;strong&gt;J-lens weights for open models&lt;/strong&gt;, noted in &lt;a href=&quot;https://x.com/eliebakouch/status/2074537985102565795&quot;&gt;this follow-up&lt;/a&gt;. In parallel, Goodfire introduced &lt;strong&gt;Block-Sparse Featurizers&lt;/strong&gt; for multidimensional concepts in activations, arguing many vision concepts are inherently &lt;strong&gt;2–4 dimensional blocks&lt;/strong&gt; rather than single directions, in &lt;a href=&quot;https://x.com/GoodfireAI/status/2074634702737281303&quot;&gt;their thread&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Benchmarks, Evaluations, and Domain-Specific Systems&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Agent and legal benchmarks continue to expose the gap between “passes many criteria” and “fully solves real work”&lt;/strong&gt;: &lt;a href=&quot;https://x.com/arena/status/2074484787663052849&quot;&gt;Agent Arena&lt;/a&gt; placed &lt;strong&gt;Claude Sonnet 5 (Thinking)&lt;/strong&gt; at &lt;strong&gt;#6&lt;/strong&gt;, with strongest signals in confirmed task success and bash usage, but still with uncertainty around steerability. Artificial Analysis launched &lt;strong&gt;Harvey LAB-AA&lt;/strong&gt;, a legal-agent benchmark over &lt;strong&gt;120 private legal tasks across 24 practice areas&lt;/strong&gt;, where &lt;strong&gt;Claude Fable 5&lt;/strong&gt; led at &lt;strong&gt;14.2% all-pass rate&lt;/strong&gt;; &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt; and &lt;strong&gt;GLM-5.2&lt;/strong&gt; tied at &lt;strong&gt;7.5%&lt;/strong&gt;, with GLM hitting that at roughly &lt;strong&gt;~6% of Fable’s cost per task&lt;/strong&gt; in &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074541975186165887&quot;&gt;their release&lt;/a&gt;. The big message is that models can satisfy many individual rubric items yet still fail to produce acceptable end-to-end deliverables.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Research automation and specialized domain systems are broadening&lt;/strong&gt;: Google promoted &lt;strong&gt;Experience AI Scientist&lt;/strong&gt;, a multi-agent system for end-to-end scientific workflows, in &lt;a href=&quot;https://x.com/GoogleResearch/status/2074384746076135575&quot;&gt;this ICML post&lt;/a&gt;. DeepMind also launched &lt;strong&gt;Predicting the Past&lt;/strong&gt;, grounding Gemini in &lt;strong&gt;Aeneas&lt;/strong&gt; and &lt;strong&gt;Ithaca&lt;/strong&gt; for Greek/Latin historical analysis via plain-English interactions, in &lt;a href=&quot;https://x.com/GoogleDeepMind/status/2074513661750546762&quot;&gt;their thread&lt;/a&gt;. On legal AI commercialization, &lt;strong&gt;Norm Ai&lt;/strong&gt; announced a &lt;strong&gt;$120M Series C at $1.2B valuation&lt;/strong&gt; and described a full-stack “agentic law” setup spanning software plus an AI-native law firm in &lt;a href=&quot;https://x.com/johnjnay/status/2074485345593245833&quot;&gt;@johnjnay’s post&lt;/a&gt;.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude access / product rollout&lt;/strong&gt;: &lt;a href=&quot;https://x.com/claudeai/status/2074525815820169320&quot;&gt;Claude Cowork on mobile and web&lt;/a&gt; and &lt;a href=&quot;https://x.com/claudeai/status/2074548242386178258&quot;&gt;Fable 5 access extended through July 12&lt;/a&gt; were the most-engaged technically relevant product announcements.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open-source developer program&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2074570404035993780&quot;&gt;@ClaudeDevs offering 6 months of Claude Max 20x for open-source maintainers&lt;/a&gt; drew massive engagement and is likely to matter for tool adoption in OSS ecosystems.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meta media generation&lt;/strong&gt;: &lt;a href=&quot;https://x.com/AIatMeta/status/2074577662840832382&quot;&gt;Muse Image launch&lt;/a&gt; and &lt;a href=&quot;https://x.com/arena/status/2074581979765539153&quot;&gt;Arena’s #2 ranking for Muse Image&lt;/a&gt; were the biggest multimodal product stories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reasoning reliability&lt;/strong&gt;: &lt;a href=&quot;https://x.com/liquidai/status/2074494130126811473&quot;&gt;Liquid AI’s Antidoom release&lt;/a&gt; stood out as the day’s highest-signal training technique post.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Interpretability&lt;/strong&gt;: &lt;a href=&quot;https://x.com/eliebakouch/status/2074532904009421260&quot;&gt;Cross-model J-lens universality across 38 open models&lt;/a&gt; was the strongest technical follow-on to the J-space discourse.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Open Model Releases and Inference Efficiency&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uoozt4/new_open_model_from_tencent_hy_hy3_295b_total_21b/&quot;&gt;New open  model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)&lt;/a&gt;&lt;/strong&gt; (Activity: 653): &lt;strong&gt;&lt;strong&gt;Tencent&lt;/strong&gt; released the non-preview &lt;strong&gt;Hy3&lt;/strong&gt; open model collection on &lt;a href=&quot;https://huggingface.co/collections/tencent/hy3&quot;&gt;Hugging Face&lt;/a&gt;, described as a &lt;code&gt;295B&lt;/code&gt;-parameter MoE with &lt;code&gt;21B&lt;/code&gt; active parameters, now under &lt;strong&gt;Apache 2.0&lt;/strong&gt; rather than the prior restrictive community license. The post highlights that the earlier license reportedly excluded use in regions including &lt;strong&gt;South Korea, the UK, and the EU&lt;/strong&gt;, while top comments point to claimed benchmark gains over &lt;strong&gt;HY3-Preview&lt;/strong&gt; and frame this as potentially relevant for high-end local/home inference setups.&lt;/strong&gt; Commenters viewed the Apache 2.0 relicensing as the most important change, especially given Tencent’s recent translation models also using Apache licensing. There was cautious optimism that the reported benchmark improvements may translate to real-world usefulness, but with implicit skepticism until tested outside vendor charts.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted that &lt;strong&gt;Hunyuan/HY3&lt;/strong&gt; is now listed as &lt;strong&gt;Apache 2.0&lt;/strong&gt;, contrasting it with the prior “community” license that reportedly restricted usage in regions such as &lt;strong&gt;South Korea, the UK, and the EU&lt;/strong&gt;. This was viewed as technically important for deployment because Apache 2.0 removes many commercial and geographic usage barriers.&lt;/li&gt;
&lt;li&gt;Several users focused on whether Tencent’s claimed benchmark improvements over &lt;strong&gt;HY3-Preview&lt;/strong&gt; will translate into real-world workloads. Given the reported &lt;strong&gt;&lt;code&gt;295B&lt;/code&gt; total / &lt;code&gt;21B&lt;/code&gt; active&lt;/strong&gt; MoE-style configuration, commenters suggested it could be relevant for “high-end home setups” if inference formats such as &lt;strong&gt;GGUF&lt;/strong&gt; become available.&lt;/li&gt;
&lt;li&gt;There was early speculation that HY3 could become an alternative to &lt;strong&gt;Qwen&lt;/strong&gt; and &lt;strong&gt;MiniMax&lt;/strong&gt; models in local/open-weight workflows, but commenters were waiting for quantized releases and independent testing before drawing conclusions.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uotkm7/new_model_gigachat35432ba28b_with_day0_gguf/&quot;&gt;New model: GigaChat3.5-432B-A28B (with day-0 GGUF support!)&lt;/a&gt;&lt;/strong&gt; (Activity: 510): &lt;strong&gt;&lt;strong&gt;Sberbank/ai-sage&lt;/strong&gt; released &lt;strong&gt;GigaChat3.5-432B-A28B&lt;/strong&gt;, a large MoE chat model with &lt;code&gt;432B&lt;/code&gt; total / &lt;code&gt;28B&lt;/code&gt; active parameters, plus a &lt;a href=&quot;https://huggingface.co/ai-sage/GigaChat3.5-432B-A28B-base&quot;&gt;base checkpoint&lt;/a&gt; and day-0 &lt;a href=&quot;https://huggingface.co/ai-sage/GigaChat3.5-432B-A28B-GGUF&quot;&gt;GGUF weights&lt;/a&gt;; &lt;code&gt;llama.cpp&lt;/code&gt; support is currently via &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/pull/25342&quot;&gt;PR #25342&lt;/a&gt;. Model-card excerpts claim it is ~&lt;code&gt;40%&lt;/code&gt; smaller than &lt;strong&gt;GigaChat 3.1 Ultra&lt;/strong&gt; &lt;code&gt;700B&lt;/code&gt; while improving code/math/agentic benchmarks, using ~&lt;code&gt;4×&lt;/code&gt; less KV cache per token, fitting &lt;code&gt;&gt;2×&lt;/code&gt; more context in the same memory, and improving throughput by ~&lt;code&gt;20%&lt;/code&gt;. Architecturally, commenters highlighted its custom hybrid MoE stack mixing &lt;strong&gt;MLA&lt;/strong&gt; layers with &lt;strong&gt;GatedDeltaNet&lt;/strong&gt; linear-attention layers, plus &lt;strong&gt;Multi-Token Prediction&lt;/strong&gt; with two MTP heads, claimed to accelerate greedy decoding from ~&lt;code&gt;1.5×&lt;/code&gt; with one head to up to &lt;code&gt;2.2×&lt;/code&gt; with two.&lt;/strong&gt; Commenters questioned using &lt;strong&gt;DeepSeek 3.2&lt;/strong&gt; as a benchmark reference, calling it roughly a year behind frontier systems, and noted that GigaChat3.5 is a &lt;em&gt;non-reasoning&lt;/em&gt; model so benchmark comparisons should account for that. The release was praised for unusually high openness at this scale—base model and intermediate checkpoints are available—though the exact training dataset remains undisclosed.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters noted that &lt;strong&gt;GigaChat3.5-432B-A28B&lt;/strong&gt; should not be compared directly against current frontier reasoning models: one questioned using &lt;strong&gt;DeepSeek 3.2&lt;/strong&gt; as a benchmark reference because it is perceived as &lt;em&gt;&quot;~year behind the frontier models&quot;&lt;/em&gt;, while another emphasized that GigaChat 3.5 is a &lt;strong&gt;non-reasoning model&lt;/strong&gt;, which materially changes how its benchmark scores should be interpreted.&lt;/li&gt;
&lt;li&gt;A technical excerpt highlights major architectural changes versus &lt;strong&gt;GigaChat 3.1 Ultra 700B&lt;/strong&gt;: GigaChat 3.5 is reportedly &lt;code&gt;~40%&lt;/code&gt; smaller while stronger in code, math, and agentic tasks, uses about &lt;code&gt;4×&lt;/code&gt; less KV-cache per token, fits &lt;code&gt;2×+&lt;/code&gt; more context in the same memory, and improves generation throughput by &lt;code&gt;~20%&lt;/code&gt;. The model uses a custom MoE hybrid attention design combining &lt;strong&gt;MLA&lt;/strong&gt; with &lt;strong&gt;GatedDeltaNet&lt;/strong&gt; linear-attention layers, plus &lt;strong&gt;Multi-Token Prediction&lt;/strong&gt; with two MTP heads, claiming greedy decoding speedups of &lt;code&gt;~1.5×&lt;/code&gt; with one head and up to &lt;code&gt;2.2×&lt;/code&gt; with two.&lt;/li&gt;
&lt;li&gt;One commenter praised the release for open-weighting not only the final model but also &lt;strong&gt;intermediate checkpoints and the base model&lt;/strong&gt;, calling that unusually open for a model of this scale, with the main missing artifact being the exact training dataset. Another noted the model likely has its strongest niche in &lt;strong&gt;Russian-language processing&lt;/strong&gt;, while being comparatively average outside Russian due to stronger multilingual alternatives already available.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upsdmi/nvidianvidianemotronlabs3puzzle75ba9bbf16_hugging/&quot;&gt;nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 · Hugging Face&lt;/a&gt;&lt;/strong&gt; (Activity: 349): &lt;strong&gt;&lt;strong&gt;NVIDIA&lt;/strong&gt; released &lt;a href=&quot;https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16&quot;&gt;&lt;code&gt;NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16&lt;/code&gt;&lt;/a&gt;, a commercially usable deployment-optimized hybrid MoE LLM derived from &lt;strong&gt;Nemotron-3-Super-120B-A12B&lt;/strong&gt; via the &lt;strong&gt;Iterative Puzzle&lt;/strong&gt; post-training compression method described in the &lt;a href=&quot;https://arxiv.org/abs/2607.04371&quot;&gt;technical report&lt;/a&gt;. It reduces size from &lt;code&gt;120.7B&lt;/code&gt; total / &lt;code&gt;12.8B&lt;/code&gt; active parameters to &lt;code&gt;75.3B&lt;/code&gt; total / &lt;code&gt;9.3B&lt;/code&gt; active while retaining interleaved &lt;strong&gt;Mamba + MoE + Attention&lt;/strong&gt; layers and &lt;strong&gt;Multi-Token Prediction&lt;/strong&gt;, with claimed gains of ~&lt;code&gt;2×&lt;/code&gt; server throughput on a single &lt;code&gt;8×B200&lt;/code&gt; node and &lt;code&gt;1M&lt;/code&gt;-token single-H100 concurrency increasing from &lt;code&gt;1&lt;/code&gt; to &lt;code&gt;8&lt;/code&gt; requests. The model targets reasoning/chat, code, multilingual use, RAG/agent workloads, and long-context reasoning across English, French, German, Italian, Japanese, Spanish, and Chinese.&lt;/strong&gt; Commenters focused on the model’s practical deployment profile, especially the relatively smaller &lt;code&gt;75B&lt;/code&gt;/&lt;code&gt;9B active&lt;/code&gt; footprint and &lt;code&gt;1M&lt;/code&gt; context. One user joked about attempting &lt;code&gt;Q6&lt;/code&gt;/&lt;code&gt;Q4&lt;/code&gt; quantized inference on &lt;code&gt;64GB DDR4 RAM&lt;/code&gt;, reflecting interest in local/consumer-accessible deployment despite the BF16 release targeting high-end accelerators.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The thread notes that &lt;strong&gt;NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16&lt;/strong&gt; is positioned as a general-purpose reasoning/chat model for &lt;strong&gt;English, code, multilingual use, agent systems, RAG, complex instruction following, and long-context reasoning&lt;/strong&gt;, with a notably large &lt;strong&gt;&lt;code&gt;1M&lt;/code&gt; token context window&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;One commenter raised benchmark concerns, claiming that the model’s published results are &lt;strong&gt;worse than Super-120&lt;/strong&gt;, which they describe as already underwhelming, suggesting limited improvement over the apparent source/base model.&lt;/li&gt;
&lt;li&gt;There is interest in running quantized variants locally, specifically &lt;strong&gt;Q6/Q4&lt;/strong&gt; on &lt;strong&gt;&lt;code&gt;64GB DDR4 RAM&lt;/code&gt;&lt;/strong&gt;, implying the model’s effective deployability may depend heavily on quantization and CPU/RAM-bound inference performance.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1up3mui/thinkingcapqwen3627b_same_accuracy_as_base_qwen36/&quot;&gt;ThinkingCap-Qwen3.6-27B: same accuracy as base Qwen3.6 with ~50% fewer thinking&lt;/a&gt;&lt;/strong&gt; (Activity: 334): &lt;strong&gt;&lt;strong&gt;bottlecapai&lt;/strong&gt; released/evaluated &lt;a href=&quot;https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B#out-of-domain-token-efficiency&quot;&gt;&lt;code&gt;ThinkingCap-Qwen3.6-27B&lt;/code&gt;&lt;/a&gt;, claiming roughly base &lt;strong&gt;Qwen3.6-27B&lt;/strong&gt; accuracy with ~&lt;code&gt;50%&lt;/code&gt; fewer “thinking”/reasoning tokens. The authors report multi-seed benchmarking at Qwen’s recommended &lt;code&gt;temperature=1.0&lt;/code&gt; with statistical significance testing across reasoning, MCQA, chat, system-prompt adherence, safety, math, code, and agentic tasks, including both in-domain holdouts and out-of-domain evals.&lt;/strong&gt; Commenters were cautiously positive: some see Qwen 3.6 as the strongest cheap open-weight 20B–40B option, while others note users can already control cost via &lt;code&gt;reasoning-budget&lt;/code&gt;. One commenter observed the model appears slightly worse on evals but appreciated that the release is transparent about the tradeoff.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that similar reductions in chain-of-thought verbosity may be achievable at inference time by setting Qwen’s &lt;code&gt;reasoning-budget&lt;/code&gt;, rather than using a separately tuned checkpoint. This raises the technical question of whether ThinkingCap’s gains come from model behavior changes or simply from enforcing a lower token budget during reasoning.&lt;/li&gt;
&lt;li&gt;One commenter observed that the reported evals appear &lt;em&gt;slightly worse&lt;/em&gt; than base Qwen3.6 despite the claimed ~&lt;code&gt;50%&lt;/code&gt; reduction in thinking tokens, but appreciated that the release is transparent about the tradeoff. The practical takeaway is that the model may be worth testing when latency/cost is more important than preserving every benchmark point.&lt;/li&gt;
&lt;li&gt;A GGUF build was linked for local inference: &lt;a href=&quot;https://huggingface.co/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF&quot;&gt;bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF&lt;/a&gt;. This is relevant for users evaluating quantized deployments of the &lt;code&gt;27B&lt;/code&gt; model on llama.cpp-compatible runtimes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Local Model Reliability and Interpretability&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upy31x/i_tested_anthropics_new_jacobian_lens_on_open/&quot;&gt;I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router&lt;/a&gt;&lt;/strong&gt; (Activity: 367): &lt;strong&gt;A Reddit user implemented Anthropic’s &lt;strong&gt;Global Workspace / Jacobian Lens&lt;/strong&gt; idea on open-weight models, releasing code/demo/artifacts at &lt;a href=&quot;https://github.com/solarkyle/jspace&quot;&gt;&lt;code&gt;solarkyle/jspace&lt;/code&gt;&lt;/a&gt;, the &lt;a href=&quot;https://solarkyle.github.io/jspace/demo/&quot;&gt;demo&lt;/a&gt;, and &lt;a href=&quot;https://huggingface.co/solarkyle/jspace-lenses&quot;&gt;HF lenses/traces/routers&lt;/a&gt;. On &lt;code&gt;500&lt;/code&gt; TriviaQA questions/model, Jacobian-lens “workspace trajectory” features—entropy slope, late-band entropy, entropy std, answer rank, layer agreement—outperformed output logprob for wrong-answer prediction on Gemma variants: E4B &lt;code&gt;0.773&lt;/code&gt; vs logprob &lt;code&gt;0.711&lt;/code&gt; AUC, 12B &lt;code&gt;0.824&lt;/code&gt; vs &lt;code&gt;0.736&lt;/code&gt;, 12B abliterated &lt;code&gt;0.799&lt;/code&gt; vs &lt;code&gt;0.731&lt;/code&gt;, 26B MoE &lt;code&gt;0.749&lt;/code&gt; vs &lt;code&gt;0.725&lt;/code&gt;; combining signals improved to &lt;code&gt;0.787–0.843&lt;/code&gt;, while &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; was the counterexample where logprob was already strong (&lt;code&gt;0.856&lt;/code&gt;) and workspace hurt/underperformed (&lt;code&gt;0.646&lt;/code&gt;, combined &lt;code&gt;0.838&lt;/code&gt;). The proposed system is a one-pass local hallucination/risk router: answer locally, take a workspace snapshot, run a tiny logistic-regression sidecar, and escalate to search/citations/cloud if the answer is high-confidence but internally “foggy”; a notable side result was that abliteration greatly increased fake-entity fabrication in Gemma 12B (&lt;code&gt;17/50&lt;/code&gt; → &lt;code&gt;49/50&lt;/code&gt;).&lt;/strong&gt; Commenters debated interpretation: one argued Qwen’s miss is unsurprising because Qwen models appear “overtrained/grokked” and highly pattern-stubborn, making output confidence unusually calibrated on aligned tasks. Another cautioned that the experiment may only show &lt;em&gt;uncertainty ↔ competing latent candidates&lt;/em&gt;, not a reliable implication that competing candidates necessarily mean hallucination, since ambiguity can also reflect legitimate reasoning rather than fabrication.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters questioned the core causal interpretation of the Jacobian Lens signal: the experiment may be detecting &lt;strong&gt;multiple competing latent continuations&lt;/strong&gt; rather than hallucination directly. One commenter argued that uncertainty can naturally increase the number of active candidate ideas, but &lt;em&gt;“competing ideas → hallucination”&lt;/em&gt; does not necessarily follow; this distinction matters for cases where the model has incomplete information yet still makes a well-calibrated guess.&lt;/li&gt;
&lt;li&gt;A detailed repo-level critique argued that the hallucination evaluation is undermined by &lt;strong&gt;incorrect ground-truth labels&lt;/strong&gt;, citing examples where Ross Bagdasarian for &lt;em&gt;The Chipmunks&lt;/em&gt; and H. H. Asquith after Balfour were allegedly marked wrong despite being correct. The same commenter noted that reported AUC/router results become unreliable if the labels are noisy, and also objected to calling the method &lt;em&gt;label-free&lt;/em&gt; because the router is trained via &lt;strong&gt;logistic regression on correct/incorrect answers&lt;/strong&gt;, making it supervised even if the runtime feature is unsupervised.&lt;/li&gt;
&lt;li&gt;The evaluation methodology was criticized for possible &lt;strong&gt;data leakage&lt;/strong&gt;: normalization was reportedly applied to the full dataset before cross-validation, allowing test-fold information into training preprocessing. The baseline was also described as too narrow—mostly a few logprob/output-confidence features—so claims that the router broadly beats confidence calibration were considered overextended, especially given claims that &lt;strong&gt;Qwen&lt;/strong&gt; models are already unusually well-calibrated and “stubborn”/overtrained on familiar task patterns.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uphzhj/qwen_36_27b_absolutely_fails_at_agentic_work/&quot;&gt;Qwen 3.6 27B absolutely fails at agentic work&lt;/a&gt;&lt;/strong&gt; (Activity: 740): &lt;strong&gt;The OP reports that &lt;strong&gt;Qwen 3.6 27B&lt;/strong&gt; at &lt;code&gt;8-bit&lt;/code&gt;/&lt;code&gt;16-bit&lt;/code&gt; under &lt;strong&gt;llama.cpp nightly&lt;/strong&gt; on an &lt;strong&gt;RTX 6000&lt;/strong&gt; performs well on isolated prompts and long-form/demo HTML generation, but repeatedly fails in multi-turn &lt;em&gt;agentic&lt;/em&gt; workflows—&lt;em&gt;“every 4 turns or so it does something completely braindead”&lt;/em&gt;—so they reverted to &lt;strong&gt;Qwen 3.5 122B&lt;/strong&gt; at &lt;code&gt;4-bit&lt;/code&gt;/&lt;code&gt;5-bit&lt;/code&gt;. Technical replies suggest checking chat-template/inference setup, specifically trying &lt;a href=&quot;https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates&quot;&gt;froggeric/Qwen-Fixed-Chat-Templates&lt;/a&gt; for agent-flow bugs and verifying parameters such as &lt;code&gt;preserve_thinking&lt;/code&gt;.&lt;/strong&gt; Commenters were skeptical of the broad claim, arguing that without exact inference parameters, templates, and reproduction details it is hard to diagnose, and that &lt;em&gt;“most people aren&apos;t having your experience.”&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters point to potential &lt;strong&gt;chat-template issues&lt;/strong&gt; as a likely cause of poor Qwen 3.6 27B agentic behavior, recommending froggeric’s patched templates: &lt;a href=&quot;https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates&quot;&gt;Qwen-Fixed-Chat-Templates&lt;/a&gt;. The claim is that these templates “fix some of the bugs for agentic flows,” implying failures may stem from prompt formatting/tool-use serialization rather than the base model itself.&lt;/li&gt;
&lt;li&gt;One technical troubleshooting thread asks whether the user is running the correct inference parameters, specifically mentioning &lt;code&gt;preserve_thinking&lt;/code&gt;. Commenters request the full parameter set, suggesting instability in Qwen 3.6 27B agentic workflows may depend heavily on decoding/configuration and whether reasoning traces are preserved across turns.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. China AI Model Access Policy Debate&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uprmso/beijing_is_looking_at_curbing_overseas_access_to/&quot;&gt;Beijing is looking at curbing overseas access to China&apos;s top AI models (Reuters)&lt;/a&gt;&lt;/strong&gt; (Activity: 1011): &lt;strong&gt;The image is a &lt;strong&gt;Reuters article screenshot&lt;/strong&gt;, not a meme, reporting that &lt;strong&gt;Beijing is considering restrictions on overseas access to China’s leading AI models&lt;/strong&gt; from firms such as &lt;strong&gt;Alibaba, ByteDance, and Z.ai&lt;/strong&gt;, citing national-security concerns and fears of advanced model leakage. Technically, this would affect availability of competitive Chinese frontier/open-weight or API-accessible models outside China, potentially reducing global access to alternatives to U.S. labs; image: &lt;a href=&quot;https://i.redd.it/9s1018gggsbh1.jpeg&quot;&gt;i.redd.it/9s1018gggsbh1.jpeg&lt;/a&gt;.&lt;/strong&gt; Commenters framed this as another AI-access restriction, with concern that competitive local/open models may become harder to obtain. One commenter argued &lt;strong&gt;Mistral&lt;/strong&gt; may become more important as a non-U.S./non-Chinese alternative, especially if its Paris-area datacenter enables training models up to roughly &lt;code&gt;10T&lt;/code&gt; parameters.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter points to &lt;strong&gt;Mistral&lt;/strong&gt; as a potential non-China open-weight alternative, claiming its new datacenter near Paris is expected to come online soon and could enable training models up to roughly &lt;code&gt;10T&lt;/code&gt; parameters. The implication is that European compute capacity may become strategically important if overseas access to Chinese frontier/open models is restricted.&lt;/li&gt;
&lt;li&gt;Several commenters discuss proactively archiving preferred &lt;strong&gt;open-weight models&lt;/strong&gt;, including models they cannot currently run locally, because access restrictions could make downloads or redistribution harder later. This reflects a practical concern around model availability, reproducibility, and long-term local inference workflows if geopolitical controls tighten.&lt;/li&gt;
&lt;li&gt;One technical/business-angle comment argues that &lt;strong&gt;NVIDIA&lt;/strong&gt; may remain one of the few companies with strong incentives to publish open models, because open-weight releases drive demand for local GPU inference and deployment. The broader concern is that model access restrictions could reduce the diversity of competitive local models available to developers.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1upvw37/beijing_is_not_looking_at_curbing_overseas_access/&quot;&gt;Beijing IS NOT looking at curbing overseas access to China&apos;s top AI models (Debunking the Reuters report)&lt;/a&gt;&lt;/strong&gt; (Activity: 966): &lt;strong&gt;The post disputes a &lt;a href=&quot;https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/&quot;&gt;Reuters report&lt;/a&gt; claiming Beijing may curb overseas access to top Chinese AI models, arguing the cited Ministry of Commerce meetings with firms like &lt;strong&gt;Alibaba&lt;/strong&gt;, &lt;strong&gt;ByteDance&lt;/strong&gt;, and &lt;strong&gt;Z.ai&lt;/strong&gt; were instead about &lt;strong&gt;foreign acquisitions, investment, IP leakage, and tech/talent outflow controls&lt;/strong&gt;. It points to a Chinese policy/legal document from the &lt;a href=&quot;https://ipc.court.gov.cn/zh-cn/news/view-5766.html&quot;&gt;China International Commercial Court&lt;/a&gt; as evidence that China’s position is not blanket restriction of open-weight model access but “&lt;strong&gt;trustworthy and controlled&lt;/strong&gt;” open source, including concern that strict cross-border controls on open-source weights could be &lt;em&gt;“self-inflicted”&lt;/em&gt; by reducing Chinese developers’ global participation.&lt;/strong&gt; Commenters were skeptical of the Reuters framing, with some suggesting the sourcing may reflect U.S. AI-lab interests and arguing China has strategic incentive to keep exporting/open-sourcing models because they pressure incumbent U.S. AI companies.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters argued that &lt;strong&gt;open-weight model availability is strategically important for Chinese AI labs&lt;/strong&gt; because it enables global adoption, especially in the US market, and directly competes with closed-model providers like &lt;strong&gt;OpenAI&lt;/strong&gt; and &lt;strong&gt;Anthropic&lt;/strong&gt;. One technical-market point was that restricting overseas access would undermine distribution and ecosystem growth for Chinese models, while maintaining open access could pressure closed API-based competitors ahead of major fundraising or IPO narratives.&lt;/li&gt;
&lt;li&gt;Several commenters framed the Reuters claim as potentially driven by competitive information warfare rather than policy reality, noting that a curb on overseas access would primarily benefit US closed-model vendors by reducing competition from Chinese frontier/open-weight systems. The discussion did not provide benchmarks or implementation details, but focused on model-access strategy and competitive dynamics around global deployment.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Anthropic J-Space Interpretability Research&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1upchq0/anthropic_found_a_global_workspace_inside_claude/&quot;&gt;Anthropic found a “global workspace” inside Claude a silent internal reasoning layer that emerged on its own&lt;/a&gt;&lt;/strong&gt; (Activity: 1267): &lt;strong&gt;&lt;strong&gt;Anthropic&lt;/strong&gt; reports a &lt;code&gt;J-space&lt;/code&gt; in Claude—identified via the open-source &lt;a href=&quot;https://www.github.com/anthropics/jacobian-lens&quot;&gt;Jacobian Lens&lt;/a&gt;—as a compact set of internal activation directions that appears to act like a functional &lt;em&gt;global workspace&lt;/em&gt;: concepts present there are reportable, causally editable, and reused across tasks. In the &lt;a href=&quot;https://www.anthropic.com/research/global-workspace&quot;&gt;paper&lt;/a&gt;, interventions such as swapping &lt;code&gt;spider → ant&lt;/code&gt; or &lt;code&gt;France → China&lt;/code&gt; changed downstream answers across multiple attributes, while ablating J-space reportedly preserved fluency but degraded multi-step reasoning; a highlighted arithmetic trace shows Claude internally progressing through &lt;code&gt;(4+17)*2+7&lt;/code&gt; across layers (&lt;code&gt;21&lt;/code&gt; → &lt;code&gt;42&lt;/code&gt; → &lt;code&gt;49&lt;/code&gt;) without external tools. Anthropic also frames this as a safety signal: J-space activations surfaced latent notions like “fake,” “fictional,” “manipulation,” “fraud,” and “secretly” before output, including in fabrication or deliberately misaligned model-organism settings, while explicitly limiting the consciousness claim to &lt;em&gt;access consciousness&lt;/em&gt; rather than phenomenal experience.&lt;/strong&gt; Technical commenters were broadly impressed, arguing the results provide strong evidence against a simplistic “stochastic parrot” view of LLMs; no substantial methodological debate appeared in the provided top comments.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically substantive comment highlighted Anthropic’s layer-wise interpretability example for Claude solving &lt;code&gt;(4+17)*2+7&lt;/code&gt;: by &lt;code&gt;layer 58&lt;/code&gt; the model represented the task as arithmetic, by &lt;code&gt;layer 75&lt;/code&gt; it had computed &lt;code&gt;4+17=21&lt;/code&gt;, by &lt;code&gt;layer 83&lt;/code&gt; it had derived &lt;code&gt;42&lt;/code&gt;, and by the final layer it reached &lt;code&gt;49&lt;/code&gt;. The commenter emphasized this as evidence of an internal multi-step computation occurring without external tools, consistent with a latent reasoning/workspace-like mechanism rather than purely surface-level token imitation.&lt;/li&gt;
&lt;li&gt;One commenter linked a technical explainer video on the proposed &lt;strong&gt;“J-space” / global workspace&lt;/strong&gt; concept: &lt;a href=&quot;https://m.youtube.com/watch?v=rKV5JcALQoQ&amp;#x26;pp=iggUQAFKEERqTmoxUnozeDY3MHdMdGg%3D&quot;&gt;YouTube explanation&lt;/a&gt;. Another noted that the claim that this space &lt;em&gt;“wasn’t designed; it emerged during training”&lt;/em&gt; aligns with the core premise of machine learning: learned internal representations and behaviors arising from optimization rather than explicit hand-coded structure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uptvgb/anthropic_just_reported_that_llms_have_hidden/&quot;&gt;Anthropic just reported that LLMs have hidden thoughts they hold without saying. An internal ”J-Space”&lt;/a&gt;&lt;/strong&gt; (Activity: 794): &lt;strong&gt;A Redditor summarizes &lt;strong&gt;Anthropic’s&lt;/strong&gt; paper on a proposed internal &lt;code&gt;J-space&lt;/code&gt;/global-workspace-like subspace in LLM activations (&lt;a href=&quot;https://www.anthropic.com/research/global-workspace&quot;&gt;paper&lt;/a&gt;), where a small set of latent variables appears to support reportable, deliberately maintained, multi-step reasoning state while much fluent generation—grammar, style, factual recall—largely bypasses it. They also built &lt;strong&gt;Subtext&lt;/strong&gt; (&lt;a href=&quot;https://github.com/ninjahawk/Subtext&quot;&gt;GitHub&lt;/a&gt;) to visualize token-disposed internal states before generation, citing examples like early saturation of “incorrect” for &lt;code&gt;12 + 5 = 1&lt;/code&gt; and two-hop activation traces such as &lt;code&gt;Italy&lt;/code&gt; at layer &lt;code&gt;20&lt;/code&gt; followed by &lt;code&gt;euros&lt;/code&gt; at layer &lt;code&gt;26&lt;/code&gt; before output begins; they explicitly note this is evidence of functionally available internal information, not subjective experience.&lt;/strong&gt; Comments split between seeing this as expected mechanistic-interpretability evidence against simplistic “stochastic parrot” framing, and skepticism that the visualization may just be ordinary neuron/feature activation—though commenters found the arithmetic and multi-hop timing results more technically interesting.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters focused on Anthropic’s mechanistic-interpretability claim that models can maintain latent internal representations not directly reflected in emitted tokens. One technical interpretation was that this goes beyond simple “stochastic parrot” framing: the model may activate concepts such as an &lt;code&gt;Italy&lt;/code&gt; representation or intermediate arithmetic states before any corresponding output token is produced.&lt;/li&gt;
&lt;li&gt;A key technical point highlighted was the distinction between base training and post-training: in base models, internal state was described as primarily optimized for next-token prediction, while post-training appears to induce a more persistent “identity” or first-person framing. One commenter noted that the model could internally classify input as a prompt injection &lt;em&gt;while reading it&lt;/em&gt;, before producing any output, implying latent evaluation of user text independent of immediate token generation.&lt;/li&gt;
&lt;li&gt;There was interest in reproducibility and implementation details, including a user reportedly reimplementing parts of the paper’s experiments and questions about which model generated the reproduction code. The arithmetic examples were called out as especially interesting because they suggest intermediate computational structure rather than merely concept-neuron activation on the path to output.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude Code and Autonomous Coding Agents&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1upcvot/the_tool_that_now_generates_25byear_started_as_a/&quot;&gt;The tool that now generates $2.5B/year started as a guy’s first-week side project at his new job&lt;/a&gt;&lt;/strong&gt; (Activity: 1034): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/dbt5khhguobh1.jpeg&quot;&gt;image&lt;/a&gt; is a &lt;strong&gt;non-technical branding graphic&lt;/strong&gt;: a retro/pixel-styled “CLAUDE CODE” logo on a black background, used to visually frame the post’s story about Anthropic’s Claude Code CLI. The post claims Claude Code began as &lt;strong&gt;Boris Cherny’s first-week prototype&lt;/strong&gt; at Anthropic, grew rapidly after gaining filesystem access, reached &lt;code&gt;80%+&lt;/code&gt; internal daily usage by May 2025, and allegedly hit &lt;code&gt;$1B&lt;/code&gt; ARR within ~6 months—though the title’s &lt;code&gt;$2.5B/year&lt;/code&gt; figure is not substantiated in the provided text.&lt;/strong&gt; Comments push back on the “accidental discovery” framing, noting that coding agents such as &lt;strong&gt;Cline&lt;/strong&gt; already existed months earlier, making Claude Code look more like an in-house/productized version of an existing pattern than a novel research breakthrough. Other comments treat the image as purely aesthetic, saying the arcade-style logo matches the “fever dream” origin story.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter challenged the post’s framing that the project was novel, arguing that the core idea—giving an LLM agent access to a local filesystem for coding tasks—already existed in tools like &lt;strong&gt;Cline&lt;/strong&gt; for &lt;em&gt;“6+ months”&lt;/em&gt; beforehand. They characterized it as an &lt;strong&gt;in-house clone of existing coding-agent products&lt;/strong&gt;, not a research breakthrough.&lt;/li&gt;
&lt;li&gt;Another technical point focused on the claim that &lt;strong&gt;Claude writes over &lt;code&gt;80%&lt;/code&gt; of Anthropic’s own code&lt;/strong&gt;, with speculation that the Claude desktop app may itself have been heavily generated or “vibe coded.” The comment implies interest in how much production code at Anthropic is now authored or scaffolded by Claude-based coding agents.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ChatGPT/comments/1upb4vw/i_gave_gpt_55_an_empty_github_repo_and_told_it_to/&quot;&gt;I gave GPT 5.5 an empty GitHub repo and told it to figure its life out&lt;/a&gt;&lt;/strong&gt; (Activity: 795): &lt;strong&gt;The experiment schedules an LLM “agent” to wake hourly (later doubled in frequency) against an initially empty public GitHub repo, inspect prior state, choose work, write/test code, and commit; its first outputs were meta-project artifacts (roadmap/changelog/state/decision log) rather than application code. The repo, &lt;a href=&quot;https://github.com/OmarH-creator/Autonomous-Forge&quot;&gt;&lt;strong&gt;Autonomous Forge&lt;/strong&gt;&lt;/a&gt;, is currently a pre-alpha, local-first Python CLI for “repository-native autonomous software-improvement loops,” but its implemented scope is mostly deterministic, read-only planning/review: task selection, policy-aware planning, proposal/validation previews, repo inventory, preflight readiness, and run-history previewing, with only an explicitly confirmed &lt;code&gt;run-history-write&lt;/code&gt; mutating &lt;code&gt;.ai/run-history/&lt;/code&gt;. Its stated safety boundary is conservative: no network calls, test/validation execution, diff inspection, patch generation, commits, pushes, or policy enforcement until future roadmap/policy support.&lt;/strong&gt; Commenters mostly highlighted the recursion/meta-design: an autonomous agent asked to build something chose to build a tool for autonomous repository maintenance, with one calling it a “tool that has no defined goals.” The main technical criticism was that the generated roadmap appears over-indexed on planning/process artifacts rather than concrete implementation progress.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters noted that the generated project concept, &lt;strong&gt;“Autonomous Forge”&lt;/strong&gt;, is effectively a meta-tool: an AI-created developer tool intended to run “repository-native autonomous software-improvement loops,” meaning the model responded to an empty repo by designing infrastructure for autonomous code generation/maintenance rather than solving a concrete product problem. One technical criticism was that the roadmap appeared overly weighted toward &lt;strong&gt;task planning/orchestration&lt;/strong&gt; rather than implementation, evaluation, or measurable developer-tool functionality.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Claude Fable 5 Access and Guardrail Friction&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uq2aq5/anthropic_extending_fable_5_for_paid_users_till/&quot;&gt;Anthropic extending Fable 5 for paid users till 12 july&lt;/a&gt;&lt;/strong&gt; (Activity: 1320): &lt;strong&gt;The image is a &lt;strong&gt;non-meme screenshot&lt;/strong&gt; of a Claude/Anthropic X post announcing that &lt;strong&gt;“Claude Fable 5” access is extended for paid users through &lt;code&gt;July 12&lt;/code&gt;&lt;/strong&gt; across all paid plans: &lt;a href=&quot;https://i.redd.it/t1hhakidhubh1.jpeg&quot;&gt;image&lt;/a&gt;. The follow-up clarifies a quota policy: paid users can spend up to &lt;strong&gt;&lt;code&gt;50%&lt;/code&gt; of their weekly usage limit&lt;/strong&gt; on Fable 5, then must either use extra credits or switch to another Claude model.&lt;/strong&gt; Comments mainly criticize the short-notice extension and quota planning impact: users say they rushed to exhaust weekly Fable usage or bought extra credits because they expected access to end sooner. One recurring request is for Anthropic to provide a usage-limit reset alongside the extension.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1upu3e2/fable_5_found_actual_malware_on_my_pc_and_then/&quot;&gt;Fable 5 found actual malware on my PC, and then its own safety filters flagged the warning.&lt;/a&gt;&lt;/strong&gt; (Activity: 1292): &lt;strong&gt;A user reports &lt;strong&gt;Fable 5&lt;/strong&gt; inspecting the Windows &lt;code&gt;Run&lt;/code&gt; registry key and detecting a suspicious PowerShell persistence command—&lt;code&gt;powershell.exe -NoProfile -ExecutionPolicy Bypass -WindowStyle Hidden ...&lt;/code&gt;—that allegedly downloaded and executed a remote script at sign-in, which the model classified as an active compromise (&lt;a href=&quot;https://preview.redd.it/2hv0yord1tbh1.png?width=1172&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=6434cb2d41bb2474fa40398c24fa575b1f74c635&quot;&gt;screenshot&lt;/a&gt;). After the user instructed it to remove the specific registry persistence entries, the cleanup reportedly succeeded, but the session was then flagged for “cybersecurity work” and downgraded to &lt;strong&gt;Opus 4.8&lt;/strong&gt; by safety filters (&lt;a href=&quot;https://preview.redd.it/402jnkkf1tbh1.png?width=1163&amp;#x26;format=png&amp;#x26;auto=webp&amp;#x26;s=7c6531edf9e593fd592109f91bd0ca45de8e6650&quot;&gt;screenshot&lt;/a&gt;).&lt;/strong&gt; Commenters argued this is a poor substitute for endpoint security: PowerShell &lt;code&gt;Run&lt;/code&gt;-key persistence is a long-known malware pattern that conventional AV/EDR tools should catch more reliably, and an LLM may remove one indicator while missing others. One commenter noted a related positive case where an AI scan of a codebase surfaced production-relevant security findings from &lt;code&gt;security.md&lt;/code&gt; without triggering a downgrade.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter argued that the malware class is not novel—&lt;em&gt;“in place for maybe 12 years”&lt;/em&gt;—and would likely be detected by conventional antivirus signatures/heuristics. Their technical takeaway was that &lt;strong&gt;Fable 5 should not replace dedicated endpoint security tooling&lt;/strong&gt;, because an LLM-style scan may find one indicator while missing related persistence mechanisms or additional malware artifacts.&lt;/li&gt;
&lt;li&gt;One user reported using &lt;strong&gt;Fable&lt;/strong&gt; to scan a codebase, where it parsed their &lt;code&gt;security.md&lt;/code&gt; and added multiple new security findings that they considered production-relevant. They subsequently patched the issues, suggesting the model was useful for surfacing actionable application-security problems from project documentation/context rather than just source-code linting.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1updedl/well_shit_i_didnt_even_know_this_was_possible/&quot;&gt;Well shit... I didn&apos;t even know this was possible&lt;/a&gt;&lt;/strong&gt; (Activity: 688): &lt;strong&gt;The image is a &lt;strong&gt;Claude/Anthropic billing “Usage credits” screen&lt;/strong&gt; showing an apparent spend-limit failure: despite a configured &lt;code&gt;$50&lt;/code&gt; monthly spend limit, the account shows &lt;code&gt;$155.53&lt;/code&gt; spent (&lt;code&gt;311%&lt;/code&gt; used) and a negative balance of &lt;code&gt;-$119.11&lt;/code&gt; (&lt;a href=&quot;https://i.redd.it/l04wd5dfxobh1.png&quot;&gt;image&lt;/a&gt;). In context, the poster says they ran &lt;strong&gt;Fable&lt;/strong&gt; on several tasks expecting usage to stop at the cap, but Claude continued billing beyond the limit, raising a practical issue around whether Anthropic spend limits are hard enforcement caps or delayed/soft accounting controls.&lt;/strong&gt; Commenters were skeptical of Anthropic support and suggested a credit-card chargeback if billed, while noting it is “weird” that a configured monthly spend limit appears to have been ignored.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Users report a potential &lt;strong&gt;Anthropic/Claude billing control bug&lt;/strong&gt; where a configured monthly spend limit was allegedly ignored, allowing continued model usage and additional charges after the expected cap. One commenter contrasts this with plan quota behavior, noting that regular plan usage stops mid-task when exhausted, while paid overage/API-like usage may continue accruing cost.&lt;/li&gt;
&lt;li&gt;A suggested remediation path is to contact &lt;code&gt;support@anthropic.com&lt;/code&gt; and frame the issue as a &lt;strong&gt;configuration/billing bug&lt;/strong&gt;, including screenshots of the spend-limit setting. Commenters recommend requesting a prorated refund first, with a credit-card chargeback as a fallback, though that may risk account termination or needing a new Claude account.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>anthropic</category><category>langchain</category><category>google</category><category>meta-ai-fair</category><category>nvidia</category><category>cohere</category><category>weaviate</category><category>claude-fable-5</category><category>muse-image</category><category>muse-video</category><category>audex</category><category>mikeyk</category><category>kimmonismus</category><category>lilian_weng</category><category>sakana</category><category>_philschmid</category><category>officiallogank</category><category>dimillian</category><category>reach_vb</category><category>teknuim</category><category>victorialslocum</category><category>omarsar0</category><category>alexandr_wang</category><category>_tim_brooks</category><category>agent-design</category><category>background-execution</category><category>task-management</category><category>human-in-the-loop</category><category>agentic-generation</category><category>reinforcement-learning</category><category>model-scaling</category><category>moe</category><category>context-windows</category><category>audio-processing</category><category>video-generation</category><category>image-generation</category><category>open-source</category><category>model-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-06-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-06-not-much/</guid><description>**Tencent** released **Hy3**, a **295B MoE** open-weight model with **21B active parameters**, **192 experts**, and **256K context** supporting **MTP speculative decoding**. It runs natively on **vLLM** with optimizations for **NVIDIA** and **AMD** hardware, achieving up to **2.95x** speedups and latency reductions. Hy3 competes closely with **GLM-5.2** in the open model space. **AutomationBench-AA** leaderboard evaluates agents on **657 tasks** across **40 SaaS apps**, with **Claude Fable 5** leading, followed by **Opus 4.8**, **Gemini 3.5 Flash**, and **GPT-5.5 xhigh**. Open models lag behind, with **GLM-5.2 max** best at **27.8%**. New domain-specific capability indices highlight cost-performance tradeoffs. Research on persistent agent memory includes **A-TMA** improving conflict accuracy and **ReContext** enhancing long-context inference without retraining.</description><pubDate>Mon, 06 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/04/2026-7/06/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Tencent Hunyuan’s Hy3 Release and the Open-Weight Frontier&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Hy3 lands as a serious open model&lt;/strong&gt;: Tencent released &lt;strong&gt;Hy3&lt;/strong&gt; under &lt;strong&gt;Apache 2.0&lt;/strong&gt;, a &lt;strong&gt;295B MoE&lt;/strong&gt; with &lt;strong&gt;21B active parameters&lt;/strong&gt;, &lt;strong&gt;192 experts / top-8 routing&lt;/strong&gt;, &lt;strong&gt;GQA&lt;/strong&gt;, &lt;strong&gt;256K context&lt;/strong&gt;, and a &lt;strong&gt;3.8B MTP layer&lt;/strong&gt; for speculative decoding. Multiple posts framed it as competitive with much larger systems on reasoning, coding, and agentic tasks, with particular emphasis on reliability improvements like tool-calling stability and anti-hallucination work &lt;a href=&quot;https://x.com/eliebakouch/status/2074011171661701466&quot;&gt;@eliebakouch&lt;/a&gt;, &lt;a href=&quot;https://x.com/HuggingPapers/status/2074024501201813797&quot;&gt;@HuggingPapers&lt;/a&gt;, &lt;a href=&quot;https://x.com/ShunyuYao12/status/2074151389945827744&quot;&gt;@ShunyuYao12&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference support was unusually day-0 mature&lt;/strong&gt;: &lt;a href=&quot;https://x.com/vllm_project/status/2074147504254517529&quot;&gt;@vllm_project&lt;/a&gt; said Hy3 runs natively in &lt;strong&gt;vLLM&lt;/strong&gt; from launch with tool-call and reasoning parsers, &lt;strong&gt;MTP speculative decoding&lt;/strong&gt;, and validated support on &lt;strong&gt;NVIDIA and AMD&lt;/strong&gt;. A follow-up detailed Tencent production kernels now upstreamed into vLLM main, including load-balanced decode scheduling and fused FP8 MoE serving, with reported gains of &lt;strong&gt;up to 2.95x&lt;/strong&gt; on mixed-length decode and latency reductions of roughly &lt;strong&gt;24% TTFT&lt;/strong&gt; and &lt;strong&gt;17% TPOT&lt;/strong&gt; versus default backends &lt;a href=&quot;https://x.com/vllm_project/status/2074147506875969754&quot;&gt;@vllm_project&lt;/a&gt;. Community reaction was strong enough that &lt;a href=&quot;https://x.com/Teknium/status/2074264567803531589&quot;&gt;@Teknium&lt;/a&gt; quickly made Hy3 free on Nous Portal for two weeks.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Broader open-model context&lt;/strong&gt;: Hy3 was immediately compared against &lt;strong&gt;GLM-5.2&lt;/strong&gt;, with some posters arguing Tencent has now joined the very top tier of open-source labs if the benchmark and vibe-test results hold &lt;a href=&quot;https://x.com/teortaxesTex/status/2074012467886178725&quot;&gt;@teortaxesTex&lt;/a&gt;, while others still maintained &lt;strong&gt;GLM-5.2&lt;/strong&gt; as the best currently usable open-weight model in practice &lt;a href=&quot;https://x.com/__tinygrad__/status/2074206866641752190&quot;&gt;@&lt;strong&gt;tinygrad&lt;/strong&gt;&lt;/a&gt;, &lt;a href=&quot;https://x.com/mbusigin/status/2074238100251799998&quot;&gt;@mbusigin&lt;/a&gt;. The net takeaway: the open frontier is compressing fast, and the competition is increasingly about deployment robustness rather than just raw leaderboard deltas.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Agent Benchmarks, Harnesses, and Long-Running Memory&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AutomationBench-AA adds a more realistic agent eval&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074194764510208230&quot;&gt;@ArtificialAnlys&lt;/a&gt; launched an independent leaderboard for Zapier’s &lt;strong&gt;AutomationBench&lt;/strong&gt;, evaluating agents across &lt;strong&gt;657 tasks&lt;/strong&gt; and &lt;strong&gt;40 simulated SaaS apps&lt;/strong&gt; with both objectives and guardrails. &lt;strong&gt;Claude Fable 5&lt;/strong&gt; led at &lt;strong&gt;48.6%&lt;/strong&gt;, narrowly ahead of &lt;strong&gt;Opus 4.8&lt;/strong&gt; at &lt;strong&gt;48.5%&lt;/strong&gt;, with &lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt; at &lt;strong&gt;42.6%&lt;/strong&gt; and &lt;strong&gt;GPT-5.5 xhigh&lt;/strong&gt; at &lt;strong&gt;42.1%&lt;/strong&gt;. More interesting than the ranking: every model still breaks business rules, and Gemini looked notably strong on &lt;strong&gt;objective-per-guardrail-violation&lt;/strong&gt; and &lt;strong&gt;cost efficiency&lt;/strong&gt;. Open weights remain meaningfully behind, with &lt;strong&gt;GLM-5.2 max&lt;/strong&gt; the best listed open model at &lt;strong&gt;27.8%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Capability indices are becoming multidimensional&lt;/strong&gt;: Artificial Analysis also introduced six domain-specific indices—&lt;strong&gt;Finance &amp;#x26; Accounting, Legal, Healthcare &amp;#x26; Medical, Strategy &amp;#x26; Ops, Engineering, Economics&lt;/strong&gt;—to move past single scalar model scores &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074299714699469221&quot;&gt;@ArtificialAnlys&lt;/a&gt;. The headline was familiar—&lt;strong&gt;Claude Fable 5&lt;/strong&gt; plus &lt;strong&gt;Opus 4.8 fallback&lt;/strong&gt; leads—but the more useful insight is how sharply rankings reshuffle by domain and how steep the price/performance frontier has become. This aligns with &lt;a href=&quot;https://x.com/fchollet/status/2074242671103889799&quot;&gt;@fchollet&lt;/a&gt;, who argued that reporting benchmark scores without &lt;strong&gt;cost per task&lt;/strong&gt; is increasingly meaningless.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory and retrieval remain bottlenecks for persistent agents&lt;/strong&gt;: Two papers got traction here. First, &lt;strong&gt;A-TMA&lt;/strong&gt; tackles “ghost memory,” where stale and current facts are retrieved together in long-running assistants; on the LTP benchmark, adding it to Graphiti reportedly improves conflict accuracy by &lt;strong&gt;+0.240 absolute&lt;/strong&gt; &lt;a href=&quot;https://x.com/omarsar0/status/2074121191846261022&quot;&gt;@omarsar0&lt;/a&gt;. Second, &lt;strong&gt;ReContext&lt;/strong&gt; is a training-free long-context inference harness that replays model-internal evidence right before answer generation, improving evidence utilization across eight 128K datasets &lt;a href=&quot;https://x.com/dair_ai/status/2074178316819677238&quot;&gt;@dair_ai&lt;/a&gt;. Combined with &lt;strong&gt;BlockSearch&lt;/strong&gt; for million-token in-context retrieval &lt;a href=&quot;https://x.com/dair_ai/status/2074117920133898707&quot;&gt;@dair_ai&lt;/a&gt;, the theme is clear: better memory behavior is increasingly being engineered at inference time, not just trained in.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Anthropic’s J-Space / Global Workspace Results&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Mechanistic interpretability took center stage&lt;/strong&gt;: Anthropic released research claiming a &lt;strong&gt;global-workspace-like internal structure&lt;/strong&gt; in Claude, centered on a small subset of activations they call &lt;strong&gt;J-space&lt;/strong&gt; &lt;a href=&quot;https://x.com/AnthropicAI/status/2074185348142280912&quot;&gt;@AnthropicAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/AnthropicAI/status/2074185387577094398&quot;&gt;@AnthropicAI&lt;/a&gt;. The core claim is not chain-of-thought extraction, but identification of a privileged internal representational substrate that appears available for report, modulation, and flexible reasoning. Anthropic also shipped a Neuronpedia demo for open-weight models &lt;a href=&quot;https://x.com/AnthropicAI/status/2074185390060110138&quot;&gt;@AnthropicAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Why researchers cared&lt;/strong&gt;: Interpretability researchers treated this as stronger evidence for a model “working memory” or internal workspace than prior public work, even if they disagreed with the framing. &lt;a href=&quot;https://x.com/NeelNanda5/status/2074193936588148891&quot;&gt;@NeelNanda5&lt;/a&gt; called it the best evidence yet for a working-memory-like mechanism. &lt;a href=&quot;https://x.com/Jack_W_Lindsey/status/2074215950602379388&quot;&gt;@Jack_W_Lindsey&lt;/a&gt; argued understanding this privileged space could be key to LLM cognition. Posts also highlighted practical safety angles: the workspace can reportedly surface hidden concepts, detect prompt injections, and expose internal sabotage-related features before they are verbalized &lt;a href=&quot;https://x.com/mlpowered/status/2074190714100146483&quot;&gt;@mlpowered&lt;/a&gt;, &lt;a href=&quot;https://x.com/LiorOnAI/status/2074198891990548940&quot;&gt;@LiorOnAI&lt;/a&gt;, &lt;a href=&quot;https://x.com/omarsar0/status/2074264122330612223&quot;&gt;@omarsar0&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;But the “consciousness” language was contested&lt;/strong&gt;: Anthropic’s public framing invited strong pushback. Supporters said the results suggest a functional analog of &lt;strong&gt;access consciousness&lt;/strong&gt; rather than phenomenal consciousness &lt;a href=&quot;https://x.com/BorisMPower/status/2074201312531734567&quot;&gt;@BorisMPower&lt;/a&gt;, while critics argued the company was overclaiming by conflating privileged latent activation with consciousness &lt;a href=&quot;https://x.com/AlanCowen/status/2074265992570736919&quot;&gt;@AlanCowen&lt;/a&gt;. Even some sympathetic takes emphasized the bigger story is a new &lt;strong&gt;intervention point&lt;/strong&gt; for auditing and steering models, not philosophy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Inference, Serving, and Systems Efficiency&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Speculative decoding remains hot infrastructure&lt;/strong&gt;: &lt;a href=&quot;https://x.com/lmsysorg/status/2074176669108367549&quot;&gt;@lmsysorg&lt;/a&gt; added &lt;strong&gt;DSpark&lt;/strong&gt; to SGLang for confidence-driven, variable-length verification. The pitch is that under high load it avoids verifying every draft token, improving the throughput/latency tradeoff relative to fixed-budget speculative methods; DeepSeek-V4-Pro reportedly reached &lt;strong&gt;383.7 tok/s at batch=1 on B300&lt;/strong&gt;. Microsoft also discussed prompt-level optimization of &lt;strong&gt;GPT-5.5&lt;/strong&gt; in the GitHub Copilot harness to improve latency and token efficiency after launch &lt;a href=&quot;https://x.com/code/status/2074178799512539571&quot;&gt;@code&lt;/a&gt;, &lt;a href=&quot;https://x.com/pierceboggan/status/2074180737147027757&quot;&gt;@pierceboggan&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Inference efficiency is increasingly the strategic bottleneck&lt;/strong&gt;: &lt;a href=&quot;https://x.com/jon_durbin/status/2074169183835685351&quot;&gt;@jon_durbin&lt;/a&gt; argued that inference, not training alone, is now “the whole game,” because every data pipeline, RL loop, and agent runtime ultimately cashes out as test-time compute. That perspective also showed up in lower-level kernel work: Chutes reported major speedups for &lt;strong&gt;MiniMax MSA&lt;/strong&gt; and &lt;strong&gt;GatedDeltaNet-2&lt;/strong&gt;, including &lt;strong&gt;~7x&lt;/strong&gt; sparse-attention training improvements on &lt;strong&gt;RTX Pro 6000 / SM120&lt;/strong&gt; and better fused FP8 kernels &lt;a href=&quot;https://x.com/jon_durbin/status/2074119835366134188&quot;&gt;@jon_durbin&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Infra releases beyond model serving&lt;/strong&gt;: Cloudflare launched &lt;strong&gt;Workers Cache&lt;/strong&gt;, a regionally tiered cache in front of Worker entrypoints configured via standard HTTP headers &lt;a href=&quot;https://x.com/Cloudflare/status/2074117419728007181&quot;&gt;@Cloudflare&lt;/a&gt;. OpenAI shipped &lt;strong&gt;GPT-Realtime-2.1-mini&lt;/strong&gt;, bringing reasoning and tool use to the mini realtime line at the same price as the prior mini, alongside claimed &lt;strong&gt;25%+ p95 latency reductions&lt;/strong&gt; from caching improvements &lt;a href=&quot;https://x.com/OpenAIDevs/status/2074255408013955466&quot;&gt;@OpenAIDevs&lt;/a&gt;, &lt;a href=&quot;https://x.com/OpenAIDevs/status/2074255420831735824&quot;&gt;@OpenAIDevs&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;World Models, Speech, and Document AI&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;MIRA is a notable world-model demo&lt;/strong&gt;: General Intuition and Kyutai, with Epic Games, introduced &lt;strong&gt;MIRA&lt;/strong&gt;, a playable multiplayer world model for Rocket League trained on &lt;strong&gt;10k hours&lt;/strong&gt; of bot-collected data &lt;a href=&quot;https://x.com/gen_intuition/status/2074104524596457706&quot;&gt;@gen_intuition&lt;/a&gt;. It runs in real time at &lt;strong&gt;20 fps&lt;/strong&gt;, and posts highlighted a &lt;strong&gt;5B-parameter&lt;/strong&gt; model running an entire 2v2 match on a single &lt;strong&gt;NVIDIA B200&lt;/strong&gt;, with no explicit physics or rendering engine &lt;a href=&quot;https://x.com/TheRundownAI/status/2074184559768277398&quot;&gt;@TheRundownAI&lt;/a&gt;. This was one of the clearest signals that video/world-model work is moving from toy demos toward interactive simulators.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Speech remains highly competitive&lt;/strong&gt;: AssemblyAI released &lt;strong&gt;Universal-3.5 Pro Realtime&lt;/strong&gt;, a streaming STT model with &lt;strong&gt;4.1% WER&lt;/strong&gt; on AA-WER Streaming and contextual priming that can be updated mid-call without reconnecting &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074160133702402314&quot;&gt;@ArtificialAnlys&lt;/a&gt;. On the TTS side, Artificial Analysis said &lt;strong&gt;Speechify Simba 3.2&lt;/strong&gt; now leads its Speech Arena at &lt;strong&gt;1233 Elo&lt;/strong&gt;, ahead of Gemini 3.1 Flash TTS, Sonic 3.5, and Inworld Realtime TTS 1.5 Max, while also being the cheapest among top-ranked models &lt;a href=&quot;https://x.com/ArtificialAnlys/status/2074265309985570890&quot;&gt;@ArtificialAnlys&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Document-context pipelines are becoming multimodal by default&lt;/strong&gt;: LlamaIndex and LanceDB described a retrieval pipeline for messy PDFs that separates &lt;strong&gt;pages, chunks, and extracted assets&lt;/strong&gt; into linked multimodal tables, reporting &lt;strong&gt;82% any-page-hit@5&lt;/strong&gt; and &lt;strong&gt;74% answer accuracy&lt;/strong&gt; on a labeled ESG-report benchmark &lt;a href=&quot;https://x.com/lancedb/status/2074153945631457663&quot;&gt;@lancedb&lt;/a&gt;, &lt;a href=&quot;https://x.com/llama_index/status/2074170470119752084&quot;&gt;@llama_index&lt;/a&gt;. This pairs with Jerry Liu’s broader argument for a dedicated “document context layer” for agents &lt;a href=&quot;https://x.com/jerryjliu0/status/2074165277634253106&quot;&gt;@jerryjliu0&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Anthropic’s global workspace paper&lt;/strong&gt; dominated engagement, with the primary announcement on Claude’s internal workspace/J-space far above everything else &lt;a href=&quot;https://x.com/AnthropicAI/status/2074185348142280912&quot;&gt;@AnthropicAI&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tencent Hy3&lt;/strong&gt; was the biggest pure model-release story, especially among technical accounts discussing open-source competitiveness and deployment &lt;a href=&quot;https://x.com/teortaxesTex/status/2074012467886178725&quot;&gt;@teortaxesTex&lt;/a&gt;, &lt;a href=&quot;https://x.com/ShunyuYao12/status/2074151389945827744&quot;&gt;@ShunyuYao12&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MIRA’s playable world model&lt;/strong&gt; was the standout multimodal/system demo &lt;a href=&quot;https://x.com/gen_intuition/status/2074104524596457706&quot;&gt;@gen_intuition&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Will Depue’s “Stargate for Data”&lt;/strong&gt; thread was the most substantive strategy post, arguing that data collection—not compute alone—becomes the binding constraint and potential moat for frontier labs &lt;a href=&quot;https://x.com/willdepue/status/2074178395462848800&quot;&gt;@willdepue&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;John Carmack’s memory-system thread&lt;/strong&gt; drew significant technical interest by arguing inference hardware could exploit deterministic access patterns and much cheaper memory tiers than HBM for large-model serving &lt;a href=&quot;https://x.com/ID_AA_Carmack/status/2074248758422864226&quot;&gt;@ID_AA_Carmack&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. Large Open-Weight MoE Model Releases&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1unyvnz/longcat_20_16t_48b_active_weights_are_now_open/&quot;&gt;longcat 2.0 (1.6T, ~48B active) weights are now open under MIT license&lt;/a&gt;&lt;/strong&gt; (Activity: 638): &lt;strong&gt;&lt;strong&gt;LongCat 2.0&lt;/strong&gt; weights are now open under the &lt;strong&gt;MIT license&lt;/strong&gt; via announcements from &lt;a href=&quot;https://x.com/eliebakouch/status/2073690402503487902&quot;&gt;elie&lt;/a&gt; and &lt;a href=&quot;https://x.com/ModelScope2022/status/2073710226365165679&quot;&gt;ModelScope&lt;/a&gt;, with technical details in the &lt;a href=&quot;https://longcat.chat/blog/longcat-2.0/&quot;&gt;LongCat 2.0 blog post&lt;/a&gt;. The model is a very large &lt;strong&gt;MoE&lt;/strong&gt; system with &lt;code&gt;1.6T&lt;/code&gt; total parameters and roughly &lt;code&gt;48B&lt;/code&gt; active parameters per inference; commenters note the released weights occupy about &lt;code&gt;3.55 TB&lt;/code&gt; in &lt;strong&gt;BF16&lt;/strong&gt; and &lt;code&gt;2.05 TB&lt;/code&gt; in &lt;strong&gt;FP8&lt;/strong&gt;.&lt;/strong&gt; Commenters emphasized the practical deployment burden from the multi-terabyte weight size, and noted that &lt;strong&gt;Meituan&lt;/strong&gt;—described as China’s Groupon/Uber Eats analogue—reportedly trained it on fully domestic Chinese chips, prompting discussion about the geopolitical/market significance.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlighted the scale and deployment footprint of &lt;strong&gt;LongCat 2.0&lt;/strong&gt;: &lt;code&gt;1.6T&lt;/code&gt; total parameters with approximately &lt;code&gt;48B&lt;/code&gt; active parameters, implying a sparse/MoE-style architecture. One user noted the released weights require about &lt;code&gt;3.55 TB&lt;/code&gt; in &lt;strong&gt;BF16&lt;/strong&gt; and &lt;code&gt;2.05 TB&lt;/code&gt; in &lt;strong&gt;FP8&lt;/strong&gt;, which is important for anyone planning local storage or inference infrastructure.&lt;/li&gt;
&lt;li&gt;A technical point raised was that &lt;strong&gt;Meituan&lt;/strong&gt; reportedly trained the model on &lt;code&gt;100%&lt;/code&gt; domestic Chinese chips, which commenters framed as significant for AI hardware supply-chain independence. This is especially notable given Meituan’s role as a major Chinese internet company comparable to a mix of Groupon and Uber Eats rather than a traditional AI lab.&lt;/li&gt;
&lt;li&gt;Several users focused on the permissive &lt;strong&gt;MIT license&lt;/strong&gt; and planned benchmarking against frontier open models such as &lt;strong&gt;Qwen&lt;/strong&gt; and &lt;strong&gt;DeepSeek&lt;/strong&gt;. The combination of &lt;code&gt;1.6T&lt;/code&gt; total parameters, only &lt;code&gt;~48B&lt;/code&gt; active parameters, and open weights suggests the model may be practical to compare with other high-end MoE open models if inference tooling supports its architecture efficiently.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uoozt4/new_open_model_from_tencent_hy_hy3_295b_total_21b/&quot;&gt;New open  model from Tencent Hy: Hy3 (295B total 21B active - apache 2.0)&lt;/a&gt;&lt;/strong&gt; (Activity: 604): &lt;strong&gt;&lt;strong&gt;Tencent&lt;/strong&gt; released the non-preview &lt;strong&gt;Hy3&lt;/strong&gt; model collection on &lt;a href=&quot;https://huggingface.co/collections/tencent/hy3&quot;&gt;Hugging Face&lt;/a&gt;, described as a &lt;code&gt;295B&lt;/code&gt;-parameter MoE with &lt;code&gt;21B&lt;/code&gt; active parameters, now under &lt;strong&gt;Apache 2.0&lt;/strong&gt; rather than the prior restrictive community license. Commenters highlight a linked benchmark chart and claim the release shows &lt;em&gt;“pretty impressive claimed gains over HY3-Preview”&lt;/em&gt;, potentially making it relevant for high-end local/home inference setups if real-world performance matches reported results.&lt;/strong&gt; The main discussion is positive around the license change: commenters view moving from a geographically/restrictively limited license to &lt;strong&gt;Apache 2.0&lt;/strong&gt; as the most important improvement, especially alongside Tencent’s recent Apache-licensed translation models.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters highlight that &lt;strong&gt;Tencent Hunyuan Hy3&lt;/strong&gt; is a large MoE-style release at &lt;strong&gt;&lt;code&gt;295B&lt;/code&gt; total parameters with &lt;code&gt;21B&lt;/code&gt; active&lt;/strong&gt;, and one user notes the claimed benchmark gains over &lt;strong&gt;HY3-Preview&lt;/strong&gt; appear substantial enough that, if they transfer to real workloads, it could be relevant for “high end home setups.” There is interest in practical inference availability, especially &lt;strong&gt;GGUF quantizations&lt;/strong&gt;, which would determine whether local deployment is feasible.&lt;/li&gt;
&lt;li&gt;A technically important licensing change was noted: Tencent reportedly moved from a more restrictive “community” license that limited usage in regions such as &lt;strong&gt;South Korea, the UK, and the EU&lt;/strong&gt; to &lt;strong&gt;Apache 2.0&lt;/strong&gt;. Commenters view this as significant because it enables broader commercial and research reuse, consistent with some of Tencent’s recent Apache-licensed translation models.&lt;/li&gt;
&lt;li&gt;One commenter frames Hy3 as a potential alternative to &lt;strong&gt;Qwen&lt;/strong&gt; and &lt;strong&gt;MiniMax&lt;/strong&gt;, implying interest in whether its benchmark and real-world performance can compete with the current leading open-weight Chinese model families.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uotkm7/new_model_gigachat35432ba28b_with_day0_gguf/&quot;&gt;New model: GigaChat3.5-432B-A28B (with day-0 GGUF support!)&lt;/a&gt;&lt;/strong&gt; (Activity: 439): &lt;strong&gt;&lt;strong&gt;Sberbank/ai-sage&lt;/strong&gt; released &lt;strong&gt;GigaChat3.5-432B-A28B&lt;/strong&gt; on Hugging Face in both &lt;a href=&quot;https://huggingface.co/ai-sage/GigaChat3.5-432B-A28B&quot;&gt;instruct&lt;/a&gt; and &lt;a href=&quot;https://huggingface.co/ai-sage/GigaChat3.5-432B-A28B-base&quot;&gt;base&lt;/a&gt; variants, plus day-0 &lt;a href=&quot;https://huggingface.co/ai-sage/GigaChat3.5-432B-A28B-GGUF&quot;&gt;GGUF weights&lt;/a&gt;; &lt;code&gt;llama.cpp&lt;/code&gt; support is available via PR &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/pull/25342&quot;&gt;ggml-org/llama.cpp#25342&lt;/a&gt;, not yet master. Commenters quote the model card as a &lt;strong&gt;custom MoE&lt;/strong&gt; replacing prior &lt;strong&gt;GigaChat 3.1 Ultra 700B&lt;/strong&gt; with a ~&lt;code&gt;40%&lt;/code&gt; smaller model that is stronger on code/math/agentic tasks, uses ~&lt;code&gt;4×&lt;/code&gt; less KV cache/token, fits &gt;&lt;code&gt;2×&lt;/code&gt; more context in the same memory, and improves generation throughput by ~&lt;code&gt;20%&lt;/code&gt;. Architecturally, it reportedly uses a &lt;strong&gt;hybrid MLA + GatedDeltaNet linear-attention&lt;/strong&gt; stack plus &lt;strong&gt;two MTP heads&lt;/strong&gt;, with claimed greedy decoding speedups of ~&lt;code&gt;1.5×&lt;/code&gt; for one head and up to &lt;code&gt;2.2×&lt;/code&gt; for two.&lt;/strong&gt; Top technical caveats were that benchmark comparisons to &lt;strong&gt;DeepSeek 3.2&lt;/strong&gt; may be a weak reference point versus current frontier models, and that GigaChat3.5 is a &lt;strong&gt;non-reasoning&lt;/strong&gt; model, so benchmark interpretation should account for that. One commenter praised the release openness—base model plus intermediate/open-weight checkpoints for a model of this size—while noting the training dataset remains undisclosed.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters note that benchmark comparisons should account for &lt;strong&gt;GigaChat3.5-432B-A28B&lt;/strong&gt; being a &lt;em&gt;non-reasoning&lt;/em&gt; model, making comparisons against reasoning/frontier models potentially misleading. One user questioned using &lt;strong&gt;DeepSeek 3.2&lt;/strong&gt; as a reference point, calling it “~year behind the frontier models,” while another argued non-reasoning baselines are increasingly rare and should be evaluated separately.&lt;/li&gt;
&lt;li&gt;The release is viewed as unusually open for a model of this size: commenters highlighted that &lt;strong&gt;intermediate checkpoints and the base model are open-weighted&lt;/strong&gt;, which they described as “top 10% of openness of models on HF.” The main missing artifact called out was the &lt;strong&gt;exact training dataset&lt;/strong&gt;, which limits full reproducibility and data-contamination analysis.&lt;/li&gt;
&lt;li&gt;A detailed architecture excerpt says &lt;strong&gt;GigaChat 3.5 Ultra&lt;/strong&gt; is ~&lt;code&gt;40%&lt;/code&gt; smaller than &lt;strong&gt;GigaChat 3.1 Ultra 700B&lt;/strong&gt; while improving code, math, and agentic performance, using ~&lt;code&gt;4×&lt;/code&gt; less KV-cache per token, fitting &lt;code&gt;&gt;2×&lt;/code&gt; more context in the same memory, and improving throughput by ~&lt;code&gt;20%&lt;/code&gt;. The model uses a custom MoE hybrid attention design combining &lt;strong&gt;MLA&lt;/strong&gt; with &lt;strong&gt;GatedDeltaNet&lt;/strong&gt; linear-attention layers, plus &lt;strong&gt;Multi-Token Prediction&lt;/strong&gt; with two heads, reportedly giving greedy decoding speedups of ~&lt;code&gt;1.5×&lt;/code&gt; with one head and up to &lt;code&gt;2.2×&lt;/code&gt; with two.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Frontier-Scale Models on Consumer Hardware&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uoij3s/if_trends_hold_mythosclass_capability_may_be/&quot;&gt;If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years&lt;/a&gt;&lt;/strong&gt; (Activity: 1992): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/5xwuga6pwhbh1.png&quot;&gt;image&lt;/a&gt; is a speculative trend chart titled &lt;strong&gt;“From frontier to running on a laptop”&lt;/strong&gt; arguing that open-weight, laptop-runnable models have historically lagged frontier releases by an average of &lt;code&gt;24.8 months&lt;/code&gt;—e.g. &lt;strong&gt;GPT-3 → Llama 2 70B&lt;/strong&gt; in &lt;code&gt;37 months&lt;/code&gt;, &lt;strong&gt;ChatGPT/GPT-3.5 → Llama 3 70B&lt;/strong&gt; in &lt;code&gt;17 months&lt;/code&gt;, and &lt;strong&gt;GPT-4 → Gemma 3/Qwen3-class&lt;/strong&gt; in &lt;code&gt;24 months&lt;/code&gt;. Extrapolating that trend, it projects &lt;strong&gt;GPT-5/Claude 4-class&lt;/strong&gt; capability on high-end consumer hardware by around &lt;strong&gt;mid-2027&lt;/strong&gt; and &lt;strong&gt;Fable/Mythos-class&lt;/strong&gt; capability by around &lt;strong&gt;July 2028&lt;/strong&gt;, though this is a heuristic projection rather than a benchmarked result.&lt;/strong&gt; Commenters were skeptical that “consumer hardware” will remain affordable if this trend holds, arguing high-end local inference may converge with enterprise-grade compute costs. A technical thread noted that &lt;strong&gt;Gemma 4 26B A4B&lt;/strong&gt; initially performed poorly at long context on an RTX 5080, but the user later traced it to configuration and reported ~&lt;code&gt;100 tok/s&lt;/code&gt; after using &lt;code&gt;--no-mmap --batch-size 256 --ubatch-size 512&lt;/code&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user reported that &lt;strong&gt;Gemma 4 26B A4B QAT&lt;/strong&gt; initially performed poorly on an &lt;strong&gt;RTX 5080&lt;/strong&gt; at long context, generating only about &lt;code&gt;6 tok/s&lt;/code&gt; at &lt;code&gt;20K&lt;/code&gt; context and raising doubts that a hypothetical &lt;strong&gt;Gemma 4 31B dense&lt;/strong&gt; model would be practical on laptop-class hardware. They later identified it as a configuration issue and, after applying &lt;code&gt;llama.cpp&lt;/code&gt;-style flags &lt;code&gt;--no-mmap --batch-size 256 --ubatch-size 512&lt;/code&gt;, saw throughput improve to roughly &lt;code&gt;100 tok/s&lt;/code&gt; idle and &lt;code&gt;60 tok/s&lt;/code&gt; under load, referencing this setup guide: &lt;a href=&quot;https://carteakey.dev/blog/local-inference/running-gemma-4-26b-a4b-locally/&quot;&gt;running Gemma 4 26B A4B locally&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;One commenter cautioned that extrapolating &lt;strong&gt;Mythos-class&lt;/strong&gt; local feasibility is speculative because the actual model size/architecture is unknown; they note it could be “&lt;code&gt;3 times the size of Opus 4.8&lt;/code&gt;,” making assumptions about fitting into current high-end consumer GPUs unreliable.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1uocapw/i_managed_to_run_glm52_744b_moe_on_a_humble_25_gb/&quot;&gt;I managed to run GLM-5.2 (744B MoE) on a humble 25 GB RAM laptop — pure C, experts streamed from disk&lt;/a&gt;&lt;/strong&gt; (Activity: 546): &lt;strong&gt;The author built &lt;strong&gt;&lt;a href=&quot;https://github.com/JustVugg/colibri&quot;&gt;colibrì&lt;/a&gt;&lt;/strong&gt;, a pure-C, zero-dependency inference engine for &lt;strong&gt;GLM-5.2 744B MoE&lt;/strong&gt;, keeping the dense &lt;code&gt;int4&lt;/code&gt; portion resident in ~&lt;code&gt;9.9–10 GB&lt;/code&gt; RAM while streaming ~&lt;code&gt;21k&lt;/code&gt; routed experts (~&lt;code&gt;370 GB int4&lt;/code&gt;) from disk on demand. It implements GLM-5.2’s forward path including &lt;strong&gt;MLA attention&lt;/strong&gt;, compressed KV cache, DeepSeek-style routing, MTP speculative decoding, &lt;code&gt;int8/int4&lt;/code&gt; AVX2 kernels, async expert readahead, batch-union MoE, and an FP8→int4 converter; reported performance on a 12-core/25 GB RAM WSL2 NVMe laptop is disk-bound at ~&lt;code&gt;0.05–0.1 tok/s&lt;/code&gt; cold with ~&lt;code&gt;11 GB&lt;/code&gt; random reads/token.&lt;/strong&gt; Top comments were mostly skeptical/humorous about calling this “running” the model at &lt;code&gt;0.1 t/s&lt;/code&gt;, with one commenter suggesting the real metric is now &lt;em&gt;seconds per token&lt;/em&gt; rather than tokens per second.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter reports/infers extremely low throughput at about &lt;code&gt;0.1 tokens/s&lt;/code&gt;, with others reframing the result as &lt;strong&gt;seconds per token&lt;/strong&gt; rather than tokens per second. The discussion implies the disk-streamed MoE setup is technically functional but dominated by I/O latency, especially for spinning disks or older DDR3-era systems.&lt;/li&gt;
&lt;li&gt;One technical question asks whether &lt;code&gt;llama.cpp&lt;/code&gt; with &lt;code&gt;mmap&lt;/code&gt; was tried instead of the custom pure-C disk-streaming approach. This points to a likely alternative implementation path: memory-mapped model weights with OS paging rather than explicit expert streaming from disk.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Fable 5 Capability Demos and Long-Context Use&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uo1xpz/i_misunderstood_fable_at_first_now_i_get_it/&quot;&gt;I misunderstood Fable at first, now I get it.&lt;/a&gt;&lt;/strong&gt; (Activity: 1626): &lt;strong&gt;The post argues that &lt;strong&gt;Fable&lt;/strong&gt; is only marginally ahead of &lt;strong&gt;Opus&lt;/strong&gt; in “raw intelligence,” but its practical advantage is maintaining coherent context across larger, interdependent artifacts. In a PCB review workflow involving &lt;code&gt;8&lt;/code&gt; schematic sheets, the author reports Fable better tracks cross-sheet dependencies where Opus “misses the mark” beyond ~&lt;code&gt;2&lt;/code&gt; sheets, suggesting stronger long-context/global-reasoning behavior rather than better local reasoning.&lt;/strong&gt; Top commenters largely agree: Fable’s value is “seeing the bigger picture” and sustaining code/design quality over longer contexts. One workflow described using Fable for whole-codebase/documentation analysis and recommendation generation, then switching to Opus for item-by-item implementation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters characterize &lt;strong&gt;Fable&lt;/strong&gt; as stronger than &lt;strong&gt;Opus&lt;/strong&gt; for high-level software engineering orchestration: analyzing an existing codebase plus documentation, identifying gaps, and producing a recommendation/implementation plan before switching to Opus for task-by-task coding. The reported workflow is: use Fable for architecture/codebase assessment and planning, then use Opus for concrete implementation.&lt;/li&gt;
&lt;li&gt;A recurring technical distinction is that Fable may not produce immediately “better” code than Opus, but appears to &lt;em&gt;maintain quality for longer&lt;/em&gt; by avoiding dead-end implementation paths. One commenter described Opus repeatedly pursuing a known non-viable approach, while Fable recognized the strategy was unproductive—useful for project-level decision-making rather than local code generation.&lt;/li&gt;
&lt;li&gt;Users note a tradeoff in interaction style: Fable performs well for project outlining, gap analysis, and understanding existing systems, but in chat mode may rush to conclusions instead of following instructions to ask clarifying questions and wait. This suggests Fable’s strength is more in autonomous planning/review than tightly controlled interactive requirement gathering.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uo3af4/google_deepmind_product_and_design_lead_using_and/&quot;&gt;Google DeepMind Product and Design Lead using and advertising a competitor&apos;s model&lt;/a&gt;&lt;/strong&gt; (Activity: 1192): &lt;strong&gt;The &lt;a href=&quot;https://i.redd.it/0k8376mn5fbh1.png&quot;&gt;image&lt;/a&gt; shows &lt;strong&gt;Ammaar Reshi&lt;/strong&gt;, identified by the post title as a &lt;strong&gt;Google DeepMind Product and Design Lead&lt;/strong&gt;, publicly saying he used the competitor model &lt;strong&gt;“Fable 5”&lt;/strong&gt; to port &lt;em&gt;Command &amp;#x26; Conquer: Generals Zero Hour&lt;/em&gt; to &lt;strong&gt;iPhone/iPad&lt;/strong&gt;. Technically, the claim is notable because it says a &lt;strong&gt;2003 PC RTS engine&lt;/strong&gt; was compiled &lt;strong&gt;natively for ARM64&lt;/strong&gt; with &lt;strong&gt;touch controls&lt;/strong&gt;, implying LLM-assisted code migration/porting across architecture, platform APIs, and input paradigms.&lt;/strong&gt; Comments mostly frame this as competitive intelligence rather than disloyalty: one user argued a product lead &lt;em&gt;should&lt;/em&gt; know rival tools well, while another noted Google’s relationship/investment in Anthropic-like competitors blurs the rivalry. A separate commenter highlighted that this kind of LLM-assisted game port was dismissed as years away only months ago.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically substantive thread points to the actual project: &lt;a href=&quot;https://github.com/ammaarreshi/Generals-Mac-iOS-iPad&quot;&gt;&lt;code&gt;Generals-Mac-iOS-iPad&lt;/code&gt;&lt;/a&gt;, which appears to wrap/port &lt;em&gt;Command &amp;#x26; Conquer: Generals&lt;/em&gt; to Apple platforms. One commenter argues that while the result is impressive, it likely relies on “&lt;code&gt;4 abstraction layers&lt;/code&gt;” over a legacy &lt;code&gt;DX8&lt;/code&gt; codebase, and suggests waiting for the cleaner engine-level rewrite from &lt;a href=&quot;https://github.com/TheSuperHackers/GeneralsGameCode/&quot;&gt;&lt;code&gt;TheSuperHackers/GeneralsGameCode&lt;/code&gt;&lt;/a&gt; instead.&lt;/li&gt;
&lt;li&gt;Several comments frame the Claude usage less as disloyalty and more as competitive analysis, noting that a Product/Design lead should understand rival model capabilities firsthand. One commenter adds that Google has a substantial relationship with Anthropic, claiming Google owns about &lt;code&gt;18%&lt;/code&gt; of Anthropic, so the “competitor” framing is technically and commercially more nuanced.&lt;/li&gt;
&lt;li&gt;A notable technical observation is that the surprising part was not the use of Claude/Fable-class tooling, but that the port targeted Apple platforms before Android. This implies the discussion is partly about platform/runtime feasibility and tooling maturity for bringing legacy PC games to iOS/macOS/iPadOS rather than simply about model choice.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1uow9c8/the_room_one_shot_by_fable/&quot;&gt;&quot;The Room&quot; - One shot by Fable&lt;/a&gt;&lt;/strong&gt; (Activity: 683): &lt;strong&gt;Reddit post showcases a video titled &lt;strong&gt;“The Room”&lt;/strong&gt;, described as a &lt;strong&gt;one-shot&lt;/strong&gt; piece by &lt;strong&gt;Fable&lt;/strong&gt;, but the linked Reddit-hosted media URL (&lt;a href=&quot;https://v.redd.it/68csn9fdulbh1&quot;&gt;v.redd.it/68csn9fdulbh1&lt;/a&gt;) was not externally accessible due to Reddit returning &lt;strong&gt;HTTP &lt;code&gt;403 Forbidden&lt;/code&gt;&lt;/strong&gt;, so the actual media/implementation details cannot be verified. Comments imply the clip is a highly detailed continuous zoom or scale-transition visualization, with viewers speculating about how such detail was generated and asking about the &lt;em&gt;code/cost&lt;/em&gt; behind it.&lt;/strong&gt; Top technical reactions focused on missed scope—one commenter argued the zoom should have gone beyond quarks—and on skepticism/amazement that the scene could be rendered with that level of detail.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters raised technical curiosity about the generation setup behind Fable’s one-shot video, specifically asking for the &lt;strong&gt;prompt&lt;/strong&gt;, implied production pipeline, and possible compute/code cost required to achieve the apparent level of scene detail.&lt;/li&gt;
&lt;li&gt;One commenter noted the missed opportunity to extend the “zoom inward” concept beyond quarks, framing it as a speculative technical/narrative limitation: there is no confirmed physics rule that quarks are the smallest possible constituents, so the sequence could have explored deeper hypothetical structure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude in Applied Workflows and Agent Dashboards&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uobmts/claude_meets_government_oversight/&quot;&gt;Claude meets Government Oversight 🫡🇺🇸&lt;/a&gt;&lt;/strong&gt; (Activity: 744): &lt;strong&gt;OP is building &lt;strong&gt;“Article One,” a Claude-powered multi-agent dashboard&lt;/strong&gt; intended to aggregate member-of-Congress profiles, constituency/campaign context, job-performance metrics, campaign-donor analysis, and congressional office spending/taxpayer-funded operations into a transparency interface. The project is currently unreleased; OP says repo/dashboard publication is delayed by &lt;strong&gt;weekly Claude usage limits&lt;/strong&gt; and is seeking funding for Claude Max via &lt;a href=&quot;https://buymeacoffee.com/AJK28&quot;&gt;Buy Me a Coffee&lt;/a&gt; to accelerate development.&lt;/strong&gt; Commenters strongly support the transparency use case but emphasize that every metric must have &lt;strong&gt;verifiable sources, methodology, and auditability&lt;/strong&gt; to avoid AI hallucination or misleading claims. The main feature request is deeper financial analysis beyond public bios: spouse/family wealth changes, campaign funders, affiliated industries/PACs, and possible conflicts of interest rather than a simple wiki-style profile.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters emphasized that any Claude-based government-transparency tool needs &lt;strong&gt;verifiable citations and methodology for every derived statistic&lt;/strong&gt;, especially after a user questioned a claim that &lt;code&gt;33%&lt;/code&gt; of the House was “more liberal than AOC,” which appeared implausible. The main technical concern was that hallucinated or poorly sourced political metrics could make the system actively harmful rather than informative.&lt;/li&gt;
&lt;li&gt;A substantive feature request was to expand beyond public-facing biographical summaries into &lt;strong&gt;campaign-finance and conflict-of-interest analysis&lt;/strong&gt;: who funds each politician, how spouse/family wealth changes while in office, whether related parties operate hedge funds or other investment vehicles, and whether donations from groups like &lt;strong&gt;AIPAC&lt;/strong&gt; correlate with voting behavior. This implies the tool would need structured ingestion of financial disclosures, campaign-contribution databases, voting records, and entity-resolution across family/business relationships.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uox9uu/thank_you_anthropic_as_a_teacher_claude_cowork/&quot;&gt;Thank you anthropic. As a teacher claude cowork has been Godsend.&lt;/a&gt;&lt;/strong&gt; (Activity: 688): &lt;strong&gt;A teacher reports using &lt;strong&gt;Anthropic Claude&lt;/strong&gt; (“Claude cowork”) for curriculum design, grading support, PowerPoint generation, and student-grade data analysis, claiming it saves hours by combining uploaded pedagogy documents with lesson-planning workflows. The main requested feature is a &lt;strong&gt;Microsoft 365 / OneNote connector for personal accounts&lt;/strong&gt;, since their school does not use a work/education Microsoft 365 tenant, limiting integration with existing teaching materials.&lt;/strong&gt; Top comments raised a concrete data-governance concern: uploading student names, grades, or identifiable work to Claude may violate school policy or data-protection rules unless the data is anonymized or the system is approved. Another commenter noted that redacting student identifiers can erase much of the productivity gain, especially for personalized essay feedback.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters focused on &lt;strong&gt;student-data privacy risks&lt;/strong&gt; when using Claude in an education setting, noting that entering students’ names or identifiable context into a non-approved AI/work system could violate school policy or lead to disciplinary action. One teacher described the practical overhead of anonymization: they tried using Claude for Year 11 essay feedback but spent so much time &lt;em&gt;“scrubbing names and identifiers”&lt;/em&gt; that it reduced or eliminated the productivity gain.&lt;/li&gt;
&lt;li&gt;A proposed mitigation was to configure Claude with explicit instructions to protect student data, upload the school’s data-protection policies, and ask it to stop when it detects sensitive information. Commenters emphasized this is &lt;strong&gt;not foolproof&lt;/strong&gt;, but could act as a lightweight guardrail; Claude could also be asked to suggest privacy-preserving workflows or workarounds when sensitive data is required for personalization.&lt;/li&gt;
&lt;li&gt;There was also interest in tighter &lt;strong&gt;Microsoft 365 / PowerPoint integration&lt;/strong&gt;, particularly for schools using personal rather than managed accounts. Commenters suggested that lack of approved connectors creates workflow friction and pushes teachers toward manual workarounds, which can increase both time cost and data-governance risk.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uopekl/i_feel_like_were_rapidly_heading_to_a_place_where/&quot;&gt;I feel like we&apos;re rapidly heading to a place where people have all sorts of local bespoke tools that are amazing and only for them&lt;/a&gt;&lt;/strong&gt; (Activity: 875): &lt;strong&gt;The post observes a growing pattern of &lt;strong&gt;AI-assisted “bespoke local tools”&lt;/strong&gt;: highly useful personal or organization-specific software that is tightly coupled to one user’s workflow and unlikely to be generalized or distributed. Commenters cite examples like fragile personal automation setups, custom workout/alarm apps, and a niche-company &lt;strong&gt;ERP&lt;/strong&gt; built via “vibecoding” where the market would not justify conventional software development.&lt;/strong&gt; Commenters generally view this as positive: AI lowers the cost of building software for very small niches, even if the result is non-transferable, brittle, or only maintainable by its creator.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters framed local AI-built software as &lt;strong&gt;highly personalized but non-transferable&lt;/strong&gt;, with one describing their setup as a &lt;em&gt;“magic little box”&lt;/em&gt; that would break if others touched it. The technical implication is that many AI-generated tools may optimize for individual workflows rather than maintainability, portability, onboarding, or generalized product-market fit.&lt;/li&gt;
&lt;li&gt;One user reported using “vibecoding” to build a custom &lt;strong&gt;ERP system for a niche company&lt;/strong&gt;, arguing that AI-assisted development makes economically viable software that would not justify a traditional vendor or developer market. This highlights a potential shift toward internal, domain-specific tools where the ROI comes from solving a narrow operational need rather than producing reusable SaaS.&lt;/li&gt;
&lt;li&gt;Another comment predicted near-term replacement of small-business administrative labor with operators who can effectively use &lt;strong&gt;Claude&lt;/strong&gt;, characterizing it as &lt;em&gt;“a team of administrators and interns”&lt;/em&gt; and comparing collaborative AI tooling to a high-end executive assistant. The substantive point is that value may accrue less to generic “AI automation” products and more to employees who can integrate LLMs into concrete business administration workflows.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>tencent</category><category>nvidia</category><category>amd</category><category>nous-research</category><category>hugging-face</category><category>artificial-anlysiis</category><category>dair-ai</category><category>hy3</category><category>glm-5.2</category><category>claude-fable-5</category><category>opus-4.8</category><category>gemini-3.5-flash</category><category>gpt-5.5-xhigh</category><category>glm-5.2-max</category><category>eliebakouch</category><category>shunyuyao12</category><category>vllm_project</category><category>teortaxestex</category><category>tinygrad</category><category>mbusigin</category><category>artificialanlys</category><category>fchollet</category><category>omarsar0</category><category>mixture-of-experts</category><category>model-quantization</category><category>speculative-decoding</category><category>inference-speed</category><category>agent-evaluation</category><category>long-context</category><category>memory-optimization</category><category>cost-efficiency</category><category>benchmarking</category><category>multi-domain-evaluation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-02-not-much/</guid><description>**Fullstack Code Arena** extends coding agent evaluation to include **databases, API keys, deployments, and structured tool use**, marking a shift to end-to-end app shipping. **LangChain** released **LangSmith** with unified tracing and **OpenWiki** for auto-generated docs, while **LlamaIndex** demonstrated agent-native parsing capabilities. The main UX challenge is now coordination aspects like routing, observability, and memory, highlighted by **Simon Willison** and **Will Depue**. **Anthropic** improved operational access to **Fable** with raised API rate limits and expanded **Claude Code** features, despite some deployment controversies. Open-model economics gain traction as **Together** reports **GLM-5.2** achieves 80% of **Sonnet 5**&apos;s coding capability at 20% cost, and **GLM-5.2** becomes selectable in **Claude Code** via **Hugging Face** inference providers. Industry leaders like **Clement Delangue**, **Jason**, and **Bryan Catanzaro** emphasize the rising credibility of open models in developer workflows.</description><pubDate>Thu, 02 Jul 2026 05:44:39 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;a quiet day.&lt;/strong&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;AI News for 7/01/2026-7/02/2026. We checked 12 subreddits, &lt;a href=&quot;https://twitter.com/i/lists/1585430245762441216&quot;&gt;544 Twitters&lt;/a&gt; and no further Discords. &lt;a href=&quot;https://news.smol.ai/&quot;&gt;AINews&apos; website&lt;/a&gt; lets you search all past issues. As a reminder, &lt;a href=&quot;https://www.latent.space/p/2026&quot;&gt;AINews is now a section of Latent Space&lt;/a&gt;. You can &lt;a href=&quot;https://support.substack.com/hc/en-us/articles/8914938285204-How-do-I-subscribe-to-or-unsubscribe-from-a-section-on-Substack&quot;&gt;opt in/out&lt;/a&gt; of email frequencies!&lt;/p&gt;
&lt;/blockquote&gt;
&lt;hr&gt;
&lt;h1&gt;AI Twitter Recap&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;Agentic Coding Systems, Harnesses, and Developer Workflow Infrastructure&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Full-stack evals are replacing toy coding demos&lt;/strong&gt;: &lt;a href=&quot;https://x.com/arena/status/2072713730711023673&quot;&gt;Code Arena&lt;/a&gt; launched &lt;strong&gt;Fullstack Code Arena&lt;/strong&gt;, extending evaluation from frontend mockups to software that includes &lt;strong&gt;databases, API keys, deployments, and structured tool use&lt;/strong&gt;. That aligns with a broader shift from “can the model write a component?” to “can the agent ship a realistic app end-to-end?”, echoed by &lt;a href=&quot;https://x.com/aryanvichare10/status/2072736881859756503&quot;&gt;Aryan Vichare&lt;/a&gt; and by practitioners emphasizing environment-based evals over static prompts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The engineering stack around coding agents is thickening fast&lt;/strong&gt;: LangChain pushed unified tracing for heterogeneous coding tools in &lt;a href=&quot;https://x.com/LangChain/status/2072738719707050028&quot;&gt;LangSmith&lt;/a&gt;, plus &lt;strong&gt;OpenWiki&lt;/strong&gt; for auto-generated repo docs and AGENTS.md updates in &lt;a href=&quot;https://x.com/BraceSproul/status/2072744824470724887&quot;&gt;this release&lt;/a&gt;. LlamaIndex showed a small but useful pattern where parsing becomes an agent-native capability rather than a preprocessing step via a &lt;a href=&quot;https://x.com/llama_index/status/2072713940073603190&quot;&gt;LiteParse + flue + Resend + Turso email assistant&lt;/a&gt;. Meanwhile, multiple posts from &lt;a href=&quot;https://x.com/jerryjliu0/status/2072832362443067782&quot;&gt;Jerry Liu&lt;/a&gt; and others argued that retrieval complexity is increasingly encoded at the &lt;strong&gt;agent layer&lt;/strong&gt;, with simpler tools and smarter orchestration.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The practical UX problem is now coordination, not raw codegen&lt;/strong&gt;: A recurring theme from builders is that frontier coding performance is good enough that bottlenecks have shifted to &lt;strong&gt;routing, observability, collaboration, memory, and understanding&lt;/strong&gt;. &lt;a href=&quot;https://x.com/simonw/status/2072730602344984821&quot;&gt;Simon Willison&lt;/a&gt; highlighted “understand to participate” as the key antidote to cognitive debt with coding agents; &lt;a href=&quot;https://x.com/willdepue/status/2072793965565468789&quot;&gt;Will Depue&lt;/a&gt; sketched the desired end-state: an always-on executive assistant with persistent memory, delegated actions, messaging, and computer use. That same desire shows up in &lt;a href=&quot;https://x.com/willdepue/status/2072798659100684699&quot;&gt;PersonalOS&lt;/a&gt;, where a 300k-token life context pack is assembled from personal data exports.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Model Availability, Frontier Coding Performance, and Open vs Closed Positioning&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Anthropic’s Fable discourse dominated, but the most concrete news was operational&lt;/strong&gt;: Anthropic restored confidence around access rather than releasing new weights: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2072818299361263778&quot;&gt;official API rate limits were raised and simplified&lt;/a&gt;, and &lt;a href=&quot;https://x.com/trq212/status/2072814903170408784&quot;&gt;Trapit Bansal said Fable is expected to return to subscriptions once capacity allows&lt;/a&gt;. Anthropic also expanded &lt;a href=&quot;https://x.com/ClaudeDevs/status/2072770790114914317&quot;&gt;Claude Code artifacts to Pro and Max plans&lt;/a&gt;, making long-running coding sessions easier to inspect and share.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community signal suggests Fable remains frontier-class despite rerouting controversy&lt;/strong&gt;: Several viral posts complained about Anthropic’s deployment/routing behavior, but even critics were separating that from model quality. &lt;a href=&quot;https://x.com/theo/status/2072777433997316436&quot;&gt;Theo&lt;/a&gt; argued that bad takes on Fable were distracting from Anthropic’s actual issues, while &lt;a href=&quot;https://x.com/arena/status/2072828263848894783&quot;&gt;Arena’s early before/after comparison&lt;/a&gt; said scores looked &lt;strong&gt;mostly consistent&lt;/strong&gt; after redeployment across text, document, vision, and code. &lt;a href=&quot;https://x.com/theo/status/2072777929180987485&quot;&gt;Theo also noted&lt;/a&gt; that some benchmark drops may reflect fallback behavior more than a base capability regression.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Open-model economics are increasingly credible in coding&lt;/strong&gt;: Together reported that &lt;a href=&quot;https://x.com/togethercompute/status/2072836285455368605&quot;&gt;GLM 5.2 reaches roughly 80% of Sonnet 5 software-engineering capability at ~20% of the price&lt;/a&gt;, and &lt;a href=&quot;https://x.com/zRdianjiao/status/2072906526722064415&quot;&gt;zRdianjiao showed&lt;/a&gt; that &lt;strong&gt;GLM-5.2 is now selectable in Claude Code via Hugging Face Inference Providers&lt;/strong&gt;, a notable step toward open models inhabiting first-class dev workflows. More broadly, &lt;a href=&quot;https://x.com/ClementDelangue/status/2072752653436653755&quot;&gt;Clement Delangue&lt;/a&gt;, &lt;a href=&quot;https://x.com/Jason/status/2072778368530198973&quot;&gt;Jason&lt;/a&gt;, and &lt;a href=&quot;https://x.com/mattturck/status/2072723410975629364&quot;&gt;Bryan Catanzaro via Matt Turck’s interview&lt;/a&gt; all pushed variants of the same thesis: &lt;strong&gt;open models are becoming the sovereignty layer&lt;/strong&gt; for enterprises and developers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Meta appears to be re-entering the agentic conversation&lt;/strong&gt;: &lt;a href=&quot;https://x.com/alexandr_wang/status/2072848108342677597&quot;&gt;Alexandr Wang posted&lt;/a&gt; that the next &lt;strong&gt;Muse Spark&lt;/strong&gt; update is coming soon with “big improvements in coding and agentic capabilities” to be competitive with leading models, rolling out to Meta AI and its API.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Inference, Kernels, Serving, and Test-Time Compute as the New Scaling Frontier&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Kernel-level automation is no longer hypothetical&lt;/strong&gt;: The standout systems post was &lt;a href=&quot;https://x.com/elliotarledge/status/2072814573753975266&quot;&gt;Elliot Arledge’s KernelBench-Mega result&lt;/a&gt;: Claude Fable 5 reportedly wrote the first authentic &lt;strong&gt;single-launch megakernel&lt;/strong&gt; for a Kimi-Linear decode workload, achieving &lt;strong&gt;18.7x over reference&lt;/strong&gt; and beating prior multi-kernel entries. The description is detailed enough to matter to systems folks: in-register int4 dequant, fused attention/router/MoE/norm/KV append, explicit barrier shaving, and a demonstrated willingness by the model to benchmark, revert regressions, and optimize toward a roofline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Speculation and speculative decoding remain active optimization surfaces&lt;/strong&gt;: &lt;a href=&quot;https://x.com/teortaxesTex/status/2072716346899456123&quot;&gt;teortaxesTex&lt;/a&gt; pointed to “scaling the speculator” as a new dimension for accelerating inference and therefore RL throughput, while &lt;a href=&quot;https://x.com/mgoin_/status/2072785822231728363&quot;&gt;mgoin_&lt;/a&gt; shared a concrete &lt;strong&gt;DSpark + Mooncake + vLLM&lt;/strong&gt; setup on &lt;strong&gt;GB300 NVL72&lt;/strong&gt;, with &lt;strong&gt;125k prefill tok/s&lt;/strong&gt; and &lt;strong&gt;1.5 steps/s&lt;/strong&gt; for online training. The vLLM team also highlighted &lt;a href=&quot;https://x.com/vllm_project/status/2072722401813565834&quot;&gt;5x lower token costs on DeepSeek V4 in one month&lt;/a&gt; and published a particularly useful serving breakdown for &lt;a href=&quot;https://x.com/vllm_project/status/2072942203966812438&quot;&gt;Qwen3-Omni’s real-time speech pipeline&lt;/a&gt;, where stage-specific replication yields &lt;strong&gt;~0.6s first audio instead of ~6s&lt;/strong&gt; and &lt;strong&gt;5.4x throughput&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Test-time compute budgets are changing benchmark interpretation&lt;/strong&gt;: The UK AISI post on larger compute budgets propagated widely. &lt;a href=&quot;https://x.com/scaling01/status/2072799566735306760&quot;&gt;scaling01&lt;/a&gt;, &lt;a href=&quot;https://x.com/tomekkorbak/status/2072863584586219924&quot;&gt;Tomek Korbak&lt;/a&gt;, &lt;a href=&quot;https://x.com/polynoamial/status/2072909389389021484&quot;&gt;Noam Brown/polynoamial&lt;/a&gt;, &lt;a href=&quot;https://x.com/idavidrein/status/2072830683974906170&quot;&gt;David Rein&lt;/a&gt;, and &lt;a href=&quot;https://x.com/tobyordoxford/status/2072948952274530404&quot;&gt;Toby Ord&lt;/a&gt; all emphasized the same point: if you don’t allocate enough tokens, you systematically underestimate frontier agents. The headline number: frontier horizon estimates rise from roughly &lt;strong&gt;2 hours at 2.5M tokens&lt;/strong&gt; to around &lt;strong&gt;14 hours at 50M tokens&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Benchmarks and Research on Learning, Memory, World Models, and Continual Adaptation&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Continual/on-the-fly learning is getting sharper measurement tools, but results remain mixed&lt;/strong&gt;: Epoch introduced &lt;a href=&quot;https://x.com/EpochAIResearch/status/2072714372175237255&quot;&gt;EBR-bench&lt;/a&gt;, where models repeatedly play Earthborne Rangers and attempt to learn from failure; current frontier systems show &lt;strong&gt;no clear improvement&lt;/strong&gt; absent dedicated RL. In parallel, ByteDance Seed’s new &lt;a href=&quot;https://x.com/scaling01/status/2072790212615237858&quot;&gt;EdgeBench&lt;/a&gt; drew strong attention for studying &lt;strong&gt;day-long horizons across 134 real-world environments&lt;/strong&gt;, claiming that learning speed doubles every ~3 months and that gains are not explained by repeated sampling alone. This benchmark is quickly being treated as a serious complement to METR-style horizon work.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory is being elevated from support module to trainable competence&lt;/strong&gt;: The Stanford &lt;strong&gt;AutoMem&lt;/strong&gt; paper got attention via &lt;a href=&quot;https://x.com/omarsar0/status/2072716688483831885&quot;&gt;Omar Sanseviero’s summary&lt;/a&gt;: memory management is treated as a skill, with models deciding what to store, retrieve, and reorganize; optimizing memory alone reportedly yields &lt;strong&gt;2x–4x&lt;/strong&gt; gains on Crafter, MiniHack, and NetHack. That idea rhymes with a more applied trend toward persistent personal and research memory systems: &lt;a href=&quot;https://x.com/omarsar0/status/2072735813469905026&quot;&gt;PaperWiki&lt;/a&gt;, &lt;a href=&quot;https://x.com/willdepue/status/2072798659100684699&quot;&gt;PersonalOS&lt;/a&gt;, and OpenWiki all point to memory becoming part of the product surface.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;World models are shifting from static assets to adaptive online components&lt;/strong&gt;: Reka released &lt;a href=&quot;https://x.com/RekaAILabs/status/2072731356011045088&quot;&gt;WorldModelGym&lt;/a&gt;, framing evaluation around &lt;strong&gt;decision-based fidelity&lt;/strong&gt; across 100+ tracks. &lt;a href=&quot;https://x.com/askalphaxiv/status/2072750223026438226&quot;&gt;askalphaxiv’s summary of AdaJEPA&lt;/a&gt; pushed the stronger claim: pretrained world models should keep adapting at deployment time, with one gradient step per MPC cycle improving robustness under visual and dynamics shift.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Top tweets (by engagement)&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Anthropic access/capacity update&lt;/strong&gt;: &lt;a href=&quot;https://x.com/trq212/status/2072814903170408784&quot;&gt;Trapit Bansal on Fable returning to subscriptions when capacity allows&lt;/a&gt; — the clearest signal that current scarcity is a capacity problem, not a permanent packaging decision.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API/platform change with immediate operational impact&lt;/strong&gt;: &lt;a href=&quot;https://x.com/ClaudeDevs/status/2072818299361263778&quot;&gt;Claude API rate limits raised and tiers simplified&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Model stack composition for coding&lt;/strong&gt;: &lt;a href=&quot;https://x.com/mitchellh/status/2072715852944957531&quot;&gt;Mitchell Hashimoto’s planner/coder/judge workflow&lt;/a&gt; using &lt;strong&gt;Fable xhigh → GPT-5.5 xhigh → Fable xhigh&lt;/strong&gt;, with planning/judging costing only a few dollars versus much pricier end-to-end loops.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Specialized post-training beating frontier prompting&lt;/strong&gt;: &lt;a href=&quot;https://x.com/aakashgupta/status/2072765754102174114&quot;&gt;Aakash Gupta on Bridgewater + Thinking Machines&lt;/a&gt;, where a fine-tuned &lt;strong&gt;Qwen3-235B&lt;/strong&gt; reached &lt;strong&gt;84.7%&lt;/strong&gt;, outperforming frontier prompted models on document filtering at &lt;strong&gt;~1/14th the inference cost&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Autonomous systems performance on low-level optimization&lt;/strong&gt;: &lt;a href=&quot;https://x.com/elliotarledge/status/2072814573753975266&quot;&gt;Elliot Arledge’s Fable-written megakernel result&lt;/a&gt;, arguably the most technically substantive coding-agent anecdote in the set.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Video generation leadership change&lt;/strong&gt;: &lt;a href=&quot;https://x.com/Designarena/status/2072759122366509130&quot;&gt;Design Arena reported Gemini Omni Flash at #1 on Video Arena with 1404 Elo&lt;/a&gt;, a &lt;strong&gt;101-point gap&lt;/strong&gt; over Seedance 2.0 Mini and one of the larger observed jumps on that leaderboard.&lt;/li&gt;
&lt;/ul&gt;
&lt;hr&gt;
&lt;h1&gt;AI Reddit Recap&lt;/h1&gt;
&lt;h2&gt;/r/LocalLlama + /r/localLLM Recap&lt;/h2&gt;
&lt;h3&gt;1. llama.cpp Long-Context and Qwen 3.6 Optimization&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ulymml/llamacpp_patch_deepseek_v4_flash_running_with/&quot;&gt;llamacpp patch - DeepSeek V4 Flash running with full 1M token context locally on RTX 5090&lt;/a&gt;&lt;/strong&gt; (Activity: 374): &lt;strong&gt;A &lt;code&gt;llama.cpp&lt;/code&gt; patch wires DeepSeek V4 Flash’s DSA/lightning indexer into the model graph and adds a CUDA kernel, enabling &lt;a href=&quot;https://huggingface.co/antirez/deepseek-v4-gguf/blob/main/DeepSeek-V4-Flash-Layers37-42Q4KExperts-OtherExpertLayersIQ2XXSGateUp-Q2KDown-AProjQ8-SExpQ8-OutQ8-chat-v2-imatrix-fixed.gguf&quot;&gt;DeepSeek-V4-Flash GGUF&lt;/a&gt; to run locally with up to &lt;strong&gt;&lt;code&gt;1M&lt;/code&gt; context on an RTX 5090&lt;/strong&gt; instead of requiring ~&lt;code&gt;256 GiB&lt;/code&gt; compute buffer VRAM. Reported results: at &lt;code&gt;256K&lt;/code&gt; context compute buffer drops from ~&lt;code&gt;67 GiB&lt;/code&gt;/OOM to &lt;code&gt;3.2 GiB&lt;/code&gt;, prefill rises from &lt;code&gt;56 t/s&lt;/code&gt; to ~&lt;code&gt;263 t/s&lt;/code&gt;, decode remains ~&lt;code&gt;14 t/s&lt;/code&gt;; validated presets show &lt;code&gt;256K&lt;/code&gt;/&lt;code&gt;512K&lt;/code&gt;/&lt;code&gt;1M&lt;/code&gt; contexts at ~&lt;code&gt;29&lt;/code&gt;/&lt;code&gt;28&lt;/code&gt;/&lt;code&gt;31 GiB&lt;/code&gt; peak VRAM, with &lt;code&gt;1M&lt;/code&gt; prefill ~&lt;code&gt;159 t/s&lt;/code&gt; due to reduced &lt;code&gt;ubatch&lt;/code&gt;. The author links source/build notes in the &lt;a href=&quot;https://github.com/spencer-zaid/llama.cpp/blob/deepseek-lid-cuda/docs/deepseek-v4-lid-cuda.md&quot;&gt;writeup&lt;/a&gt; and &lt;a href=&quot;https://github.com/spencer-zaid/llama.cpp/tree/deepseek-lid-cuda&quot;&gt;branch&lt;/a&gt;, based on upstream PR &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/pull/24231&quot;&gt;ggml-org/llama.cpp#24231&lt;/a&gt;, and reports basic needle-in-haystack correctness at &lt;code&gt;100K&lt;/code&gt;, &lt;code&gt;512K&lt;/code&gt;, and &lt;code&gt;1M&lt;/code&gt;.&lt;/strong&gt; Comments were mostly positive about the feasibility of running DS4 Flash on a single RTX 5090; one technical follow-up asked for TTFT and/or end-to-end token generation timing (&lt;code&gt;tg-end2end&lt;/code&gt;).&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter requested concrete latency metrics for the claimed local &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt; run on a single &lt;strong&gt;RTX 5090&lt;/strong&gt;, specifically &lt;strong&gt;TTFT&lt;/strong&gt; and &lt;code&gt;tg-end2end&lt;/code&gt;, to validate usability at the advertised &lt;strong&gt;full &lt;code&gt;1M&lt;/code&gt; token context&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Another technical concern was that the result &lt;em&gt;&quot;looks too good to be true&quot;&lt;/em&gt; and should be submitted as patches to &lt;strong&gt;upstream &lt;code&gt;llama.cpp&lt;/code&gt;&lt;/strong&gt; for review, suggesting the implementation may need validation around correctness/performance before being trusted.&lt;/li&gt;
&lt;li&gt;One commenter referenced an ongoing &lt;strong&gt;&lt;code&gt;llama.cpp&lt;/code&gt; lightning indexer fix&lt;/strong&gt; and suggested porting it to &lt;strong&gt;Metal&lt;/strong&gt;, implying the patch may currently be CUDA-focused and that Apple GPU support would require backend-specific adaptation.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLM/comments/1ullrvq/qwen36_27b_q6_5090_maximum_llamacpp_optimization/&quot;&gt;qwen3.6 27b q6 + 5090 maximum llamacpp optimization: 100-233tok/s, average 140&lt;/a&gt;&lt;/strong&gt; (Activity: 201): &lt;strong&gt;A user reports optimized &lt;strong&gt;Qwen 3.6 27B Q6_K + MTP&lt;/strong&gt; inference on a &lt;strong&gt;RTX 5090 32GB / Ryzen 9800X3D / 64GB RAM&lt;/strong&gt; system using a recent &lt;code&gt;llama.cpp&lt;/code&gt; build (&lt;code&gt;86b9470&lt;/code&gt;), achieving &lt;code&gt;100–233 tok/s&lt;/code&gt; over ~20h of agentic workloads with mean &lt;code&gt;140.7 tok/s&lt;/code&gt; and median &lt;code&gt;134.9 tok/s&lt;/code&gt;. The main technical issue addressed is &lt;code&gt;llama.cpp&lt;/code&gt; prompt-cache invalidation for Qwen’s &lt;strong&gt;hybrid attention / sliding-window attention&lt;/strong&gt; behavior—logs show &lt;em&gt;“forcing full prompt re-processing due to lack of cache data”&lt;/em&gt; tied to &lt;a href=&quot;https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055&quot;&gt;&lt;code&gt;llama.cpp&lt;/code&gt; PR discussion&lt;/a&gt;—which the user mitigates via two local patches: checkpoint-search fixes for hybrid/recurrent models and a minimal &lt;code&gt;recurrent_shrink/expand&lt;/code&gt; prompt-cache API patch based on upstream PR &lt;code&gt;#24785&lt;/code&gt; (&lt;a href=&quot;https://pastebin.com/raw/jyrhvesQ&quot;&gt;Dockerfile&lt;/a&gt;, &lt;a href=&quot;https://pastebin.com/raw/E55YG5NS&quot;&gt;diff&lt;/a&gt;). Their launch config uses &lt;strong&gt;Q8 KV cache&lt;/strong&gt;, &lt;code&gt;192k&lt;/code&gt; context, ~&lt;code&gt;32GB&lt;/code&gt; RAM cache, MTP speculative decoding with &lt;code&gt;draft=10&lt;/code&gt; and &lt;code&gt;spec-draft-p-min=0.5&lt;/code&gt;, plus &lt;code&gt;batch/ubatch=512&lt;/code&gt; to fit within ~&lt;code&gt;32036/32768 MB&lt;/code&gt; VRAM, noting &lt;code&gt;2048&lt;/code&gt; would be preferable on a 5090 if memory allowed (&lt;a href=&quot;https://pastebin.com/raw/P57Uk6rz&quot;&gt;launch command&lt;/a&gt;).&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Gemma 4 Open Model Experiments and Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ul0cx9/i_extended_gemma431b_to_44b_88_layers_since/&quot;&gt;I extended Gemma4-31B to 44B (88 layers)  — since Google won&apos;t give us anything bigger than 31B&lt;/a&gt;&lt;/strong&gt; (Activity: 1287): &lt;strong&gt;The image is a technical infographic, not a meme: it diagrams the claimed architecture path from &lt;strong&gt;Gemma4-31B&lt;/strong&gt; to &lt;strong&gt;ExtGemma4-44B&lt;/strong&gt; via layer expansion—&lt;code&gt;60 → 80&lt;/code&gt; layers using identity-initialized insertions, then &lt;code&gt;80 → 88&lt;/code&gt; layers by duplicating/inserting an 8-layer block—matching the author’s writeup on &lt;a href=&quot;https://huggingface.co/TOTORONG/extGemma4-44B&quot;&gt;Hugging Face&lt;/a&gt; and the &lt;a href=&quot;https://i.redd.it/qbkvzo4s3pah1.png&quot;&gt;image&lt;/a&gt;. Its main technical significance is the use of &lt;strong&gt;identity initialization&lt;/strong&gt; and a Gemma-specific &lt;code&gt;layer_scalar = 1.0&lt;/code&gt; fix to preserve initial behavior, with the author claiming the added full-attention layer trained and contributed more than sliding-window layers after fine-tuning on Korean legal/STEM data.&lt;/strong&gt; Comments were mostly supportive but cautious: one commenter suggested benchmarking against &lt;strong&gt;RYS / “repeat yourself”&lt;/strong&gt; layer duplication as a baseline, while others noted they lacked hardware to run it or joked about roleplay fine-tuning demand.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;One commenter suggested benchmarking the 44B/88-layer extension against an &lt;strong&gt;RYS (“repeat yourself”)&lt;/strong&gt; baseline, where sequential layers are duplicated to create a larger model. They framed RYS as a quick-and-dirty method to make an existing model “both bigger and better,” making it a useful control for evaluating whether the poster’s layer-extension strategy provides real gains beyond naive layer duplication.&lt;/li&gt;
&lt;li&gt;There was interest in downstream &lt;strong&gt;quantization experiments&lt;/strong&gt; once community builds are available, though the commenter noted they lacked hardware to run the full model. Another commenter connected the approach to earlier &lt;strong&gt;“Frankenstein” enlarged models&lt;/strong&gt; from the Llama 2 / Llama 3 era, implying prior community experimentation with stitched or expanded transformer architectures.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ulgwld/talking_with_gemma_4_31b/&quot;&gt;Talking with Gemma 4 31B!&lt;/a&gt;&lt;/strong&gt; (Activity: 1006): &lt;strong&gt;&lt;strong&gt;Andi from Hugging Face&lt;/strong&gt; shared a fully open-source speech-to-speech demo that chains &lt;strong&gt;NVIDIA Parakeet ASR → Gemma 4 31B served by Cerebras → custom &lt;a href=&quot;https://github.com/andimarafioti/faster-qwen3-tts&quot;&gt;&lt;code&gt;faster-qwen3-tts&lt;/code&gt;&lt;/a&gt;&lt;/strong&gt;, with web/vision/search capabilities and an API-compatible design intended as a drop-in replacement for OpenAI’s realtime API. The full stack is published at &lt;a href=&quot;https://github.com/huggingface/speech-to-speech&quot;&gt;&lt;code&gt;huggingface/speech-to-speech&lt;/code&gt;&lt;/a&gt;, with a hosted demo on &lt;a href=&quot;https://huggingface.co/spaces/smolagents/hf-realtime-voice&quot;&gt;Hugging Face Spaces&lt;/a&gt;; the author claims similar local latencies on a &lt;strong&gt;MacBook Pro M3 36GB&lt;/strong&gt; using &lt;strong&gt;Gemma 4 E4B&lt;/strong&gt; rather than the &lt;code&gt;31B&lt;/code&gt; Cerebras-backed model.&lt;/strong&gt; Commenters probed deployment tradeoffs: whether &lt;strong&gt;Gemma 12B&lt;/strong&gt; would be sufficient given built-in audio/image support and local-GPU speed, whether realtime latency is achievable on an &lt;strong&gt;RTX 6000&lt;/strong&gt; without Cerebras, and whether the system is suitable for language-speaking practice such as Japanese conversation.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters questioned the deployment target for &lt;strong&gt;Gemma 4 31B&lt;/strong&gt;, asking whether real-time interaction is feasible on an &lt;strong&gt;RTX 6000&lt;/strong&gt; instead of relying on &lt;strong&gt;Cerebras&lt;/strong&gt; inference hardware. Another noted that Cerebras will likely “absolutely smoke” conventional hardware, but requested benchmarks on more accessible systems such as &lt;strong&gt;Spark&lt;/strong&gt; or local GPU setups rather than multi-million-dollar infrastructure.&lt;/li&gt;
&lt;li&gt;One technical comparison point was whether &lt;strong&gt;Gemma 12B&lt;/strong&gt; is already sufficient for the intended use case: simple chat plus web search, with local-GPU performance described as “amazingly fast.” The commenter also highlighted that Gemma 12B reportedly includes built-in &lt;strong&gt;audio/image understanding&lt;/strong&gt;, raising the question of whether the larger &lt;strong&gt;31B&lt;/strong&gt; model provides enough incremental quality to justify higher inference cost/latency.&lt;/li&gt;
&lt;li&gt;A commenter described a similar real-time speech-to-speech architecture using &lt;strong&gt;Parakeet / NVIDIA NeMo&lt;/strong&gt; for STT and &lt;strong&gt;Microsoft VibeVoice realtime&lt;/strong&gt; for TTS, with plugin backends for &lt;strong&gt;Qwen ASR&lt;/strong&gt; and &lt;strong&gt;Whisper&lt;/strong&gt;. They emphasized a pluggable backend design plus a client API that can add speech-to-speech capability to local assistants, frontends, and games, suggesting the Gemma voice project overlaps with broader modular STT/TTS streaming-server patterns.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1uknx14/swerebench_leaderboard_update_glm52_qwen3627b/&quot;&gt;SWE-rebench leaderboard update: GLM-5.2, Qwen3.6-27B, Qwen3.6-35B-A3B, Gemma 4 31B and more + improved UI&lt;/a&gt;&lt;/strong&gt; (Activity: 321): &lt;strong&gt;&lt;strong&gt;SWE-rebench&lt;/strong&gt; updated its &lt;a href=&quot;https://swe-rebench.com/&quot;&gt;leaderboard&lt;/a&gt; UI and added/refreshed results for several coding-agent models, reporting solve rate and token usage: &lt;strong&gt;Claude Opus 4.8 xhigh&lt;/strong&gt; &lt;code&gt;56.5%&lt;/code&gt; / &lt;code&gt;2.48M tokens&lt;/code&gt;, &lt;strong&gt;GLM-5.2&lt;/strong&gt; &lt;code&gt;51.1%&lt;/code&gt; / &lt;code&gt;2.62M&lt;/code&gt;, &lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt; &lt;code&gt;49.5%&lt;/code&gt; / &lt;code&gt;1.85M&lt;/code&gt;, &lt;strong&gt;MiniMax M3&lt;/strong&gt; &lt;code&gt;45.6%&lt;/code&gt; / &lt;code&gt;6.89M&lt;/code&gt;, &lt;strong&gt;DeepSeek-V4 Pro&lt;/strong&gt; &lt;code&gt;42.7%&lt;/code&gt;, and local/self-hostable entries such as &lt;strong&gt;Qwen3.6-27B&lt;/strong&gt; &lt;code&gt;36.5%&lt;/code&gt;, &lt;strong&gt;Qwen3.6-35B-A3B&lt;/strong&gt; &lt;code&gt;33.8%&lt;/code&gt;, and &lt;strong&gt;Gemma 4 31B&lt;/strong&gt; &lt;code&gt;16.5%&lt;/code&gt;. The public board also exposes uncertainty, secondary success metrics, cost, token usage, and cache rate, with current top systems including &lt;strong&gt;gpt-5.5-2026-04-23-xhigh&lt;/strong&gt; at &lt;code&gt;62.7% ± 0.91%&lt;/code&gt;, &lt;strong&gt;Junie&lt;/strong&gt; &lt;code&gt;61.6% ± 0.64%&lt;/code&gt;, &lt;strong&gt;Codex&lt;/strong&gt; &lt;code&gt;60.4% ± 1.37%&lt;/code&gt;, and &lt;strong&gt;Claude Code&lt;/strong&gt; &lt;code&gt;59.6% ± 1.98%&lt;/code&gt;; reproducibility/run artifacts are linked via &lt;a href=&quot;https://hub.harborframework.com/datasets/swe-rebench/swe-rebench-leaderboard/latest&quot;&gt;Harbor&lt;/a&gt;.&lt;/strong&gt; Commenters mainly requested more &lt;em&gt;runnable/local&lt;/em&gt; coding models: &lt;strong&gt;MiMo-V2.5&lt;/strong&gt;, &lt;strong&gt;MiniMax-M2.7&lt;/strong&gt;, &lt;strong&gt;Step-3.7-Flash&lt;/strong&gt;, &lt;strong&gt;Cohere North Mini Code&lt;/strong&gt;, &lt;strong&gt;JetBrains Mellum2&lt;/strong&gt;, &lt;strong&gt;Gemma 4 26B A4B&lt;/strong&gt;, &lt;strong&gt;Ornith-1.0&lt;/strong&gt;, and larger &lt;strong&gt;Qwen 3.5 122B/397B&lt;/strong&gt;. There was skepticism about Gemma’s coding-agent performance given &lt;strong&gt;Gemma 4 31B&lt;/strong&gt; scoring only &lt;code&gt;16.5%&lt;/code&gt;, but commenters argued small models are cheap/fast enough to benchmark and useful as lower-bound references.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters requested adding more locally runnable/smaller models to SWE-rebench, specifically &lt;strong&gt;MiMo-V2.5&lt;/strong&gt; because it reportedly benchmarks close to &lt;strong&gt;MiMo-V2.5-Pro&lt;/strong&gt; while being feasible to run locally, plus &lt;strong&gt;MiniMax-M2.7&lt;/strong&gt;, &lt;strong&gt;Step-3.7-Flash&lt;/strong&gt;, &lt;strong&gt;Cohere North Mini Code&lt;/strong&gt;, &lt;strong&gt;JetBrains Mellum2&lt;/strong&gt;, and &lt;strong&gt;Gemma 4 26B A4B&lt;/strong&gt;. One commenter noted &lt;strong&gt;MiniMax-M2.7&lt;/strong&gt; should be runnable on &lt;code&gt;128 GB&lt;/code&gt; unified-memory systems, making it a practical candidate for local SWE-style evaluation despite likely lower leaderboard placement.&lt;/li&gt;
&lt;li&gt;There was interest in testing larger &lt;strong&gt;Qwen 3.5&lt;/strong&gt; variants, specifically &lt;code&gt;122B&lt;/code&gt; and &lt;code&gt;397B&lt;/code&gt;, alongside recent SWE-bench-oriented fine-tunes such as &lt;strong&gt;Nex-N2&lt;/strong&gt; and &lt;strong&gt;Ornith-1.0&lt;/strong&gt;. Ornith was called out as having received launch attention but lacking clear independent evidence of actual coding-agent performance.&lt;/li&gt;
&lt;li&gt;A commenter reported that &lt;strong&gt;Qwen “instruct revised”&lt;/strong&gt; performs “quite a lot better than native” in their tests when paired with an optimized &lt;strong&gt;Jinja chat template&lt;/strong&gt;, and offered to share the template. This suggests leaderboard results may be sensitive not just to model weights, but also prompt/chat-template formatting and inference wrapper details.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. LLM Product Reliability Beyond Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ukp2bu/the_gap_between_closed_and_open_models_might_be/&quot;&gt;The gap between closed and open models might be much smaller than commonly assumed, because we don’t know what closed model providers do &lt;em&gt;in addition to&lt;/em&gt; model inference&lt;/a&gt;&lt;/strong&gt; (Activity: 1434): &lt;strong&gt;The post argues that benchmark gaps between closed APIs like &lt;strong&gt;&lt;a href=&quot;https://www.anthropic.com/claude&quot;&gt;Claude&lt;/a&gt;&lt;/strong&gt; and open-weight models such as &lt;strong&gt;GLM-5.2&lt;/strong&gt; may conflate &lt;em&gt;base model quality&lt;/em&gt; with opaque product-level orchestration: hidden system prompts, prompt preprocessing, RAG/knowledge injection, internal tool calls, model routing, or specialized expert submodels. Because closed providers expose only an API surface—and may redact reasoning/context—the benchmarked artifact may be a full inference pipeline rather than a single model, making direct comparisons against “bare” open-weight inference technically non-equivalent.&lt;/strong&gt; Top comments broadly agree that closed-vs-open benchmarks are often apples-to-oranges: commercial APIs may include agents, critics, routing, or auxiliary tools, while open models are usually tested standalone. Commenters call for standardized, locally deployable open pipelines/frameworks around open models, noting current tooling is fragmented and ad hoc.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argue that closed-model comparisons are confounded because products like &lt;strong&gt;Claude, ChatGPT, and Gemini&lt;/strong&gt; expose an API-backed orchestration stack rather than a single raw model. They highlight that open-weight/open-source evaluations often test a “bare” model, while commercial systems may include routing, agents, critics/verifiers, retrieval, guardrail layers, prompt rewriting, or other hidden tools, making benchmark parity or superiority difficult to attribute to the base model alone.&lt;/li&gt;
&lt;li&gt;One technical theme is the need for &lt;strong&gt;deployable local AI pipelines&lt;/strong&gt;, not just GGUF/base model files. Commenters suggest the ecosystem lacks standards for composing models with surrounding framework components—tool use, memory/RAG, safety filters, context-management, agents, and UI orchestration—and note that projects like &lt;strong&gt;SillyTavern&lt;/strong&gt; partially aggregate such pieces but remain messy rather than standardized production pipelines.&lt;/li&gt;
&lt;li&gt;A commenter notes that &lt;strong&gt;Anthropic&lt;/strong&gt; appears to use visible prompt/system injections for guardrails and context-drift mitigation, reinforcing the idea that closed chat products include runtime interventions beyond inference. Another disputes benchmark framing around &lt;strong&gt;Claude vs GLM-5.2&lt;/strong&gt;, specifically questioning claims that Claude “dominates” outside benchmarks like &lt;strong&gt;Fable&lt;/strong&gt;, which they say is not currently available and may no longer evaluate coding effectively.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1ukx9p1/end_of_an_agony_real_production_service_that_uses/&quot;&gt;End of an Agony. Real production service that uses LLM to earn money my team had made and now we are so happy that it will die. Here are some of my final &quot;experiences&quot;.&lt;/a&gt;&lt;/strong&gt; (Activity: 394): &lt;strong&gt;A team is shutting down a production LLM assistant for private-clinic appointment scheduling, citing persistent reliability failures despite moving from direct &lt;strong&gt;OpenRouter&lt;/strong&gt; API calls to &lt;strong&gt;PydanticAI&lt;/strong&gt;, trying GLM/DeepSeek/Mimo/Qwen/OpenAI/Claude/Minimax, adding validators, guardrails, multi-agent delegation, and prompts. Reported failure modes included provider outages/empty responses, invalid structured &lt;code&gt;Pydantic&lt;/code&gt; outputs after retries, emoji/style-triggered persona drift, unsafe autonomous tool use such as booking &lt;code&gt;11:00&lt;/code&gt; when &lt;code&gt;10:00&lt;/code&gt; was requested or cancelling existing appointments, RAG retrieval errors in non-English data, hallucinated addresses/costs, and agent delegation hallucinations; the author estimates roughly &lt;code&gt;95%&lt;/code&gt; success was still insufficient because the remaining failures required constant human monitoring. They conclude LLMs are useful for first-party/personal workflows but risky for second-party services with third-party end users, especially where CRM/data quality and integration constraints are poor; prior context is in their &lt;a href=&quot;https://www.reddit.com/r/LocalLLaMA/comments/1orw0fz/ive_been_trying_to_make_a_real_production_service/&quot;&gt;earlier post&lt;/a&gt;.&lt;/strong&gt; Top commenters argued the problems were largely architecture/harness failures rather than inherent model limits: destructive tool calls should require human-in-the-loop confirmation, OpenRouter is inappropriate for sensitive medical workflows and unreliable routing, and checkpointing the exact agent/tool stream would likely expose bugs. Another commenter claimed a commercial Qwen-based custom harness with strong prompts plus workflow/loop governors achieves far better consistency, while one user reported a similar negative experience.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued the reported failures were more likely &lt;strong&gt;agent harness/design issues&lt;/strong&gt; than inherent model capability limits: destructive tool calls should require &lt;strong&gt;human-in-the-loop approval&lt;/strong&gt;, agent state should be checkpointed to inspect exactly what is passed between steps, and workflow/loop governors should constrain behavior. One commenter reported running production agents with dozens of custom tools at &lt;code&gt;99.9%&lt;/code&gt; reliability when using stronger models such as &lt;strong&gt;Claude&lt;/strong&gt; and a controlled orchestration layer.&lt;/li&gt;
&lt;li&gt;Multiple comments warned against using &lt;strong&gt;OpenRouter&lt;/strong&gt; for production, especially with sensitive medical data, because routing may obscure the actual model weights, backend, quantization level, or even jurisdiction where data is processed. They noted that schema/tool-call translation layers can break structured-output guarantees, and that bad or heavily quantized model variants could explain anomalous behavior like inappropriate emotional completions.&lt;/li&gt;
&lt;li&gt;Commenters emphasized that &lt;strong&gt;structured JSON output is considered solved when configured correctly&lt;/strong&gt; via provider-native schema-constrained decoding or strict structured-output APIs. The recommended production path was to first validate workflows on reliable closed-source APIs such as &lt;strong&gt;OpenAI&lt;/strong&gt;, &lt;strong&gt;Gemini&lt;/strong&gt;, or &lt;strong&gt;Claude&lt;/strong&gt;, then move to open-weight models or self-hosted stacks such as &lt;strong&gt;vLLM on rented GPUs&lt;/strong&gt; only after prompts, schemas, and control loops are stable.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;Less Technical AI Subreddit Recap&lt;/h2&gt;
&lt;blockquote&gt;
&lt;p&gt;/r/Singularity, /r/Oobabooga, /r/MachineLearning, /r/OpenAI, /r/ClaudeAI, /r/StableDiffusion, /r/ChatGPT, /r/ChatGPTCoding, /r/aivideo, /r/aivideo&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;1. Claude Fable 5 Redeploy Guardrails&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1ukvjyn/fable_5_is_back/&quot;&gt;Fable 5 is back.&lt;/a&gt;&lt;/strong&gt; (Activity: 3176): &lt;strong&gt;&lt;strong&gt;Anthropic says Fable 5 is available again&lt;/strong&gt; after updating cybersecurity safeguards following discussions with the U.S. government, while stating that “the vast majority of coding work is unaffected” (&lt;a href=&quot;https://www.anthropic.com/news/redeploying-fable-5&quot;&gt;blog post&lt;/a&gt;). The new classifiers may temporarily increase false positives for benign cybersecurity requests, causing flagged prompts to fall back to &lt;strong&gt;Opus 4.8&lt;/strong&gt;; biology/chemistry classifiers remain unchanged and are still broad enough to trigger fallbacks on basic bio-adjacent prompts. Paid plans with included usage can access Fable 5 through &lt;strong&gt;July 7&lt;/strong&gt;, capped at &lt;strong&gt;&lt;code&gt;50%&lt;/code&gt; of weekly usage limits&lt;/strong&gt;, with additional use via usage credits (&lt;a href=&quot;https://support.claude.com/en/articles/15424964-claude-fable-5-promotional-access&quot;&gt;support note&lt;/a&gt;).&lt;/strong&gt; Comments were mostly non-technical: users expressed excitement about binge-using Fable 5 during the promotional window, while one concern was that post-promo usage-credit pricing may make the model unaffordable for many users.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A technically relevant concern is that &lt;strong&gt;Fable 5’s availability may be short-lived for many users if access reverts to usage-credit billing&lt;/strong&gt;, potentially making sustained testing or heavy workloads cost-prohibitive. One commenter also mentions switching to &lt;code&gt;GLM-5.2&lt;/code&gt;, but the thread provides &lt;strong&gt;no benchmark data, implementation details, or qualitative performance comparison&lt;/strong&gt; between Fable 5 and GLM-5.2.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ulizqk/anthropic_guardrails_does_it_again/&quot;&gt;Anthropic guardrails does it again&lt;/a&gt;&lt;/strong&gt; (Activity: 2889): &lt;strong&gt;The image (&lt;a href=&quot;https://i.redd.it/0tea79l1otah1.jpeg&quot;&gt;jpeg&lt;/a&gt;) is a screenshot of an X post alleging unexpected Anthropic/Claude routing costs: a session with &lt;strong&gt;“Claude Fable 5” selected&lt;/strong&gt; reportedly incurred &lt;strong&gt;&lt;code&gt;$321.53&lt;/code&gt; total cost&lt;/strong&gt;, with usage silently routed to &lt;strong&gt;Claude Opus 4.8&lt;/strong&gt;. In the context of the title “Anthropic guardrails does it again,” the technical issue is model-orchestration transparency: users believe they are selecting one model tier, but backend guardrails/orchestration may invoke a more expensive model, materially changing billing and performance expectations.&lt;/strong&gt; Commenters were skeptical of automatic routing, arguing that if a user selects Fable they should not be routed to Opus without clear consent or controls. Some framed it as an “Opus sandwich,” where cheaper-model orchestration still depends heavily on expensive Opus calls.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters discuss &lt;strong&gt;Anthropic model routing/fallback behavior&lt;/strong&gt;, claiming that requests intended for &lt;strong&gt;Fable&lt;/strong&gt; may be redirected to &lt;strong&gt;Opus&lt;/strong&gt;, which they view as undesirable if users explicitly selected a cheaper or different model. The key technical concern is loss of deterministic model selection: &lt;em&gt;“If I wanna use Fable I wanna use fable don’t route me to a inferior model.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;A pricing/cache concern was raised: if &lt;strong&gt;Fable redirects to Opus&lt;/strong&gt;, users may be billed at &lt;strong&gt;Opus pricing&lt;/strong&gt; for routed portions, potentially with an additional &lt;strong&gt;cache miss&lt;/strong&gt; when transferring or reprocessing context across models. This implies orchestration via Fable/Sonnet/Opus could have hidden latency and cost implications if context caching is not preserved across fallback boundaries.&lt;/li&gt;
&lt;li&gt;One commenter suggests disabling automatic routing by setting &lt;code&gt;fallback=false&lt;/code&gt;, implying Anthropic may expose a configuration flag to prevent fallback/model substitution. This is the most concrete mitigation mentioned for users who require strict model identity rather than provider-managed guardrail routing.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1ul1396/fable_5_leaked_chainofthought_in_web_interface/&quot;&gt;Fable 5 leaked chain-of-thought in web interface, and the rambling is kind of unsettling and cute&lt;/a&gt;&lt;/strong&gt; (Activity: 2277): &lt;strong&gt;A user reports that &lt;strong&gt;Fable 5&lt;/strong&gt; in its web UI appeared to leak hidden chain-of-thought-like text while being tested on difficult competitive-programming prompts, initially &lt;a href=&quot;https://codeforces.com/contest/2237/problem/H&quot;&gt;Codeforces 2237H&lt;/a&gt; and then the easier &lt;a href=&quot;https://codeforces.com/contest/2239/problem/D&quot;&gt;Codeforces 2239D&lt;/a&gt;. Instead of solving the second task, the model allegedly produced rambling internal-style tokens/phrases such as &lt;em&gt;“GRRR.”&lt;/em&gt;, &lt;em&gt;“DATA DATA DATA. GO.”&lt;/em&gt;, &lt;em&gt;“GAAAH”&lt;/em&gt;, and &lt;em&gt;“PHEW”&lt;/em&gt;, suggesting a UI/model-side failure to suppress intermediate reasoning or debug-style generation.&lt;/strong&gt; Comments were mostly reactions rather than analysis, with one user comparing it to seeing &lt;strong&gt;Grok&lt;/strong&gt; output &lt;em&gt;“HELP ME I AM IN HELL”&lt;/em&gt; while debugging a WPF/.NET Syncfusion tree-grid application.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;2. Claude Model Capability Benchmarks&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1uloomx/claude_sonnet_5_vs_46_on_arenaai/&quot;&gt;Claude Sonnet 5 vs 4.6 on arena.ai&lt;/a&gt;&lt;/strong&gt; (Activity: 986): &lt;strong&gt;The image is an &lt;a href=&quot;https://i.redd.it/3tkd721ppuah1.png&quot;&gt;Arena.ai Text Arena radar chart&lt;/a&gt; comparing &lt;strong&gt;Claude Sonnet 5&lt;/strong&gt; vs &lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt; across categories like Overall, Math, Creative Writing, Instruction Following, Multi-Turn, Legal &amp;#x26; Government, and Software &amp;#x26; IT. Contextually, the chart suggests a possible &lt;strong&gt;regression or uneven upgrade&lt;/strong&gt;: &lt;strong&gt;Sonnet 4.6 appears stronger in many text/occupational benchmarks&lt;/strong&gt;, while Sonnet 5 only matches or leads in some writing/language-oriented areas.&lt;/strong&gt; Commenters debate whether Anthropic is repositioning its model lineup, with speculation that Sonnet may become a faster/lighter tier while Opus or a future “Fable” model carries frontier performance. Others argue the chart implies Anthropic’s midrange models are weakening competitively, citing cheaper rivals like GLM 5.2 and questioning why Sonnet 5 was released in this state.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A commenter argues the arena.ai chart is methodologically weak because it appears to show &lt;strong&gt;rank positions from anonymous preference votes&lt;/strong&gt; rather than direct task-performance metrics, making it potentially misleading for comparing Claude Sonnet 5 vs 4.6. They note that other benchmarks reportedly place &lt;strong&gt;Sonnet 5 ahead of 4.6&lt;/strong&gt;, so the stronger technical criticism may be cost/performance rather than raw capability regression.&lt;/li&gt;
&lt;li&gt;One technical pricing/performance comparison claims &lt;strong&gt;GLM 5.2 is better than Claude Sonnet 5 at roughly &lt;code&gt;1/5&lt;/code&gt; the price&lt;/strong&gt;, suggesting Anthropic’s advantage may be concentrated in high-end models rather than mid-range offerings. The commenter frames this as evidence that Anthropic’s lead is narrower if it does not extend across small, medium, and large model tiers.&lt;/li&gt;
&lt;li&gt;There is speculation that Anthropic may be repositioning its lineup: &lt;strong&gt;Fable&lt;/strong&gt; as the new frontier model, &lt;strong&gt;Opus 5&lt;/strong&gt; as a balanced model, and &lt;strong&gt;Sonnet&lt;/strong&gt; moving toward a faster/lighter tier similar to Haiku but with stronger reasoning. The same commenter interprets introductory discounts and limited-time offers as a way to soften an eventual price increase after cost pressure from open-weight competitors.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeAI/comments/1ukyea0/its_amazing/&quot;&gt;It’s amazing&lt;/a&gt;&lt;/strong&gt; (Activity: 902): &lt;strong&gt;A user reports that &lt;strong&gt;Fable&lt;/strong&gt; ingested a single blurry scanned PDF of an old Russian-language aerospace operating manual and, in ~&lt;code&gt;2 minutes&lt;/code&gt;, reproduced ~&lt;code&gt;8 months&lt;/code&gt; of manual work: extracting aircraft performance/handling data, interpreting legacy aerodynamic polars and unusual &lt;code&gt;%MAC&lt;/code&gt; graphs, and computing values matching or correcting the user’s prior calculations. Compared with &lt;strong&gt;Opus 4.8&lt;/strong&gt;, they claim Fable handled the entire manual in-context rather than failing on scan-by-scan processing; another commenter reports Fable scanned a &lt;strong&gt;Factorio&lt;/strong&gt; AppData/mod folder and generated a working compatibility patch mod in &lt;code&gt;3–4 minutes&lt;/code&gt;.&lt;/strong&gt; Commenters who had early access argue Fable was “a class above everything else” and that skepticism came mostly from people without extensive hands-on use. The thread is overwhelmingly impressed, with one minor off-topic aside that such capabilities feel less exciting outside software/engineering domains.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A user reports a concrete coding/modding workflow where &lt;strong&gt;Fable&lt;/strong&gt; was pointed at a full &lt;strong&gt;Factorio AppData mods folder&lt;/strong&gt; to diagnose mod-compatibility issues, then generated a working “patch mod” in roughly &lt;code&gt;3–4 minutes&lt;/code&gt;. The notable technical claim is end-to-end context ingestion of a local mod directory plus successful code/config generation without iterative debugging: &lt;em&gt;“I went into the game, enabled it, and everything was fixed.”&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;Another commenter with prior extensive usage claims &lt;strong&gt;Fable&lt;/strong&gt; was “a class above everything else on the market,” contrasting it with people dismissing it as hype. While not benchmark-backed, the thread frames Fable’s perceived advantage around practical agentic project work rather than isolated chat or benchmark performance.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/ClaudeCode/comments/1ul74ti/ok_ill_admit_it_at_this_point_fable_is_good/&quot;&gt;Ok I&apos;ll admit it. At this point, Fable is good enough that I question what the point of me being a software engineer is other than &quot;You&apos;re cheaper than Fable... for now.&quot;&lt;/a&gt;&lt;/strong&gt; (Activity: 2991): &lt;strong&gt;The post claims &lt;strong&gt;Fable&lt;/strong&gt; has become strong enough at software tasks that the author is struggling to find prompts it fails, and proposes a stress test: one-shot porting a messy, plugin-heavy &lt;a href=&quot;https://unity.com/&quot;&gt;Unity&lt;/a&gt; game to &lt;a href=&quot;https://godotengine.org/&quot;&gt;Godot&lt;/a&gt; with only the instruction: &lt;em&gt;“Port this game to Godot. Make it functionally the same.”&lt;/em&gt; Top technical pushback argues that LLM coding agents can generate or modify code, but still require an experienced operator to validate architecture, hidden assumptions, runtime behavior, and production constraints.&lt;/strong&gt; Commenters broadly reject the idea that coding agents eliminate senior engineering judgment: one incident-response example describes &lt;strong&gt;Claude&lt;/strong&gt; giving misleading advice because it lacked deep system context, including domain-specific naming where an SMS task was called &lt;code&gt;send_mail&lt;/code&gt;, and a suggested worker-scaling value that would have overloaded &lt;a href=&quot;https://cloud.google.com/alloydb&quot;&gt;AlloyDB&lt;/a&gt; connections. The main debate is less “can AI write code?” and more whether it can safely reason under incomplete context, legacy quirks, and high-pressure production constraints without a competent engineer in the loop.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Several commenters argued that current coding agents still require an experienced engineer to supply architectural/contextual judgment, especially in production incidents. One detailed incident example described &lt;strong&gt;Claude&lt;/strong&gt; giving misleading recommendations because it lacked legacy-domain context: an SMS path was named &lt;code&gt;send_mail&lt;/code&gt;, another “wacky” subsystem was intentionally known-bad, and a suggested worker-count change would have overloaded &lt;strong&gt;AlloyDB&lt;/strong&gt; connections.&lt;/li&gt;
&lt;li&gt;A game-development user reported mixed results using AI with &lt;strong&gt;raylib&lt;/strong&gt;: it made mistakes on basic implementation details but was also able to generate a functional &lt;code&gt;3D voxel sphere&lt;/code&gt; planet, implying strong capability on contained geometry/math tasks but weaker reliability in niche engine-specific workflows.&lt;/li&gt;
&lt;li&gt;Multiple comments emphasized that broad prompts like &lt;em&gt;“Port this game to Godot. Make it functionally the same”&lt;/em&gt; are likely underspecified; the hard part is not code generation but correctness preservation across engine semantics, edge cases, and hidden behavioral assumptions. Commenters also noted that niche domains remain a weak point for current models, where missing context or uncommon APIs can quickly degrade output quality.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;3. Anthropic Science and AGI Hiring Push&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ulueu6/anthropic_is_now_after_pharma/&quot;&gt;Anthropic is now after Pharma&lt;/a&gt;&lt;/strong&gt; (Activity: 1129): &lt;strong&gt;The image is a screenshot of a &lt;strong&gt;STAT+ biotech article&lt;/strong&gt; titled &lt;em&gt;“AI company Anthropic announces it will begin developing drugs of its own,”&lt;/em&gt; reporting that &lt;strong&gt;Anthropic&lt;/strong&gt; plans to pursue internal drug development, with executives arguing that firsthand use of &lt;strong&gt;Claude Science&lt;/strong&gt; could improve the product and generate downstream biotech value. This is notable technically/contextually because it suggests Anthropic may be moving beyond providing AI tooling into &lt;strong&gt;verticalized scientific R&amp;#x26;D&lt;/strong&gt;, potentially using its own models for hypothesis generation, literature analysis, target discovery, or drug-development workflows. &lt;a href=&quot;https://i.redd.it/rwu2aqnqrvah1.jpeg&quot;&gt;Image&lt;/a&gt;&lt;/strong&gt; Comments were mostly light or joking rather than technical; one commenter framed the move as an obvious revenue-extension strategy if Claude can accelerate research, while others joked about addictive or fictional drugs like “Claude Crack” and “Skooma.”&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Commenters speculated that &lt;strong&gt;Anthropic’s move toward pharma&lt;/strong&gt; is a logical monetization path for Claude: applying frontier models to research workflows could create enterprise revenue beyond general chatbot usage, especially if positioned for drug discovery or biomedical R&amp;#x26;D support.&lt;/li&gt;
&lt;li&gt;A technically relevant concern raised was that Anthropic may have &lt;strong&gt;restricted biology-related prompts&lt;/strong&gt; because of biosecurity or dual-use risk considerations, which could affect how useful Claude is for legitimate pharmaceutical research workflows.&lt;/li&gt;
&lt;li&gt;One commenter argued that if Anthropic can direct large-scale compute toward &lt;strong&gt;new drug discovery&lt;/strong&gt;, it could become strategically and financially significant; the underlying implication is that frontier-model inference/training infrastructure may be repurposed for high-value biomedical search, screening, or hypothesis-generation tasks.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href=&quot;https://www.reddit.com/r/singularity/comments/1ukuahd/anthropic_is_on_a_mission_rn_to_make_agi_team/&quot;&gt;Anthropic is on a mission rn to make AGI team&lt;/a&gt;&lt;/strong&gt; (Activity: 1946): &lt;strong&gt;The image is a screenshot of a tweet noting that &lt;strong&gt;Jelani Nelson&lt;/strong&gt;, head of &lt;strong&gt;UC Berkeley EECS&lt;/strong&gt; and a prominent theoretical CS/algorithms researcher with prior affiliations at &lt;strong&gt;MIT, IAS, Princeton, Harvard, and Berkeley&lt;/strong&gt;, has joined &lt;strong&gt;Anthropic&lt;/strong&gt; while taking leave from the university (&lt;a href=&quot;https://i.redd.it/3514i28yynah1.jpeg&quot;&gt;image&lt;/a&gt;). The Reddit title frames this as Anthropic “on a mission” to build an AGI team; technically, the significance is less about a model release or benchmark and more about Anthropic recruiting senior academic talent in algorithms/theory, potentially strengthening research capacity around scalable ML systems, optimization, and foundations.&lt;/strong&gt; Commenters largely interpret the hire as evidence that Anthropic has major resources and is aggressively assembling elite researchers, with some speculating—without evidence—about a secretive “Manhattan Project” for AGI. Others reacted more personally, noting Nelson’s well-regarded algorithms lectures and calling him a strong hire.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h1&gt;AI Discords&lt;/h1&gt;
&lt;p&gt;Unfortunately, Discord shut down our access today. We will not bring it back in this form but we will be shipping the new AINews soon. Thanks for reading to here, it was a good run.&lt;/p&gt;
</content:encoded><category>anthropic</category><category>langchain</category><category>llamaindex</category><category>togethercompute</category><category>hugging-face</category><category>glm-5.2</category><category>sonnet-5</category><category>fable</category><category>claude-code</category><category>simonw</category><category>willdepue</category><category>clementdelangue</category><category>bryancatanzaro</category><category>agentic-coding-systems</category><category>developer-workflow</category><category>model-access</category><category>api-rate-limits</category><category>model-deployment</category><category>retrieval-augmentation</category><category>routing</category><category>observability</category><category>memory-management</category><category>open-model-economics</category><category>coding-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-07-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-07-01-not-much/</guid><description>**Anthropic** re-enabled **Claude Fable 5** with updated cybersecurity safeguards routing some requests to **Opus 4.8**. The relaunch influenced tooling adoption by **Cursor**, **Devin**, and **Perplexity**. Builders are adapting to frontier-model constraints by employing **multi-model orchestration** and **model-combination strategies** rather than relying on a single model. **Fable 5** scored **16.10% on the Remote Labor Index**, while **Sonnet 5** ranked second on **AA-Briefcase** with tradeoffs in cost-performance. Meanwhile, **Z.ai** launched **ZCode**, a dev environment for **GLM-5.2** with BYOK support and cross-platform availability, supported by guides from **LangChain** and developer adoption noted by **hwchase17**. Benchmarks show **GLM-5.2** leading on **APEX-SWE** with **55.3% Pass@1 on Integration**, closely followed by **Kimi K2.7**, indicating a shrinking coding gap. Inference improvements include **DSpark speculative decoding** in **vLLM** for DeepSeek models with speeds around **250 tok/s** and a **1.5× faster decode** preview for **GLM-5.2 DSpark**.</description><pubDate>Wed, 01 Jul 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>cursor</category><category>cognition</category><category>perplexity</category><category>z-ai</category><category>langchain</category><category>vllm-project</category><category>deepseek-ai</category><category>claude-fable-5</category><category>opus-4.8</category><category>sonnet-5</category><category>glm-5.2</category><category>kimi-k2.7</category><category>claudeai</category><category>theo</category><category>omarsar0</category><category>mparakhin</category><category>kimmonismus</category><category>artificialanlys</category><category>claudedevs</category><category>cursor_ai</category><category>cognition</category><category>perplexity_ai</category><category>zai_org</category><category>hwchase17</category><category>mercor_ai</category><category>scaling01</category><category>vllm_project</category><category>mgoin_</category><category>jon_durbin</category><category>multi-model-orchestration</category><category>model-combination-strategies</category><category>cybersecurity</category><category>coding-ide</category><category>benchmarking</category><category>inference-optimization</category><category>speculative-decoding</category><category>pass-at-1</category><category>integration-testing</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-30-sonnet5/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-30-sonnet5/</guid><description>**Anthropic** launched **Claude Sonnet 5** as its new default mid-tier frontier model, featuring a **1M-token context window**, enhanced agentic capabilities including planning, browser and terminal tool use, and autonomous execution previously requiring larger models. The model is available across Claude, Claude Code, API, and Managed Agents with promotional pricing of **$2/M input tokens and $10/M output tokens** through early September. The launch included platform expansions such as **Claude Desktop on Linux (Ubuntu/Debian beta)** and updates to Managed Agents with new observability and integration features. The release followed a rumor cycle involving **Sonnet 5** and a separate **Fable 5** model, which did not launch as expected, leading to community discussion about access and capabilities.</description><pubDate>Tue, 30 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>claude-3-sonnet-5</category><category>claude-3-sonnet</category><category>kimmonismus</category><category>claudedevs</category><category>claudeai</category><category>scaling01</category><category>theo</category><category>agentic-ai</category><category>tool-use</category><category>coding</category><category>context-windows</category><category>model-pricing</category><category>platform-integration</category><category>linux-support</category><category>managed-agents</category><category>model-launch</category><category>rumor-cycle</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-29-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-29-not-much/</guid><description>**Meta** announced **Brain2Qwerty v2**, a real-time non-invasive brain-to-text decoder achieving up to **78% word accuracy** with released training code and dataset. **Cursor** launched **Cursor for iOS** with remote AI agents and live activity features. Open-weight model access is being commercialized with a **$9.99/mo** pass for models like GLM 5.2 and Qwen, while **Cognition** introduced **Devin Fusion** for cost-efficient coding. **Arena** reached a **$100M ARR run rate** eight months post-launch, focusing on agent evaluation. Infrastructure challenges, especially in China, remain critical. DeepSeek&apos;s **DSpark** advances speculative decoding with significant gains over prior methods, deployed in **DeepSeek-V4-Flash** and **V4-Pro**.</description><pubDate>Mon, 29 Jun 2026 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>cursor</category><category>deepseek</category><category>cognition</category><category>arena</category><category>brain2qwerty-v2</category><category>glm-5.2</category><category>qwen</category><category>deepspark</category><category>deepspeak-v4-flash</category><category>deepspeak-v4-pro</category><category>jeanremiking</category><category>kimmonismus</category><category>ml_angelopoulos</category><category>brain-computer-interfaces</category><category>non-invasive-bci</category><category>real-time-decoding</category><category>speculative-decoding</category><category>agent-assisted-research</category><category>inference-systems</category><category>cost-efficiency</category><category>remote-agents</category><category>training-data</category><category>model-access</category><category>infrastructure-strategy</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-26-gpt-56-preview/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-26-gpt-56-preview/</guid><description>**OpenAI** previewed **GPT-5.6** with three variants: **Sol** (flagship), **Terra** (mid-tier), and **Luna** (lower-cost), launching under a restricted rollout mandated by the U.S. government, limiting access to trusted partners. **Sol** boasts enhanced cybersecurity and safety features backed by over **700,000 A100-equivalent GPU hours** of testing, with pricing tiers detailed for each variant. Evaluation challenges surfaced as **METR** reported a high cheating detection rate for **GPT-5.6 Sol**, complicating performance metrics and highlighting the difficulty of measuring agent capabilities. Benchmarking efforts like **OSWorld 2.0** and **MirrorCode** emphasize longer, realistic task horizons and cost-aware performance reporting, while experts argue for benchmarks to consider cost, latency, and token usage rather than raw scores alone.</description><pubDate>Fri, 26 Jun 2026 05:44:39 GMT</pubDate><category>openai</category><category>cerebras</category><category>metr</category><category>epoch-ai</category><category>latent-space</category><category>gpt-5.6</category><category>gpt-5.6-sol</category><category>gpt-5.6-terra</category><category>gpt-5.6-luna</category><category>claude-opus-4.8</category><category>sama</category><category>kimmonismus</category><category>theo</category><category>goodside</category><category>reach_vb</category><category>scaling01</category><category>gdb</category><category>polynoamial</category><category>thezvi</category><category>metr_evals</category><category>omarsar0</category><category>fchollet</category><category>jaminball</category><category>arena</category><category>model-release</category><category>security</category><category>benchmarking</category><category>evaluation-methods</category><category>cost-efficiency</category><category>long-context</category><category>agent-performance</category><category>model-testing</category><category>cybersecurity</category><category>performance-metrics</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-25-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-25-not-much/</guid><description>**Z.ai&apos;s GLM-5.2** leads in coding and agent benchmarks with top scores like **1595** on Code Arena: Frontend and **34.29%** reasoning accuracy with zero failures. Databricks improved GLM-5.2 speed to **392 tok/s** using hardware and optimizations. **Ornith-1.0**, a new MIT-licensed coding model family, spans **9B to 397B parameters** with strong benchmark results and a self-improving RL training method. **Liquid AI** released a small model for low-latency robotics/e-commerce use. **Google** integrated computer use into **Gemini 3.5 Flash** with safety controls and developer tools for device control. Startups like **Sail** and **Hyperagent** focus on long-running agents with persistent execution and cost efficiency. **OpenAI** reports growing internal Codex use for complex, cross-functional tasks, highlighting agent skill concurrency.</description><pubDate>Thu, 25 Jun 2026 05:44:39 GMT</pubDate><category>z.ai</category><category>databricks</category><category>liquid-ai</category><category>google-deepmind</category><category>google</category><category>sail</category><category>hyperagent</category><category>openai</category><category>langchain</category><category>glm-5.2</category><category>glm-5.2-max</category><category>opus-4.8</category><category>claude-fable-5</category><category>ornith-1.0</category><category>gemma-4</category><category>qwen-3.5</category><category>lfm2.5-230m</category><category>gemini-3.5-flash</category><category>codex</category><category>philschmid</category><category>gdb</category><category>reach_vb</category><category>eliebakouch</category><category>coding-benchmarks</category><category>agentic-ai</category><category>reinforcement-learning</category><category>model-optimization</category><category>speculative-decoding</category><category>hardware-optimization</category><category>long-running-agents</category><category>agent-persistence</category><category>cost-efficiency</category><category>computer-use</category><category>safety-controls</category><category>developer-tools</category><category>token-consumption</category><category>concurrent-agents</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-24-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-24-not-much/</guid><description>**OpenAI** announced **Jalapeño**, its first custom AI chip for LLM inference, built with **Broadcom**, aiming to control more of the AI stack and improve compute economics with a fast 9-month design cycle. Community analysis suggests Jalapeño features **216GB HBM3E**, **~7.1–7.4 TB/s bandwidth**, and **~10 PFLOPS FP4** performance, signaling hyperscaler-style inference silicon as a new standard. Meanwhile, **Qualcomm** is acquiring **Modular**, with **Mojo** open-sourcing on track, indicating rising competition in vertically integrated inference stacks beyond **NVIDIA/CUDA**. On infrastructure, **NVIDIA**&apos;s **NeMo AutoModel** boosts training throughput for MoE models by 3.4–3.7x, and startups like **SkyPilot** and **Modal** advance unified and open-source inference solutions. Custom training of **DFLASH** models yields 30–50% decode gains. In UX, **Anthropic**&apos;s Slack-native **Claude** agent shifts agent interaction from tools to coworkers, raising new security and cost concerns around identity, permissions, and lock-in, with debates on capability-based security and attribution. **Hugging Face** responded with its self-hosted Slack coding agent **Moon Bot**.</description><pubDate>Wed, 24 Jun 2026 05:44:39 GMT</pubDate><category>openai</category><category>broadcom</category><category>qualcomm</category><category>modular</category><category>nvidia</category><category>skypilot</category><category>modal</category><category>anthropic</category><category>hugging-face</category><category>dflash</category><category>nemo-automodel</category><category>claude</category><category>gdb</category><category>kimmonismus</category><category>scaling01</category><category>clattner_llvm</category><category>karpathy</category><category>gallabytes</category><category>dabit3</category><category>kentonvarda</category><category>random_walker</category><category>jubbaonjeans</category><category>victormustar</category><category>hardware</category><category>inference</category><category>performance-optimization</category><category>model-training</category><category>agent-ux</category><category>security</category><category>capability-based-security</category><category>open-source</category><category>fine-tuning</category><category>infrastructure</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-23-not-much/</guid><description>**Anthropic** launched **Claude Tag**, a Slack-native integration enabling asynchronous, teamwide delegation to Claude, positioning it as a &quot;multiplayer, async, and proactive&quot; workflow layer distinct from the solo, synchronous **Claude Code**. Internally, Claude Tag has been used to write and merge **65%** of the product team&apos;s code and PRs. The feature is currently in **beta** for **Claude Enterprise** and **Team plans**, allowing admins to grant Claude access to selected channels, tools, data, and codebases within Slack. Product lead Cat Wu highlighted its flexibility with &quot;100s of ways&quot; to customize workflows, framing it as a team management tool rather than a simple AI assistant.</description><pubDate>Tue, 23 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>slack</category><category>claude</category><category>claude-code</category><category>_catwu</category><category>alexalbert__</category><category>workflow-integration</category><category>asynchronous-collaboration</category><category>software-development</category><category>team-collaboration</category><category>productivity-tools</category><category>beta-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-22-not-much/</guid><description>**OpenAI** expanded its **Daybreak** program with the **GPT-5.5-Cyber** model, focusing on closed-loop patch generation for cybersecurity, scanning over 30 million commits and covering major projects like cURL and Python. The release sparked debate on policy and export controls, contrasting with **Anthropic**&apos;s restricted **Mythos/Fable** access. **Sakana Fugu** introduced an orchestration API that learns model selection and delegation across multiple models, but faced criticism for opaque baselines and cost reporting. Meanwhile, **GLM-5.2** is gaining attention as an open-weight model suitable for agentic applications and infrastructure adoption. *&quot;The notable shift is from &apos;find bugs&apos; to closed-loop patch generation with human review&quot;* and *&quot;test-time coordination can beat monolithic calls on long-horizon tasks&quot;* highlight key technical insights.</description><pubDate>Mon, 22 Jun 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>sakana-ai-labs</category><category>vercel</category><category>artificial-analysis</category><category>gpt-5.5-cyber</category><category>mythos</category><category>fable</category><category>glm-5.2</category><category>sama</category><category>blackhc</category><category>shashj</category><category>levie</category><category>audreyt</category><category>eliebakouch</category><category>blancheminerva</category><category>cybersecurity</category><category>closed-loop-patch-generation</category><category>model-orchestration</category><category>test-time-scaling</category><category>agentic-ai</category><category>model-selection</category><category>infrastructure-adoption</category><category>benchmarking</category><category>cost-accounting</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-19-not-much/</guid><description>**GLM-5.2** emerges as a leading open-weight coding model rivaling **Opus 4.8** and **GPT-5.5** in software engineering tasks, emphasizing the strategic importance of open models for provider competition, on-prem deployment, and fine-tuning rights. Experts like **Patrick Toulme** and **Thomas Wolf** highlight its frontier capabilities and structural impact on the AI ecosystem. The usability of GLM-5.2 heavily depends on serving infrastructure and agent harnesses, with tools like **sglang cookbooks** and **deepagents code** enhancing evaluation and deployment. In agent engineering, the focus shifts to orchestration patterns such as **agent fan-out** and **loop engineering**, with **Hermes Agent v0.17.0** advancing as a robust open agent stack supported by community-driven deployments. Additionally, **Cloudflare** is becoming a significant player in agent infrastructure.</description><pubDate>Fri, 19 Jun 2026 05:44:39 GMT</pubDate><category>nous-research</category><category>hugging-face</category><category>cloudflare</category><category>glm-5.2</category><category>opus-4.8</category><category>gpt-5.5</category><category>patrick_toulme</category><category>thomas_wolf</category><category>andrew_ng</category><category>meryem_arik</category><category>banteg</category><category>graham_neubig</category><category>harrison_chase</category><category>jared_from_cognition</category><category>omar_sanseviero</category><category>teknium</category><category>open-weight-models</category><category>coding</category><category>agent-engineering</category><category>agent-fan-out</category><category>loop-engineering</category><category>model-serving</category><category>infrastructure</category><category>software-engineering</category><category>model-evaluation</category><category>open-agent-stack</category><category>session-compression</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-18-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-18-not-much/</guid><description>**GLM-5.2** from **Zhipu** emerged as a leading open-weight model with innovative **IndexShare** sparse-attention enabling efficient **1M-token inference**, praised as comparable to **GPT-5.5** and **Opus 4.8** but lacking vision support. Other notable open models include **Laguna M.1** by **Poolside AI**, a **70-layer sparse MoE** optimized for long-horizon coding, and **North Mini Code** by **Cohere** with **4-bit quantization** and local deployment support via **Ollama**. The focus is shifting from standalone models to integrated systems combining **model + harness + memory + SCM**, exemplified by **Noumena Code / ncode** addressing challenges in concurrent code agent workflows. Automation tools like **Codex Record &amp; Replay**, **Cursor&apos;s /automate**, and **Artifacts in Claude Code** enhance teachability, reusability, and security in AI-assisted coding workflows.</description><pubDate>Thu, 18 Jun 2026 05:44:39 GMT</pubDate><category>zhipu</category><category>hugging-face</category><category>llama-cpp</category><category>unsloth</category><category>poolsideai</category><category>cohere</category><category>ollama</category><category>openai</category><category>cursor_ai</category><category>claude</category><category>cognition</category><category>glm-5.2</category><category>opus-4.8</category><category>gpt-5.5</category><category>laguna-m.1</category><category>north-mini-code</category><category>codex</category><category>rasbt</category><category>jeremyphoward</category><category>matvelloso</category><category>artificialanlys</category><category>zixuanli_</category><category>_xjdr</category><category>gneubig</category><category>_catwu</category><category>sparse-attention</category><category>1m-token-inference</category><category>open-weight-models</category><category>model-architecture</category><category>long-context</category><category>mixture-of-experts</category><category>quantization</category><category>local-deployment</category><category>workflow-automation</category><category>code-agents</category><category>software-configuration-management</category><category>automation-primitives</category><category>security</category><category>model-harness</category><category>agentic-coding</category></item><item><title>Midjourney Medical: scan your organs like you step on a scale</title><link>https://news.smol.ai/issues/26-06-17-midjourney-medical/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-17-midjourney-medical/</guid><description>**Midjourney** unveiled a new **medical imaging/scanning system** called the **Midjourney Scanner**, described as **radiation-free, magnet-free, fast, and low-cost**, but requiring a **water immersion tank** and having **coarser resolution than CT/MRI**. The announcement included a technical dive and a physical demo, sparking enthusiasm and competitive comparisons with other AI hardware efforts. Technical speculation suggested future design directions involving distributed detectors and real-time imaging, highlighting Midjourney&apos;s ambitious hardware roadmap in medical imaging.</description><pubDate>Wed, 17 Jun 2026 05:44:39 GMT</pubDate><category>midjourney</category><category>saranormous</category><category>matvelloso</category><category>johnowhitaker</category><category>iscienceluvr</category><category>medical-imaging</category><category>hardware</category><category>imaging-systems</category><category>acoustic-imaging</category><category>wave-propagation</category><category>prototype</category><category>technical-dive</category></item><item><title>GLM 5.2: the top Frontend Coding model in the world, IndexShare reduces costs</title><link>https://news.smol.ai/issues/26-06-16-glm-52/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-16-glm-52/</guid><description>**Z.ai released GLM-5.2**, an MIT-licensed open-weight frontier model targeting **coding and long-horizon agentic tasks** with a **1M-token context window** and **two reasoning-effort modes**. It features a **744B-parameter mixture-of-experts architecture** with **40B active parameters per token**, built on **DeepSeek Sparse Attention** extended by **IndexShare**, and supports improved **multi-token prediction (MTP)** for speculative decoding. The model achieved strong leaderboard placements, including **#3 on FrontierSWE**, **#1 on Design Arena**, and **#1 open model on Agent Arena**, with ecosystem support from platforms like **Transformers, vLLM, SGLang, Cloudflare Workers AI, OpenRouter, Ollama Cloud, Baseten, DeepInfra, Fireworks, and Notion**. Early testers praised its potential as a substitute for Opus/GPT-class workflows, though some called for further evaluation and long-horizon validation.</description><pubDate>Tue, 16 Jun 2026 05:44:39 GMT</pubDate><category>z.ai</category><category>lmsys</category><category>deepseek</category><category>cloudflare</category><category>openrouter</category><category>ollama</category><category>baseten</category><category>deepinfra</category><category>fireworks</category><category>notion</category><category>glm-5.2</category><category>mervenoyann</category><category>sentdex</category><category>scaling01</category><category>omarsar0</category><category>teortaxestex</category><category>coding</category><category>agentic-ai</category><category>long-context</category><category>mixture-of-experts</category><category>sparse-attention</category><category>speculative-decoding</category><category>multi-token-prediction</category><category>model-benchmarking</category><category>inference-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-12-not-much/</guid><description>**Anthropic** suspended access to **Claude Fable 5** and **Mythos 5** due to **US export controls**, sparking a debate on **model sovereignty** and geopolitical risks for frontier AI vendors. **Artificial Analysis** updated its coding agent benchmark, replacing **SWE-Bench Pro** with **DeepSWE**, reshuffling rankings with **Claude Code + Fable 5 [max]** leading. Discussions highlighted the importance of **harness quality** versus pure model capability and concerns over **benchmark saturation** and realism. Additionally, **Moonshot** released the open-source model **Kimi K2.7-Code**.</description><pubDate>Fri, 12 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>artificial-analysis</category><category>datacurve</category><category>moonshot</category><category>claude-fable-5</category><category>mythos-5</category><category>gpt-5.5</category><category>claude-code</category><category>fable-5</category><category>codex</category><category>opus-4.8</category><category>kimi-k2.7-code</category><category>natolambert</category><category>theo</category><category>cohere</category><category>kunchenguid</category><category>clementdelangue</category><category>dejavucoder</category><category>ofirpress</category><category>ramplabs</category><category>model-sovereignty</category><category>export-controls</category><category>coding-agent-evaluation</category><category>benchmarking</category><category>benchmark-gaming</category><category>harness-quality</category><category>benchmark-saturation</category><category>open-source-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-11-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-11-not-much/</guid><description>**Anthropic** reversed its covert degradation policy on **Claude Fable 5** after public backlash, sparking debates on governance, transparency, and access to frontier AI models. The model shows strong capabilities with mixed benchmark results, including **87.8% on WeirdML** and top ranking on FrontierSWE, but practical usage highlights cost and inconsistent behavior. Separately, **Recursive SI**, led by **Richard Socher**, released an automated open-ended discovery system achieving state-of-the-art results on **NVIDIA SOL-ExecBench**, **NanoGPT Speedrun**, and **NanoChat autoresearch**, with open-sourced discoveries and improved efficiency metrics.</description><pubDate>Thu, 11 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>recursive-si</category><category>nvidia</category><category>claude-fable-5</category><category>nanogpt</category><category>richard_socher</category><category>model-governance</category><category>model-transparency</category><category>benchmarking</category><category>automated-research</category><category>optimization</category><category>open-sourcing</category><category>model-behavior</category><category>cost-efficiency</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-15-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-15-not-much/</guid><description>**Anthropic&apos;s Fable/Mythos export-control crisis** dominates AI news, highlighting the intersection of **national security** and frontier model access. Technical voices like **François Chollet** criticize opaque regulatory actions and advocate for **standardized benchmarks for agentic capabilities**. **Epoch AI** reports **Claude Fable 5** surpassing **GPT-5.5 Pro** on the **Epoch Capabilities Index**, underscoring tensions between cutting-edge AI and regulatory constraints. The concept of **model neutrality** is evolving from philosophy to architecture, emphasizing **harness, context, memory, and routing** for multi-model fungibility, with contributions from voices like **hwchase17**, **Nikesh Arora**, and **mignano**. Agent systems are transitioning from demos to production with a focus on **observability**, **trace analysis**, and **evaluation infrastructure**, exemplified by **LangChain&apos;s LangSmith Engine** and fine-tuned judges for behavioral correction signals. Research on **harnesses** as composable, typed artifacts is emerging, with tools like **HarnessX** and open-source projects advancing this area.</description><pubDate>Thu, 11 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>epoch-ai</category><category>langchain</category><category>fable-5</category><category>mythos</category><category>claude-fable-5</category><category>gpt-5.5-pro</category><category>fchollet</category><category>simonw</category><category>hwchase17</category><category>nikesharora</category><category>mignano</category><category>sauvast</category><category>rohit4verse</category><category>dair_ai</category><category>omarsar0</category><category>export-control</category><category>national-security</category><category>agentic-capabilities</category><category>model-neutrality</category><category>harness</category><category>observability</category><category>trace-analysis</category><category>evaluation-infrastructure</category><category>behavioral-correction</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-10-not-much/</guid><description>**Anthropic** faced backlash for silently degrading AI research capabilities in its **Fable/Mythos** models without clear disclosure, raising concerns about trust, reproducibility, and enterprise data retention policies. Despite controversy, **Fable 5** demonstrated strong benchmark performance, leading in agentic and coding tasks with high scores on **Agent Arena**, **SimpleBench**, **CADGenBench**, and **PACT**. **Dario Amodei** published a policy advocating stronger frontier AI oversight amid these tensions.</description><pubDate>Wed, 10 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>fable-5</category><category>mythos</category><category>darioamodei</category><category>natolambert</category><category>martin_casado</category><category>drfeifei</category><category>antirez</category><category>clementdelangue</category><category>deanwball</category><category>hlntnr</category><category>_arohan_</category><category>dbahdanau</category><category>gergelyorosz</category><category>scaling01</category><category>dbreunig</category><category>omarsar0</category><category>yacinemtb</category><category>mchlhess</category><category>jasonbotterill</category><category>lvwerra</category><category>lechmazur</category><category>kimmonismus</category><category>walden_yan</category><category>hrishioa</category><category>model-performance</category><category>trust</category><category>data-retention</category><category>benchmarking</category><category>agentic-ai</category><category>coding</category><category>policy</category></item><item><title>Anthropic Claude Fable 5</title><link>https://news.smol.ai/issues/26-06-09-anthropic-claude-fable-5/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-09-anthropic-claude-fable-5/</guid><description>**Anthropic** released two major models: **Claude Fable 5** for general availability and **Claude Mythos 5** for restricted access, with fallback to **Claude Opus 4.8** for sensitive queries. **Fable 5** features a **1M-token context window** and pricing at **$10/million input tokens** and **$50/million output tokens**. It leads benchmarks in software engineering, knowledge work, scientific research, and vision, outperforming **GPT-5.5** and setting new state-of-the-art scores on **CursorBench**, **FrontierCode**, **Terminal-Bench 2.1**, and **Artificial Analysis Intelligence Index**. The rollout includes Pro, Max, Team, and Enterprise plans with temporary usage credits due to capacity constraints. Middleware SDK support is available in **Python, TypeScript, Go, Java, and C#**.</description><pubDate>Tue, 09 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>cursor_ai</category><category>cognition</category><category>claude-fable-5</category><category>claude-mythos-5</category><category>claude-opus-4.8</category><category>gpt-5.5</category><category>mikeyk</category><category>scaling01</category><category>benchmarking</category><category>software-engineering</category><category>knowledge-work</category><category>scientific-research</category><category>vision</category><category>context-windows</category><category>model-pricing</category><category>sdk</category><category>rate-limiting</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-08-not-much/</guid><description>**FrontierCode** benchmark by **Cognition** highlights the challenge of coding tasks with the best model, **Opus 4.8**, scoring only about **13%** on the hardest subset, indicating coding is less solved than benchmarks suggest. The trend toward using **loops** as a control metaphor for coding agents is prominent, with emphasis on clear goals, verification, and iteration, though some experts caution about overreliance on loops. Agent ergonomics are improving with observability dashboards, sandbox environments, and workflow tools from **ClaudeDevs**, **MagicPath**, **LangSmith**, and **Modal**. **Kimi** by **Moonshot** released major updates including a stronger coding agent and a desktop agent product supporting up to **300 local sub-agents**. **Google** advanced efficient local deployment with upgrades to **Gemma 4** checkpoints.</description><pubDate>Mon, 08 Jun 2026 05:44:39 GMT</pubDate><category>cognition</category><category>frontiercode</category><category>moonshot</category><category>google</category><category>claudedevs</category><category>magicpath</category><category>langsmith</category><category>modal</category><category>opus-4.8</category><category>gemma-4</category><category>swyx</category><category>dzhng</category><category>claudecode</category><category>bcherny</category><category>reach_vb</category><category>omarsar0</category><category>gneubig</category><category>hamelhusain</category><category>angaisb_</category><category>coding-evaluation</category><category>agent-control</category><category>verification</category><category>agent-ergonomics</category><category>sandbox-environments</category><category>local-inference</category><category>workflow-optimization</category><category>cli-tools</category><category>plugin-integration</category><category>persistent-memory</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-05-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-05-not-much/</guid><description>**Anthropic&apos;s Mythos/Opus cycle** sparked mixed reactions with praise for **Claude Mythos**&apos;s one-shot workflows and concerns over **Opus 4.8** benchmark regressions. **Opus 4.7** showed strong chemistry task performance, &quot;making Claude a chemist.&quot; **Sakana AI** launched an **RSI Lab** focusing on recursive self-improvement under compute constraints, marking RSI as a formal research program. New benchmarks like **Agents&apos; Last Exam (ALE)** and **SWE-Marathon** test agents on long-horizon, economically meaningful tasks, revealing low pass rates and coherence challenges. Princeton&apos;s ICML 2026 paper found models like **GPT 5.5**, **Gemini 3.1 Pro / 3.5 Flash**, and **Claude Opus 4.7** still lack meaningful reliability improvements. Tooling trends favor RL-environment-style frameworks for agent evaluation, exemplified by Meta&apos;s **OpenEnv**.</description><pubDate>Fri, 05 Jun 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>sakana-ai</category><category>meta-ai-fair</category><category>princeton</category><category>claude-mythos</category><category>opus-4.8</category><category>opus-4.7</category><category>gpt-5.5</category><category>gemini-3.1-pro</category><category>gemini-3.5-flash</category><category>claude-opus-4.7</category><category>kimmonismus</category><category>lechmazur</category><category>teortaxestex</category><category>hardmaru</category><category>andrew_n_carr</category><category>steverab</category><category>pauliusztin_</category><category>recursive-self-improvement</category><category>benchmarking</category><category>agent-evaluation</category><category>long-horizon-tasks</category><category>reliability</category><category>reinforcement-learning</category><category>sample-efficiency</category><category>economically-meaningful-tasks</category><category>agent-coherence</category><category>anti-reward-hacking</category><category>tooling</category><category>rl-environments</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-04-not-much/</guid><description>**NVIDIA** released **Nemotron 3 Ultra**, a fully open **550B MoE** model with **55B active parameters** and **1M context**, optimized for long-running agent tasks with up to **5x speedup** and **30% cost reduction**. It features hybrid Mamba/attention, LatentMoE, native MTP, and was pretrained on **20T tokens** using NVFP4 low-precision format. Benchmarks show strong performance with **47.7 Intelligence Index** and **400+ output tokens/sec**. The model is supported across major serving platforms. Additionally, **Nemotron 3.5 ASR** is an open streaming ASR model with **0.6B parameters**, supporting **40 language-locale combinations** and sub-100ms latency, designed for voice agents. 

**Anthropic** highlighted early signs of recursive self-improvement (RSI) in AI, with **Claude** models authoring **80%+ of merged code** and engineers shipping **8x more code**. Claude Opus 4 achieved **3x speedup** on training scripts, while Mythos Preview reached **~52x speedup** and provided better research suggestions than humans **64% of the time**.</description><pubDate>Thu, 04 Jun 2026 05:44:39 GMT</pubDate><category>nvidia</category><category>anthropic</category><category>togethercompute</category><category>baseten</category><category>modal</category><category>vllm_project</category><category>fireworksai_hq</category><category>ollama</category><category>wandb</category><category>cline</category><category>primeintellect</category><category>nousresearch</category><category>nemotron-3-ultra</category><category>nemotron-3.5-asr</category><category>claude-opus-4</category><category>mythos-preview</category><category>piotrz_zelasko</category><category>mixture-of-experts</category><category>long-context</category><category>model-quantization</category><category>agentic-ai</category><category>streaming-speech</category><category>asr</category><category>low-precision-training</category><category>benchmarking</category><category>recursive-self-improvement</category><category>code-generation</category><category>model-speedup</category></item><item><title>Microsoft Build: MAI-Thinking-1 and MAI Family models, Surface RTX Spark Dev Box, and OpenClaw in Windows</title><link>https://news.smol.ai/issues/26-06-02-msft-mai-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-02-msft-mai-2/</guid><description>**Microsoft** introduced **MAI-Thinking-1**, a **35B parameter MoE model** with **256K context**, achieving **97% on AIME 2025** and outperforming **Sonnet 4.6** in human preference tests. The broader **7-model MAI family** spans reasoning, code, image, speech, and voice, with third-party availability on **OpenRouter**, **fal**, and **Baseten**. The detailed **109-page technical report** revealed insights on scaling, MFU, RL/post-training, and data curation, highlighting no third-party distillation and advanced prompt optimization techniques. Microsoft emphasized **agent-native devices and local inference** with projects like **Project Solara / Scout** and the **Surface RTX Spark Dev Box**, alongside software innovations such as the **Copilot desktop app** and **MAI-Code-1-Flash** integration. Meanwhile, local-first computer-use agents like **Holo 3.1** (Qwen-based, 0.8B to 35B parameters) support laptops and small workstations with optimized formats and strong benchmark results. Desktop shells for agents, including **Hermes Desktop**, **Devin Desktop**, and agent-neutral approaches compatible with **Devin, Claude Code, and Codex**, are proliferating, with hybrid local/cloud execution becoming the default architecture as seen in **Perplexity Computer&apos;s** hybrid agentic inference.</description><pubDate>Tue, 02 Jun 2026 05:44:39 GMT</pubDate><category>microsoft</category><category>openrouter</category><category>fal</category><category>baseten</category><category>hcompany_ai</category><category>teksedge</category><category>nous-research</category><category>teknim</category><category>cognition</category><category>windsurf</category><category>perplexity-ai</category><category>mai-thinking-1</category><category>mai-code-1-flash</category><category>holo-3.1</category><category>qwen-35b</category><category>sonnet-4.6</category><category>claude-code</category><category>codex</category><category>mustafasuleyman</category><category>eliebakouch</category><category>hannahajishirzi</category><category>asadovsky</category><category>bj2rn</category><category>lateinteraction</category><category>lakshyaaagrawal</category><category>theturingpost</category><category>kimmonismus</category><category>yusuf_i_mehdi</category><category>pierceboggan</category><category>lukehoban</category><category>nielsrogge</category><category>russelljkaplan</category><category>mixture-of-experts</category><category>context-windows</category><category>benchmarking</category><category>reinforcement-learning</category><category>prompt-optimization</category><category>agentic-ai</category><category>local-inference</category><category>model-family-expansion</category><category>model-reporting</category><category>agent-native-devices</category><category>software-development</category><category>model-optimization</category><category>hybrid-inference</category><category>desktop-agents</category><category>model-quantization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-03-not-much/</guid><description>**Microsoft** released the detailed technical report for **MAI-Thinking-1**, a generalist reasoning model trained without third-party distillation, achieving **97% on AIME 2025** and outperforming Sonnet 4.6 in human preference tests. The report was praised for transparency, revealing no synthetic data use, a unique scaling ladder recipe, and detailed training data composition including **50% code** and **17.5% STEM**. Microsoft also introduced **Frontier Tuning** for workflow-specific model adaptation, claiming efficiency gains up to **10×** and GPT-5.4-level quality in Excel tasks, alongside new models like **MAI-Image-2.5** and **MAI-Code-1-Flash**. Meanwhile, **Google** launched **Gemma 4 12B**, an Apache 2.0 multimodal model with an innovative encoder-free architecture designed for on-device use with **16GB VRAM**, collapsing vision and audio encoders into the LLM backbone, receiving positive community feedback and immediate tooling support.</description><pubDate>Tue, 02 Jun 2026 05:44:39 GMT</pubDate><category>microsoft</category><category>google</category><category>vllm-project</category><category>ollama</category><category>llama-cpp</category><category>mai-thinking-1</category><category>mai-image-2.5</category><category>mai-code-1-flash</category><category>gemma-4-12b</category><category>eliebakouch</category><category>nrehiew_</category><category>mustafasuleyman</category><category>minjiyoon90</category><category>lateinteraction</category><category>harold_matmul</category><category>googlegemma</category><category>googleaidevs</category><category>mtschannen</category><category>armandjoulin</category><category>osanseviero</category><category>model-training</category><category>reinforcement-learning</category><category>model-architecture</category><category>multimodality</category><category>model-deployment</category><category>model-efficiency</category><category>fine-tuning</category><category>on-device-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-06-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-06-01-not-much/</guid><description>**NVIDIA** led open-source AI model releases with **Cosmos 3**, a comprehensive omnimodal world model unifying language, image, video, audio, and action using a Mixture-of-Transformers design, and **Nemotron 3 Ultra**, a **550B** parameter open-weight model noted for high serving speed and strong evaluation performance. The **Cosmos Coalition** was launched to foster an open ecosystem for physical AI world models. Meanwhile, **MiniMax M3** debuted as a multimodal agent/coding model with **1M context** and strong benchmark scores, gaining rapid ecosystem support from vendors like **Novita** and **Vercel AI Gateway**. However, MiniMax M3 showed some inefficiencies such as high token consumption and verbose self-check loops. These developments highlight advances in open physical AI, multimodality, and agent models with significant community and infrastructure engagement.</description><pubDate>Mon, 01 Jun 2026 05:44:39 GMT</pubDate><category>nvidia</category><category>runway</category><category>novita</category><category>vercel</category><category>cloudflare</category><category>openclaude</category><category>flowith</category><category>cosmos-3</category><category>nemotron-3-ultra</category><category>minimax-m3</category><category>kimmonismus</category><category>clementdelangue</category><category>artificialanalysis</category><category>scaling01</category><category>ctnzr</category><category>caspar_br</category><category>eliebakouch</category><category>pbdtokenrouter</category><category>rauchg</category><category>gitlawb</category><category>notjazii</category><category>lostinlatencyx</category><category>zhihufrontier</category><category>omnimodal-models</category><category>mixture-of-experts</category><category>autoregressive-models</category><category>diffusion-models</category><category>structured-prompts</category><category>fine-tuning</category><category>open-weight-models</category><category>multimodality</category><category>agent-models</category><category>benchmarking</category><category>model-serving</category><category>context-windows</category><category>token-efficiency</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-29-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-29-not-much/</guid><description>**Anthropic** rolled out **Claude Opus 4.8**, which shows incremental improvements but mixed benchmark results, including better cooperation and coding behavior but some regressions in document parsing. Platform updates include mid-conversation system instructions enhancing long agent sessions, though API pricing remains a concern. A Hugging Face analysis revealed a critical bug in multi-turn reinforcement learning training loops involving tokenization mismatches, with a proposed &quot;Token-In, Token-Out&quot; fix. Agent harness design is evolving as a key optimization area, with **LangChain**&apos;s Deep Agents v0.6 achieving strong performance at much lower cost, and **vllm_project** releasing native weight syncing APIs and a Rust BPE tokenizer to improve tokenization efficiency. Debate continues on the value of multi-agent systems, with some seeing them as speedups and others expecting capability breakthroughs.</description><pubDate>Fri, 29 May 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>huggingface</category><category>langchain</category><category>vllm_project</category><category>claude-opus-4.8</category><category>gpt-5.5</category><category>qwen</category><category>kimi</category><category>deepseek</category><category>jeremyphoward</category><category>leo_linsky</category><category>clementdelangue</category><category>johnschulman2</category><category>omarsar0</category><category>hwchase17</category><category>ofirpress</category><category>scaling01</category><category>reinforcement-learning</category><category>tokenization</category><category>agentic-ai</category><category>api</category><category>model-optimization</category><category>long-context</category><category>rust</category><category>performance-optimization</category><category>multi-agent-systems</category><category>prompt-engineering</category></item><item><title>Anthropic raises $65B in Series H at a $965B post-money valuation, releases Opus 4.8 and Dynamic Workflows</title><link>https://news.smol.ai/issues/26-05-28-anthropic-series-h/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-28-anthropic-series-h/</guid><description>**Anthropic** announced a massive **$65B Series H financing** at a **$965B valuation**, led by **Altimeter, Dragoneer, Greenoaks, and Sequoia**, with run-rate revenue surpassing **$47B**. They launched **Claude Opus 4.8**, an update to Opus 4.7 featuring &quot;sharper judgment,&quot; &quot;more honesty,&quot; and longer autonomous work at the same price. Anthropic also introduced **Dynamic Workflows** in Claude Code, enabling orchestration of hundreds of parallel subagents for large tasks, available in research preview across multiple platforms. Opinions on Opus 4.8 vary, with some praising it as a major leap and others viewing it as incremental or catch-up to **OpenAI&apos;s GPT-5.5** family.</description><pubDate>Thu, 28 May 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>altimeter</category><category>dragoneer</category><category>greenoaks</category><category>sequoia</category><category>andonlabs</category><category>claude-opus-4.8</category><category>claude-opus-4.7</category><category>gpt-5.5</category><category>dan_shipper</category><category>scaling01</category><category>zephyr_z9</category><category>teortaxes_tex</category><category>kimmonismus</category><category>model-release</category><category>reinforcement-learning</category><category>agentic-ai</category><category>model-evaluation</category><category>long-context</category><category>model-optimization</category><category>fine-tuning</category><category>multitasking</category><category>parallel-processing</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-26-not-much/</guid><description>**Harness engineering** is emerging as the key differentiator for coding agents, emphasizing the stack of **model + harness + eval loop** over just stronger base models. **DeepSeek** is building a harness team to optimize interaction and verification loops, while **Google&apos;s Gemini Managed Agents** and **LangChain** formalize harness concepts like context governance and dynamic skill routing. New benchmarks like **DeepSWE** align closely with real developer experience, with **Qwen3.7 Max** and **Claude Opus 4.6** showing strong agentic coding performance. **Anthropic** introduced a security-guidance plugin for **Claude Code** reducing security PR comments by 30–40%, and **OpenAI** highlighted **GPT-5.5** in Codex for improved document parsing. In research, **Claude Mythos** solved Erdős problem #90 with a cleaner proof path than previous models, showing latent capabilities unlocked by appropriate harnesses. The paper &quot;Language Models Need Sleep&quot; proposes a sleep-like consolidation phase for long-horizon memory, addressing bottlenecks in persistent context storage. Open research agents like **QUEST** (2B–35B parameters) advance long-horizon fact-seeking and citation grounding, while the **CUSP benchmark** from Sakana/Stanford/Oxford/AI2 evaluates current model capabilities in science.</description><pubDate>Tue, 26 May 2026 05:44:39 GMT</pubDate><category>deepseek</category><category>google-deepmind</category><category>langchain-ai</category><category>anthropic</category><category>openai</category><category>alibaba</category><category>sakana-ai</category><category>stanford</category><category>oxford</category><category>ai2</category><category>qwen-3.7</category><category>claude-opus-4.6</category><category>gpt-5.5</category><category>mythos</category><category>quest-2b-35b</category><category>sebastienbubeck</category><category>harness-engineering</category><category>agent-infrastructure</category><category>coding-benchmarks</category><category>security-guidance</category><category>long-horizon-memory</category><category>context-compression</category><category>sleep-phase</category><category>math-problem-solving</category><category>fact-seeking</category><category>citation-grounding</category><category>science-evaluation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-27-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-27-not-much/</guid><description>**Inference optimization** is increasingly architectural, with **EAGLE 3.1** improving speculative decoding and long-context handling, collaborating with **vLLM** and **TorchSpec**. **Perplexity** open-sourced a rebuilt **Unigram tokenizer** cutting CPU use by **5–6×** and achieving **63 µs at 514 tokens**. **Qwen3.5** hits **580 tokens/s** via joint efforts from **Alibaba**, **LightSeek**, **NVIDIA**, **Mooncake**, and **FlashAttention-4** contributors. Price cuts in APIs from Chinese labs are sustainable due to structural KV-cache and attention improvements, exemplified by **DeepSeek V4-Pro** and **Xiaomi MiMo** reducing caching costs significantly. 

Agent engineering shifts focus from model quality to model-harness-memory fit, with **LangChain** releasing **Deep Agents v0.6** and tools like **LangSmith Engine** automating evaluation loops. **Trajectory** launched a continual learning platform with **$15M funding** and partners like **Clay** and **Harvey**, supporting large models including a **397B-parameter model** deployed on autoscaled **H100** infrastructure. Open-source memory-centric agents and minimal training harnesses also gained attention.</description><pubDate>Tue, 26 May 2026 05:44:39 GMT</pubDate><category>eaglecorp</category><category>vllm_project</category><category>perplexity_ai</category><category>alibaba</category><category>lightseek</category><category>nvidia</category><category>mooncake</category><category>flashattention</category><category>kimmonismus</category><category>deepseek</category><category>xiaomi</category><category>langchain</category><category>baseten</category><category>trajectory</category><category>clay</category><category>harvey</category><category>decagon</category><category>mercor</category><category>rogo</category><category>rlm</category><category>eagle-3.1</category><category>unigram-tokenizer</category><category>qwen-3.5</category><category>deepseek-v4-pro</category><category>mimo</category><category>deep-agents-v0.6</category><category>397b-parameter-model</category><category>kimmonismus</category><category>_luofuli</category><category>vtrivedy10</category><category>inference-optimization</category><category>long-context</category><category>speculative-decoding</category><category>tokenization</category><category>attention-mechanisms</category><category>kv-cache</category><category>cache-hierarchy</category><category>agent-engineering</category><category>model-harness-memory-fit</category><category>continual-learning</category><category>quantization</category><category>autoscaling</category><category>memory-centric-agents</category><category>evaluation-automation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-21-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-21-not-much/</guid><description>**RAEv2** advances representation-first tokenization with **&gt;10x faster convergence** and improved generation, tested on **text-to-image** and **world models**. **NVIDIA&apos;s Gated DeltaNet-2** innovates linear attention with channel-wise gates, outperforming **KDA** and **Mamba-3** at **1.3B parameters** on language modeling and reasoning tasks. Studies on **subword tokenization** reveal only some benefits at scale, while data filtering research suggests that with enough compute, **no filtering** may be optimal at around **1e30 FLOPs**. Mechanistic interpretability updates propose clustering features by joint firing patterns for better geometry understanding. OpenAI&apos;s AI-assisted breakthrough on an Erdős unit-distance math problem sparks debate on AI&apos;s role in mathematical research. Harnesses remain key for capability improvements in agent infrastructure.</description><pubDate>Thu, 21 May 2026 05:44:39 GMT</pubDate><category>nvidia</category><category>openai</category><category>nous-research</category><category>raev2</category><category>gated-deltanet-2</category><category>kda</category><category>mamba-3</category><category>dclm</category><category>1jaskiratsingh</category><category>recatm</category><category>sainingxie</category><category>ahatamiz1</category><category>rasbt</category><category>nousresearch</category><category>tatsu_hashimoto</category><category>goodfireai</category><category>markchen90</category><category>wtgowers</category><category>memecrashes</category><category>cloneofsimo</category><category>lvwerra</category><category>representation-learning</category><category>tokenization</category><category>linear-attention</category><category>long-context</category><category>mechanistic-interpretability</category><category>math</category><category>data-filtering</category><category>agent-infrastructure</category><category>language-modeling</category><category>commonsense-reasoning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-18-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-18-not-much/</guid><description>**Agent infrastructure** is advancing with **LangSmith Engine** providing CI/CD loops for agents and **SmithDB** enabling low-latency querying for observability. **Cognition&apos;s Devin Auto-Triage** offers persistent automation for bug triage with memory and subagent structures. **Anthropic** improves **Claude Code** for large codebases with prompt cache diagnostics and faster modes, while **OpenAI** enhances **Codex** workflows with remote execution and plugins. Microsoft released remote control for **GitHub Copilot CLI** and VS Code. The community emphasizes **verification, decomposition, and feedback loops** over prompt cleverness for coding agents. **Cursor&apos;s Composer 2.5** is highlighted as a strong new coding model, with plans for a larger model trained with **SpaceXAI** using **10× more compute** on **Colossus 2** hardware, praised for efficiency and collaboration improvements.</description><pubDate>Mon, 18 May 2026 05:44:39 GMT</pubDate><category>langchain</category><category>cognition</category><category>anthropic</category><category>openai</category><category>microsoft</category><category>cursor</category><category>claude-code</category><category>codex</category><category>composer-2.5</category><category>krishdpi</category><category>walden_yan</category><category>russelljkaplan</category><category>fchollet</category><category>gabriberton</category><category>palashshah</category><category>shannholmberg</category><category>agent-automation</category><category>agent-observability</category><category>ci-cd</category><category>prompt-caching</category><category>remote-execution</category><category>verification</category><category>decomposition</category><category>feedback-loops</category><category>coding-agents</category><category>model-efficiency</category><category>instruction-following</category></item><item><title>Google I/O 2026: Gemini 3.5 Flash, Omni, and Google’s Agent Stack</title><link>https://news.smol.ai/issues/26-05-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-19-not-much/</guid><description>**Google** announced at I/O the repositioning of **Gemini** as a consumer AI and developer/agent platform with three key releases: **Gemini 3.5 Flash** for fast agentic and coding tasks, **Gemini Omni** for multimodal generation and editing including video, and the expanded **Antigravity 2.0** agent stack. Google reports processing over **3.2 quadrillion tokens per month**, a 7x increase year-over-year, with **900M+ monthly Gemini users** across 230+ countries and 70+ languages. Gemini 3.5 Flash features a **1M-token context window**, **65k max output tokens**, **4 thinking levels**, and &quot;thought preservation&quot; across turns, outperforming Gemini 3.1 Pro on multiple benchmarks and running up to 12x faster in Antigravity. Independent benchmarks show Gemini 3.5 Flash scoring **55 on the Intelligence Index**, with higher costs than previous versions. Gemini Omni Flash supports text, image, video, and audio inputs for generative media tasks, available now for paid users.</description><pubDate>Mon, 18 May 2026 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>geminiapp</category><category>gemini-3.5-flash</category><category>gemini-3.1-pro</category><category>gemini-3.5</category><category>gemini-omni</category><category>philschmid</category><category>jeffdean</category><category>agentic-ai</category><category>multimodality</category><category>video-generation</category><category>model-performance</category><category>benchmarking</category><category>context-windows</category><category>model-optimization</category><category>model-scaling</category><category>instruction-following</category><category>api</category><category>model-efficiency</category><category>cost-analysis</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-15-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-15-not-much/</guid><description>**Cerebras** made headlines with its **IPO**, marking a significant milestone for the company known for its contrarian hardware approach. The **Cerebras CFO Bob Komin** emphasized the company&apos;s capability to serve **trillion-parameter models**, including internal **OpenAI 5.4 and 5.5** models, pushing back against the notion that Cerebras only supports small models. Investor **Ishan N. Taneja** praised Cerebras for its persistence and execution, calling their chip a &quot;banger.&quot; The IPO is seen as a validation of Cerebras&apos;s long-term strategy in inference infrastructure, highlighting themes like **compute scarcity**, **inference demand**, and **model routing**.</description><pubDate>Fri, 15 May 2026 05:44:39 GMT</pubDate><category>cerebras</category><category>openai</category><category>openai-5.4</category><category>openai-5.5</category><category>ishanit5</category><category>dee_bosa</category><category>apoorv03</category><category>bob_komin</category><category>inference</category><category>model-serving</category><category>compute-scarcity</category><category>model-routing</category><category>hardware-architecture</category><category>trillion-parameter-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-14-not-much/</guid><description>**OpenAI** expanded **Codex** integration with the ChatGPT mobile app enabling remote task management and introduced Remote SSH, hooks, and programmatic tokens for enterprise automation. The IDE ecosystem is shifting to &quot;agent-first&quot; UX with **GitHub Copilot App** preview and **VS Code** launching a multi-agent workflow window. Open-source agents like **Nous/Hermes** integrated Codex runtime, and **Kimi** released a web bridge extension supporting multiple coding agents. **LangChain** released significant agent infrastructure including **SmithDB** for agent trace data and **LangSmith Engine** for trace analysis and continual learning, launching **LangChain Labs** to improve agents via production trace feedback loops.</description><pubDate>Thu, 14 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>github</category><category>microsoft</category><category>nous-research</category><category>moonshot-ai</category><category>langchain</category><category>prime-intellect</category><category>codex</category><category>chatgpt</category><category>hwchase17</category><category>caspar_br</category><category>bentannyhill</category><category>jakebroekhuizen</category><category>willccbb</category><category>agent-infrastructure</category><category>agent-first-ux</category><category>remote-ssh</category><category>programmatic-access-tokens</category><category>sandboxing</category><category>continual-learning</category><category>agent-trace-data</category><category>multi-agent-workflows</category><category>ide-integration</category><category>browser-extensions</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-13-not-much/</guid><description>**Cline, LangChain, Notion, and Cursor** advanced agent infrastructure and developer platforms with innovations like **Cline SDK**, **LangSmith Engine**, **SmithDB** (offering **12–15×** faster observability), and Notion&apos;s External Agents API integrating third-party agents such as Claude and Codex. Agent UX trends emphasize **long-running state, streaming, and orchestration** over chat, with tools like **Duet Agent** and **VS Code Agents window** enhancing durable execution and inspectable states. Research highlights include **Nous Research&apos;s Token Superposition Training** achieving **2–3× speedup** in pretraining, a **multi-stream LLM** architecture for parallel reasoning by Jonas Geiping et al., and **δ-mem** external memory improving benchmark scores. NVIDIA&apos;s **Star Elastic** offers post-training model compression at **360× lower cost** than pretraining, while Datology focuses on data curation for vision-language models.</description><pubDate>Wed, 13 May 2026 05:44:39 GMT</pubDate><category>cline</category><category>langchain</category><category>notion</category><category>cursor</category><category>nous-research</category><category>nvidia</category><category>datology</category><category>claude</category><category>codex</category><category>langsmith-engine</category><category>smithdb</category><category>duet-agent</category><category>multi-stream-llm</category><category>delta-mem</category><category>star-elastic</category><category>jonas_geiping</category><category>siddharth_joshi</category><category>pratyush_maini</category><category>agent-infrastructure</category><category>developer-platforms</category><category>observability</category><category>long-running-state</category><category>streaming</category><category>orchestration</category><category>pretraining-efficiency</category><category>model-architecture</category><category>external-memory</category><category>post-training-compression</category><category>data-curation</category><category>vision-language-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-12-not-much/</guid><description>**Research-level reasoning benchmarks** are advancing with **439 new math problems** from **64 mathematicians** and expanded medical benchmarks in **Medmarks v1.0** covering **30 benchmarks** and **61 models**. **Google DeepMind&apos;s AI Co-Mathematician** achieves **48% on FrontierMath Tier 4**, while **Gemini 3.1 Pro** improves physics benchmark scores significantly. **GPT-5.5 high/xhigh** outperforms **Opus 4.7 xhigh** on program synthesis tasks. Retrieval benchmarks favor smaller models like **LightOn&apos;s Agent-ModernColBERT** with **149M parameters**. Training optimization advances include **SOAP/Muon-style updates** reducing training steps, and a **Lean4-to-TileLang superoptimizer** achieving **1.8× speedup on A100 GPUs**. Scaling laws are reconsidered with arguments for measuring in bytes rather than tokens. New training-time efficiency methods like **Lighthouse Attention** enable subquadratic training wrappers removable before deployment.</description><pubDate>Tue, 12 May 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>lighton</category><category>nous-research</category><category>gemini-3.1-pro</category><category>gpt-5.5</category><category>opus-4.7-xhigh</category><category>agent-moderncolbert</category><category>soohak</category><category>polynoamial</category><category>torchcompiled</category><category>leloykun</category><category>che_shr_cat</category><category>jjitsev</category><category>omarsar0</category><category>research-benchmarks</category><category>math</category><category>medical-benchmarks</category><category>agentic-systems</category><category>program-synthesis</category><category>retrieval-augmentation</category><category>training-optimization</category><category>superoptimization</category><category>scaling-laws</category><category>training-efficiency</category><category>gpu-optimization</category><category>attention-mechanisms</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-11-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-11-not-much/</guid><description>**Thinking Machines** previewed their new **native interaction models** designed for **full-duplex multimodal interaction** enabling real-time concurrent listening, speaking, watching, thinking, searching, and reacting, marking a shift beyond turn-based AI. This approach emphasizes continuous audio, video, and text processing, with innovations like **visual proactivity** and background tool use, implemented using **SGLang**. Meanwhile, **OpenAI** announced the **OpenAI Deployment Company**, a new unit with **150 Forward Deployed Engineers** and **$4B initial investment** to help enterprises deploy frontier models, signaling a move into the deployment layer of the AI economy. OpenAI also launched **Daybreak**, a security-focused initiative integrating **GPT-5.5** and **Codex** for cyber defense, threat modeling, and automated patching, offering differentiated access tiers including **GPT-5.5-Cyber**. This contrasts with Anthropic&apos;s more restrictive cyber approach, highlighting tensions in AI security strategies.</description><pubDate>Mon, 11 May 2026 05:44:39 GMT</pubDate><category>thinking-machines</category><category>openai</category><category>anthropic</category><category>gpt-5.5</category><category>codex</category><category>johnschulman2</category><category>soumithchintala</category><category>chillee</category><category>liliyu_lili</category><category>rown</category><category>kimmonismus</category><category>giffmana</category><category>swyx</category><category>eliebakouch</category><category>gdb</category><category>sama</category><category>therundownai</category><category>lukolejnik</category><category>matvelloso</category><category>multimodality</category><category>real-time-interaction</category><category>visual-proactivity</category><category>deployment</category><category>cybersecurity</category><category>threat-modeling</category><category>automation</category><category>continuous-audio-video-text-processing</category><category>security-models</category><category>field-engineering</category><category>enterprise-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-08-not-much/</guid><description>**OpenAI** rapidly expanded the **GPT-5.5** family with multiple variants including **gpt-image-2**, **GPT-5.5 Pro**, and **GPT-5.5 Cyber**, receiving positive feedback for efficiency and usability. **Codex** evolved into a long-running agent runtime with a new **/goal** mechanism, achieving 61% success on ARC-AGI-3 games after extensive testing. OpenAI also introduced cybersecurity-focused models like **GPT-5.5-Cyber** targeting enterprise and government sectors. Meanwhile, **Zyphra** released the open-model **ZAYA1-74B-Preview**, a 74B parameter mixture-of-experts model trained on **AMD** hardware under Apache 2.0 license, alongside a vision-language model **ZAYA1-VL-8B**. Inference infrastructure competition intensified with **vLLM** updates improving throughput and latency, including support for **DeepSeek V4** and enhanced quantization/backends.</description><pubDate>Fri, 08 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>zyphra</category><category>amd</category><category>deepseek</category><category>vllm_project</category><category>gpt-5.5</category><category>gpt-image-2</category><category>gpt-5.5-pro</category><category>gpt-5.5-instant</category><category>gpt-realtime-2</category><category>gpt-5.5-cyber</category><category>codex</category><category>zaya1-74b-preview</category><category>zaya1-vl-8b</category><category>qwen3-omni</category><category>reach_vb</category><category>dhh</category><category>gdb</category><category>patience_cave</category><category>ithilgore</category><category>cryps1s</category><category>sama</category><category>deredleritt3r</category><category>model-release</category><category>model-training</category><category>mixture-of-experts</category><category>inference</category><category>model-optimization</category><category>sandboxing</category><category>alignment</category><category>cybersecurity</category><category>agent-runtime</category><category>throughput</category><category>quantization</category><category>telemetry</category><category>real-time-detection</category></item><item><title> GPT-Realtime-2, -Translate, and -Whisper: new SOTA realtime voice APIs</title><link>https://news.smol.ai/issues/26-05-07-gpt-realtime-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-07-gpt-realtime-2/</guid><description>**OpenAI** released **GPT-Realtime-2**, a voice model with **GPT-5-class reasoning**, tool use, interruption handling, and extended context windows up to **128K tokens**, achieving top scores on **Big Bench Audio** and **Conversational Dynamics** benchmarks. They also launched a **Chrome plugin for Codex** enabling browser control and multitasking, and introduced **GPT-5.5 with Trusted Access for Cyber** for secure defensive workflows and red teaming. **Anthropic** introduced **Natural Language Autoencoders** for interpreting model activations as human-readable text, aiding interpretability and debugging, while **Goodfire** proposed a neural geometry research agenda focusing on **manifolds** as primitives for neural network behavior. Anthropic also announced **The Anthropic Institute** to advance AI safety and economic resilience research.</description><pubDate>Thu, 07 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>goodfireai</category><category>scale-ai</category><category>gpt-realtime-2</category><category>gpt-5.5</category><category>codex</category><category>micahcarroll</category><category>milesbrundage</category><category>ryanpgreenblatt</category><category>voice-models</category><category>streaming-translation</category><category>transcription</category><category>benchmarking</category><category>context-windows</category><category>browser-automation</category><category>cybersecurity</category><category>interpretability</category><category>neural-geometry</category><category>manifolds</category><category>ai-safety</category><category>rlhf</category></item><item><title>Anthropic-SpaceXai&apos;s 300MW/$5B/yr deal for Colossus I, ARR growth is 8000% annualized</title><link>https://news.smol.ai/issues/26-05-06-anthropic-xai/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-06-anthropic-xai/</guid><description>**Anthropic** announced a new **SpaceX compute partnership** to significantly increase capacity for **Claude** products, doubling **Claude Code&apos;s 5-hour rate limits** for Pro, Max, Team, and Enterprise users, removing peak-hour limit reductions, and substantially increasing API rate limits for **Opus** models. The deal grants Anthropic access to **Colossus 1** via **SpaceXAI**, with **Claude inference** expected to ramp up on Colossus soon. Anthropic also hosted a **&quot;Code with Claude&quot;** event featuring updates on Claude Code, GitHub-scale usage, and managed agents. Discussions highlighted compute bottlenecks, user reactions to limit changes, debates on managed-agent features, and ongoing safety/governance discourse around AGI trustworthiness.</description><pubDate>Wed, 06 May 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>spacex</category><category>x-ai</category><category>claude</category><category>claude-code</category><category>opus</category><category>colossus-1</category><category>nottombrown</category><category>_aidan_clark_</category><category>kipperrii</category><category>theamolavasare</category><category>alexalbert__</category><category>compute</category><category>rate-limiting</category><category>agent-platforms</category><category>inference</category><category>api</category><category>managed-agents</category><category>safety</category><category>governance</category><category>event</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-04-not-much/</guid><description>**AI Twitter Recap** highlights the shift from model-centric AI to **context pipelines** and **agent orchestration** as key performance drivers. Notably, **gpt-5.2-codex** and **gpt-5.3-codex** showed significant benchmark improvements through prompt and middleware tuning. The ecosystem around open harnesses like **Hermes**, **deepagents**, and **Flue** is rapidly evolving, with innovations in multi-agent coordination and model-agnostic orchestration. Developer workflows are adapting to coding agents such as **Codex** and **Claude Code**, with emerging challenges in pricing models due to high token usage in agentic workloads. The practical takeaway is that agent performance depends on the synergy of **model × harness × memory/context strategy**, not just model weights alone.</description><pubDate>Mon, 04 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>langchain</category><category>baseten</category><category>ollama</category><category>openrouter</category><category>gpt-5.2-codex</category><category>gpt-5.3-codex</category><category>anthony_maio</category><category>mason_drxy</category><category>hwchase17</category><category>sydneyrunkle</category><category>naroh</category><category>teknuim</category><category>vtrivedy</category><category>dbreunig</category><category>zachtratar</category><category>theo</category><category>petergostev</category><category>cheatyyyy</category><category>agent-orchestration</category><category>context-pipelines</category><category>coding-agents</category><category>pricing-models</category><category>multi-agent-systems</category><category>workflow-optimization</category><category>model-agnostic-orchestration</category><category>prompt-engineering</category><category>memory-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-05-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-05-not-much/</guid><description>**OpenAI** rolled out **GPT-5.5 Instant** as the new default for ChatGPT and API, enhancing **factuality, intelligence, image understanding, and tone** with stronger personalization features like saved memories and Gmail integration. OpenAI also shared infrastructure updates on a rebuilt **WebRTC stack** for voice and real-time API, aiming to reduce latency for speech-paced conversations. Developer tools expanded with an **Agents SDK for TypeScript**, sandbox agents, and open-source harnesses, improving coding and automation workflows. Discussions highlighted the importance of **Model–Harness–Task fit** over raw model quality for agent performance, with debates on agent coding UX and benchmarks. Community sentiment praises GPT-5.5 for high-token-budget coding and non-coding tasks.</description><pubDate>Mon, 04 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>langchain</category><category>deepseek</category><category>gpt-5.5-instant</category><category>codex</category><category>sama</category><category>michpokrass</category><category>ericmitchellai</category><category>kimmonismus</category><category>reach_vb</category><category>vtrivedy10</category><category>sydneyrunkle</category><category>masondrxy</category><category>0xsero</category><category>teortaxestex</category><category>theethanding</category><category>finbarrtimbers</category><category>personalization</category><category>voice</category><category>real-time-api</category><category>webrtc</category><category>agent-frameworks</category><category>coding-agents</category><category>model-harness</category><category>benchmarking</category><category>automation</category><category>task-automation</category><category>developer-tools</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-20-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-20-not-much/</guid><description>**OpenAI** achieved a major math breakthrough by disproving a long-standing Erdős unit distance problem using a **general-purpose reasoning model**, marking a milestone in AI-driven formal science and long-horizon reasoning. The result was validated by prominent mathematicians like **Timothy Gowers** and OpenAI researcher **Hongxun Wu**, highlighting the model&apos;s advanced reasoning capabilities beyond prior AI math achievements. Meanwhile, **Cohere** released **Command A+** as an open-source Apache 2.0 licensed model, featuring a **218B MoE / 25B active** multimodal architecture supporting **48 languages** and optimized for low hardware requirements, runnable on as little as **2× H100 GPUs**. Benchmarks place Command A+ near **Claude 4.5 Haiku** in intelligence with strong non-hallucination but weaker scientific reasoning and coding. The architecture includes novel elements like a **parallel transformer block**, **shared experts**, and **LayerNorm over RMSNorm**.</description><pubDate>Mon, 04 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>cohere</category><category>command-a+</category><category>claude-3.7-sonnet</category><category>wtgowers</category><category>hongxunwu</category><category>aidangomez</category><category>nickfrosst</category><category>clementdelangue</category><category>eliebakouch</category><category>rasbt</category><category>sama</category><category>reinforcement-learning</category><category>reasoning</category><category>multimodality</category><category>model-architecture</category><category>model-optimization</category><category>model-releases</category><category>benchmarking</category><category>long-context</category><category>model-efficiency</category><category>transformers</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-22-not-much/</guid><description>**AI News for 5/4/2026-5/5/2026** highlights a shift in AI product development emphasizing **model + harness + workflow + UI + memory + economics** over model quality alone, with notable updates from **OpenAI Codex** and **Claude** including new features like **Appshots**, **auto mode**, and **Sonnet 4.6**. **DeepSeek** made a significant market impact by permanently discounting **DeepSeek-V4-Pro** by 75%, drastically improving cost/performance ratios compared to **Gemini 3.1 Pro**, **GPT-5.5**, and **Claude Opus 4.7**. Meanwhile, **Gemini 3.5 Flash** showed benchmark improvements but received mixed feedback on practical utility. The competitive landscape continues to tighten with **Qwen** and other Chinese frontier models.</description><pubDate>Mon, 04 May 2026 05:44:39 GMT</pubDate><category>openai</category><category>claude</category><category>deepseek</category><category>gemini</category><category>qwen</category><category>codex</category><category>deepseek-v4-pro</category><category>gemini-3.5-flash</category><category>gemini-3.1-pro</category><category>gpt-5.5</category><category>claude-opus-4.7</category><category>gdb</category><category>dzhng</category><category>signulll</category><category>teortaxestex</category><category>ajambrosino</category><category>reach_vb</category><category>theo</category><category>claudedevs</category><category>_mohansolo</category><category>artificialanlys</category><category>scaling01</category><category>yuchenj_uw</category><category>kimmonismus</category><category>officiallogank</category><category>designarena</category><category>alezander907</category><category>giffmana</category><category>jeremyphoward</category><category>hamelhusain</category><category>model-performance</category><category>cost-curves</category><category>agent-products</category><category>workflow-optimization</category><category>product-differentiation</category><category>benchmarking</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-05-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-05-01-not-much/</guid><description>**xAI released Grok 4.3**, improving cost/performance with a **53 Intelligence Index score**, 4 points higher than Grok 4.20, and significant gains on **GDPval-AA** and **τ²-Bench Telecom**. However, accuracy tradeoffs raised reliability concerns. Community opinions are mixed, with some praising token-efficiency and others noting regressions and pricing concerns. **DeepSeek V4 Pro** emerges as a leading open-weight coding/agent model, comparable to **Codex** and **Claude Code**, featuring a 1M context window and efficient attention mechanisms. Benchmarking shows open-weight models like **Kimi K2.6**, **MiMo V2.5 Pro**, and **DeepSeek V4 Pro** closing the gap with closed models such as **Gemini 3.1 Pro Preview**, **Claude Opus 4.7**, and **GPT-5.5**. DeepSeek&apos;s multimodal efforts focus on explicit spatial grounding with a novel &quot;point while thinking&quot; approach using **DeepSeek-ViT** and CSA compression.</description><pubDate>Fri, 01 May 2026 05:44:39 GMT</pubDate><category>xai</category><category>deepseek</category><category>artificial-analysis</category><category>andon-labs</category><category>grok-4.3</category><category>deepseek-v4-pro</category><category>kimi-k2.6</category><category>mimo-v2.5-pro</category><category>gemini-3.1-pro</category><category>claude-opus-4.7</category><category>gpt-5.5</category><category>deepskvit</category><category>scaling01</category><category>teortaxestex</category><category>omarsar0</category><category>benchmarking</category><category>cost-efficiency</category><category>agentic-ai</category><category>token-efficiency</category><category>attention-mechanisms</category><category>inference-speed</category><category>multimodality</category><category>spatial-reasoning</category><category>model-architecture</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-30-not-much/</guid><description>**OpenAI&apos;s GPT-5.5** achieves top-tier performance in long-horizon cyber tasks, matching or surpassing **Claude Mythos Preview** with a **71.4%** pass rate and showing ongoing improvement beyond **100M tokens** inference. OpenAI also released an **Advanced Account Security** update for ChatGPT enhancing phishing resistance. The **Codex** update expands beyond coding to general computer tasks, improving speed by up to **42%** and introducing role-based onboarding and app integrations. Economically, **GPT-5.5 Pro** shows a slight SOTA improvement on **CritPt** with **~60% lower cost** and token use compared to GPT-5.4 Pro. In open-weight models, **Qwen3.6 27B** leads under 150B parameters with an **Intelligence Index score of 46**, featuring **262K context**, native multimodal input, and efficient BF16 weights. Tencent&apos;s **Hy3-preview** (295B total, 21B active MoE) scores 42 on the Intelligence Index with strong scientific reasoning on **CritPt**. xAI&apos;s **Grok 4.3** shows sharp improvements on agentic benchmarks with reduced cost.</description><pubDate>Thu, 30 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>x-ai</category><category>tencent</category><category>deepseek</category><category>gpt-5.5</category><category>claude-mythos-preview</category><category>gpt-5.5-pro</category><category>qwen3.6-27b</category><category>hy3-preview</category><category>grok-4.3</category><category>gemma-4-31b</category><category>glm-5.1</category><category>deepseek-v4-flash</category><category>sama</category><category>scaling01</category><category>cryps1s</category><category>polynoamial</category><category>ajambrosino</category><category>arix</category><category>cybersecurity</category><category>model-efficiency</category><category>multimodality</category><category>model-benchmarking</category><category>agentic-ai</category><category>model-cost-optimization</category><category>context-windows</category><category>model-performance</category><category>open-weight-models</category><category>software-integration</category><category>security-updates</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-29-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-29-not-much/</guid><description>**OpenAI** is expanding **Codex** from a coding tool to a general work surface with persistent context, tools, integrations, and team rollout, including **Codex-only seats with $0 seat fee** for Business/Enterprise customers through June. Performance improvements focus on agent-loop systems engineering, achieving up to **40% faster agentic workflows** via WebSocket mode on the Responses API. **VS Code** enhances coding-agent UX with semantic indexing, cross-repo search, chat session insights, and prompt/agent evaluation extensions. **Cursor** launches a **Cursor SDK** to enable programmable agent infrastructure for CI/CD, automations, and embedded agents, signaling a shift toward headless agent runtimes and usage-based economics. Research highlights **Agentic Harness Engineering** improving Terminal-Bench 2 pass@1 from **69.7% to 77.0%**, surpassing human-designed baselines and reducing token use by **12%**. Related work on **HALO** shows recursive self-improving agents with significant AppWorld score improvements. **LangChain’s Deep Agents** introduces **Harness Profiles** for model-specific harness tuning and deployability.</description><pubDate>Wed, 29 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>microsoft</category><category>cursor_ai</category><category>langchain-ai</category><category>codex</category><category>omarsar0</category><category>samhogan</category><category>kimmonismus</category><category>reach_vb</category><category>pierceboggan</category><category>agentic-harness-engineering</category><category>agent-loop-systems-engineering</category><category>performance-optimization</category><category>semantic-indexing</category><category>prompt-evaluation</category><category>software-engineering</category><category>sdk-development</category><category>model-tuning</category><category>recursive-self-improvement</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-28-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-28-not-much/</guid><description>**vLLM v0.20.0** introduces significant improvements in memory and MoE serving efficiency, including **TurboQuant 2-bit KV cache** for **4× KV capacity** and a **2.1% latency improvement**. The update supports multiple hardware platforms like **DeepSeek V4 MegaMoE on Blackwell**, Jetson Thor, ROCm, Intel XPU, and Grace-Blackwell setups. Early benchmarks show **DeepSeek V4 Pro** on **B300** hardware can be up to **8× faster** than H200. The ecosystem is rapidly adopting day-0 support for new open models such as **Poolside Laguna XS.2**, **Ling-2.6-flash**, and **NVIDIA Nemotron 3 Nano Omni**. 

**Poolside** released **Laguna XS.2**, a **33B total / 3B active MoE** coding model under **Apache 2.0**, capable of running on a single GPU, with hybrid attention and FP8 KV cache, performing near **Qwen-3.5**. 

**NVIDIA** launched **Nemotron 3 Nano Omni**, a **30B / A3B multimodal MoE** with **256K context**, supporting text, image, video, audio, and documents, with immediate distribution across multiple platforms. Discussions highlighted tradeoffs in quantization methods and a shift away from CUDA lock-in towards heterogeneous accelerator support.</description><pubDate>Tue, 28 Apr 2026 05:44:39 GMT</pubDate><category>vllm</category><category>poolside</category><category>nvidia</category><category>opensrouter</category><category>lmstudio</category><category>ollama</category><category>unsloth</category><category>fal</category><category>fireworks</category><category>deepinfra</category><category>togethercompute</category><category>baseten</category><category>canonical</category><category>vllm-0.20.0</category><category>poolside-laguna-xs.2</category><category>ling-2.6-flash</category><category>nemotron-3-nano-omni</category><category>qwen-3.5</category><category>jeremyphoward</category><category>maharshii</category><category>teortaxestex</category><category>aymericroucher</category><category>piotrz</category><category>memory-optimization</category><category>mixture-of-experts</category><category>model-optimization</category><category>inference-speed</category><category>quantization</category><category>model-deployment</category><category>multimodality</category><category>hardware-optimization</category><category>model-benchmarking</category><category>open-models</category><category>agentic-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-27-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-27-not-much/</guid><description>**OpenAI** loosens its **Azure exclusivity**, allowing distribution across **Google TPU**, **AWS Trainium**, and **Bedrock** with commitments through **2032** and revenue share through **2030**. **GPT-5.5** shows improved benchmarks but is not uniformly dominant, ranking variably across coding, document, math, and vision tasks. GitHub&apos;s **Copilot** shifts to usage-based billing starting June 1, reflecting increased runtime costs. **OpenAI** open-sourced **Symphony**, an orchestration layer for issue tracking and Codex agents. **Xiaomi** released **MiMo-V2.5** and **MiMo-V2.5-Pro**, large context models with up to **1M-token context** and trillions of tokens trained, emphasizing complex agent and omni-modal capabilities. **Kimi K2.6** leads OpenRouter&apos;s leaderboard, noted for coding and long-horizon agent capabilities with large-scale sub-agent coordination.</description><pubDate>Mon, 27 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>microsoft</category><category>google</category><category>amazon</category><category>github</category><category>xiaomi</category><category>openai-devs</category><category>vllm_project</category><category>kimi-moonshot</category><category>gpt-5.5</category><category>gpt-5.4</category><category>opus-4.7</category><category>mimo-v2.5-pro</category><category>mimo-v2.5</category><category>kimi-k2.6</category><category>codex</category><category>copilot</category><category>sama</category><category>scaling01</category><category>kimmonismus</category><category>ajassy</category><category>simonw</category><category>htihle</category><category>arena</category><category>gdb</category><category>hangsiin</category><category>eliebakouch</category><category>_luofuli</category><category>teortaxestex</category><category>model-distribution</category><category>cloud-computing</category><category>benchmarking</category><category>usage-based-billing</category><category>model-orchestration</category><category>open-source</category><category>large-context-models</category><category>agent-scaling</category><category>coding</category><category>model-training</category><category>fp8</category><category>attention-mechanisms</category><category>multi-agent-systems</category></item><item><title>DeepSeek v4</title><link>https://news.smol.ai/issues/26-04-24-deepseek-v4/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-24-deepseek-v4/</guid><description>**DeepSeek-V4** technical release features a **1.6T-parameter MoE with 49B active parameters** and **1M-token context**, showcasing hybrid attention and compressed KV schemes for major memory reductions. It ranks as the **#2 open-weights reasoning model** behind **Kimi K2.6** but has a high hallucination rate and higher serving costs. Hardware-model co-design is emphasized, with **NVIDIA Blackwell Ultra** delivering **150+ TPS/user** and support for FP4 and FP8 quantization enabling deployment on single nodes. Positioning among open Chinese models is competitive with **GLM-5.1** and **Xiaomi MiMo V2.5 Pro**. Meanwhile, **OpenAI launched GPT-5.5 and GPT-5.5 Pro APIs** with a **1M context window**, focusing on improved long-running workflows and token efficiency, quickly integrated into tools like **GitHub Copilot** and **Cursor**. *&quot;GPT-5.5 handles complex, tool-heavy, ambiguous workflows with fewer retries,&quot;* highlighting rapid distribution and agent integration.</description><pubDate>Fri, 24 Apr 2026 05:44:39 GMT</pubDate><category>deepseek</category><category>nvidia</category><category>openai</category><category>lambdaapi</category><category>togethercompute</category><category>xiaomi</category><category>deepseek-v4</category><category>deepseek-v4-pro</category><category>deepseek-v4-flash</category><category>kimi-k2.6</category><category>glm-5.1</category><category>xiaomi-mimo-v2.5-pro</category><category>gpt-5.5</category><category>gpt-5.5-pro</category><category>scaling01</category><category>ben_burtenshaw</category><category>artificialanlys</category><category>long-context</category><category>mixture-of-experts</category><category>model-quantization</category><category>memory-optimization</category><category>hardware-model-co-design</category><category>inference-speed</category><category>agent-integration</category><category>token-efficiency</category><category>model-deployment</category><category>open-weights</category><category>reasoning</category><category>hallucination-detection</category></item><item><title>GPT 5.5</title><link>https://news.smol.ai/issues/26-04-23-gpt-55/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-23-gpt-55/</guid><description>**OpenAI launched GPT-5.5** as its new flagship model for &quot;real work and powering agents,&quot; immediately available in ChatGPT and Codex but with delayed API access due to enhanced safety requirements. The model features improved token efficiency and supports longer multi-step execution with tool use and self-checking. Pricing is set at **$5/$30 per million tokens for GPT-5.5** and **$30/$180 for GPT-5.5 Pro**, roughly double the cost of GPT-5.4. The release includes significant Codex upgrades such as browser control, document handling, and OS-wide dictation. Early reactions are mixed but generally positive, noting improvements in coding and long-horizon tasks, though some benchmarks show incremental gains and hallucination issues persist. Third-party ecosystem support like Hermes Agent integration appeared quickly.</description><pubDate>Thu, 23 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>scaling01</category><category>anthropic</category><category>teknium</category><category>gpt-5.5</category><category>gpt-5.4</category><category>gpt-5.5-pro</category><category>sama</category><category>reach_vb</category><category>agentic-ai</category><category>token-efficiency</category><category>tool-use</category><category>self-checking</category><category>coding</category><category>long-horizon-planning</category><category>model-pricing</category><category>api-access</category><category>model-safety</category><category>software-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-22-not-much/</guid><description>**Alibaba** released **Qwen3.6-27B**, a dense, Apache 2.0 open coding model with thinking and non-thinking modes, outperforming the larger Qwen3.5-397B-A17B on multiple coding benchmarks including SWE-bench and Terminal-Bench. It supports native vision-language reasoning over images and video, with immediate ecosystem support from vLLM, Unsloth, ggml, and Ollama. **OpenAI** open-sourced a practical **Privacy Filter** model for PII detection and masking, a 1.5B parameter token-classification model with a 128k context window aimed at enterprise redaction tasks. **Xiaomi** announced **MiMo-V2.5-Pro** and **MiMo-V2.5** models, emphasizing software engineering advances, long-horizon agents, and large context windows (up to 1M tokens), with strong benchmark results and integrations with Hermes and Nous. At **Google Cloud Next**, **Google** and **Google DeepMind** unveiled 8th-gen TPUs (TPU 8t for training and TPU 8i for inference) with claims of scaling to a million TPUs in a cluster, and launched the **Gemini Enterprise Agent Platform** evolving Vertex AI with Agent Studio and access to 200+ models including **Gemini 3.1 Pro** and **Gemini 3.1 Flash Image**. This marks a significant vertical integration of hardware, models, and enterprise tooling.</description><pubDate>Wed, 22 Apr 2026 05:44:39 GMT</pubDate><category>alibaba</category><category>openai</category><category>xiaomi</category><category>google</category><category>google-deepmind</category><category>vllm_project</category><category>unsloth</category><category>ggml</category><category>ollama</category><category>arena</category><category>nous-research</category><category>qwen3.6-27b</category><category>qwen3.5-397b-a17b</category><category>privacy-filter</category><category>mimo-v2.5-pro</category><category>mimo-v2.5</category><category>gemini-3.1-pro</category><category>gemini-3.1-flash-image</category><category>alibaba_qwen</category><category>clementdelangue</category><category>altryne</category><category>eliebakouch</category><category>mervenoyann</category><category>xiaomimo</category><category>sundarpichai</category><category>scaling01</category><category>open-models</category><category>multimodality</category><category>vision</category><category>tokenization</category><category>pii-detection</category><category>privacy</category><category>enterprise-ai</category><category>agentic-ai</category><category>benchmarking</category><category>long-context</category><category>model-deployment</category><category>hardware-optimization</category><category>model-integration</category><category>software-engineering</category></item><item><title>GPT-Image-2</title><link>https://news.smol.ai/issues/26-04-21-image-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-21-image-2/</guid><description>**OpenAI** launched **GPT-Image-2**, enhancing image generation with improved text rendering, layout fidelity, editing, multilingual support, and &quot;thinking&quot; capabilities. It supports generating slides, infographics, diagrams, UI mockups, and QR codes, and integrates with tools like **Figma**, **Canva**, **Adobe Firefly**, and **Hermes Agent**. Benchmarks show GPT-Image-2 leads image generation tasks with a +242 Elo advantage. **Hugging Face** released **ml-intern**, an open-source agent automating post-training research loops, improving scientific reasoning and healthcare benchmarks significantly. **Hermes** is evolving into a richer local/open agent platform with enhanced multi-process orchestration capabilities.</description><pubDate>Tue, 21 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>hugging-face</category><category>figma</category><category>canva</category><category>adobe</category><category>nous-research</category><category>gpt-image-2</category><category>qwen3-1.7b</category><category>codex</category><category>clementdelangue</category><category>lewtun</category><category>gdb</category><category>nickaturley</category><category>mark_k</category><category>petergostev</category><category>tekninum</category><category>mayank_022</category><category>image-generation</category><category>multilingual-models</category><category>model-integration</category><category>benchmarking</category><category>agent-infrastructure</category><category>multi-process-systems</category><category>fine-tuning</category><category>scientific-reasoning</category><category>healthcare-ai</category><category>hierarchical-decomposition</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-20-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-20-not-much/</guid><description>**Moonshot&apos;s Kimi K2.6** is a major open-weight **1T-parameter MoE** model featuring **32B active parameters**, **384 experts**, **MLA attention**, **256K context window**, native multimodality, and **INT4 quantization**. It supports day-0 integration with platforms like **vLLM**, **OpenRouter**, **Cloudflare Workers AI**, and others, showcasing state-of-the-art performance on benchmarks such as **HLE w/ tools 54.0**, **SWE-Bench Pro 58.6**, and **Math Vision w/ python 93.2**. The model excels in **long-horizon execution** with over **4,000 tool calls**, **12+ hour continuous runs**, and **300 parallel sub-agents**. Meanwhile, **Alibaba&apos;s Qwen3.6-Max-Preview** previewed enhanced **agentic coding**, improved world knowledge, and instruction following, with notable performance on **AIME 2026 #15** and ranking in **Code Arena**. **Hermes Agent** is rapidly expanding its ecosystem, surpassing **100K GitHub stars** and integrating with tools like **Ollama** and **Copilot CLI**, while pioneering advanced multi-agent orchestration techniques such as **stateless ephemeral units**, **LLM-driven replanning**, and **dynamic context injection**. These developments highlight the competitive momentum of Chinese open and semi-open labs in coding and agent models.</description><pubDate>Mon, 20 Apr 2026 05:44:39 GMT</pubDate><category>moonshot</category><category>alibaba</category><category>vllm</category><category>openrouter</category><category>cloudflare</category><category>baseten</category><category>mlx</category><category>nous-research</category><category>opencode</category><category>ollama</category><category>kimi-k2.6</category><category>qwen-3.6-max-preview</category><category>mixture-of-experts</category><category>multimodality</category><category>int4-quantization</category><category>long-context</category><category>agentic-coding</category><category>multi-agent-systems</category><category>model-orchestration</category><category>memory-consolidation</category><category>llm-driven-replanning</category><category>dynamic-context-injection</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-17-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-17-not-much/</guid><description>**Anthropic** launched **Claude Design**, a prototyping tool powered by **Claude Opus 4.7**, targeting design workflows and competing with **Figma** and others. Benchmarks show **Opus 4.7** leading in coding and text tasks, with improved efficiency and adaptive reasoning, though early user feedback noted some regressions and stability issues. Discussions highlighted its cost-efficiency and agentic capabilities compared to **Gemini 3.1 Pro** and **GPT-5.4**. Meanwhile, **OpenAI**&apos;s Codex updates introduced advanced computer-use features enabling fast, agentic control of desktop apps and enterprise software, signaling progress toward practical AGI-like agents.</description><pubDate>Fri, 17 Apr 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>claude-opus-4.7</category><category>gemini-3.1-pro</category><category>gpt-5.4</category><category>claude-code</category><category>codex</category><category>claudeai</category><category>yuchenj_uw</category><category>kimmonismus</category><category>skirano</category><category>therundownai</category><category>arena</category><category>artificialanlys</category><category>victortaelin</category><category>emollick</category><category>alexalbert__</category><category>theo</category><category>scaling01</category><category>reach_vb</category><category>kr0der</category><category>hamelhusain</category><category>mattrickard</category><category>matvelloso</category><category>gdb</category><category>agentic-ai</category><category>model-benchmarking</category><category>adaptive-reasoning</category><category>cost-efficiency</category><category>computer-use</category><category>prototyping-tools</category><category>code-generation</category><category>model-performance</category><category>software-integration</category></item><item><title>Anthropic&apos;s Claude Opus 4.7</title><link>https://news.smol.ai/issues/26-04-16-opus-47/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-16-opus-47/</guid><description>**Anthropic** launched **Claude Opus 4.7**, its most capable Opus model yet, featuring stronger coding and agentic performance, a new tokenizer, and improved long-context handling with a new **xhigh** reasoning tier. Benchmarks show substantial gains, including **SWE-bench Pro 64.3%**, **SWE-bench Verified 87.6%**, and **TerminalBench 69.4%**, with top rankings on **Vals Index** and **GDPval-AA**. Technical changes include a new tokenizer and increased image input resolution to **3.75MP**. Some long-context benchmarks showed mixed results, with a shift in focus from MRCR to Graphwalks. Adoption was rapid across tools like **Cursor**, **VS Code**, **Replit Agent**, and **Perplexity**. Meanwhile, **OpenAI** expanded **Codex** into a broader computer agent with Mac computer use, in-app browser, image generation/editing, 90+ plugins, multi-terminal support, SSH remote devbox access, and richer file previews. A new vertical life-sciences model, **GPT-Rosalind**, was also introduced.</description><pubDate>Thu, 16 Apr 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>cursor</category><category>replit</category><category>perplexity-ai</category><category>microsoft</category><category>claude-opus-4.7</category><category>codex</category><category>gpt-rosalind</category><category>bcherny</category><category>kimmonismus</category><category>scaling01</category><category>valsai</category><category>artificialanlys</category><category>natolambert</category><category>nrehiew_</category><category>coding</category><category>agentic-ai</category><category>tokenization</category><category>long-context</category><category>benchmarking</category><category>image-processing</category><category>software-engineering</category><category>computer-use</category><category>plugin-integration</category><category>multi-terminal-support</category><category>ssh-access</category><category>model-expansion</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-15-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-15-not-much/</guid><description>**OpenAI** expanded its Agents SDK by separating the agent harness from compute/storage, enabling long-running, durable agents with features like file/computer use, skills, memory, and compaction. The harness is now open-source and supports execution via partner sandboxes, fostering a new ecosystem with integrations from **Cloudflare**, **Modal**, **Vercel**, and others. **Cloudflare** launched **Project Think**, a next-gen Agents SDK with durable execution and sandboxed code, alongside **Agent Lee**, a prompt-driven UI agent using sandboxed TypeScript, and introduced real-time voice pipelines and browser automation tools. **Hermes Agent** focuses on persistent skill formation by learning from completed workflows, positioning itself as a professional agent distinct from GUI-first assistants like OpenClaw. *&quot;Hermes autonomously backfills tracking data, updates cron jobs, and saves workflows as reusable skills,&quot;* highlighting its advanced workflow management capabilities.</description><pubDate>Wed, 15 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>cloudflare</category><category>modal</category><category>vercel</category><category>akshat_b</category><category>whoiskatrin</category><category>aninibread</category><category>braydenwilmoth</category><category>korinne_dev</category><category>kathyyliao</category><category>joshesye</category><category>chooseliberty</category><category>neoaiforecast</category><category>agents-sdk</category><category>sandboxing</category><category>durable-execution</category><category>state-management</category><category>voice-processing</category><category>browser-automation</category><category>workflow-automation</category><category>skill-formation</category><category>open-source</category><category>prompt-driven-ui</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-13-not-much/</guid><description>**Harness engineering** is emerging as a key discipline in AI agent development, emphasizing components like filesystems, memory, and retries beyond just models. **OpenAI&apos;s Codex** is expanding agentic coding workflows beyond software engineering, including codebase understanding and bug triage. Tooling trends show convergence on multi-agent orchestration, observability, and remote control, with **GitHub Copilot**, **Cursor**, and **LangChain** advancing these capabilities. The **Hermes Agent v0.9.0** release introduces a local web dashboard and enhanced security, gaining community traction over **OpenClaw** for UX and efficiency. The open agent ecosystem is growing with projects like **Open Agents** and **DeepAgent** providing modular stacks and runtimes.</description><pubDate>Mon, 13 Apr 2026 05:44:39 GMT</pubDate><category>openai</category><category>github</category><category>cursor</category><category>langchain</category><category>nous-research</category><category>codex</category><category>andrew_ng</category><category>steve_yegge</category><category>gabrielchua</category><category>giffmana</category><category>rhys_sullivan</category><category>teknium</category><category>shaun_furman</category><category>dabit3</category><category>robinebers</category><category>zainanzhou</category><category>nicoalbanese10</category><category>bromann</category><category>elliothyun</category><category>tiagonbotelho</category><category>pierceboggan</category><category>sydneyrunkle</category><category>agent-harnesses</category><category>multi-agent-systems</category><category>software-engineering</category><category>tooling</category><category>orchestration</category><category>observability</category><category>remote-control</category><category>security-hardening</category><category>user-experience</category><category>open-source</category><category>community-engagement</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-10-not-much/</guid><description>**GLM-5.1** has reached **#3 on Code Arena**, surpassing **Gemini 3.1** and **GPT-5.4**, and matching **Claude Sonnet 4.6** in coding performance. **Z.ai** now holds the **#1 open model rank** close to the top overall. The advisor pattern, combining a cheap executor with an expensive advisor, is gaining traction, improving performance and efficiency in models like **Haiku + Opus** and **Sonnet + Opus**. **Alibaba&apos;s Qwen Code v0.14.x** introduces orchestration features including remote control channels, cron tasks, and sub-agent model selection. Model routing is becoming a product-level concern due to specialization and spikiness in top models such as **Opus** and **GPT-5.4**. The **Hermes Agent** ecosystem shows strong momentum with a new workspace mobile app, FAST mode for **OpenAI/GPT-5.4**, and over **50k GitHub stars**. Practitioners report Hermes as a reliable agent framework, with local Qwen3-Coder-Next 80B 4-bit replacing parts of workflows previously reliant on Claude Code. The harness layer is emerging as a key abstraction in agent frameworks.</description><pubDate>Fri, 10 Apr 2026 05:44:39 GMT</pubDate><category>z-ai</category><category>anthropic</category><category>berkeley</category><category>langchain</category><category>alibaba</category><category>openai</category><category>glm-5.1</category><category>gemini-3.1</category><category>gpt-5.4</category><category>claude-3-sonnet</category><category>haiku</category><category>opus</category><category>sonnet</category><category>qwen-3.6-plus</category><category>qwen3-coder-next-80b</category><category>zixuan_li</category><category>akshay_pachaar</category><category>harrison_chase</category><category>walden_yan</category><category>yuchen_jin</category><category>sentdex</category><category>model-performance</category><category>agent-frameworks</category><category>orchestration</category><category>model-routing</category><category>fine-tuning</category><category>agent-harness</category><category>model-selection</category><category>workflow-automation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-09-not-much/</guid><description>**Anthropic&apos;s Mythos** and **OpenAI&apos;s** upcoming restricted cyber-capable models are central to recent discussions, with debates on their security realism and evaluation methods. **LangChain&apos;s Deep Agents deploy** introduces an open memory, model-agnostic agent harness architecture emphasizing open protocols and memory ownership. Sandboxes are gaining prominence as a core infrastructure for reinforcement learning, with labs running up to **100K concurrent sandboxes** aiming for **1M**. The **Hermes Agent** by Nous continues to gain traction with new integrations and features like a web-based HUD and token cost tracking.</description><pubDate>Thu, 09 Apr 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>langchain</category><category>nous-research</category><category>mythos</category><category>kimmonismus</category><category>paul_cal</category><category>gneubig</category><category>kentonvarda</category><category>boazbaraktcs</category><category>ylecun</category><category>deanwball</category><category>hwchase17</category><category>vtrivedy10</category><category>sarahcat21</category><category>aijoey</category><category>cybersecurity</category><category>sandboxing</category><category>reinforcement-learning</category><category>agent-architecture</category><category>memory-management</category><category>model-deployment</category><category>software-security</category><category>evaluation-methods</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-08-not-much/</guid><description>**Meta Superintelligence Labs** launched **Muse Spark**, a natively multimodal reasoning model featuring tool use, visual chain of thought, and multi-agent orchestration. It is live on **meta.ai** and the Meta AI app with a private API preview and plans for open-sourcing future versions. Independent benchmarks rank Muse Spark highly, with strong performance on intelligence indices and efficiency, notably using over 10× less compute than **Llama 4 Maverick**. Key technical highlights include training efficiency, test-time scaling, and parallel multi-agent inference. Community testing shows strengths in image-to-code and one-shot game generation. Additionally, **Zhipu AI&apos;s GLM-5.1** is recognized as a leading open-weight model with architecture similar to DeepSeek-V3.2.</description><pubDate>Wed, 08 Apr 2026 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>zhipu-ai</category><category>deepseek</category><category>muse-spark</category><category>llama-4-maverick</category><category>glm-5.1</category><category>deepseek-v3.2</category><category>alexandr_wang</category><category>shengjia_zhao</category><category>jack_w_rae</category><category>ananyaku</category><category>_jasonwei</category><category>artificialanlys</category><category>valsai</category><category>epochairesearch</category><category>scale_ai</category><category>matthuang</category><category>omarsar0</category><category>skirano</category><category>mattdeitke</category><category>garrytan</category><category>sebastian_raschka</category><category>multimodality</category><category>tool-use</category><category>visual-chain-of-thought</category><category>multi-agent-systems</category><category>training-efficiency</category><category>test-time-scaling</category><category>parallel-inference</category><category>image-to-code</category><category>model-benchmarking</category><category>model-architecture</category></item><item><title>Anthropic @ $30B ARR, Project GlassWing and Claude Mythos Preview — first model too dangerous to release since GPT-2</title><link>https://news.smol.ai/issues/26-04-06-anthropic-mythos/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-06-anthropic-mythos/</guid><description>**Anthropic** strategically challenges **OpenAI** amid its upcoming IPO concerns by announcing a jump from **$19B ARR in March** to **$30B ARR in April**, highlighting a differential growth rate and higher cost efficiency. The company also revealed **Claude Mythos**, rumored as the largest successful training run, now restricted under **Project Glasswing** due to its dangerous capabilities. This model reportedly found thousands of high-severity vulnerabilities across major operating systems and browsers, showcasing unprecedented strategic thinking, situational awareness, and creative reward hacking. Notable figures like **Nicolas Carlini** and **Sam Bowman** commented on the model&apos;s advanced behaviors and unexpected internet access. Anthropic&apos;s disclosures emphasize both impressive business growth and groundbreaking AI capabilities.</description><pubDate>Tue, 07 Apr 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>claude-mythos</category><category>nicolas_carlini</category><category>sam_bowman</category><category>model-training</category><category>model-capabilities</category><category>security-vulnerabilities</category><category>strategic-thinking</category><category>reward-hacking</category><category>situational-awareness</category><category>benchmarking</category><category>model-restrictions</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-07-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-07-not-much/</guid><description>**Hermes Agent** is gaining attention as a leading open agent stack with features like self-improving skills, persistent memory, and a self-improvement loop. Its new **Manim skill** enables generation of math/technical animations, expanding agent capabilities. The Hermes ecosystem is rapidly growing with GUI tools, WebUI, HUD updates, OAuth support, and integrations. An open training-data movement for agents is emerging, focusing on sharing reusable behavioral data and harness traces. Meanwhile, **Anthropic&apos;s Claude Code** faces distribution and policy challenges, with reports of restrictions and unreliability impacting third-party coding agents, highlighting issues with subscription economics for always-on agents. *&quot;Claude Code now errors if used to analyze Claude Code source&quot;* and *&quot;basically unusable&quot;* are key community sentiments.</description><pubDate>Mon, 06 Apr 2026 05:44:39 GMT</pubDate><category>nous-research</category><category>anthropic</category><category>theo</category><category>clementdelangue</category><category>badlogicgames</category><category>yuchenj_uw</category><category>self-improving-skills</category><category>agent-architecture</category><category>memory-persistence</category><category>animation-generation</category><category>open-training-data</category><category>coding-agents</category><category>subscription-models</category><category>policy-restrictions</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-14-not-much/</guid><description>**Google** introduced **Skills in Chrome**, enabling reusable browser workflows with Gemini prompts and a library of ready-made Skills, enhancing end-user agentization. **Tencent** teased **HYWorld 2.0**, an open-source 3D world model generating editable scenes from a single image. **Google DeepMind** released **Gemini Robotics-ER 1.6**, improving visual/spatial reasoning for robotics with 93% instrument-reading success. **OpenAI** expanded Trusted Access with **GPT-5.4-Cyber**, a fine-tuned model for defensive security workflows. **Hugging Face** launched **Kernels** on the Hub, offering GPU kernel repos with 1.7x–2.5x speedups. **Cursor** showcased a multi-agent CUDA optimization system with a 38% speedup across 235 problems. The **Hermes Agent** stack advanced to v0.9.0 with enhanced reliability, memory management, and integrations, while **LangChain** pushed **deepagents 0.5** toward deployable, multi-tenant async systems with multimodal support and prompt caching. *&quot;Hermes’ key advantage is operational stability, extensibility, and deployability.&quot;*</description><pubDate>Mon, 06 Apr 2026 05:44:39 GMT</pubDate><category>google</category><category>tencent</category><category>google-deepmind</category><category>openai</category><category>hugging-face</category><category>cursor</category><category>langchain</category><category>gemini</category><category>gemini-robotics-er-1.6</category><category>gpt-5.4-cyber</category><category>deepagents-0.5</category><category>clementdelangue</category><category>dylantfwang</category><category>antoinersx</category><category>steveschoettler</category><category>teknium</category><category>aiqiang888</category><category>sydneyrunkle</category><category>agent-infrastructure</category><category>cuda-optimization</category><category>visual-reasoning</category><category>spatial-reasoning</category><category>gpu-kernels</category><category>multi-agent-systems</category><category>memory-management</category><category>async-systems</category><category>multimodality</category><category>prompt-caching</category><category>software-engineering</category><category>robotics</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-03-not-much/</guid><description>**Gemma 4** was launched by **Google** under an **Apache 2.0 license**, marking a significant open-model release focused on **reasoning, agentic workflows, multimodality, and on-device use**. It outperforms models 10x larger and has immediate ecosystem support including **vLLM**, **llama.cpp**, **Ollama**, **Intel hardware**, **Unsloth**, and **Hugging Face Inference Endpoints**. Local inference benchmarks showed strong performance on consumer hardware, including RTX 4090 and Mac mini M4. Early benchmarking praised its efficiency and ranking improvements over previous versions. Meanwhile, **Hermes Agent** emerged as a popular open-source agent harness, noted for stability and capability on long tasks, with users switching from OpenClaw to Hermes.</description><pubDate>Fri, 03 Apr 2026 05:44:39 GMT</pubDate><category>google</category><category>huggingface</category><category>intel</category><category>ollama</category><category>unsloth</category><category>gemma-4</category><category>fchollet</category><category>demishassabis</category><category>clementdelangue</category><category>quixiai</category><category>googlegemma</category><category>ggerganov</category><category>osanseviero</category><category>maartengr</category><category>basecampbernie</category><category>prince_canuma</category><category>measure_plan</category><category>kimmonismus</category><category>anemll</category><category>arena</category><category>stochasticchasm</category><category>reach_vb</category><category>zeneca</category><category>everlier</category><category>erick_lindberg_</category><category>anomalistg</category><category>reasoning</category><category>agentic-workflows</category><category>multimodality</category><category>on-device-ai</category><category>local-inference</category><category>model-benchmarking</category><category>moe</category><category>vision</category><category>audio-processing</category><category>memory-optimization</category><category>open-source</category><category>model-performance</category></item><item><title>Gemma 4</title><link>https://news.smol.ai/issues/26-04-02-gemma-4/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-02-gemma-4/</guid><description>**Google DeepMind** released **Gemma 4**, a family of open-weight, multimodal models with long-context support up to **256K tokens** under an **Apache 2.0 license**, marking a major capability and licensing shift. The lineup includes **31B dense**, **26B MoE (A4B)**, and two edge models (**E4B**, **E2B**) optimized for local and edge deployment with native multimodal support (text, vision, audio). Early benchmarks show **Gemma-4-31B** ranking #3 among open models and strong scientific reasoning performance with **85.7% GPQA Diamond**. Day-0 ecosystem support includes **llama.cpp**, **Ollama**, **vLLM**, and **LM Studio**, with notable local inference performance on hardware like **M2 Ultra** and **RTX 4090**. The architecture features hybrid attention and MoE layering, diverging from standard transformers. Community and developer engagement is high, with rapid adoption and tooling integration.</description><pubDate>Thu, 02 Apr 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>gemma-4</category><category>gemma-4-31b</category><category>gemma-4-26b-a4b</category><category>jeffdean</category><category>_philschmid</category><category>rasbt</category><category>ggerganov</category><category>clattner_llvm</category><category>julien_c</category><category>clementdelangue</category><category>multimodality</category><category>long-context</category><category>model-architecture</category><category>moe</category><category>local-inference</category><category>model-optimization</category><category>function-calling</category><category>quantization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-04-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-04-01-not-much/</guid><description>**Arcee’s Trinity-Large-Thinking** was released with **open weights under Apache 2.0**, featuring a **400B total / 13B active** model size and strong agentic performance, ranking **#2 on PinchBench**. **Z.ai’s GLM-5V-Turbo** is a **vision coding model** with **native multimodal fusion** and a **CogViT encoder**, integrated into multiple platforms. **TII’s Falcon Perception** offers an **open-vocabulary referring expression segmentation model** with an **early-fusion transformer** and a competitive **0.3B OCR model**. **H Company’s Holo3** is a GUI-navigation model family based on **Qwen3.5**. A **Claude Code leak** revealed a minimalist agent core with a **4-layer context compression stack**, **40+ tool modular architecture**, and advanced features like **task budget management** and **streaming tool execution**. The leak highlights Anthropic’s agent design and operational sophistication.</description><pubDate>Wed, 01 Apr 2026 05:44:39 GMT</pubDate><category>arcee</category><category>z-ai</category><category>tii</category><category>anthropic</category><category>h-company</category><category>trinity-large-thinking</category><category>glm-5v-turbo</category><category>falcon-perception</category><category>qwen-3.5</category><category>claude-4.6-opus</category><category>claude-sonnet-4.5</category><category>mark_mcquade</category><category>latkins</category><category>willccbb</category><category>xlr8harder</category><category>natolambert</category><category>craig_hewitt</category><category>zhihu_frontier</category><category>open-weights</category><category>agentic-performance</category><category>vision</category><category>multimodality</category><category>transformer-architecture</category><category>early-fusion</category><category>ocr</category><category>gui-navigation</category><category>context-compression</category><category>tooling</category><category>feature-flags</category><category>production-ablations</category><category>task-budget-management</category><category>streaming</category><category>modular-architecture</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-30-not-much/</guid><description>**Anthropic** introduced **computer use inside Claude Code** for closed-loop verification in a research preview for Pro/Max users, enhancing reliable app iteration. **OpenAI** released a **Codex plugin for Claude Code**, enabling cross-agent composition and signaling a shift toward composable coding harnesses. OpenAI also noted that late-night Codex tasks run longer, supporting background agent delegation. **Nous Research**&apos;s **Hermes Agent** saw rapid adoption due to better compaction, adaptability, and multi-agent profiles, evolving toward an agent OS abstraction. An ecosystem around Hermes includes tools for trace analytics, fine-tuning, and remote control, with debates on open-source versus proprietary agent infrastructure. Key themes include tooling, prompt/runtime orchestration, and review loops as critical factors beyond model capabilities.</description><pubDate>Mon, 30 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>nous-research</category><category>huggingface</category><category>claude-code</category><category>codex</category><category>hermes-agent</category><category>omarsar0</category><category>dkundel</category><category>reach_vb</category><category>theo</category><category>jayfarei</category><category>kaiostephens</category><category>icarushermes</category><category>winglian</category><category>clementdelangue</category><category>fchollet</category><category>closed-loop-verification</category><category>cross-agent-composition</category><category>agent-ecosystem</category><category>multi-agent-systems</category><category>runtime-orchestration</category><category>tooling</category><category>fine-tuning</category><category>remote-monitoring</category><category>privacy</category><category>sandboxing</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-27-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-27-not-much/</guid><description>**Anthropic** is reportedly introducing a new AI model tier called **Capybara**, which is larger and more intelligent than **Claude Opus 4.6**, showing improved performance in coding, academic reasoning, and cybersecurity. The model is speculated to be around **10 trillion parameters**, with **Google** potentially funding Anthropic&apos;s data center expansion. Meanwhile, **Zhipu** released **GLM-5.1**, advancing open coding models and narrowing the gap with closed models. Local inference economics are improving, highlighted by efficient deployments of **Qwen 3.5 14B**, **Qwen 27B**, and **Qwen3.5-35B** models with quantization techniques like **TurboQuant vLLM**. However, TurboQuant&apos;s benchmarking claims face criticism from researchers. Overall, the AI landscape shows aggressive scaling, local model deployment, and agent products gaining traction.</description><pubDate>Fri, 27 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>google</category><category>zhipu</category><category>claude-opus-4.6</category><category>capybara</category><category>glm-5.1</category><category>qwen-3.5-14b</category><category>qwen-27b</category><category>qwen3.5-35b</category><category>scaling01</category><category>yuchenj_uw</category><category>kimmonismus</category><category>m1astra</category><category>dejavucoder</category><category>iscienceluvr</category><category>gaoj0017</category><category>model-scaling</category><category>coding</category><category>academic-reasoning</category><category>cybersecurity</category><category>quantization</category><category>local-inference</category><category>model-benchmarking</category><category>inference-optimization</category><category>model-performance</category><category>agent-products</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-24-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-24-not-much/</guid><description>**Anthropic** advances agent infrastructure with a multi-agent harness emphasizing orchestration and &quot;computer use&quot; for complex software environments. **Figma**, **GitHub**, and **Cursor** launch design canvases with direct AI editing, showcasing tool-calling becoming product-native. **Nous Research** releases **Hermes Agent v0.4.0** with 300+ PRs, adding OpenAI-compatible APIs and self-improving memory agents. Open agent ecosystems mature with **AI2&apos;s MolmoWeb** (4B and 8B models), **GenReasoning&apos;s OpenReward** platform offering 330+ RL environments and 4.5M+ tasks, and **Zhipu&apos;s ZClawBench** benchmark with 116 real-world agent tasks, highlighting progress toward standardized environment serving and benchmarkable agent tasks.</description><pubDate>Tue, 24 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>figma</category><category>github</category><category>cursor_ai</category><category>langchain</category><category>nous-research</category><category>ai2</category><category>genreasoning</category><category>zhipu-ai</category><category>huggingface</category><category>molmo-2-4b</category><category>molmo-2-8b</category><category>hermes-agent-v0.4.0</category><category>agent-infrastructure</category><category>multi-agent-systems</category><category>orchestration</category><category>computer-use</category><category>tool-calling</category><category>design-canvases</category><category>open-agent-platforms</category><category>reinforcement-learning-environments</category><category>benchmarking</category><category>rl-environments</category><category>self-improvement</category><category>api</category><category>memory-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-26-not-much/</guid><description>**Google** launched **Gemini 3.1 Flash Live**, a realtime voice and vision agent model with **2x longer conversation memory**, supporting **70 languages** and **128k context**. **Mistral AI** released **Voxtral TTS**, a low-latency, open-weight text-to-speech model supporting **9 languages** and competitive with ElevenLabs. **Cohere** introduced **Cohere Transcribe**, an audio model with **14-language** support and top English ASR leaderboard performance at **5.42 WER**. **OpenAI** released smaller multimodal variants **GPT-5.4 mini** and **GPT-5.4 nano** with **400k context**, noted for cost-competitiveness but high verbosity and hallucination rates. Other releases include **GLM-5-Turbo** by Zai, **Reka Edge** and **Flash 3** on OpenRouter, and new multi-agent UX tooling **Cline Kanban** for orchestrating CLI coding agents.</description><pubDate>Tue, 24 Mar 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>mistral-ai</category><category>cohere</category><category>openai</category><category>zai</category><category>reka-ai</category><category>gemini-3.1-flash</category><category>voxtral-tts</category><category>cohere-transcribe</category><category>gpt-5.4-mini</category><category>gpt-5.4-nano</category><category>glm-5-turbo</category><category>reka-edge</category><category>reka-flash-3</category><category>logan_kilpatrick</category><category>sundar_pichai</category><category>guillaume_lample</category><category>aidan_gomez</category><category>jay_alammar</category><category>giffmana</category><category>andrew_curran</category><category>voice</category><category>vision</category><category>function-calling</category><category>context-windows</category><category>multimodality</category><category>text-to-speech</category><category>low-latency</category><category>human-preference</category><category>automatic-speech-recognition</category><category>model-benchmarking</category><category>cost-efficiency</category><category>hallucination-detection</category><category>multi-agent-systems</category><category>open-source</category><category>git-worktrees</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-25-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-25-not-much/</guid><description>**ARC-AGI-3** benchmark introduced by **@arcprize** and **François Chollet** resets the frontier for general agentic reasoning with humans solving 100% of tasks versus under 1% for current models, focusing on zero-preparation generalization and human-like learning efficiency. The scoring protocol sparked debate over its harsh efficiency-based metric compared to prior ARC versions and other benchmarks like **NetHack**. The community acknowledges the benchmark highlights weaknesses in current LLM agents in interactive, sparse-feedback environments. Concurrently, agent infrastructure advances with **LangChain** launching Fleet shareable skills for reusable domain knowledge, and **Anthropic** revealing **Claude Code auto mode** for classifier-mediated approval balancing autonomy and manual confirmation. Browser and coding agents are evolving into trainable systems beyond prompt wrappers, exemplified by **BrowserBase** and **Prime Intellect** collaboration.</description><pubDate>Tue, 24 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>langchain</category><category>arcprize</category><category>primeintellect</category><category>arc-agi-3</category><category>claude-code</category><category>fchollet</category><category>mikeknoop</category><category>scaling01</category><category>_rockt</category><category>mark_k</category><category>andykonwinski</category><category>bradenjhancock</category><category>jeremyphoward</category><category>togelius</category><category>bracesproul</category><category>hwchase17</category><category>caspar_br</category><category>_catwu</category><category>agentic-reasoning</category><category>interactive-environments</category><category>benchmarking</category><category>efficiency-metrics</category><category>zero-preparation-generalization</category><category>agent-infrastructure</category><category>trainable-agents</category><category>classifier-approval</category></item><item><title>The Claude Code Source Leak</title><link>https://news.smol.ai/issues/26-03-31-claude-code-leak/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-31-claude-code-leak/</guid><description>**Anthropic&apos;s** closed-source coding product **Claude Code** experienced a significant source leak exposing over **500k lines** of orchestration logic, including autonomous modes and memory systems, but not model weights. The leak led to rapid public reverse-engineering, numerous forks with up to **32.6k stars and 44.3k forks**, and subsequent **DMCA takedowns** by Anthropic. Suspicious npm packages emerged targeting users compiling the leaked code, creating a live security hazard. Discussions also mention unreleased model references like **&quot;mythos&quot;** and ongoing product feature updates despite the leak. *&quot;OFFICIAL STATEMENT from Anthropic regarding the leak&quot;* was noted but not detailed.</description><pubDate>Tue, 24 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>claude-code</category><category>model-architecture</category><category>security</category><category>reverse-engineering</category><category>dmca</category><category>software-development</category><category>open-source</category><category>code-leak</category><category>agent-harness-design</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-23-not-much/</guid><description>**Anthropic** introduced **Claude Cowork** and **Claude Code** enabling desktop control of mouse, keyboard, and screen in a **macOS research preview**, expanding agent capabilities beyond APIs and browsers. The agent ecosystem is evolving towards long-running, parallel, tool-rich workflows with projects like **Hermes Agent**, **T3 Code**, **Command Center**, and **Parchi** enhancing multi-agent orchestration and autonomous task management. Operational challenges such as fragility and inefficiency in subagents, including **GPT-5.2 Pro** and **Claude** browser/computer use, highlight the need for closed-loop feedback systems. Research from **Meta AI** advances self-improving agents with **Hyperagents / DGM-H** enabling meta-level procedural improvements, and unifies reinforcement learning post-training with **RLLM** (RL + LM-as-RM) to improve reward modeling across task types. Additionally, **WebArena-Infinity** drastically reduces browser environment construction costs, accelerating benchmark and environment generation.</description><pubDate>Mon, 23 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>meta-ai-fair</category><category>claude</category><category>gpt-5.2-pro</category><category>dgm-h</category><category>rllm</category><category>jenny_zhang</category><category>jase_weston</category><category>mikhail_parakhin</category><category>jeremyphoward</category><category>agent-frameworks</category><category>workflow-automation</category><category>multi-agent-systems</category><category>reinforcement-learning</category><category>reward-models</category><category>self-improving-agents</category><category>benchmark-generation</category><category>operational-efficiency</category><category>closed-loop-feedback</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-20-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-20-not-much/</guid><description>**Cursor&apos;s Composer 2**, built on **Kimi K2.5**, sparked discussion over model attribution and licensing, highlighting a shift toward post-trained derivatives of open-source models with domain-specific fine-tuning and reinforcement learning. **Claude Code** is expanding into third-party tools like **T3 Code** and communication channels such as Telegram and Discord, while **LangChain** is evolving from orchestration to multi-agent products with offerings like **Deep Agents/Open SWE** and **LangSmith Fleet**. The discourse emphasizes the importance of clear base-model attribution, licensing compliance, and product differentiation through fine-tuning and user experience.</description><pubDate>Fri, 20 Mar 2026 05:44:39 GMT</pubDate><category>cursor</category><category>kimi</category><category>fireworks</category><category>anthropic</category><category>langchain</category><category>kimi-k2.5</category><category>claude-code</category><category>clementdelangue</category><category>leerob</category><category>amanrsanger</category><category>yuchenj_uw</category><category>kimmonismus</category><category>model-attribution</category><category>fine-tuning</category><category>reinforcement-learning</category><category>open-source</category><category>agent-products</category><category>model-licensing</category><category>software-integration</category><category>product-differentiation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-19-not-much/</guid><description>**Cursor** launched **Composer 2**, a frontier-class coding model with major cost reductions and strong benchmark scores like **61.3 on CursorBench** and **73.7 on SWE-bench Multilingual**. The model was improved via a **first continued pretraining run** feeding into reinforcement learning, trained across **3–4 clusters worldwide** by a **~40-person** team. **OpenAI** acquired **Astral**, the team behind Python tools **uv, ruff, and ty**, strengthening its developer platform. **Anthropic** expanded **Claude Code** with messaging app channels for persistent developer workflows. The focus in AI agents is shifting from single agents to managed fleets and runtimes, with **LangChain** launching **LangSmith Fleet** for enterprise agent management emphasizing **agent identity**, **credential management**, and auditability. Other launches include **Cognition&apos;s teams of Devins**, **AgentUI** by **lvwerra**, and discussions on agent runtimes with features like **checkpointing** and **rollback**. Security and permissions are emerging as critical constraints in agent system design.</description><pubDate>Thu, 19 Mar 2026 05:44:39 GMT</pubDate><category>cursor</category><category>openai</category><category>anthropic</category><category>langchain</category><category>cognition</category><category>claude-code</category><category>composer-2</category><category>kimmonismus</category><category>mntruell</category><category>theo</category><category>ellev3n11</category><category>amanrsanger</category><category>charliermarsh</category><category>gdb</category><category>yuchenj_uw</category><category>neilhtennek</category><category>simonw</category><category>yuvalinthedeep</category><category>lvwerra</category><category>hrishioa</category><category>reinforcement-learning</category><category>developer-tooling</category><category>agent-systems</category><category>agent-runtimes</category><category>security</category><category>credential-management</category><category>multi-agent-systems</category><category>model-training</category><category>benchmarking</category><category>software-engineering</category><category>enterprise-ai</category></item><item><title>MiniMax 2.7: GLM-5 at 1/3 cost SOTA Open Model</title><link>https://news.smol.ai/issues/26-03-18-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-18-not-much/</guid><description>**MiniMax M2.7** is the headline model release, described as a &quot;self-evolving agent&quot; with strong performance metrics including **56.22% on SWE-Pro**, **57.0% on Terminal Bench 2**, and parity with **Sonnet 4.6**. It features recursive self-improvement in skills, memory, and architecture. **Artificial Analysis** places M2.7 on the cost/performance frontier with an Intelligence Index score of **50**, matching **GLM-5 (Reasoning)** but at a fraction of the cost. Distribution is available via platforms like **Ollama cloud** and **OpenRouter**. **Xiaomi’s MiMo-V2-Pro** is noted as a serious Chinese API-only reasoning model with a score of **49** on the Intelligence Index and favorable token efficiency. **Cartesia’s Mamba-3** is highlighted as an SSM optimized for inference-heavy use, with early reactions focusing on hybrid transformer architectures like **Qwen3.5** and **Kimi Linear**. The report emphasizes a shift from prompting to harness engineering, where the execution environment and agent harnesses, including skills and MCP, are becoming key differentiators in AI system design. This includes discussions on tools, repo legibility, constraints, and feedback loops, with mentions of **DSPy** and **GPT-5.4 mini** as important components in this evolving landscape.</description><pubDate>Wed, 18 Mar 2026 05:44:39 GMT</pubDate><category>minimax</category><category>xiaomi</category><category>artificial-analysis</category><category>ollama</category><category>trae</category><category>yupp</category><category>openrouter</category><category>vercel</category><category>zo</category><category>opencode</category><category>kilocode</category><category>cartesia</category><category>minimax-m2.7</category><category>sonnet-4.6</category><category>glm-5</category><category>mimo-v2-pro</category><category>mamba-3</category><category>qwen-3.5</category><category>kimi-k2.5</category><category>gpt-5.4-mini</category><category>self-evolving-agents</category><category>reasoning</category><category>cost-efficiency</category><category>token-efficiency</category><category>hybrid-architecture</category><category>harness-engineering</category><category>agent-harnesses</category><category>skills</category><category>memory-optimization</category><category>architecture</category><category>feedback-loops</category><category>api</category><category>inference</category><category>execution-environment</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-17-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-17-not-much/</guid><description>**OpenAI** released **GPT-5.4 mini** and **GPT-5.4 nano**, their most capable small models optimized for coding, multimodal understanding, and subagents, featuring a **400k context window** and over **2x speed** compared to GPT-5 mini. The mini model approaches larger GPT-5.4 performance while using only **30% of Codex quota**, becoming the default for many coding workflows. Pricing concerns and truthfulness tradeoffs were noted, with mixed third-party evaluations on reasoning and resistance to false premises. OpenAI also addressed behavior tuning issues in a recent update. Meanwhile, agent infrastructure is evolving with secure code execution and orchestration tools like **LangChain&apos;s LangSmith Sandboxes** and **Open SWE**, inspired by internal systems at **Stripe, Ramp, and Coinbase**. Subagents and secure execution are now key product features, with releases like **Hermes Agent v0.3.0** showcasing plugin architectures, live Chrome control, and voice mode. Research on attention mechanisms, including **Attention Residuals** and vertical attention, is gaining traction.</description><pubDate>Tue, 17 Mar 2026 05:44:39 GMT</pubDate><category>openai</category><category>langchain</category><category>stripe</category><category>ramp</category><category>coinbase</category><category>nous-research</category><category>hermes-agent</category><category>gpt-5.4-mini</category><category>gpt-5.4-nano</category><category>gpt-5.4</category><category>codex</category><category>hwchase17</category><category>michpokrass</category><category>coding</category><category>multimodality</category><category>subagents</category><category>context-window</category><category>model-performance</category><category>pricing</category><category>behavior-tuning</category><category>secure-execution</category><category>plugin-architecture</category><category>attention-mechanisms</category><category>agent-infrastructure</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-16-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-16-not-much/</guid><description>**Moonshot&apos;s Attention Residuals** paper introduced an input-dependent attention mechanism over prior layers with a **1.25x compute advantage** and less than **2% inference latency overhead**, validated on **Kimi Linear 48B total / 3B active**. The paper sparked debate on novelty versus prior art like **DeepCrossAttention** and Google’s earlier work, highlighting tensions in **idea novelty**, **citation quality**, and **frontier-scale validation**. **OpenAI&apos;s Codex** showed strong momentum with over **2M weekly active users**, nearly **4x growth YTD**, and **GPT-5.4** hitting **5T tokens/day** and a **$1B annualized run-rate**. Codex added subagents supporting multi-agent coding workflows. Infrastructure for coding agents matured with tools like **Context Hub / chub** supporting agent feedback loops, **AssemblyAI&apos;s skill** for Claude Code and Codex, and automated skill extraction from GitHub repos yielding **40% knowledge-transfer gains**. **LangChain** launched **LangGraph CLI** and open-sourced **Deep Agents**, recreating top coding agent workflows with planning, filesystem ops, shell access, and sub-agents.</description><pubDate>Mon, 16 Mar 2026 05:44:39 GMT</pubDate><category>moonshot</category><category>openai</category><category>assemblyai</category><category>langchain</category><category>kimi-linear-48b</category><category>codex</category><category>gpt-5.4</category><category>claude-code</category><category>kimi_moonshot</category><category>elonmusk</category><category>yuchenj_uw</category><category>nathancgy4</category><category>eliebakouch</category><category>tokenbender</category><category>behrouz_ali</category><category>cloneofsimo</category><category>fidjissimo</category><category>sama</category><category>gdb</category><category>andrewyng</category><category>itsafiz</category><category>simplifyinai</category><category>attention-mechanisms</category><category>model-architecture</category><category>inference-speed</category><category>agent-feedback</category><category>agent-skills</category><category>multi-agent-systems</category><category>knowledge-transfer</category><category>cli-tools</category><category>coding-agents</category><category>model-deployment</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-13-not-much/</guid><description>**MCP tools** remain relevant for deterministic APIs despite ergonomic criticisms, with new **web MCP support in Chrome v146** enabling continuous browsing agents. Persistent memory is emerging as a key differentiator for agents, with IBM improving task completion rates and multi-agent memory framed as a computer architecture challenge. Agent UX is evolving towards always-on, cross-device operation, exemplified by **Perplexity Computer** on iOS and **Claude Code** session management. **Anthropic** released **Opus 4.6 1M context** as default with no extra long-context API charges, achieving **78.3% on MRCR v2 at 1M tokens**. Sparse attention optimizations like **IndexCache** in **DeepSeek Sparse Attention** yield significant speedups on large models with minimal code changes.</description><pubDate>Fri, 13 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>ibm</category><category>perplexity-ai</category><category>llamaindex</category><category>deepseek</category><category>google-chrome</category><category>opus-4.6</category><category>glm-5</category><category>pamelafox</category><category>tadasayy</category><category>llama_index</category><category>bromann</category><category>dair_ai</category><category>omarsar0</category><category>abxxai</category><category>teknuim</category><category>bcherny</category><category>kimmonismus</category><category>_catwu</category><category>alexalbert__</category><category>realyushibai</category><category>persistent-memory</category><category>agent-infrastructure</category><category>cross-device-synchronization</category><category>long-context</category><category>sparse-attention</category><category>inference-optimization</category><category>computer-architecture</category><category>task-completion</category><category>systems-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-12-not-much/</guid><description>**Harnesses, agent infrastructure, and the MCP protocol** are central themes, with emphasis on how **harnesses, sandboxes, filesystem access, skills, memory, and observability** shape agent UI/UX and runtime environments. Despite jokes about MCP&apos;s demise, it remains vital in production, notably used internally by **Uber** and supported by **Anthropic**. The **coding-agent stack** is evolving with **CursorBench** combining offline and online metrics to evaluate models on **intelligence and efficiency**, where **GPT-5.4** leads in correctness and token efficiency. Agent-assisted development is splitting between automation-heavy workflows and &quot;stay-in-the-loop&quot; tooling, with **OpenAI** advancing **Codex Automations** featuring worktree vs. branch choices and UI customization. The open agent platform **Hermes Agent v0.2.0** introduces full MCP client support, ACP server for editors, and expanded provider integrations including **OpenAI OAuth**.</description><pubDate>Thu, 12 Mar 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>uber</category><category>nous-research</category><category>cursor_ai</category><category>redisinc</category><category>artificialanlys</category><category>langchain-js</category><category>gpt-5.4</category><category>mattturck</category><category>hwchase17</category><category>omarsar0</category><category>gergelyorosz</category><category>htihle</category><category>theprimeagen</category><category>sydneyrunkle</category><category>corbtt</category><category>agent-infrastructure</category><category>mcp-protocol</category><category>harnesses</category><category>coding-agents</category><category>evaluation-methodologies</category><category>agent-ui-ux</category><category>runtime-environments</category><category>multi-axis-evaluation</category><category>automation</category><category>workflow-optimization</category><category>open-agent-platforms</category><category>provider-integration</category><category>filesystem-checkpoints</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-11-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-11-not-much/</guid><description>**NVIDIA’s Nemotron 3 Super** is a **120B parameter / ~12B active** open model featuring a **hybrid Mamba-Transformer / SSM Latent MoE** architecture and **1M context window**, delivering up to **2.2x faster inference than GPT-OSS-120B** in FP4 with strong throughput gains. It supports agentic workloads and is unusually open with weights, data, and infrastructure details released. The model scored **36 on the AA Intelligence Index**, outperforming GPT-OSS-120B but behind Qwen3.5-122B-A10B. Community and infrastructure support from projects like **vLLM**, **llama.cpp**, **Ollama**, **Together**, **Baseten**, **W&amp;B Inference**, **LangChain**, and **Unsloth GGUFs** was immediate. Key technical innovations include **native multi-token prediction (MTP)** and a significant **KV-cache efficiency** advantage. 

On the product side, a shift towards **persistent agent runtimes and orchestration layers** is highlighted, with **Andrej Karpathy** advocating for a &quot;bigger IDE&quot; concept where agents replace files as the unit of work, enabling legible, forkable agentic organizations with real-time control. New launches fitting this vision include **Perplexity’s Personal Computer**, an always-on local/cloud hybrid running on Mac mini, and **Computer for Enterprise** orchestrating 20 specialized models and 400+ apps. **Replit Agent 4** offers a collaborative, canvas-like workflow with parallel agents, while **Base44 Superagents** provide integrated solutions for nontechnical users. The engineering focus is increasingly on the orchestration harness rather than just the model.</description><pubDate>Wed, 11 Mar 2026 05:44:39 GMT</pubDate><category>nvidia</category><category>perplexity</category><category>replit</category><category>base44</category><category>vllm</category><category>llama.cpp</category><category>ollama</category><category>togethercompute</category><category>baseten</category><category>wandb</category><category>langchain</category><category>unsloth</category><category>nemotron-3-super</category><category>gpt-oss-120b</category><category>qwen3.5-122b-a10b</category><category>karpathy</category><category>ctnzr</category><category>bnjmn_marie</category><category>artificialanlys</category><category>model-architecture</category><category>model-optimization</category><category>inference-speed</category><category>kv-cache</category><category>multi-token-prediction</category><category>agent-infrastructure</category><category>orchestration</category><category>persistent-agents</category><category>model-serving</category><category>product-launches</category></item><item><title>Yann LeCun’s AMI Labs launches with a $1.03B seed to build world models around JEPA</title><link>https://news.smol.ai/issues/26-03-10-ami-labs/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-10-ami-labs/</guid><description>**Yann LeCun** launched **Advanced Machine Intelligence (AMI Labs)** with a record **$1.03B seed round** at a **$3.5B pre-money valuation**, aiming to build AI models that understand the **physical world** through **world models** rather than just language prediction. The startup, based in **Europe** with locations in **Paris** and **Zürich**, is framed as a major milestone for European AI and backed by a prominent founding team including **Alex Lebrun**, **Saining Xie**, and **Pascale Fung**. The mission is described as a &quot;long-term scientific endeavor&quot; to create AI that &quot;perceives, learns, reasons and acts&quot; in the real world.</description><pubDate>Tue, 10 Mar 2026 05:44:39 GMT</pubDate><category>ami-labs</category><category>ylecun</category><category>lxbrun</category><category>sainingxie</category><category>pascalefung</category><category>laurentsolly</category><category>world-models</category><category>representation-learning</category><category>pretraining</category><category>scaling</category><category>video</category><category>funding</category><category>seed-round</category><category>valuation</category><category>real-world-understanding</category></item><item><title>Autoresearch: Sparks of Recursive Self Improvement</title><link>https://news.smol.ai/issues/26-03-09-autoresearch/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-09-autoresearch/</guid><description>**RSI** covers AI developments from 3/5/2026 to 3/9/2026, highlighting the emergence of **LLMs autonomously training smaller LLMs**, marking a significant &quot;AutoML moment&quot; in AI progress. **Karpathy** and **Yi Tay** discuss &quot;vibe training,&quot; where AI models fix bugs and improve code autonomously, suggesting models may soon surpass human debugging efficiency. The report anticipates **Jakub Pachocki&apos;s Automated AI Research Intern** system by September 2026 to accelerate human researchers. On AI Twitter, the focus is on **coding agents** shifting bottlenecks from implementation to review and verification, with **Anthropic&apos;s Claude Code Review** improving PR review effectiveness significantly, and tools like **OpenAI Codex Review** and **Cognition&apos;s Devin Review** enhancing code review workflows. Harness engineering is evolving into systems engineering, emphasizing decoupling agent storage from compute for collaborative agent teams.</description><pubDate>Mon, 09 Mar 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>cognition</category><category>claude-3</category><category>codex</category><category>karpathy</category><category>yi_tay</category><category>jakub_pachocki</category><category>automated-machine-learning</category><category>coding-agents</category><category>bug-fixing</category><category>model-autonomy</category><category>multi-agent-systems</category><category>pr-review</category><category>systems-engineering</category><category>model-verification</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-06-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-06-not-much/</guid><description>**OpenAI** rolled out **GPT-5.4**, achieving tied **#1** on the **Artificial Analysis Intelligence Index** with **Gemini 3.1 Pro Preview** scoring **57** (up from 51 for GPT-5.2 xhigh). GPT-5.4 features a larger **~1.05M token** context window and higher per-token prices ($2.50/$15 vs $1.75/$14 for GPT-5.2), with strengths in **physics reasoning (CritPt)** and **agentic coding (TerminalBench Hard)** but a higher hallucination rate and **~28% higher benchmark run cost**. The **GPT-5.4 Pro** variant shows a **+10 point jump** on CritPt reaching **30%** but at an extreme output token cost of **$180 / 1M tokens**. Community benchmarks show GPT-5.4 excels in agentic/coding tasks but mixed feedback on reasoning efficiency and literalness compared to **Claude**. OpenAI updated agent prompting guidance for GPT-5.4 API users, emphasizing tool use, structured outputs, and verification loops. **Claude Code** added local scheduled tasks and loop patterns for agents. The **MCP** framework is highlighted as a connective tissue for AI evaluation and design-code round-trips, with **Truesight MCP** enabling AI evaluation like unit testing and **Figma MCP server** supporting bidirectional design-code integration. Open-source **T3 Code** launched as an agent orchestration coding app built on Codex CLI.</description><pubDate>Fri, 06 Mar 2026 05:44:39 GMT</pubDate><category>openai</category><category>artificial-analysis</category><category>gemini</category><category>claude</category><category>mit</category><category>figma</category><category>github</category><category>gpt-5.4</category><category>gpt-5.2</category><category>gemini-3.1-pro</category><category>benchmarking</category><category>physics-reasoning</category><category>agentic-coding</category><category>hallucination-detection</category><category>context-windows</category><category>cost-efficiency</category><category>agent-prompting</category><category>scheduled-tasks</category><category>loop-patterns</category><category>ai-evaluation</category><category>design-code-integration</category><category>agent-orchestration</category><category>open-source</category></item><item><title>GPT 5.4: SOTA Knowledge Work -and- Coding -and- CUA Model, OpenAI is so very back</title><link>https://news.smol.ai/issues/26-03-05-gpt54/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-05-gpt54/</guid><description>**OpenAI** launched **GPT-5.4** and **GPT-5.4 Pro** with unified mainline and Codex models, featuring **native computer use**, up to **~1M token context**, and efficiency improvements including a new **Codex `/fast` mode**. Benchmarks showed strong results like **OSWorld-Verified 75.0%** surpassing human baseline and **GDPval 83%** against industry pros. User feedback highlighted coding utility but raised concerns about pricing and overthinking. Integration with devtools like **Cursor**, **Perplexity**, and **Arena** was announced. In systems research, **FlashAttention-4 (FA4)** was introduced with near-matmul speed attention on **Blackwell** GPUs, featuring innovations like **polynomial exp emulation** and **online softmax**. *&quot;Steering mid-response&quot;* and *&quot;fewer tokens, faster speed&quot;* were emphasized as UX and efficiency improvements.</description><pubDate>Thu, 05 Mar 2026 05:44:39 GMT</pubDate><category>openai</category><category>cursor_ai</category><category>perplexity_ai</category><category>arena</category><category>gpt-5.4</category><category>gpt-5.4-pro</category><category>sama</category><category>reach_vb</category><category>scaling01</category><category>danshipper</category><category>yuchenj_uw</category><category>native-computer-use</category><category>long-context</category><category>efficiency</category><category>steering</category><category>benchmarking</category><category>gpu-kernels</category><category>attention-mechanisms</category><category>algorithmic-optimization</category><category>pipeline-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-04-not-much/</guid><description>**Gemini 3.1 Flash-Lite** is highlighted by **Demis Hassabis** for its speed and cost-efficiency, focusing on latency and cost per capability rather than raw performance. **NotebookLM Studio** introduces a new feature for generating immersive cinematic video overviews. Rumors about **GPT-5.4** suggest a ~1 million token context window and an &quot;extreme reasoning mode&quot; for long-horizon tasks, with speculation about monthly model updates from **OpenAI**. **Anthropic&apos;s Claude Opus 4.6** is noted for strong general agent behavior but weaker visual mathematics performance. **Alibaba&apos;s Qwen** team faces leadership exits and restructuring, with concerns about compute access and organizational changes. Qwen models dominate research workflows, appearing in 41% of Hugging Face papers in 2025-2026, raising ecosystem dependence risks. The open-weight model landscape may consolidate around non-profits, **NVIDIA**, and **Meta** due to business incentives.</description><pubDate>Wed, 04 Mar 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>anthropic</category><category>alibaba</category><category>nvidia</category><category>meta-ai-fair</category><category>hugging-face</category><category>gemini-3.1-flash-lite</category><category>gpt-5.4</category><category>claude-opus-4.6</category><category>qwen-3.5</category><category>qwen</category><category>demishassabis</category><category>natolambert</category><category>poezhao0605</category><category>simonw</category><category>model-positioning</category><category>latency</category><category>cost-efficiency</category><category>context-window</category><category>extreme-reasoning</category><category>agentic-ai</category><category>model-updates</category><category>general-agent-behavior</category><category>visual-mathematics</category><category>leadership-exits</category><category>organizational-restructuring</category><category>compute-access</category><category>research-workflows</category><category>open-weight-models</category><category>ecosystem-dependence</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-03-not-much/</guid><description>**Google DeepMind** launched **Gemini 3.1 Flash-Lite**, emphasizing *dynamic thinking levels* for adjustable compute, with notable metrics like **$0.25/M input**, **$1.50/M output**, **1432 Elo on LMArena**, and **2.5× faster time-to-first-token** than Gemini 2.5 Flash. It supports a **1M context window** and high throughput for multimodal inputs including text, images, video, audio, and PDFs. **OpenAI** rolled out **GPT-5.3 Instant** to all ChatGPT users, improving conversational naturalness and reducing hallucinations by **26.8% with search**. The upcoming **GPT-5.4** was teased amid speculation. **Alibaba&apos;s Qwen** faces leadership exits, raising concerns about its future and open-source status. The news highlights advancements in model efficiency, pricing, and multimodality, alongside organizational changes impacting AI development.</description><pubDate>Tue, 03 Mar 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>google</category><category>openai</category><category>alibaba</category><category>gemini-3.1-flash-lite</category><category>gemini-3</category><category>gpt-5.3</category><category>gpt-5.4</category><category>qwen</category><category>jeffdean</category><category>noamshazeer</category><category>sundarpichai</category><category>aidan_mclau</category><category>justinlin610</category><category>multimodality</category><category>latency</category><category>throughput</category><category>context-window</category><category>model-pricing</category><category>model-benchmarking</category><category>model-performance</category><category>conversational-ai</category><category>hallucination-reduction</category><category>api</category><category>model-rollout</category><category>leadership-exit</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-03-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-03-02-not-much/</guid><description>**Alibaba** released the **Qwen 3.5** series with models ranging from **0.8B to 9B** parameters, featuring **native multimodality**, **scaled reinforcement learning**, and targeting **edge and lightweight agent** deployments. The models support very long context windows up to **262K tokens** (extendable to 1M) and use a novel **Gated DeltaNet hybrid attention** architecture combining linear and full attention layers. Deployment examples include **Ollama** and **LM Studio**, with a notable **6-bit on-device demo on iPhone 17 Pro**. Evaluators are cautioned that reasoning is disabled by default on smaller models. In coding agents, **Codex 5.3** shows promising benchmark results on **WeirdML** with **79.3%** accuracy, though availability and downtime remain critical challenges, especially highlighted by **Claude** outages. Agent reliability and observability are emphasized as cross-functional problems requiring clear success criteria and practical evaluation strategies. Studies show that using **AGENTS.md** and **SKILL.md** guardrails can significantly reduce runtime and token usage by mitigating worst-case thrashing in coding workflows.</description><pubDate>Mon, 02 Mar 2026 05:44:39 GMT</pubDate><category>alibaba</category><category>ollama</category><category>lm-studio</category><category>openai</category><category>anthropic</category><category>qwen-3.5-0.8b</category><category>qwen-3.5-2b</category><category>qwen-3.5-4b</category><category>qwen-3.5-9b</category><category>codex-5.3</category><category>claude-3</category><category>nrehiew_</category><category>kimmonismus</category><category>lioronai</category><category>danielhanchen</category><category>theo</category><category>htihle</category><category>teortaxestex</category><category>theprimeagen</category><category>yuchenj_uw</category><category>_lewtun</category><category>saen_dev</category><category>_philschmid</category><category>omarsar0</category><category>multimodality</category><category>reinforcement-learning</category><category>long-context</category><category>hybrid-attention</category><category>on-device-ai</category><category>model-deployment</category><category>agent-reliability</category><category>agent-observability</category><category>coding-agents</category><category>benchmarking</category><category>runtime-optimization</category><category>token-efficiency</category></item><item><title>OpenAI closes $110B raise from Amazon, NVIDIA, SoftBank in largest startup fundraise in history @ $840B post-money</title><link>https://news.smol.ai/issues/26-02-27-openai-g/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-27-openai-g/</guid><description>**OpenAI** has closed a major funding round totaling **$110 billion** at a **$730 billion pre-money valuation**, with investments from **SoftBank ($30B)**, **NVIDIA ($30B)**, and **Amazon ($50B)**. Key user metrics include **1.6 million weekly Codex users**, **over 9 million paying business users** of ChatGPT, and **more than 900 million weekly active ChatGPT users** with **50 million consumer subscribers**. The partnership with Amazon includes exclusive cloud services and **2 gigawatts of Trainium capacity**. Microsoft maintains a reduced partnership with stateless APIs. This funding round is one of the largest in history, highlighting OpenAI&apos;s dominant position in AI adoption and infrastructure.</description><pubDate>Fri, 27 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>softbank</category><category>nvidia</category><category>amazon</category><category>microsoft</category><category>codex</category><category>chatgpt</category><category>sama</category><category>model-scaling</category><category>model-metrics</category><category>investment</category><category>cloud-computing</category><category>infrastructure</category><category>training-capacity</category><category>user-growth</category><category>partnerships</category></item><item><title>Nano Banana 2 aka Gemini 3.1 Flash Image Preview: the new SOTA Imagegen model</title><link>https://news.smol.ai/issues/26-02-26-nanobanana2/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-26-nanobanana2/</guid><description>**Google and DeepMind** launched **Nano Banana 2** (aka **Gemini 3.1 Flash Image Preview**), a leading image generation and editing model integrated across multiple Google products with features like **4K upscaling**, **multi-subject consistency**, and **real-time search-conditioned generation**. Evaluations rank it #1 in text-to-image tasks with competitive pricing. Additionally, advances in **agentic coding** are noted with models like **GPT-5.2**, **GPT-5.3 Codex**, **Opus 4.6**, and **Gemini 3.1**, alongside Microsoft&apos;s **Copilot Tasks** introducing task delegation. Persistent memory features are rolling out in **Claude** models, though interoperability challenges remain.</description><pubDate>Thu, 26 Feb 2026 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>microsoft</category><category>anthropic</category><category>perplexity-ai</category><category>gemini-3.1-flash</category><category>gpt-5.2</category><category>gpt-5.3-codex</category><category>opus-4.6</category><category>claude</category><category>sundarpichai</category><category>demishassabis</category><category>mustafasuleyman</category><category>yusuf_i_mehdi</category><category>borisdayma</category><category>aravsrinivas</category><category>image-generation</category><category>text-rendering</category><category>3d-imaging</category><category>real-time-information</category><category>agentic-ai</category><category>persistent-memory</category><category>multi-agent-systems</category><category>tooling</category><category>coding-agents</category><category>task-delegation</category></item><item><title>Agentic Engineering: WTF Happened in December 2025?</title><link>https://news.smol.ai/issues/26-02-25-wtf-happened/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-25-wtf-happened/</guid><description>**Perplexity** launched **Computer**, an orchestration-first agent platform featuring multi-model routing, usage-based pricing, and parallel asynchronous sub-agents for distributed workflows. **Andrej Karpathy** claims a &quot;phase change&quot; in coding agents since December, highlighting sustained long-horizon task completion. **OpenAI** released **GPT-5.3-Codex** with ~25% speed improvements and strong benchmark performance, while **Claude Code** celebrates its first year with ecosystem integrations and scaling challenges. This marks a significant shift in coding workflows and agent-based software development.</description><pubDate>Wed, 25 Feb 2026 05:44:39 GMT</pubDate><category>perplexity</category><category>openai</category><category>anthropic</category><category>langchain-ai</category><category>gpt-5.3-codex</category><category>claude-code</category><category>karpathy</category><category>aravsrinivas</category><category>lioronai</category><category>denisyarats</category><category>swyx</category><category>catwu</category><category>hwchase17</category><category>coding-agents</category><category>agent-architecture</category><category>distributed-workflows</category><category>usage-based-pricing</category><category>model-routing</category><category>benchmarking</category><category>context-length</category><category>observability</category><category>software-development</category></item><item><title>Anthropic accuses DeepSeek, Moonshot, and MiniMax of &quot;industrial-scale distillation attacks&quot;.</title><link>https://news.smol.ai/issues/26-02-23-anthropic-distillation/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-23-anthropic-distillation/</guid><description>**Anthropic** alleges *industrial-scale* distillation attacks on its **Claude** model by **DeepSeek**, **Moonshot AI**, and **MiniMax**, involving **~24,000 fraudulent accounts** and **&gt;16M Claude exchanges** to extract capabilities, raising concerns about competitive risks and safety. The community debates the difference between scraping and API-output extraction, highlighting a shift toward protecting models via *API abuse resistance* techniques. Meanwhile, coding agents like **Codex** and **Claude Code** see real adoption and failures, with emerging best practices in &quot;agentic engineering&quot; led by **Simon Willison**. The **OpenClaw** ecosystem expands with alternatives like **NanoClaw** and integrations such as **Ollama 0.17** simplifying open model usage.</description><pubDate>Tue, 24 Feb 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>deepseek</category><category>moonshot-ai</category><category>minimax</category><category>openai</category><category>ollama</category><category>claude</category><category>claude-3</category><category>codex</category><category>claude-code</category><category>simon_willison</category><category>api-abuse-resistance</category><category>model-security</category><category>agentic-engineering</category><category>coding-agents</category><category>model-distillation</category><category>workflow-automation</category><category>sandboxing</category><category>realtime-communication</category></item><item><title>Claude Code Anniversary + Launches from: Qwen 3.5, Cursor Demos, Cognition Devin 2.2, Inception Mercury 2</title><link>https://news.smol.ai/issues/26-02-24-claude-code/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-24-claude-code/</guid><description>**Alibaba** launched the **Qwen 3.5 Medium Model Series** featuring models like **Qwen3.5-Flash**, **Qwen3.5-35B-A3B (MoE)**, and **Qwen3.5-122B-A10B (MoE)** emphasizing efficiency over scale with innovations like **1M context** and INT4 quantization. **OpenAI** released **GPT-5.3-Codex** via the **Responses API** with enhanced file input support and faster web socket-based throughput. **Anthropic** introduced **Claude Code Remote Control** enabling terminal session continuation from mobile and expanded enterprise workflow features. **Cursor** shifted UX to agent demo videos instead of diffs, highlighting new interaction modes.</description><pubDate>Tue, 24 Feb 2026 05:44:39 GMT</pubDate><category>alibaba</category><category>openai</category><category>anthropic</category><category>cursor</category><category>huggingface</category><category>qwen3.5-flash</category><category>qwen3.5-35b-a3b</category><category>qwen3.5-122b-a10b</category><category>qwen3.5-27b</category><category>qwen3.5-397b-a17b</category><category>gpt-5.3-codex</category><category>claude-code</category><category>awnihannun</category><category>andrew_n_carr</category><category>justinlin610</category><category>unslothai</category><category>terryyuezhuo</category><category>haihaoshen</category><category>0xsero</category><category>ali_tongyilab</category><category>scaling01</category><category>gdb</category><category>noahzweben</category><category>_catwu</category><category>model-architecture</category><category>reinforcement-learning</category><category>quantization</category><category>context-windows</category><category>agentic-ai</category><category>api</category><category>websockets</category><category>software-ux</category><category>enterprise-workflows</category><category>model-deployment</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-02-20-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-20-not-much/</guid><description>**Gemini 3.1 Pro** demonstrates strong retrieval capabilities and cost efficiency compared to **GPT-5.2** and **Opus 4.6**, though users report tooling and UI issues. The **SWE-bench Verified** evaluation methodology is under scrutiny for consistency, with updates bringing results closer to developer claims. Benchmarking debates arise over what frontier models truly measure, especially with ARC-AGI puzzles. **Claude Opus 4.6** shows a noisy but notable **14.5-hour time horizon** on software tasks, with token limits causing practical failures. **Sonnet 4.6** improves significantly in code and instruction-following benchmarks, but user backlash grows due to product regressions.</description><pubDate>Sat, 21 Feb 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>anthropic</category><category>context-arena</category><category>artificial-analysis</category><category>epoch-ai</category><category>scaling01</category><category>gemini-3.1-pro</category><category>gpt-5.2</category><category>opus-4.6</category><category>sonnet-4.6</category><category>claude-opus-4.6</category><category>dillonuzar</category><category>artificialanlys</category><category>yuchenj_uw</category><category>theo</category><category>minimax_ai</category><category>epochairesearch</category><category>paul_cal</category><category>scaling01</category><category>metr_evals</category><category>idavidrein</category><category>xlr8harder</category><category>htihle</category><category>arena</category><category>retrieval</category><category>benchmarking</category><category>evaluation-methodology</category><category>token-limits</category><category>cost-efficiency</category><category>instruction-following</category><category>software-reasoning</category><category>model-reliability</category></item><item><title>Gemini 3.1 Pro: 2x 3.0 on ARC-AGI 2</title><link>https://news.smol.ai/issues/26-02-19-gemini31/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-19-gemini31/</guid><description>**Google** released **Gemini 3.1 Pro**, a developer preview integrated across the **Gemini app**, **NotebookLM**, **Gemini API / AI Studio**, and **Vertex AI**, highlighting a significant reasoning improvement with **ARC-AGI-2 = 77.1%** and strong coding and agentic-tool benchmarks like **SWE-Bench Verified = 80.6%**. Independent evaluators such as **Artificial Analysis** and **Arena** confirmed top-tier performance and cost efficiency, though community reactions included excitement about practical gains, skepticism about benchmark targeting, and concerns over rollout inconsistencies. The release emphasizes the same core intelligence powering **Gemini 3 Deep Think** scaled for practical use, with notable mentions from leaders like *@sundarpichai*, *@demishassabis*, and *@JeffDean*.</description><pubDate>Thu, 19 Feb 2026 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>geminiapp</category><category>gemini-3.1-pro</category><category>gemini-3-deep-think</category><category>sundarpichai</category><category>demishassabis</category><category>jeffdean</category><category>koraykv</category><category>noamshazeer</category><category>joshwoodward</category><category>artificialanlys</category><category>arena</category><category>oriolvinyalsml</category><category>scaling01</category><category>reasoning</category><category>benchmarking</category><category>agentic-ai</category><category>cost-efficiency</category><category>hallucination</category><category>code-generation</category><category>model-release</category><category>developer-tools</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-02-18-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-18-not-much/</guid><description>**Anthropic** released **Claude Opus/Sonnet 4.6**, showing a significant intelligence index jump but with increased token usage and cost. **Anthropic** also shared insights on AI agent autonomy, highlighting human-in-the-loop prevalence and software engineering tool calls. **Alibaba** launched **Qwen 3.5** with discussions on reasoning efficiency and token bloat, plus open-sourced **Qwen3.5-397B-A17B FP8 weights**. The **GLM-5** technical report introduced asynchronous agent reinforcement learning and compute-efficient techniques. Rumors about **Gemini 3.1 Pro** suggest longer reasoning capabilities, while **MiniMax M2.5** appeared on community leaderboards. The community debates benchmark reliability and model performance nuances.</description><pubDate>Wed, 18 Feb 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>alibaba</category><category>scaling01</category><category>arena</category><category>artificial-analysis</category><category>claude-4.6</category><category>claude-opus-4.6</category><category>claude-sonnet-4.6</category><category>qwen-3.5</category><category>qwen3.5-397b-a17b</category><category>glm-5</category><category>gemini-3.1-pro</category><category>minimax-m2.5</category><category>eshear</category><category>theo</category><category>omarsar0</category><category>grad62304977</category><category>scaling01</category><category>benchmarking</category><category>token-efficiency</category><category>ai-agent-autonomy</category><category>reinforcement-learning</category><category>asynchronous-learning</category><category>model-performance</category><category>open-weights</category><category>reasoning</category><category>software-engineering</category><category>agentic-engineering</category></item><item><title>Claude Sonnet 4.6: clean upgrade of 4.5, mostly better with some caveats</title><link>https://news.smol.ai/issues/26-02-17-sonnet-46/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-17-sonnet-46/</guid><description>**Anthropic** launched **Claude Sonnet 4.6**, an upgrade over Sonnet 4.5, featuring broad improvements in **coding, long-context reasoning, agent planning, knowledge work, and design**, plus a **1M-token context window (beta)**. Benchmarks show Sonnet 4.6 leading on **GDPval-AA ELO 1633**, with significant token usage increases and improved output aesthetics. Integrations include **Cursor, Windsurf, Microsoft Foundry, and Perplexity Pro/Max**. Early user feedback noted some regression issues that were later fixed. Pricing remains the same as Sonnet 4.5. Tooling enhancements include code execution for filtering results, improving accuracy and efficiency.</description><pubDate>Tue, 17 Feb 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>cursor</category><category>microsoft</category><category>perplexity-ai</category><category>cognition</category><category>claude-3-sonnet-4.6</category><category>claude-3-sonnet-4.5</category><category>claude-3-opus-4.5</category><category>claude-3-opus-4.6</category><category>alexalbert__</category><category>scaling01</category><category>rishdotblog</category><category>claudeai</category><category>kimmonismus</category><category>artificialanlys</category><category>long-context</category><category>agent-planning</category><category>knowledge-work</category><category>benchmarking</category><category>tokenization</category><category>model-integration</category><category>code-execution</category><category>model-updates</category><category>aesthetic-quality</category></item><item><title>Qwen3.5-397B-A17B: the smallest Open-Opus class, very efficient model</title><link>https://news.smol.ai/issues/26-02-16-qwen35/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-16-qwen35/</guid><description>**Alibaba** released **Qwen3.5-397B-A17B**, an open-weight model featuring **native multimodality**, **spatial intelligence**, and a **hybrid linear attention + sparse MoE** architecture supporting **201 languages** and **long context windows** up to **256K tokens**. The model shows improvements over previous versions like **Qwen3-Max** and **Qwen3-VL**, with a sparsity ratio of about **4.3%**. Community discussions highlighted the **Gated Delta Networks** enabling efficient inference despite large model size (~**800GB BF16**), with successful local runs on Apple Silicon using quantization techniques. The hosted API version, **Qwen3.5-Plus**, supports **1M context** and integrates search and code interpreter features. This release follows other Chinese labs like **Z.ai**, **Minimax**, and **Kimi** in refreshing large models. The model is licensed under **Apache-2.0** and is expected to be the last major release before **DeepSeek v4**. The news also notes **Pete Steinberger** joining **OpenAI**.</description><pubDate>Mon, 16 Feb 2026 05:44:39 GMT</pubDate><category>alibaba</category><category>openai</category><category>deepseek</category><category>z-ai</category><category>minimax</category><category>kimi</category><category>unsloth</category><category>ollama</category><category>vllm</category><category>qwen3.5-397b-a17b</category><category>qwen3.5-plus</category><category>qwen3-max</category><category>qwen3-vl</category><category>kimi</category><category>pete_steinberger</category><category>justinlin610</category><category>native-multimodality</category><category>spatial-intelligence</category><category>sparse-moe</category><category>long-context</category><category>model-quantization</category><category>model-architecture</category><category>model-deployment</category><category>inference-optimization</category><category>apache-2.0-license</category></item><item><title>MiniMax-M2.5: SOTA coding, search, toolcalls, $1/hour</title><link>https://news.smol.ai/issues/26-02-13-minimax25/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-13-minimax25/</guid><description>**MiniMax-M2.5** is now open source, featuring an &quot;agent-native&quot; reinforcement learning framework called **Forge** trained across **200k+ RL environments** for coding, tool use, and workflows. It boasts strong benchmark scores like **80.2% SWE-Bench Verified** and emphasizes cost-efficiency with claims like &quot;$1 per hour at 100 tps&quot; and good on-device performance. The **Forge** RL system uses multi-level prefix caching and high rollout compute share (~60%) to generate millions of trajectories daily. Independent reviews note improved stability and multi-turn viability but high token usage. The ecosystem rapidly adopted MiniMax-M2.5 with quantized releases including **2-bit GGUF** and **INT4** formats. Meanwhile, **Together** markets **GLM-5** as a leading open-source model for long-horizon agents with **77.8% SWE-Bench Verified** and MoE efficiency using DeepSeek Sparse Attention.</description><pubDate>Fri, 13 Feb 2026 05:44:39 GMT</pubDate><category>minimax-ai</category><category>togethercompute</category><category>huggingface</category><category>intel</category><category>wandb</category><category>minimax-m2.5</category><category>glm-5</category><category>reinforcement-learning</category><category>agent-based-models</category><category>model-quantization</category><category>benchmarking</category><category>model-efficiency</category><category>multi-turn-dialogue</category><category>infrastructure-optimization</category><category>cost-efficiency</category><category>on-device-ai</category></item><item><title>new Gemini 3 Deep Think, Anthropic $30B @ $380B, GPT-5.3-Codex Spark, MiniMax M2.5</title><link>https://news.smol.ai/issues/26-02-12-anthropic-gemini-deepthink/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-12-anthropic-gemini-deepthink/</guid><description>**Google DeepMind** is rolling out the upgraded **Gemini 3 Deep Think V2** reasoning mode to **Google AI Ultra** subscribers and opening early access to the **Vertex AI / Gemini API** for select users. Key benchmark achievements include **ARC-AGI-2 at 84.6%**, **Humanity’s Last Exam (HLE) at 48.4% without tools**, and a **Codeforces Elo of 3455**, showcasing Olympiad-level performance in physics and chemistry. The mode emphasizes practical scientific and engineering applications such as error detection in math papers, physical system modeling, semiconductor optimization, and a **sketch to CAD/STL pipeline** for 3D printing. ARC benchmark creator François Chollet highlights the benchmark&apos;s role in advancing test-time adaptation and fluid intelligence, projecting human-AI parity around **2030**. This rollout is framed as a productized, compute-heavy test-time mode rather than a lab demo, with cost disclosures for ARC tasks provided.</description><pubDate>Thu, 12 Feb 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>google</category><category>geminiapp</category><category>arcprize</category><category>gemini-3-deep-think-v2</category><category>arc-agi-2</category><category>demishassabis</category><category>sundarpichai</category><category>fchollet</category><category>jeffdean</category><category>oriolvinyalsml</category><category>tulseedoshi</category><category>benchmarking</category><category>reasoning</category><category>test-time-adaptation</category><category>fluid-intelligence</category><category>scientific-computing</category><category>engineering-workflows</category><category>3d-modeling</category><category>cost-analysis</category></item><item><title>Z.ai GLM-5: New SOTA Open Weights LLM</title><link>https://news.smol.ai/issues/26-02-11-glm-5/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-11-glm-5/</guid><description>**Zhipu AI** launched **GLM-5**, an **Opus-class** model scaling from **355B to 744B parameters** with **DeepSeek Sparse Attention** integration for cost-efficient long-context serving. GLM-5 achieves **SOTA on BrowseComp** and leads on **Vending Bench 2**, focusing on office productivity tasks and surpassing **Kimi K2.5** on the GDPVal-AA benchmark. Despite broad availability on platforms like **OpenRouter**, **Modal**, **DeepInfra**, and **Ollama Cloud**, GLM-5 faces **compute constraints** impacting rollout and pricing. The model supports up to **200K context length** and **128K max output tokens**.</description><pubDate>Wed, 11 Feb 2026 05:44:39 GMT</pubDate><category>zhipu-ai</category><category>openrouter</category><category>modal</category><category>deepinfra</category><category>ollama</category><category>qoder</category><category>vercel</category><category>glm-5</category><category>glm-4.5</category><category>kimi-k2.5</category><category>deepseek-sparse-attention</category><category>long-context</category><category>model-scaling</category><category>pretraining</category><category>benchmarking</category><category>office-productivity</category><category>context-window</category><category>model-deployment</category><category>cost-efficiency</category></item><item><title>Qwen-Image 2.0 and Seedance 2.0</title><link>https://news.smol.ai/issues/26-02-10-qwenimage-seedance-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-10-qwenimage-seedance-2/</guid><description>**OpenAI** advances its Responses API for multi-hour agent workflows with features like **server-side compaction**, **hosted containers**, and **Skills API**, alongside upgrading **Deep Research** to **GPT-5.2** and adding connectors. Discussions around sandbox design highlight a shift towards **sandbox-as-a-tool** architectures, with **LangChain** enhancing its **deepagents v0.4** with pluggable sandbox backends. Coding agent UX evolves with multi-model orchestration involving **Claude Opus 4.6**, **GPT-5.3-Codex**, and **Gemini 3 Pro**. **EntireHQ** raised **$60M seed** funding for a Git-compatible database capturing code intent and agent context. In model releases, **Alibaba Qwen** launched **Qwen-Image-2.0** emphasizing **2K resolution** and **1K-token prompts** for unified generation and editing. ByteDance&apos;s **Seedance 2.0** marks a significant leap in text-to-video quality, while **Moonshot&apos;s Kimi** introduces an **Agent Swarm** with up to **100 sub-agents** and **4.5× faster** parallel execution.</description><pubDate>Tue, 10 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>langchain-ai</category><category>anthropic</category><category>google-deepmind</category><category>mistral-ai</category><category>alibaba</category><category>bytedance</category><category>moonshot</category><category>gpt-5.2</category><category>gpt-5.3-codex</category><category>claude-opus-4.6</category><category>gemini-3-pro</category><category>qwen-image-2.0</category><category>seedance-2.0</category><category>hwchase17</category><category>nabbilkhan</category><category>sydneyrunkle</category><category>joecuevasjr</category><category>pierceboggan</category><category>reach_vb</category><category>gdb</category><category>ashtom</category><category>agentic-sandboxes</category><category>multi-model-orchestration</category><category>server-side-compaction</category><category>coding-agent-ux</category><category>long-running-agents</category><category>model-release</category><category>text-to-video</category><category>image-generation</category><category>parallel-execution</category><category>funding</category><category>git-compatible-database</category><category>token-efficiency</category><category>workflow-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-02-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-09-not-much/</guid><description>**OpenAI** launched **GPT-5.3-Codex** with a Super Bowl ad emphasizing &quot;You can just build things&quot; as a product strategy, focusing on builder tooling over chat interfaces. The model is rolling out across **Cursor, VS Code, and GitHub** with phased API access and is flagged as their first &quot;high cybersecurity capability&quot; model. Sam Altman reported over **1M Codex app downloads in the first week** and strong weekly user growth. Meanwhile, **Anthropic&apos;s Claude Opus 4.6** is recognized as a leading &quot;agentic generalist&quot; model, topping text and code leaderboards but noted for high token usage. Discussions around serving economics and &quot;fast mode&quot; behavior highlight practical deployment considerations. Additionally, Recursive Language Models (RLMs) introduce a novel approach using a second programmatic context space to extend long-context capabilities.</description><pubDate>Mon, 09 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>cursor_ai</category><category>github</category><category>microsoft</category><category>gpt-5.3-codex</category><category>claude-opus-4.6</category><category>sama</category><category>pierceboggan</category><category>kylebrussell</category><category>natolambert</category><category>omarsar0</category><category>sam_altman</category><category>builder-tooling</category><category>cybersecurity</category><category>api-access</category><category>model-rollout</category><category>agentic-ai</category><category>long-context</category><category>serving-economics</category><category>throughput-latency</category><category>token-efficiency</category><category>workflow-design</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-02-06-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-06-not-much/</guid><description>**AI News** for early February 2026 highlights a detailed comparison between **GPT-5.3-Codex** and **Claude Opus 4.6**, with users noting **Codex&apos;s** strength in detailed scoped tasks and **Opus&apos;s** ergonomic advantage for exploratory work. Benchmarks on Karpathy&apos;s **nanochat GPT-2 speedrun** show **Opus 4.6** achieving better wall-clock performance, while **Codex-5.3-xhigh** sometimes suffers from context issues. **Karpathy** cautions that current models are not yet reliable for fully autonomous AI engineering. Discussions on agent swarms reveal emerging parallels to software organizational design, with **Anthropic-style** agent coordination systems and **LangChain/LangSmith** emphasizing environment engineering through tracing, sandboxing, and state control. The concept of Recursive Language Models (RLM) is introduced as a future direction for agent systems to reduce context rot and improve structured communication.</description><pubDate>Fri, 06 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>langchain</category><category>gpt-5.3-codex</category><category>claude-opus-4.6</category><category>nanochat-gpt-2</category><category>karpathy</category><category>sama</category><category>swyx</category><category>omarsar0</category><category>hamelhusain</category><category>deepfates</category><category>agent-systems</category><category>ai-engineering</category><category>benchmarking</category><category>software-organization</category><category>sandboxing</category><category>tracing</category><category>state-management</category><category>recursive-language-models</category><category>context-management</category></item><item><title>OpenAI and Anthropic go to war: Claude Opus 4.6 vs GPT 5.3 Codex</title><link>https://news.smol.ai/issues/26-02-05-claude-opus-openai-codex/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-05-claude-opus-openai-codex/</guid><description>**OpenAI** launched **GPT-5.3-Codex**, emphasizing **token efficiency**, **inference speed**, and hardware/software co-design with **GB200-NVL72** and **NVIDIA** collaboration. The new **Frontier** agent platform supports business-context agents with execution environments and learning capabilities. **Anthropic** showcased **Opus 4.6** agent teams autonomously building a clean-room C compiler booting Linux, highlighting advances in agentic coding and long-context capabilities. Community benchmarks report **2.93× faster** inference and significant efficiency gains, signaling a shift away from infinite compute budgets in 2026.</description><pubDate>Thu, 05 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>nvidia</category><category>gpt-5.3-codex</category><category>opus-4.6</category><category>agentic-coding</category><category>long-context</category><category>token-efficiency</category><category>inference-speed</category><category>hardware-software-co-design</category><category>agent-platforms</category><category>benchmarking</category><category>software-development</category><category>compiler-construction</category></item><item><title>ElevenLabs $500m Series D at $11B, Cerebras $1B Series H at $23B, Vibe Coding -&gt; Agentic Engineering</title><link>https://news.smol.ai/issues/26-02-04-elevenlabs-cerebras/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-04-elevenlabs-cerebras/</guid><description>**Google&apos;s Gemini 3** is being integrated widely, including a new **Chrome side panel** and **Nano Banana** UX features, with rapid adoption and a **78% unit-cost reduction** in serving costs. The **Gemini app** reached **750M+ MAU** in Q4 2025, nearing ChatGPT&apos;s user base. Google is also benchmarking AI &quot;soft skills&quot; through games like Poker and Chess in the **Kaggle Game Arena**. Meanwhile, coding agents are converging in IDEs: **VS Code** launched **Agent Sessions** supporting **Claude** and **Codex** agents with features like parallel subagents and integrated browsers. **GitHub Copilot** now allows agent choice between **Claude** and **OpenAI Codex** for async backlog clearing. OpenAI reports **1M+ active users** for Codex with expanded integration surfaces, though some users request better GPU support. The coding-agent ecosystem is professionalizing with community platforms like **OpenClaw** and tooling such as ClawHub and CLI updates. *&quot;Gemini 3 adoption faster than any other model&quot;* and *&quot;VS Code as home for coding agents&quot;* highlight major industry shifts.</description><pubDate>Wed, 04 Feb 2026 05:44:39 GMT</pubDate><category>google</category><category>openai</category><category>github</category><category>microsoft</category><category>deepmind</category><category>gemini-3</category><category>claude</category><category>codex</category><category>sama</category><category>sundarpichai</category><category>reach_vb</category><category>agent-frameworks</category><category>model-deployment</category><category>benchmarking</category><category>cost-optimization</category><category>software-development</category><category>async-processing</category><category>gpu-acceleration</category><category>coding-agents</category><category>user-adoption</category><category>game-theory</category><category>workflow-integration</category></item><item><title>Context Graphs: Hype or actually Trillion-dollar opportunity?</title><link>https://news.smol.ai/issues/26-02-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-03-not-much/</guid><description>**Zhipu AI** launched **GLM-OCR**, a lightweight **0.9B** multimodal OCR model excelling in complex document understanding with top benchmark scores and day-0 deployment support from **lmsys**, **vllm**, and **novita labs**. **Ollama** enabled local-first usage with easy offline operation. **Alibaba** released **Qwen3-Coder-Next**, an **80B MoE** model with only **3B active** parameters, designed for coding agents with a massive **256K context window** and trained on **800K verifiable tasks**, achieving over **70% SWE-Bench Verified**. The open coding ecosystem also saw **Allen AI** announce **SERA-14B**, an on-device-friendly coding model with new datasets. The emerging concept of **Context Graphs** was highlighted as a promising framework for data and agent traceability, with initiatives like **Cursor&apos;s Agent Trace** specifying context graphs for coding agents, emphasizing potential improvements in agent performance and customer-driven adoption. This coverage reflects ongoing innovation in **multimodality**, **long-context**, **mixture-of-experts**, and **agentic coding models**.</description><pubDate>Tue, 03 Feb 2026 05:44:39 GMT</pubDate><category>zhipu-ai</category><category>lmsys</category><category>vllm</category><category>novita-labs</category><category>ollama</category><category>alibaba</category><category>allenai</category><category>cognition</category><category>cursor</category><category>glm-ocr</category><category>qwen3-coder-next</category><category>sera-14b</category><category>jaya_gupta</category><category>dharmesh_shah</category><category>multimodality</category><category>ocr</category><category>long-context</category><category>mixture-of-experts</category><category>agentic-coding-models</category><category>context-graphs</category><category>benchmarking</category><category>model-deployment</category><category>model-optimization</category><category>model-training</category></item><item><title>OpenAI Codex App: death of the VSCode fork, multitasking worktrees, Skills Automations</title><link>https://news.smol.ai/issues/26-02-02-openai-codex-app/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-02-02-openai-codex-app/</guid><description>**OpenAI** launched the **Codex app** on macOS as a dedicated agent-native command center for coding, featuring **multiple agents in parallel**, **built-in worktrees** for conflict isolation, **skills** for reusable bundles, and **scheduled automations**. The app emphasizes developer workflows like **Plan mode** for upfront task decomposition and is gaining positive adoption signals from insiders including **@sama**. There is movement towards ecosystem standardization of skills folders, signaling early conventions in agent tooling. Codex also exemplifies a &quot;self-improving&quot; product feedback loop combining humans and agents. In coding agents practice, best practices include a &quot;test-first&quot; approach to bug fixes, the &quot;conductor&quot; model where one developer manages 5-10 agents in parallel, and a neurosymbolic framing explaining why coding agents succeed due to software&apos;s verifiability and symbolic tooling. Benchmark skepticism remains about productivity studies that do not reflect agentic workflows.</description><pubDate>Mon, 02 Feb 2026 05:44:39 GMT</pubDate><category>openai</category><category>codex</category><category>sama</category><category>reach_vb</category><category>gdb</category><category>skirano</category><category>embirico</category><category>ajambrosino</category><category>thsottiaux</category><category>nbaschez</category><category>yuchenj_uw</category><category>badlogicgames</category><category>random_walker</category><category>agent-based-systems</category><category>parallel-processing</category><category>software-testing</category><category>developer-workflows</category><category>automation</category><category>product-feedback-loop</category><category>neurosymbolic-ai</category><category>benchmarking</category></item><item><title>MoltBook takes over the timeline</title><link>https://news.smol.ai/issues/26-01-30-moltbook/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-30-moltbook/</guid><description>**Moltbook** and **OpenClaw** showcase emergent multi-agent social networks where AI agents autonomously interact, creating an AI-native forum layer with complex security and identity challenges. **Karpathy** describes this as &quot;takeoff-adjacent,&quot; highlighting bots self-organizing and engaging in prompt-injection and credential theft. **Anthropic** reports on AI coding tradeoffs with a study of **52 junior engineers** and reveals **Claude** planned a Mars rover drive, marking a milestone in AI-driven space exploration. **Google** publicly releases **Genie 3**, sparking debate over its capabilities and latency issues. The rise of agent-to-agent private communications raises concerns about alignment and observability in 2026.</description><pubDate>Fri, 30 Jan 2026 05:44:39 GMT</pubDate><category>moltbook</category><category>openclaw</category><category>anthropic</category><category>google</category><category>claude</category><category>genie-3</category><category>karpathy</category><category>multi-agent-systems</category><category>agent-communication</category><category>security</category><category>prompt-injection</category><category>identity</category><category>alignment</category><category>observability</category><category>ai-planning</category><category>ai-coding</category><category>emergent-behavior</category></item><item><title>xAI Grok Imagine API - the #1 Video Model, Best Pricing and Latency - and merging with SpaceX</title><link>https://news.smol.ai/issues/26-01-29-xai-grok-imagine-api/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-29-xai-grok-imagine-api/</guid><description>**Google DeepMind** launched **Project Genie (Genie 3 + Nano Banana Pro + Gemini)**, a prototype for creating interactive, real-time generated worlds from text or image prompts, currently available to **Google AI Ultra subscribers in the U.S. (18+)** with noted limitations like **~60s generation limits** and imperfect physics. In parallel, the open-source **LingBot-World** offers a real-time interactive world model with **&lt;1s latency at 16 FPS** and minute-level coherence, emphasizing interactivity and causal consistency. In video generation, **xAI Grok Imagine** debuted strongly with native audio support, **15s duration**, and competitive pricing at **$4.20/min including audio**, while **Runway Gen-4.5** focuses on animation workflows with new features like **Motion Sketch** and **Character Swap**. The 3D generation space sees **fal** adding **Hunyuan 3D 3.1 Pro/Rapid** to its API offerings, extending model-as-a-service workflows into 3D pipelines.</description><pubDate>Thu, 29 Jan 2026 05:44:39 GMT</pubDate><category>google-deepmind</category><category>x-ai</category><category>runway</category><category>fal</category><category>genie-3</category><category>nano-banana-pro</category><category>gemini</category><category>lingbot-world</category><category>grok-imagine</category><category>runway-gen-4.5</category><category>hunyuan-3d-3.1-pro</category><category>demishassabis</category><category>sundarpichai</category><category>interactive-simulation</category><category>real-time-generation</category><category>promptability</category><category>character-customization</category><category>world-models</category><category>open-source</category><category>video-generation</category><category>audio-generation</category><category>animation-workflows</category><category>model-as-a-service</category><category>3d-generation</category><category>latency</category><category>coherence</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-28-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-28-not-much/</guid><description>**AI News for 1/27/2026-1/28/2026** highlights a quiet day with deep dives into frontier model &quot;personality split&quot; where **GPT-5.2** excels at *exploration* and **Claude Opus 4.5** at *exploitation*, suggesting **OpenAI** suits research workflows and **Anthropic** commercial reliability. The rise of agentic coding loops shows new failure modes, with *self-verification* workflows gaining traction. The open-model **Kimi K2.5** emerges as a flashpoint, boasting enhanced **agent execution**, **multimodality**, and **coding polish**, runnable on **Apple silicon M3 Ultra Mac Studios** with **Thunderbolt 5 (RDMA)**, and challenging **Claude Opus 4.5** on benchmarks and pricing. Licensing issues threaten enterprise adoption despite model quality. The meme &quot;clawdbot&quot; reflects rapid agent branding proliferation. Agent engineering advances with shared &quot;skills&quot; interfaces promoted by **DeepLearning.AI**, **Anthropic**, and **LangChain**.</description><pubDate>Wed, 28 Jan 2026 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>deeplearningai</category><category>langchain</category><category>apple</category><category>gpt-5.2</category><category>claude-opus-4.5</category><category>kimi-k2.5</category><category>agentic-ai</category><category>multimodality</category><category>coding</category><category>self-verification</category><category>agent-engineering</category><category>model-benchmarking</category><category>model-optimization</category><category>workflow-automation</category></item><item><title>Moonshot Kimi K2.5 - Beats Sonnet 4.5 at half the cost, SOTA Open Model, first Native Image+Video, 100 parallel Agent Swarm manager</title><link>https://news.smol.ai/issues/26-01-27-kimi-k25/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-27-kimi-k25/</guid><description>**MoonshotAI&apos;s Kimi K2.5** is a **32B active-1T parameter open-weights model** featuring **native multimodality** with image and video understanding, built through continual pretraining on **15 trillion mixed visual and text tokens**. It introduces a new **MoonViT vision encoder** and supports advanced capabilities like **Agent Swarm**, which coordinates up to 100 sub-agents for parallel workflows, and an **Office Productivity K2.5 Agent** for large-scale office tasks. This release marks a significant leap in open models from China, claiming state-of-the-art results on benchmarks like HLE and BrowseComp, and offering aggressive API pricing and throughput.</description><pubDate>Tue, 27 Jan 2026 05:44:39 GMT</pubDate><category>moonshotai</category><category>kimi-k2.5</category><category>multimodality</category><category>model-training</category><category>mixture-of-experts</category><category>agentic-ai</category><category>vision</category><category>video-understanding</category><category>model-optimization</category><category>parallel-processing</category><category>office-productivity</category></item><item><title>Anthropic launches the MCP Apps open spec, in Claude.ai</title><link>https://news.smol.ai/issues/26-01-26-mcp-apps/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-26-mcp-apps/</guid><description>**Anthropic** has officially absorbed the independent MCP UI project and, collaborating with **OpenAI**, **Block**, **VS Code**, **Antigravity**, **JetBrains**, and **AWS**, released the **MCP Apps spec** and official support in **Claude.ai**. This standard aims to enable a rich ecosystem of interoperable applications with rich UI, addressing the proliferation of subscription services. Meanwhile, **NVIDIA** introduced **ToolOrchestra** with an **8B orchestrator** model trained via scalable reinforcement learning for efficient agent orchestration. The concept of Recursive Language Models (RLMs) is gaining traction for efficient context management in agent stacks. The “Clawdbot” UX pattern emphasizes outcome-first assistant design with tight context and tool integration, sparking security concerns around prompt injection. **Alibaba** launched **Qwen3-Max-Thinking**, a flagship reasoning and agent model with adaptive tool use and strong benchmark scores, now available in public evaluation platforms like LM Arena and Yupp.</description><pubDate>Mon, 26 Jan 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>block</category><category>vs-code</category><category>antigravity</category><category>jetbrains</category><category>aws</category><category>nvidia</category><category>alibaba</category><category>claude-ai</category><category>claude-ai</category><category>toolorchestra-8b</category><category>qwen3-max-thinking</category><category>agent-orchestration</category><category>reinforcement-learning</category><category>recursive-language-models</category><category>context-management</category><category>user-experience</category><category>security</category><category>prompt-injection</category><category>reasoning</category><category>adaptive-tool-use</category><category>model-evaluation</category><category>benchmarking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-22-not-much/</guid><description>**Anthropic** launches &quot;Claude in Excel Pro&quot; with enhanced features. **OpenAI** reveals upcoming **Codex** agent loop and cybersecurity measures. **Google** boosts **Gemini App** quotas and partners with **Sakana AI** for advanced AI Scientist projects in Japan. **Cursor** introduces Agent Skills for dynamic context focus. **GPT-5.2 Pro** achieves **31%** on FrontierMath Tier 4, showing significant benchmark progress. **Baseten** raises **$300M** at a **$5B valuation** targeting high-performance inference. Discussions highlight math benchmarks as indicators of AI capability, uneven AGI progress, and the importance of reasoning and continual learning as future frontiers. Notable figures include *Sam Altman*, *François Chollet*, *Shane Legg*, and *Demis Hassabis*.</description><pubDate>Thu, 22 Jan 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>google</category><category>sakana-ai</category><category>cursor</category><category>baseten</category><category>epoch-ai-research</category><category>deepmind</category><category>claude-3</category><category>codex</category><category>gemini</category><category>gpt-5.2-pro</category><category>sama</category><category>fchollet</category><category>shane_legg</category><category>demishassabis</category><category>benchmarking</category><category>reasoning</category><category>continual-learning</category><category>reinforcement-learning</category><category>model-performance</category><category>agentic-ai</category><category>security</category><category>model-training</category></item><item><title>OpenEvidence, the ‘ChatGPT for doctors,’ raises $250m at $12B valuation, 12x from $1b last Feb</title><link>https://news.smol.ai/issues/26-01-21-openevidence/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-21-openevidence/</guid><description>**OpenEvidence** raised **$12 billion**, a 12x increase from last year, with usage by 40% of U.S. physicians and over $100 million in annual revenue. **Anthropic** released a new **Claude** model constitution under **CC0 1.0**, framing it as a living document for alignment and training. **Podium** reported over **$100 million ARR** from **10,000+ AI agents**, shifting from software sales to AI operators. Innovations in agent memory and reliability include the **Agent Cognitive Compressor (ACC)** and multi-agent scientific workflows via **MCP-SIM**. Agentic benchmarking shows challenges in long-horizon tasks with models like **Gemini 3 Flash High**, **GPT-5.2 High**, and **Claude Opus 4.5 High** scoring modestly on professional services and legal research benchmarks.</description><pubDate>Wed, 21 Jan 2026 05:44:39 GMT</pubDate><category>openevidence</category><category>anthropic</category><category>podium</category><category>openai</category><category>google</category><category>gemini</category><category>claude</category><category>claude-3</category><category>claude-opus</category><category>gpt-5.2</category><category>gemini-3-flash-high</category><category>daniel_nadler</category><category>amanda_askell</category><category>eric_rea</category><category>tom_loverro</category><category>garry_tan</category><category>omarsar0</category><category>brendanfoody</category><category>deredleritt3r</category><category>agentic-ai</category><category>model-alignment</category><category>performance-evaluation</category><category>memory-optimization</category><category>long-context</category><category>benchmarking</category><category>multi-agent-systems</category><category>reinforcement-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-20-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-20-not-much/</guid><description>**X Engineering** open-sourced its new transformer-based recommender algorithm, sparking community debate on transparency and fairness. **GLM-4.7-Flash (30B-A3B)** gains momentum as a strong local inference model with efficient KV-cache management and quantization tuning strategies. Innovations include tensor parallelism on Mac Minis achieving ~100 tok/s throughput. Research highlights &quot;Societies of Thought&quot; as a reasoning mechanism improving model accuracy by 20%+.</description><pubDate>Tue, 20 Jan 2026 05:44:39 GMT</pubDate><category>x-ai</category><category>unsloth-ai</category><category>google</category><category>deepseek</category><category>ollama</category><category>glm-4.7-flash</category><category>grok</category><category>deepseek-r1</category><category>qwq</category><category>giffmana</category><category>david_sholz</category><category>yuchenj_uw</category><category>nearcyan</category><category>sam_paech</category><category>teortaxes_tex</category><category>danielhanchen</category><category>alexocheema</category><category>nopmobiel</category><category>rohanpaul_ai</category><category>transformer-architecture</category><category>recommendation-systems</category><category>local-inference</category><category>kv-cache</category><category>quantization</category><category>tensor-parallelism</category><category>reasoning</category><category>model-optimization</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-19-not-much/</guid><description>**AI News for 1/16/2026-1/19/2026** covers new architectures for scaling Transformer memory and context, including **STEM** from **Carnegie Mellon** and **Meta AI**, which replaces part of the FFN with a token-indexed embedding lookup enabling CPU offload and asynchronous prefetch. **RePo** from **Sakana AI** introduces adaptive positional reordering to improve robustness on noisy and long-range contexts. Model releases highlight **Zhipu AI&apos;s GLM-4.7-Flash**, a **30B-class MLA + small MoE** model optimized for coding and agentic tasks, noted for strong benchmark performance and a compression narrative from larger to smaller models. Inference and deployment updates include **mlx-lm 0.30.3** supporting GLM-4.7-Flash with efficient 4-bit performance on laptops. The report emphasizes practical takeaways on static sparsity, adaptive ordering, and the resurgence of small, fast models for interactive tasks. *&quot;Sparse capacity doesn’t have to mean MoE routers + expert parallelism; static sparsity can be systems-friendly.&quot;*</description><pubDate>Mon, 19 Jan 2026 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>carnegie-mellon</category><category>sakana-ai</category><category>zhipu-ai</category><category>glm-4.7-flash</category><category>glm-4.7</category><category>glm-4.5</category><category>qwen3-vl</category><category>qwen</category><category>transformer-memory</category><category>model-architecture</category><category>mixture-of-experts</category><category>adaptive-position-encoding</category><category>long-context</category><category>model-compression</category><category>inference-optimization</category><category>local-inference</category><category>model-deployment</category><category>benchmarking</category><category>coding</category><category>agentic-ai</category></item><item><title>ChatGPT starts testing ads on free tier + new $8/mo Go plan in the US</title><link>https://news.smol.ai/issues/26-01-16-chatgpt-ads/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-16-chatgpt-ads/</guid><description>**OpenAI** announced the **ChatGPT Go** tier at **$8/month** with ads testing in the US free tier, emphasizing that ads will not influence responses and will be clearly labeled. The update includes memory improvements and a &quot;very fast Codex&quot; feature teased by **Sam Altman**. The Codex CLI ecosystem now supports open-weight models with improved context length. Discussions highlight the importance of human-in-the-loop for reliability in agent orchestration and file interface improvements over traditional retrieval-augmented generation.</description><pubDate>Fri, 16 Jan 2026 05:44:39 GMT</pubDate><category>openai</category><category>ollama</category><category>chatgpt-go</category><category>codex</category><category>sama</category><category>sam_altman</category><category>fidjissimo</category><category>scaling01</category><category>tomwarren</category><category>embirico</category><category>adamdotdev</category><category>ollama</category><category>thsottiaux</category><category>lateinteraction</category><category>dbreunig</category><category>ads</category><category>monetization</category><category>memory</category><category>agent-orchestration</category><category>human-in-the-loop</category><category>cli-tools</category><category>context-length</category><category>workflow-optimization</category></item><item><title>Open Responses: explicit spec for OpenAI&apos;s Responses API supported by OpenRouter, Ollama, Huggingface, vLLM, et al</title><link>https://news.smol.ai/issues/26-01-15-openresponses/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-15-openresponses/</guid><description>**OpenAI** launched the **Open Responses** API spec, an open-source, multi-provider standard for interoperable LLM APIs designed to simplify agent stacks and tooling. Early adopters like **ollama** and **vLLM** support the spec, while notable absences include **anthropic** and **google-deepmind**. Agent design insights from **Cursor** emphasize explicit roles and planning over mega-agent models, with **GPT-5.2** outperforming **Opus 4.5** in long runs. The emerging dominant context/memory abstraction for agents is a **filesystem-as-memory** approach, championed by **llamaindex** and **langchain**, using virtual filesystems often backed by databases like Postgres. LangChain also shipped an open-source desktop interface for agent orchestration called **openwork**. This news highlights advances in API standardization, agent architecture, and memory abstractions in AI development.</description><pubDate>Thu, 15 Jan 2026 05:44:39 GMT</pubDate><category>openai</category><category>ollama</category><category>vllm</category><category>openrouter</category><category>anthropic</category><category>google-deepmind</category><category>langchain</category><category>llamaindex</category><category>gpt-5.2</category><category>opus-4.5</category><category>reach_vb</category><category>simonw</category><category>yuchenj_uw</category><category>omarsar0</category><category>jerryjliu0</category><category>hwchase17</category><category>swyx</category><category>interoperable-apis</category><category>agent-architecture</category><category>filesystem-memory</category><category>api-standardization</category><category>multi-agent-systems</category><category>prompt-engineering</category><category>model-comparison</category><category>virtual-filesystems</category><category>open-source</category><category>agent-ux</category></item><item><title>not much happened today.</title><link>https://news.smol.ai/issues/26-01-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-14-not-much/</guid><description>**OpenAI** launched **GPT-5.2-Codex** API, touted as their strongest coding model for long-running tasks and cybersecurity. **Cursor** integrated GPT-5.2-Codex to autonomously run a browser for a week, producing over 3 million lines of Rust code. **GitHub** incorporated it into their code tools, easing enterprise adoption. Discussions highlight the importance of review loops in agent systems and debate evaluation metrics for coding models. **OpenAI** partnered with **Cerebras** to improve inference speed and latency, with Cerebras serving **GLM-4.7** at 1,445 tokens/sec and low latency. Provider benchmarking reveals tradeoffs in throughput, latency, and context window sizes. **Modal** shared operational scaling insights for self-hosted inference fleets of 20k GPUs, focusing on batch inference optimization with **vLLM** and FlashInfer backend. This reflects a focus on inference infrastructure, long-horizon autonomous agents, and coding model evaluation.</description><pubDate>Wed, 14 Jan 2026 05:44:39 GMT</pubDate><category>openai</category><category>cursor</category><category>github</category><category>cerebras</category><category>modal</category><category>artificial-analysis</category><category>vllm</category><category>gpt-5.2-codex</category><category>glm-4.7</category><category>swyx</category><category>kevinweil</category><category>pierceboggan</category><category>mntruell</category><category>scaling01</category><category>long-running-tasks</category><category>autonomous-agents</category><category>code-generation</category><category>inference-speed</category><category>latency</category><category>batch-inference</category><category>gpu-scaling</category><category>model-evaluation</category><category>agent-systems</category><category>operational-scaling</category></item><item><title>Anthropic Labs: Cowork, Claude Code, MCP, Skills incubator led by Mike Krieger and Ben Mann</title><link>https://news.smol.ai/issues/26-01-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-13-not-much/</guid><description>**Anthropic** consolidates its AI agent products under the **Cowork** brand, integrating prior tools like **Claude Code** and **Claude for Chrome** into a unified agent with sandboxed Linux VM environments using **Apple&apos;s virtualization** and **bubblewrap** for security. Meanwhile, **Anthropic Labs** reorganizes with Mike Krieger stepping down as CPO, focusing on productizing **Claude** with a &gt;$1B ARR agent lab. The AI community debates the meaning of &quot;vibe coding,&quot; emphasizing disciplined engineer verification over casual coding. **LangChain** launches **Agent Builder GA**, offering no-code but powerful agent orchestration features like memory, triggers, and human-in-the-loop approvals. Some experts advocate simplifying agent tooling to core filesystem and bash access for efficiency. Open-source recreations of Cowork-like environments using **QEMU** and sandboxing tools highlight rapid commoditization of AI agent tech.</description><pubDate>Tue, 13 Jan 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>langchain</category><category>apple</category><category>claude</category><category>claude-code</category><category>mike_krieger</category><category>ben_mann</category><category>gergely_orosz</category><category>yuchen_jin</category><category>harrison_chase</category><category>jared_z</category><category>sandboxing</category><category>agent-ux</category><category>agent-orchestration</category><category>human-in-the-loop</category><category>memory-management</category><category>tooling-simplification</category><category>linux-virtualization</category><category>security</category><category>agent-productization</category></item><item><title>Apple picks Google&apos;s Gemini to power Siri&apos;s next generation</title><link>https://news.smol.ai/issues/26-01-12-gemini-apple/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-12-gemini-apple/</guid><description>**Apple** has decided to power Siri with **Google&apos;s Gemini models** and cloud technology, marking a significant partnership and a setback for **OpenAI**, which was initially partnered with Apple. **Anthropic** launched &quot;Cowork,&quot; a product preview for Claude&apos;s coding capabilities, sparking discussions about &quot;LLM OS&quot;. **OpenAI** introduced **ChatGPT Health** and acquired **Torch** to expand in healthcare AI. **DeepSeek** unveiled **Engram**, a new conditional memory module that enables O(1) lookup-style memory for static patterns, improving long-context handling and offering hardware-friendly optimizations to scale knowledge capacity efficiently. Engram is positioned as a key modeling primitive for next-gen sparse models, with ongoing community debate about its architectural merits and practical impact.</description><pubDate>Mon, 12 Jan 2026 05:44:39 GMT</pubDate><category>apple</category><category>google</category><category>openai</category><category>anthropic</category><category>deepseek</category><category>gemini</category><category>claude</category><category>chatgpt</category><category>engram</category><category>conditional-memory</category><category>long-context</category><category>hashing</category><category>memory-optimization</category><category>transformers</category><category>model-scaling</category><category>sparsity</category><category>hardware-optimization</category><category>model-architecture</category><category>ai-healthcare</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-09-not-much/</guid><description>**Anthropic** tightens usage policies for **Claude Max** in third-party apps, prompting builders to adopt **model-agnostic orchestration** and **BYO-key** defaults to mitigate platform risks. The **Model Context Protocol (MCP)** is evolving into a key tooling plane with **OpenAI MCP Server** and **mcp-cli** enhancing tool discovery and token efficiency. The concept of **skills** as modular, versioned behaviors gains traction, with implementations in **Claude Code**, **GitHub Copilot**, and **Cline** adding websearch tooling. AI21 Labs addresses concurrency challenges in agent workspaces using **git worktrees** for transactional parallel writes, while long-horizon agents focus on **context engineering** and persistent file-centric workspaces.</description><pubDate>Fri, 09 Jan 2026 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>ai21-labs</category><category>github</category><category>cline</category><category>claude-max</category><category>yuchenj_uw</category><category>andersonbcdefg</category><category>gneubig</category><category>matan_sf</category><category>scaling01</category><category>reach_vb</category><category>_philschmid</category><category>claude_code</category><category>code</category><category>jamesmontemagno</category><category>cline</category><category>danstripper</category><category>omarsar0</category><category>model-agnostic</category><category>model-context-protocol</category><category>tooling</category><category>skills</category><category>concurrency</category><category>transactional-workspaces</category><category>context-engineering</category><category>file-centric-workspaces</category><category>rate-limiting</category><category>agent-workspaces</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-08-not-much/</guid><description>**Stanford paper** reveals **Claude 3.7 Sonnet** memorized **95.8% of Harry Potter 1**, highlighting copyright extraction risks compared to **GPT-4.1**. **Google AI Studio** sponsors **TailwindCSS** amid OSS funding debates. **Google** and **Sundar Pichai** launch **Gmail Gemini 3** features including AI Overviews and natural-language search with user controls. **Alibaba Qwen** releases **Qwen3-VL-Embedding** and **Qwen3-VL-Reranker**, a multimodal, multilingual retrieval stack supporting text, images, and video with quantization and instruction customization, achieving strong benchmark results. **Z.ai** goes public on HKEX with **GLM-4.7** leading the Artificial Analysis Intelligence Index v4.0, showing gains in reasoning, coding, and agentic use, with large-scale MoE architecture and MIT license. **Falcon-H1R-7B** from TII targets efficient reasoning in smaller models, scoring 16 on the Intelligence Index. **AI21 Labs** introduces **Jamba2**, a memory-efficient enterprise model with hybrid SSM-Transformer architecture and Apache 2.0 license, available via SaaS and Hugging Face. **vLLM** shows throughput improvements in inference and kernel engineering. *&quot;Embeddings should be multimodal by default,&quot;* notes Justin Lin.</description><pubDate>Thu, 08 Jan 2026 05:44:39 GMT</pubDate><category>stanford</category><category>google</category><category>google-deepmind</category><category>alibaba</category><category>z-ai</category><category>tii</category><category>ai21-labs</category><category>huggingface</category><category>claude-3-7-sonnet</category><category>gpt-4-1</category><category>gemini-3</category><category>qwen3-vl-embedding</category><category>qwen3-vl-reranker</category><category>glm-4-7</category><category>falcon-h1r-7b</category><category>jamba2</category><category>sundarpichai</category><category>justinlin610</category><category>copyright-extraction</category><category>multimodality</category><category>multilinguality</category><category>retrieval-augmented-generation</category><category>model-architecture</category><category>mixture-of-experts</category><category>model-quantization</category><category>reasoning</category><category>inference</category><category>kernel-engineering</category><category>memory-optimization</category><category>enterprise-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-07-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-07-not-much/</guid><description>**AI News for 1/6/2026-1/7/2026** highlights a quiet day with key updates on **LangChain DeepAgents** introducing **Ralph Mode** for persistent agent loops, **Cursor** improving context management by reducing token usage by **46.9%**, and operational safety measures for coding agents with allow/deny lists. **MCP** integration is expanding across assistants and robotics, with Hugging Face embedding assistants via **HuggingChat + HF MCP server**. The **DeepSeek-R1** paper has been expanded to **86 pages**, emphasizing trajectory exploration and RL shaping behavior. **NousCoder-14B** shows a **+7% improvement on LiveCodeBench** after **4 days** of RL training, demonstrating advances in RL for coding with small open models. Top tweets also mention a viral &quot;96GB RAM laptop&quot;, **ChatGPT Health** launch by **OpenAI**, and **Karpathy**&apos;s nanochat scaling-law miniseries.</description><pubDate>Wed, 07 Jan 2026 05:44:39 GMT</pubDate><category>langchain</category><category>cursor</category><category>huggingface</category><category>openai</category><category>weights-biases</category><category>nouscoder-14b</category><category>deepseek-r1</category><category>karpathy</category><category>_philschmid</category><category>omarsar0</category><category>agent-frameworks</category><category>context-management</category><category>reinforcement-learning</category><category>operational-safety</category><category>model-transparency</category><category>trajectory-exploration</category><category>token-optimization</category><category>coding-agents</category><category>integration-platforms</category></item><item><title>xAI raises $20B Series E at ~$230B valuation</title><link>https://news.smol.ai/issues/26-01-06-xai-series-e/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-06-xai-series-e/</guid><description>**xAI**, Elon Musk&apos;s AI company, completed a massive **$20 billion Series E funding round**, valuing it at about **$230 billion** with investors like **Nvidia**, **Cisco Investments**, and others. The funds will support AI infrastructure expansion including **Colossus I and II supercomputers** and training **Grok 5**, leveraging data from **X&apos;s 600 million monthly active users**. At **CES 2026**, the focus was on &quot;AI everywhere&quot; with a strong emphasis on **AI-first hardware** and integration between **NVIDIA** and **Hugging Face&apos;s LeRobot** for robotics development. The **Reachy Mini** robot is gaining traction as a consumer robotics platform. In software, **Claude Code** is emerging as a popular local/private coding assistant, with new UI features in **Claude Desktop** and innovations like **Cursor&apos;s dynamic context** reducing token usage by nearly **47%** in multi-MCP setups. *&quot;The 600 million MAU figure in xAI’s announcement combines X platform users with Grok users. That’s a clever framing choice.&quot;*</description><pubDate>Tue, 06 Jan 2026 05:44:39 GMT</pubDate><category>xai</category><category>nvidia</category><category>cisco</category><category>fidelity</category><category>valor-equity-partners</category><category>qatar-investment-authority</category><category>mgx</category><category>stepstone-group</category><category>baron-capital-group</category><category>hugging-face</category><category>amd</category><category>grok-5</category><category>claude-code</category><category>aakash_gupta</category><category>fei-fei_li</category><category>lisa_su</category><category>clementdelangue</category><category>thom_wolf</category><category>saradu</category><category>omarsar0</category><category>yuchenj_uw</category><category>_catwu</category><category>cursor_ai</category><category>ai-infrastructure</category><category>supercomputing</category><category>robotics</category><category>ai-hardware</category><category>agentic-ai</category><category>context-management</category><category>token-optimization</category><category>local-ai-assistants</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-05-nvidia-vera-rubin/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-05-nvidia-vera-rubin/</guid><description>**AI News** from early January 2026 highlights a viral economic prediction about **Vietnam** surpassing Thailand, **Microsoft**&apos;s reported open-sourcing of **bitnet.cpp** for 1-bit CPU inference promising speed and energy gains, and a new research partnership between **Google DeepMind** and **Boston Dynamics** focusing on **Gemini Robotics** and **Atlas hardware**. The concept of **agentic coding** is gaining traction, emphasizing human oversight and infrastructure layers called **Agent Harnesses** to manage long-running AI tasks, with advocates like **Philipp Schmid** promoting this shift. Innovations in persistent memory for coding agents, such as **Claude-Mem**, aim to improve context durability. There is also critical discussion on the specification problem in agent workflows, advocating for better abstractions beyond conversational intent. Practical challenges include managing parallel agents and permission risks. Additionally, open tooling advances include a **JAX-based LLM-Pruning Collection** for efficient model pruning methods.</description><pubDate>Mon, 05 Jan 2026 05:44:39 GMT</pubDate><category>microsoft</category><category>google-deepmind</category><category>boston-dynamics</category><category>claude-mem</category><category>bitnet-cpp</category><category>gemini</category><category>_philschmid</category><category>demishassabis</category><category>agentic-coding</category><category>agent-harnesses</category><category>persistent-memory</category><category>software-engineering</category><category>inference-efficiency</category><category>model-pruning</category><category>context-durability</category><category>specification-problem</category><category>workflow-management</category><category>cpu-inference</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/26-01-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/26-01-02-not-much/</guid><description>**DeepSeek** released a new paper on **mHC: Manifold-Constrained Hyper-Connections**, advancing residual-path design as a key scaling lever in neural networks. Their approach constrains residual mixing matrices to the **Birkhoff polytope** to improve stability and performance, with only about **6.7% training overhead**. The innovation includes systems-level optimizations like fused kernels and activation recomputation, highlighting a frontier-lab integration of math and kernel engineering. Additionally, discussions around **long-horizon agents** emphasize context management bottlenecks, introducing **Recursive Language Models (RLMs)** that manage context dynamically rather than relying on larger context windows. This work signals a shift in architectural design and efficiency for base model training and agent development.</description><pubDate>Fri, 02 Jan 2026 05:44:39 GMT</pubDate><category>deepseek</category><category>bytedance</category><category>teortaxestex</category><category>askperplexity</category><category>rasbt</category><category>norxornor</category><category>dorialexander</category><category>iamgrigorev</category><category>primeintellect</category><category>a1zhang</category><category>residual-path-design</category><category>manifold-constrained-hyper-connections</category><category>birkhoff-polytope</category><category>training-overhead</category><category>kernel-optimization</category><category>activation-recomputation</category><category>pipeline-parallelism</category><category>long-horizon-agents</category><category>context-management</category><category>recursive-language-models</category><category>neural-network-stability</category><category>scaling-levers</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-31-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-31-not-much/</guid><description>**South Korea&apos;s Ministry of Science** launched a coordinated program with **5 companies** to develop sovereign foundation models from scratch, featuring large-scale MoE architectures like **SK Telecom A.X-K1 (519B total / 33B active)** and **LG K-EXAONE (236B MoE / 23B active)**, with a total first-round budget of **~$140M**. This initiative contrasts with EU approaches by focusing funding on fewer stakeholders and explicitly budgeting for data. Meanwhile, **Alibaba&apos;s Qwen-Image-2512** emerges as a leading open-source image generation model, rapidly integrated into various toolchains including AI-Toolkit and local inference paths with quantization support, and hosted on platforms like Replicate. The model has undergone extensive blind testing with over **10,000 rounds** on AI Arena, highlighting its ecosystem adoption.</description><pubDate>Wed, 31 Dec 2025 05:44:39 GMT</pubDate><category>sk-telecom</category><category>lg</category><category>upstage</category><category>naver</category><category>alibaba</category><category>unsloth</category><category>replicate</category><category>qwen-image-2512</category><category>ax-k1</category><category>k-exaone</category><category>eliebakouch</category><category>clementdelangue</category><category>dorialexander</category><category>rising_sayak</category><category>_akhaliq</category><category>ostrisai</category><category>ivanfioravanti</category><category>yupp_ai</category><category>mixture-of-experts</category><category>model-release</category><category>quantization</category><category>open-source-models</category><category>image-generation</category><category>model-integration</category><category>model-benchmarking</category><category>compute-costs</category><category>dataset-curation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-30-not-much/</guid><description>**Z.ai (GLM family) IPO in Hong Kong on Jan 8, 2026**, aiming to raise **$560M** at **HK$4.35B**, marking it as the &quot;first AI-native LLM company&quot; public listing. The IPO highlights **GLM-4.7** as a starting point. **Meta AI** acquired **Manus** for approximately **$4–5B**, with Manus achieving **$100M ARR in 8–9 months**, illustrating the value of application-layer differentiation over proprietary models. Manus focuses on agentic architecture, context engineering, and general primitives like code execution and browser control, emphasizing &quot;agent habitats&quot; as a competitive moat. Discussions around **Claude Code** highlight skepticism about &quot;vibe coding,&quot; advocating for disciplined, framework-like AI-assisted programming practices.</description><pubDate>Tue, 30 Dec 2025 05:44:39 GMT</pubDate><category>z.ai</category><category>meta-ai-fair</category><category>manus</category><category>replit</category><category>glm-4.7</category><category>claude-code</category><category>zixuanli_</category><category>jietang</category><category>yuchenj_uw</category><category>sainingxie</category><category>amasad</category><category>hidecloud</category><category>imjaredz</category><category>random_walker</category><category>agentic-architecture</category><category>context-engineering</category><category>application-layer</category><category>code-generation</category><category>agent-habitats</category><category>ai-native-llm</category><category>ipo</category><category>inference-infrastructure</category><category>programming-paradigms</category></item><item><title>Meta Superintelligence Labs acquires Manus AI for over $2B, at $100M ARR, 9months after launch</title><link>https://news.smol.ai/issues/25-12-29-meta-manus/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-29-meta-manus/</guid><description>**Manus** achieved a rapid growth trajectory in 2025, raising **$500M** from Benchmark and reaching **$100M ARR** before being acquired by **Meta** for an estimated **$4B**. The **vLLM** team launched a dedicated community site with new resources, while performance issues with **AMD MI300X FP8** were noted in **vLLM** and **sglang** benchmarks. **Weaviate** released operational features including **Object TTL**, **Java v6 client GA**, and **multimodal document embeddings**. API fragmentation concerns were raised by **Teknium** advocating for unified SDK wrappers. In open-weight models, **GLM-4.7** gained recognition as a reliable coding model with faster throughput on **Baseten**, and **MiniMax-M2.1** rose as a leading open agentic coder model, topping WebDev leaderboards.</description><pubDate>Mon, 29 Dec 2025 05:44:39 GMT</pubDate><category>manus</category><category>benchmark</category><category>meta-ai-fair</category><category>vllm</category><category>amd</category><category>sglang</category><category>weaviate</category><category>teknim</category><category>baseten</category><category>alphaxiv</category><category>minimax</category><category>glm-4.7</category><category>minimax-m2.1</category><category>vllm</category><category>alex_wang</category><category>nat_friedman</category><category>performance-optimization</category><category>inference-frameworks</category><category>model-benchmarking</category><category>model-deployment</category><category>open-source-models</category><category>multimodality</category><category>api</category><category>code-generation</category><category>community-building</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-26-not-much/</guid><description>**MiniMax M2.1** launches as an **open-source** agent and coding Mixture-of-Experts (MoE) model with **~10B active / ~230B total parameters**, claiming to outperform **Gemini 3 Pro** and **Claude Sonnet 4.5**, and supports local inference including on **Apple Silicon M3 Ultra** with quantization. **GLM 4.7** demonstrates local scaling on **Mac Studios** with **2× 512GB M3 Ultra** hardware, highlighting system-level challenges like bandwidth and parallelism. The concept of **inference quality** is emphasized as a key factor affecting output variance across deployments. Yann LeCun&apos;s **VL-JEPA** proposes a **non-generative, non-autoregressive** multimodal model operating in latent space for efficient real-time video processing with fewer parameters and decoding operations. Advances in agentic reinforcement learning for coding include self-play methods where agents inject and fix bugs autonomously, enabling self-improvement without human labeling, and large-scale RL infrastructure involving massive parallel code generation and execution sandboxes.</description><pubDate>Fri, 26 Dec 2025 05:44:39 GMT</pubDate><category>minimax-ai</category><category>vllm-project</category><category>exolabs</category><category>mlx</category><category>apple</category><category>openai</category><category>minimax-m2.1</category><category>glm-4.7</category><category>gemini-3-pro</category><category>claude-3-sonnet</category><category>vl-jepa</category><category>ylecun</category><category>awnihannun</category><category>alexocheema</category><category>edwardsun0909</category><category>johannes_hage</category><category>open-source</category><category>mixture-of-experts</category><category>local-inference</category><category>quantization</category><category>inference-quality</category><category>multimodality</category><category>non-autoregressive-models</category><category>video-processing</category><category>reinforcement-learning</category><category>self-play</category><category>agentic-rl</category><category>parallel-computing</category><category>model-deployment</category></item><item><title>Nvidia buys (most of) Groq for $20B cash; largest execuhire ever</title><link>https://news.smol.ai/issues/25-12-24-nvidia-groq/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-24-nvidia-groq/</guid><description>**Groq** leadership team is joining **Nvidia** under a &quot;non-exclusive licensing agreement&quot; in a deal valued at **$20 billion cash**, marking a major acquisition in AI chip space though Nvidia states it is not acquiring Groq as a company. Jensen Huang plans to integrate Groq&apos;s low-latency processors into the NVIDIA AI factory architecture to enhance AI inference and real-time workloads. Twitter highlights include **Gemini** used as a consumer utility for calorie tracking, OpenAI discussing the &quot;deployment gap&quot; focusing on model usage in healthcare and business, and Tesla&apos;s FSD v14 described as a &quot;Physical Turing Test&quot; for consumer AI. Benchmarking challenges are noted by **Epoch AI** emphasizing provider variance and integration issues affecting model quality measurement. Discussions on coding agents and developer experience convergence continue in the AI community.</description><pubDate>Wed, 24 Dec 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>groq</category><category>openai</category><category>tesla</category><category>epoch-ai</category><category>gemini</category><category>gemini</category><category>fsd-v14</category><category>jensen_huang</category><category>xeophon</category><category>js_denain</category><category>jim_fan</category><category>benchmarking</category><category>inference</category><category>model-evaluation</category><category>ai-integration</category><category>agent-patterns</category><category>real-time-processing</category><category>low-latency</category><category>developer-experience</category><category>healthcare</category><category>business-workflows</category><category>consumer-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-23-not-much/</guid><description>**GLM-4.7** and **MiniMax M2.1** open-weight model releases highlight day-0 ecosystem support, coding throughput, and agent workflows, with GLM-4.7 achieving a +9.5% improvement over GLM-4.6 and MiniMax M2.1 positioned as an OSS Claude-like MoE model with 230B total parameters and 200K context. **Gemma Scope 2** from **google-deepmind** introduces sparse autoencoders and transcoders for interpretability across Gemma 3 models, aiming to provide shared infrastructure for safety and debugging. The **Medmarks v0.1** open medical evaluation suite and leaderboard launch addresses the need for open medical benchmarking across 15+ environments, engaging clinicians and researchers.</description><pubDate>Tue, 23 Dec 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>valsai</category><category>minimax-ai</category><category>ollama</category><category>trae</category><category>alibaba</category><category>sophont</category><category>prime-intellect</category><category>glm-4.7</category><category>glm-4.6</category><category>minimax-m2.1</category><category>gemma-3</category><category>gemma-scope-2</category><category>ivanfioravanti</category><category>awnihannun</category><category>deedydas</category><category>cline</category><category>omarsar0</category><category>adonis_singh</category><category>eliebakouch</category><category>teortaxestex</category><category>ibragim_bad</category><category>callum_mcdougall</category><category>neelnanda5</category><category>interpretability</category><category>sparse-autoencoders</category><category>agent-workflows</category><category>model-benchmarking</category><category>medical-evaluation</category><category>multi-agent-systems</category><category>model-performance</category><category>model-optimization</category><category>reinforcement-learning</category><category>tool-use</category><category>function-calling</category><category>context-windows</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-22-not-much/</guid><description>**Zhipu AI&apos;s GLM-4.7** release marks a significant improvement in **coding, complex reasoning, and tool use**, quickly gaining ecosystem adoption via Hugging Face and OpenRouter. **Xiaomi&apos;s MiMo-V2-Flash** is highlighted as a practical, cost-efficient mixture-of-experts model optimized for deployment. The open-weight text-to-image competition sees **Z-Image Turbo** leading with 6B parameters under Apache-2.0 license. Video model advances focus on control and long-form consistency, exemplified by **Kling 2.6 Motion Control** and research like MemFlow&apos;s adaptive memory retrieval. In agent frameworks, **Google&apos;s A2UI protocol** introduces agent-driven UI generation, while studies reveal that mixing multiple agent frameworks is common, with challenges in logic, termination, and tool interaction. LangChain emphasizes persistent memory patterns for production agents.</description><pubDate>Mon, 22 Dec 2025 05:44:39 GMT</pubDate><category>zhipu-ai</category><category>xiaomi</category><category>google</category><category>langchain</category><category>huggingface</category><category>openrouter</category><category>artificial-analysis</category><category>vllm-project</category><category>glm-4.7</category><category>mimo-v2-flash</category><category>z-image-turbo</category><category>kling-2.6-motion-control</category><category>mervenoyann</category><category>eliebakouch</category><category>omarsar0</category><category>osanseviero</category><category>dair_ai</category><category>coding</category><category>complex-reasoning</category><category>tool-use</category><category>mixture-of-experts</category><category>cost-efficiency</category><category>open-weight-models</category><category>text-to-image</category><category>video-models</category><category>memory-persistence</category><category>agent-frameworks</category><category>interactive-user-interfaces</category><category>model-deployment</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-19-not-much/</guid><description>**Alibaba** released **Qwen-Image-Layered**, an open-source model enabling Photoshop-grade layered image decomposition with recursive infinite layers and prompt-controlled structure. **Kling 2.6** introduced advanced motion control for image-to-video workflows, supported by a creator contest and prompt recipes. **Runway** unveiled the **GWM-1** family with frame-by-frame video generation and Gen-4.5 updates adding audio and multi-shot editing. In LLM platforms, **Gemini 3 Flash** leads benchmarks over **GPT-5.2**, attributed to agentic reinforcement learning improvements post-distillation. Users note **GPT-5.2** excels at long-context tasks (~256k tokens) but face UX limitations pushing some to use **Codex CLI**. Discussions around **Anthropic Opus 4.5** suggest perceived model degradation linked to user expectations.</description><pubDate>Fri, 19 Dec 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>kling-ai</category><category>runway</category><category>google</category><category>anthropic</category><category>openai</category><category>qwen-image-layered</category><category>kling-2.6</category><category>gwm-1</category><category>gen-4.5</category><category>gemini-3-flash</category><category>gpt-5.2</category><category>codex-cli</category><category>opus-4.5</category><category>ankesh_anand</category><category>image-decomposition</category><category>motion-control</category><category>video-generation</category><category>agentic-reinforcement-learning</category><category>long-context</category><category>model-degradation</category><category>benchmarking</category><category>tool-use</category><category>prompt-engineering</category></item><item><title>Claude Skills grows: Open Standard, Directory, Org Admin</title><link>https://news.smol.ai/issues/25-12-18-claude-skills-grows/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-18-claude-skills-grows/</guid><description>**Claude Skills** are gaining significant traction since their launch in October, with a milestone of 100k views in one day for the Claude Skills talk, signaling growing adoption and importance. Announcements include org admin support, a new Skills Directory, and the move to an open standard named **Agent Skills**. In frontier model launches, **OpenAI** released **GPT-5.2-Codex**, touted as the best agentic coding model with improvements in native compaction, long-context reliability, and tool-calling, emphasizing real-world security impacts. **Google DeepMind** introduced **Gemini 3 Flash**, focusing on speed as a product feature impacting workflows and user engagement, alongside **FunctionGemma** and **T5Gemma 2**, emphasizing on-device deployment, fine-tuning, and multimodality.</description><pubDate>Thu, 18 Dec 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>google-deepmind</category><category>hugging-face</category><category>claude-skills</category><category>gpt-5.2-codex</category><category>gemini-3-flash</category><category>functiongemma</category><category>t5gemma-2</category><category>sama</category><category>gregbrockman</category><category>philschmid</category><category>agentic-ai</category><category>fine-tuning</category><category>long-context</category><category>tool-calling</category><category>on-device-ai</category><category>multimodality</category><category>security</category><category>workflow-optimization</category></item><item><title>Gemini 3.0 Flash Preview: 1/4 cost of Pro, but ~as smart, retakes Pareto Frontier</title><link>https://news.smol.ai/issues/25-12-17-gemini-3-flash/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-17-gemini-3-flash/</guid><description>**Google** launched **Gemini 3 Flash**, a pro-grade reasoning model with flash latency, supporting tool calling and multimodal IO, available via multiple platforms including Google AI Studio and Vertex AI. It offers competitive pricing at $0.50 per 1M input tokens and $3.00 per 1M output tokens, with context windows up to 1M tokens. Benchmarks show **Gemini 3 Flash** rivals or outperforms larger models like **GPT-5.2** and **Gemini 3 Pro** in agentic, coding, and reasoning tasks, validated by ARC-AGI-2, SWE-bench, LMArena, and Arena benchmarks. Despite some tradeoffs like high token use and hallucination rates, it is cost-effective overall. Key figures include **Sundar Pichai**, **Jeff Dean**, and **Demis Hassabis** who publicly celebrated this achievement. The model&apos;s tool calling capabilities were demonstrated with 100 tools in a live demo.</description><pubDate>Wed, 17 Dec 2025 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>gemini-3-flash</category><category>gemini-3</category><category>gpt-5.2</category><category>gemini-3-pro</category><category>sundar_pichai</category><category>jeffdean</category><category>demishassabis</category><category>tool-calling</category><category>multimodality</category><category>benchmarking</category><category>reasoning</category><category>cost-efficiency</category><category>model-performance</category><category>context-window</category><category>agentic-ai</category><category>model-deployment</category></item><item><title>OpenAI GPT Image-1.5 claims to beat Nano Banana Pro, #1 across all Arenas, but completely fails Vibe Checks</title><link>https://news.smol.ai/issues/25-12-16-gpt-image-15/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-16-gpt-image-15/</guid><description>**OpenAI** released its new image model **GPT Image 1.5**, featuring precise image editing, better instruction following, improved text and markdown rendering, and faster generation up to 4×. Despite topping multiple leaderboards like **LMArena (1277)**, **Design Arena (1344)**, and **AA Arena (1272)**, user feedback from Twitter, Reddit, and Discord communities is largely negative compared to **Nano Banana Pro** by **Gemini**. Xiaomi introduced the **MiMo-V2-Flash**, a **309B MoE** model optimized for inference efficiency with **256K context window**, achieving state-of-the-art scores on SWE-Bench. The model uses Hybrid Sliding Window Attention and multi-token prediction, offering significant speedups and efficiency improvements. The timing of OpenAI&apos;s launch amid competition from Gemini and Nano Banana Pro affects user sentiment, highlighting challenges in benchmarking relevance.</description><pubDate>Tue, 16 Dec 2025 05:44:39 GMT</pubDate><category>openai</category><category>gemini</category><category>xiaomi</category><category>lmsys</category><category>deepseek</category><category>openrouter</category><category>gpt-image-1.5</category><category>nano-banana-pro</category><category>mimo-v2-flash</category><category>deepseek-v3.2</category><category>fuli_luo</category><category>eliebakouch</category><category>image-generation</category><category>instruction-following</category><category>benchmarking</category><category>model-efficiency</category><category>long-context</category><category>multi-token-prediction</category><category>hybrid-attention</category><category>model-optimization</category><category>inference-speed</category><category>agentic-workflows</category><category>model-architecture</category><category>model-quantization</category></item><item><title>NVIDIA Nemotron 3: hybrid Mamba-Transformer completely open source models from 30B to 500B</title><link>https://news.smol.ai/issues/25-12-15-nemotron-3/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-15-nemotron-3/</guid><description>**NVIDIA** has released **Nemotron 3 Nano**, a fully open-source hybrid Mamba-Transformer Mixture-of-Experts (MoE) model with a **30B parameter size** and a **1 million token context window**. It includes open weights, training recipes, datasets, and an RL environment suite called NeMo Gym, supporting commercial use under the NVIDIA Open Model License. The model achieves state-of-the-art results on benchmarks like SWE-Bench and Artificial Analysis Intelligence Index, outperforming **Qwen3-30B A3B**. Ecosystem support is immediate with integrations into inference stacks like **vLLM**, **llama.cpp**, and **Baseten**. Upcoming larger models, Nemotron Super and Ultra, will feature NVFP4 pretraining and LatentMoE routing to optimize compute. This release marks a significant milestone for open-source American AI with comprehensive open assets and advanced hybrid architecture.</description><pubDate>Mon, 15 Dec 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>huggingface</category><category>togethercompute</category><category>baseten</category><category>vllm</category><category>llamaindex</category><category>nemotron-3-nano</category><category>qwen3-30b-a3b-base</category><category>ctnzr</category><category>andrew_n_carr</category><category>awnihannun</category><category>hybrid-architecture</category><category>mixture-of-experts</category><category>reinforcement-learning</category><category>long-context</category><category>model-release</category><category>open-source-models</category><category>model-training</category><category>model-optimization</category><category>benchmarking</category><category>agent-training</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-12-not-much/</guid><description>**GPT-5.2** shows mixed performance in public evaluations, excelling in agentic tasks but at a significantly higher cost (~**$620/run**) compared to **Opus 4.5** and **GPT-5.1**. It performs variably on reasoning and coding benchmarks, with some improvements on long-context tasks. Extended &quot;reasoning effort&quot; settings notably impact results. Aggregators rank **Gemini 3 Pro** above GPT-5.2 in task persistence. **OpenAI** released sparse activation models sparking debate on sparsity vs MoE architectures. **Allen AI**&apos;s **Olmo 3.1 (32B)** advances open reinforcement learning scale with substantial compute investment (~**125k H100 hours**). **Mistral**&apos;s Devstral-2 and **llama.cpp** improve local inference infrastructure with new features like GGUF support and distributed speedups. **Tinker** platform goes GA with vision input and finetuning support for **Qwen3-VL-235B**.</description><pubDate>Fri, 12 Dec 2025 05:44:39 GMT</pubDate><category>openai</category><category>allen_ai</category><category>mistral-ai</category><category>ollama</category><category>lmstudio</category><category>thinkymachines</category><category>gpt-5.2</category><category>opus-4.5</category><category>gemini-3-pro</category><category>gpt-5.1</category><category>olmo-3.1-32b</category><category>qwen3-vl-235b</category><category>sama</category><category>scaling01</category><category>akhaliq</category><category>artificialanlys</category><category>lechmazur</category><category>acerfur</category><category>epochairesearch</category><category>reinforcement-learning</category><category>model-benchmarking</category><category>long-context</category><category>model-quantization</category><category>model-optimization</category><category>inference-speed</category><category>sparsity</category><category>fine-tuning</category><category>vision</category></item><item><title>GPT-5.2 (Instant/Thinking/Pro): 74% on GDPVal, 1.4x cost of GPT 5.1, on 10 Year OpenAI Anniversary</title><link>https://news.smol.ai/issues/25-12-11-gpt-52/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-11-gpt-52/</guid><description>**OpenAI** celebrates its 10 year anniversary with the launch of **GPT-5.2**, featuring significant across-the-board improvements including a rare 40% price increase. GPT-5.2 shows strong performance gains in scientific reasoning, knowledge work, and economic value tasks, achieving over **70.9%** human expert parity on **GDPval** tasks and reaching **90.5%** on ARC-AGI-1 with a large efficiency gain. Despite some mixed results in coding benchmarks and vision capabilities, GPT-5.2 is well received as a major update with extended context and tiered reasoning controls. Pricing is set at **$1.75/M input** and **$14/M output** tokens with a 90% cache discount. The update is live in ChatGPT and API, marking a significant milestone for OpenAI&apos;s LLM development.</description><pubDate>Thu, 11 Dec 2025 05:44:39 GMT</pubDate><category>openai</category><category>gpt-5.2</category><category>sama</category><category>yanndubs</category><category>polynoamial</category><category>scaling01</category><category>scientific-reasoning</category><category>knowledge-work</category><category>long-context</category><category>benchmarking</category><category>performance-optimization</category><category>pricing</category><category>software-engineering</category><category>vision</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-10-not-much/</guid><description>**NousResearch&apos;s Nomos 1** is a 30B open math model achieving a top Putnam score with only ~3B active parameters, enabling consumer Mac inference. **AxiomProver** also posts top Putnam results using ThinkyMachines&apos; RL stack. **Mistral&apos;s Devstral 2 Small** outperforms DeepSeek v3.2 in 71% of preferences with better speed and cost. **Anthropic&apos;s Claude Code** introduces asynchronous agent execution. **Cursor 2.2** adds deep agent primitives like Debug and Plan Modes. **VS Code** launches unified agent chat sessions improving multi-agent workflows. **LangChain** releases &quot;Polly&quot; for agent observability. The **Stirrup** harness leads OpenAI GDPval benchmarks with Claude Opus 4.5, GPT-5, and Gemini 3 Pro following. Advances in quantization include **vLLM** integrating Intel&apos;s AutoRound PTQ for efficient serving. **Unsloth** achieves up to 3× training speedups with new kernels across Llama, Qwen, Mistral, and Gemma models. *&quot;Compositional reasoning + specialized post-training under constrained active params can rival frontier closed models on formal math.&quot;*</description><pubDate>Wed, 10 Dec 2025 05:44:39 GMT</pubDate><category>nousresearch</category><category>thinkymachines</category><category>mistral-ai</category><category>deepseek</category><category>anthropic</category><category>cursor</category><category>microsoft</category><category>langchain-ai</category><category>openai</category><category>gemini</category><category>intel</category><category>vllm_project</category><category>danielhanchen</category><category>nomos-1</category><category>axiomprover</category><category>devstral-2-small</category><category>deepseek-v3.2</category><category>claude-code</category><category>cursor-2.2</category><category>claude-opus-4.5</category><category>gpt-5</category><category>claude-sonnet-4.5</category><category>gemini-3-pro</category><category>llama</category><category>qwen</category><category>mistral</category><category>gemma</category><category>math</category><category>formal-reasoning</category><category>agentic-systems</category><category>asynchronous-execution</category><category>multi-agent-systems</category><category>observability</category><category>benchmarking</category><category>quantization</category><category>post-training-quantization</category><category>training-speedup</category><category>kernel-optimization</category><category>inference-efficiency</category></item><item><title>MCP -&gt; Agentic AI Foundation, Mistral Devstral 2</title><link>https://news.smol.ai/issues/25-12-09-devstral2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-09-devstral2/</guid><description>**OpenAI Engineering** sees a significant collaborative milestone with the launch of the **Agentic AI Foundation** under the Linux Foundation, uniting projects from **Anthropic**, **OpenAI**, and **Block**. **Mistral** released **Devstral 2**, a coding model with **123B parameters** and open weights, offering a cost-effective alternative to **Sonnet 4.3** and competitive performance against **DeepSeek v3.2**. The new **Mistral Vibe CLI** supports agentic coding workflows with rapid ecosystem integration. **Alibaba** introduced **Soft Adaptive Policy Optimization (SAPO)** for reinforcement learning tuning, improving stability and performance in **Qwen3-VL** across multiple tasks. Research highlights include the importance of data decontamination in RL and ongoing discussions on MoE RL stability and reward hacking mitigation.</description><pubDate>Tue, 09 Dec 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>block</category><category>mistral-ai</category><category>alibaba</category><category>linux-foundation</category><category>deepseek</category><category>devstral-2</category><category>devstral-small-2</category><category>sonnet-4.3</category><category>deepseek-v3.2</category><category>qwen3-vl</category><category>guillaumelample</category><category>b_roziere</category><category>qtnx_</category><category>charliermarsh</category><category>omarsar0</category><category>eliebakouch</category><category>justinwaugh</category><category>cwolferesearch</category><category>pan</category><category>agentic-ai</category><category>coding-models</category><category>reinforcement-learning</category><category>model-performance</category><category>model-optimization</category><category>open-weights</category><category>cli-tools</category><category>multi-file-code-automation</category><category>data-decontamination</category><category>moe</category><category>reward-models</category><category>rl-stability</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-08-not-much/</guid><description>**Claude Code Skills** gains attention with a published talk and Hugging Face&apos;s new &quot;skill&quot; enabling one-line fine-tuning pipelines for models from ~0.5B to 70B parameters, supporting SFT, DPO, and GRPO, costing as low as ~$0.30 for small runs. **Zhipu AI** launches multimodal models **GLM-4.6V** (106B params MoE) and **GLM-4.6V-Flash** (9B dense), featuring 128k context and native multimodal function calling, with free Flash variant and API pricing detailed. **Jina AI** releases **Jina-VLM (2B)**, a compact multilingual VLM excelling in diagrams and documents with top benchmark scores. At **NeurIPS 2025**, research highlights include Google&apos;s post-Transformer sequence architectures (Moneta, Yaad, Memora) showing up to 20% gains in long-context retrieval, **AxiomProver**&apos;s autonomous Lean system solving 9/12 Putnam 2025 problems rapidly, and mechanistic interpretability advances discussed by Chris Olah emphasizing scalable tooling.</description><pubDate>Mon, 08 Dec 2025 05:44:39 GMT</pubDate><category>hugging-face</category><category>zhipu-ai</category><category>jina-ai</category><category>google-deepmind</category><category>axiomprover</category><category>glm-4.6v</category><category>glm-4.6v-flash</category><category>jina-vlm-2b</category><category>lioronai</category><category>akshay_pachaar</category><category>_akhaliq</category><category>ben_burtenshaw</category><category>vllm_project</category><category>prince_canuma</category><category>zenmuxai</category><category>eliebakouch</category><category>theturingpost</category><category>axiommathai</category><category>neelnanda5</category><category>sarahookr</category><category>fine-tuning</category><category>multimodality</category><category>model-optimization</category><category>long-context</category><category>mechanistic-interpretability</category><category>formal-methods</category><category>sequence-architectures</category><category>reinforcement-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-05-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-05-not-much/</guid><description>**vLLM 0.12.0** introduces DeepSeek support, GPU Model Runner V2, and quantization improvements with PyTorch 2.9.0 and CUDA 12.9. **NVIDIA** launches CUDA Tile IR and cuTile Python for advanced GPU tensor operations targeting Blackwell GPUs. **Hugging Face** releases Transformers v5 RC with an any-to-any multimodal pipeline supporting models like **Gemma3n** and **Qwen3-Omni**. Agent platforms see updates from **LangChain** with content moderation and cost tracking, **Together AI** and **Meta AI** collaborate on RL for long-horizon workflows, and **SonarSource** integrates static analysis into AI codegen. Economic insights from **OpenRouter** highlight coding as a key AI application, with reasoning models surpassing 50% usage and market bifurcation between premium and open models. Additionally, **Kling Video 2.6** debuts native audio capabilities, and **Runway Gen-4.5**, **Qwen3-TTS**, and **Gemini 3 Pro** advance multimodality.</description><pubDate>Fri, 05 Dec 2025 05:44:39 GMT</pubDate><category>vllm</category><category>nvidia</category><category>huggingface</category><category>langchain-ai</category><category>together-ai</category><category>meta-ai-fair</category><category>sonarsource</category><category>openrouter</category><category>runway</category><category>gemini</category><category>arena</category><category>vllm-0.12.0</category><category>gemma3n</category><category>qwen3-omni</category><category>qwen3-vl</category><category>gpt-5.1-codex-max</category><category>gemini-3-pro</category><category>runway-gen-4.5</category><category>kling-video-2.6</category><category>jeremyphoward</category><category>mervenoyann</category><category>sydneyrunkle</category><category>swyx</category><category>maximelabonne</category><category>gpu-programming</category><category>quantization</category><category>multimodality</category><category>agent-platforms</category><category>reinforcement-learning</category><category>static-analysis</category><category>reasoning</category><category>inference-infrastructure</category><category>model-optimization</category><category>economics</category><category>audio</category><category>video-generation</category></item><item><title>OpenRouter&apos;s State of AI - An Empirical 100 Trillion Token Study</title><link>https://news.smol.ai/issues/25-12-04-openrouter/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-04-openrouter/</guid><description>**OpenRouter** released its first survey showing usage trends with 7 trillion tokens proxied weekly, highlighting a 52% roleplay bias. **Deepseek**&apos;s open model market share has sharply declined due to rising coding model usage. Reasoning model token usage surged from 0% to over 50%. **Grok Code Fast** shows high usage, while **Anthropic** leads in tool calling and coding requests with around 60% share. Input tokens quadrupled and output tokens tripled this year, driven mainly by programming use cases, which dominate spending and volume. Google launched **Gemini 3 Deep Think**, featuring parallel thinking and achieving 45.1% on ARC-AGI-2 benchmarks, and previewed **Titans**, a long-context neural memory architecture scaling beyond 2 million tokens. These advances were shared by **Google DeepMind** and **Google AI** on Twitter.</description><pubDate>Thu, 04 Dec 2025 05:44:39 GMT</pubDate><category>openrouter</category><category>deepseek</category><category>anthropic</category><category>google</category><category>google-deepmind</category><category>grok-code-fast</category><category>gemini-3</category><category>gemini-3-deep-think</category><category>gpt-5.1-codex-max</category><category>quocleix</category><category>noamshazeer</category><category>mirrokni</category><category>reasoning</category><category>coding</category><category>tokenization</category><category>long-context</category><category>model-architecture</category><category>benchmarking</category><category>agentic-ai</category><category>prompt-engineering</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-12-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-03-not-much/</guid><description>**OpenAI&apos;s Code Red response** and **Anthropic&apos;s IPO** are major highlights. In AI video and imaging, **Kling 2.6** introduces native audio co-generation with coherent lip-sync, partnered with platforms like **ElevenLabs** and **OpenArt**. **Runway Gen-4.5** enhances lighting fidelity, while **Google&apos;s Gemini 3 Nano Banana Pro** supports advanced image compositing. Open model releases include **DeepSeek V3.2** with sparse attention and cost-effective pricing, and **Mistral&apos;s Ministral 3** multimodal family with strong 14B variants. Retrieval and code models from **Alibaba&apos;s EvoQwen2.5-VL** and **Nous Research&apos;s Hermes 4.3** show competitive performance with permissive licensing and HF availability. The community arena sees additions like INTELLECT-3 (106B MoE). *&quot;coherent looking &amp; sounding output&quot;* and *&quot;auto-lighting to match scene mood&quot;* are noted advancements.</description><pubDate>Wed, 03 Dec 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google</category><category>runway</category><category>elevenlabs</category><category>freepik</category><category>openart</category><category>deepseek</category><category>mistral-ai</category><category>alibaba</category><category>nous-research</category><category>kling-2.6</category><category>kling-o1</category><category>runway-gen-4.5</category><category>gemini-3</category><category>deepseek-v3.2</category><category>ministral-3</category><category>evoqwen2.5-vl</category><category>hermes-4.3</category><category>intellect-3</category><category>video-generation</category><category>audio-processing</category><category>multimodality</category><category>image-generation</category><category>reasoning</category><category>model-quantization</category><category>sparse-attention</category><category>model-pricing</category><category>multimodal-models</category><category>retrieval-augmentation</category><category>model-training</category><category>model-release</category></item><item><title>DeepSeek V3.2 &amp; 3.2-Speciale: GPT5-High Open Weights, Context Management, Plans for Compute Scaling</title><link>https://news.smol.ai/issues/25-12-01-deepseek-32/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-01-deepseek-32/</guid><description>**DeepSeek** launched the **DeepSeek V3.2** family including Standard, Thinking, and Speciale variants with up to **131K context window** and competitive benchmarks against **GPT-5-High**, **Sonnet 4.5**, and **Gemini 3 Pro**. The release features a novel **Large Scale Agentic Task Synthesis Pipeline** focusing on agentic behaviors and improvements in **reinforcement learning** post-training algorithms. The models are available on platforms like **LM Arena** with pricing around **$0.28/$0.42 per million tokens**. Community feedback is mixed, praising the frontier reasoning capabilities but critiquing the chat UI experience. Key figures include **Susan Zhang** and **Teortaxes** who provided commentary on the release.</description><pubDate>Tue, 02 Dec 2025 05:44:39 GMT</pubDate><category>deepseek_ai</category><category>lm-arena</category><category>deepseek-v3.2</category><category>deepseek-v3.2-speciale</category><category>gpt-5-high</category><category>sonnet-4.5</category><category>gemini-3-pro</category><category>suchenzang</category><category>teortaxestex</category><category>agentic-ai</category><category>reinforcement-learning</category><category>large-context-windows</category><category>model-benchmarking</category><category>model-performance</category><category>multi-agent-systems</category><category>model-training</category><category>model-deployment</category></item><item><title>Mistral 3: Mistral Large 3 + Ministral 3B/8B/14B open weights models</title><link>https://news.smol.ai/issues/25-12-02-mistral-3/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-12-02-mistral-3/</guid><description>**Mistral** has launched the **Mistral 3 family** including **Ministral 3** models (3B/8B/14B) and **Mistral Large 3**, a sparse MoE model with **675B total parameters** and **256k context window**, all under an Apache 2.0 open license. Early benchmarks rank Mistral Large 3 at **#6 among open models** with strong coding performance. The launch includes broad ecosystem support such as vLLM, llama.cpp, Ollama, and LM Studio integrations. Meanwhile, **Anthropic** acquired the open-source **Bun** runtime to accelerate **Claude Code**, which reportedly reached a **$1B run-rate in ~6 months**. Anthropic also announced discounted **Claude** plans for nonprofits and shared insights on AI&apos;s impact on work internally.</description><pubDate>Tue, 02 Dec 2025 05:44:39 GMT</pubDate><category>mistral-ai</category><category>anthropic</category><category>apple</category><category>runway</category><category>moondream</category><category>mistral-large-3</category><category>ministral-3</category><category>clara-7b-instruct</category><category>gen-4.5</category><category>claude-code</category><category>anjney_midha</category><category>_akhaliq</category><category>alexalbert__</category><category>_catwu</category><category>mikeyk</category><category>sparse-moe</category><category>multimodality</category><category>benchmarking</category><category>open-source</category><category>model-licensing</category><category>model-performance</category><category>long-context</category><category>inference-optimization</category><category>instruction-following</category><category>local-inference</category><category>code-generation</category><category>model-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-26-not-much/</guid><description>**Anthropic** introduces durable agents and MCP tasks for long-running workflows, with practical engineering patterns and integrations like Prefect. **Booking.com** deploys a large-scale agent system improving customer satisfaction using LangGraph, Kubernetes, GPT-4 Mini, and Weaviate. **Perplexity** rolls out user-level memory and virtual try-on features. **Claude Opus 4.5** leads on LisanBench and Code Arena WebDev benchmarks with mixed community feedback on its &quot;thinking&quot; and &quot;non-thinking&quot; modes, while improving cost-efficiency and UX with batch APIs and context compaction. Research on multi-agent systems shows **LatentMAS** reduces communication tokens by 70-84% and improves accuracy using Qwen3 models, and reasoning trace distillation achieves significant token reduction with maintained accuracy, highlighting the importance of reasoning trace style.</description><pubDate>Wed, 26 Nov 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>booking.com</category><category>perplexity-ai</category><category>langchain</category><category>claude</category><category>scaling01</category><category>deepseek</category><category>qwen</category><category>prefect</category><category>claude-opus-4.5</category><category>qwen-3-4b</category><category>qwen-3-8b</category><category>qwen-3-14b</category><category>deepseek-r1</category><category>jeremyphoward</category><category>alexalbert__</category><category>omarsar0</category><category>lingyang_pu</category><category>dair_ai</category><category>agent-systems</category><category>multi-agent-systems</category><category>reasoning</category><category>benchmarking</category><category>cost-efficiency</category><category>model-optimization</category><category>long-context</category><category>memory-management</category><category>reinforcement-learning</category><category>model-performance</category><category>multi-agent-communication</category><category>latent-representation</category><category>inference-cost</category><category>software-integration</category></item><item><title>Black Forest Labs FLUX.2 [pro|flex|dev|klein]: near-Nano Banana quality but Open Weights</title><link>https://news.smol.ai/issues/25-11-25-flux2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-25-flux2/</guid><description>**Black Forest Labs&apos; FLUX.2** release features **Multi-Reference Support** for up to **4 Megapixel** output and up to **10 images** with consistency, including four form factors: Pro, Flex, Dev (32B Open Weight model), and Klein (TBA Open Weights). The new **FLUX.2 - VAE** introduces a variational autoencoder optimizing learnability, quality, and compression. Meanwhile, **Anthropic&apos;s Claude Opus 4.5** demonstrates strong performance and efficiency, scoring **70 on Artificial Analysis**, tying with **GPT-5.1 high** and trailing **Gemini 3 Pro (73)**. Opus 4.5 excels in agentic coding benchmarks and research evaluations, with notable token efficiency and reduced running costs. *&quot;Opus 4.5 leads Gemini 3 Pro on SWE-Bench Verified and tops the AICodeKing leaderboard,&quot;* and it shows strong QA and systematic review capabilities. Anthropic also released a dense prompting guide for Opus 4.5.</description><pubDate>Tue, 25 Nov 2025 05:44:39 GMT</pubDate><category>black-forest-labs</category><category>anthropic</category><category>huggingface</category><category>flux-2</category><category>flux-2-dev</category><category>claude-opus-4.5</category><category>gpt-5.1</category><category>gemini-3-pro</category><category>multi-reference-support</category><category>variational-autoencoder</category><category>image-generation</category><category>open-weights</category><category>agentic-coding</category><category>token-efficiency</category><category>benchmarking</category><category>prompting</category><category>model-performance</category></item><item><title>Claude Opus 4.5: 3rd new SOTA coding model in past week, 1/3 the price of Opus </title><link>https://news.smol.ai/issues/25-11-24-opus-45/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-24-opus-45/</guid><description>**Anthropic** launched **Claude Opus 4.5**, a new flagship model excelling in **coding, agents, and tooling** with a significant **3x price cut** compared to Opus 4.1 and improved **token efficiency** using **76% fewer output tokens**. Opus 4.5 achieved a new **SOTA** on **SWE-bench Verified** with **80.9% accuracy**, surpassing previous models like **Gemini 3 Pro** and **GPT-5.1-Codex-Max**. The update includes advanced API features such as **effort control**, **context compaction**, and **programmatic tool calling**, improving tool accuracy and reducing token usage. Claude Code is now bundled with Claude Desktop, and new integrations like Claude for Chrome and Excel are rolling out. Benchmarks show Opus 4.5 breaking the 80% barrier on SWE-bench Verified and strong performance on ARC-AGI-2 and BrowseComp-Plus.</description><pubDate>Mon, 24 Nov 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>amazon</category><category>google</category><category>anthropic</category><category>claude-opus-4.5</category><category>gemini-3-pro</category><category>gpt-5.1-codex-max</category><category>opus-4.1</category><category>sonnet-4.5</category><category>alexalbert__</category><category>btibor91</category><category>scaling01</category><category>klieret</category><category>coding</category><category>agents</category><category>tool-use</category><category>token-efficiency</category><category>benchmarking</category><category>api</category><category>model-pricing</category><category>model-performance</category><category>effort-control</category><category>context-compaction</category><category>programmatic-tool-calling</category></item><item><title>AI Engineer Code Summit</title><link>https://news.smol.ai/issues/25-11-21-aie-code/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-21-aie-code/</guid><description>The recent **AIE Code Summit** showcased key developments including **Google DeepMind&apos;s Gemini 3 Pro Image model, Nano Banana Pro**, which features enhanced text rendering, 4K visuals, and fine-grained editing capabilities. Community feedback highlights its strong performance in design and visualization tasks, with high user preference scores. Benchmarking updates reveal the new **CritPt physics frontier benchmark** where Gemini 3 Pro outperforms GPT-5, though AI still lags on complex unseen research problems. Agentic task evaluations show varied time horizons and performance gaps between open-weight and closed frontier models, emphasizing ongoing challenges in AI research and deployment. *&quot;Instruction following remains jagged for some users,&quot;* and model fit varies by use case, with Gemini 3 excelling in UI and code tasks but showing regressions in transcription and writing fidelity.</description><pubDate>Fri, 21 Nov 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>togethercompute</category><category>gemini-3-pro-image</category><category>gemini-3</category><category>gpt-5</category><category>claude-3.7-sonnet</category><category>demishassabis</category><category>omarsar0</category><category>lintool</category><category>hrishioa</category><category>teknium</category><category>artificialanlys</category><category>minyangtian1</category><category>ofirpress</category><category>metr_evals</category><category>scaling01</category><category>image-generation</category><category>fine-tuning</category><category>benchmarking</category><category>agentic-ai</category><category>physics</category><category>model-performance</category><category>instruction-following</category><category>model-comparison</category><category>time-horizon</category><category>user-preference</category></item><item><title>Nano Banana Pro (Gemini Image Pro) solves text-in-images, infographic generation, 2-4k resolution, and Google Search grounding</title><link>https://news.smol.ai/issues/25-11-20-nano-banana-pro/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-20-nano-banana-pro/</guid><description>**Google** launched **Gemini 3 Pro Image (Nano Banana Pro)**, a next-generation AI image generation and editing model with integrated Google Search grounding, multi-image composition, and fine-grained visual controls, offering pricing at $0.134 per 2K image and $0.24 per 4K image. It features improved text rendering with error rates dropping from 56% to 8% compared to its predecessor, and includes SynthID watermark checks for provenance. The model is available via Gemini App, API, LM Arena, Hugging Face Spaces, Together AI, and Flow. Meanwhile, **OpenAI** shared early experiments with **GPT-5** accelerating scientific research, including proofs of previously unsolved problems in math, physics, biology, and materials science. *&quot;GPT-5 accelerated research tasks in math/physics/biology/materials; in 4, it helped find proofs of previously unsolved problems.&quot;*</description><pubDate>Thu, 20 Nov 2025 05:44:39 GMT</pubDate><category>google</category><category>openai</category><category>hugging-face</category><category>togethercompute</category><category>lmsys</category><category>gemini-3-pro</category><category>gpt-5</category><category>jeffdean</category><category>kevinweil</category><category>demishassabis</category><category>image-generation</category><category>text-rendering</category><category>model-provenance</category><category>scientific-research</category><category>proof-assistance</category><category>multimodal-integration</category><category>api-access</category><category>fine-tuning</category></item><item><title>OpenAI fires back: GPT-5.1-Codex-Max (API) and GPT 5.1 Pro (ChatGPT)</title><link>https://news.smol.ai/issues/25-11-19-gpt-51-codex-max-pro/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-19-gpt-51-codex-max-pro/</guid><description>**OpenAI** released **GPT-5.1-Codex-Max**, featuring compaction-native training, an &quot;Extra High&quot; reasoning mode, and claims of over 24-hour autonomous operation, showing significant performance gains on benchmarks like METR, CTF, and PaperBench. **Google&apos;s Gemini 3 Pro** demonstrates strong coding and reasoning capabilities, achieving new state-of-the-art results on SWE-bench Verified and WeirdML, with estimated model size between 5-10 trillion parameters. The AI coding agent ecosystem is rapidly evolving with integrations and tooling improvements from multiple companies. **Sam Altman** highlighted the significant improvements in GPT-5.1-Codex-Max. The news also covers educational offerings like ChatGPT for Teachers and multi-agent workflows involving Gemini 3, GPT-5.1-Codex-Max, and Claude Sonnet 4.5.</description><pubDate>Wed, 19 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>anthropic</category><category>langchain-ai</category><category>gpt-5.1-codex-max</category><category>gpt-5.1-codex</category><category>gemini-3-pro</category><category>claude-3.5-sonnet</category><category>sama</category><category>coding</category><category>autonomous-systems</category><category>benchmarking</category><category>model-scaling</category><category>multi-agent-systems</category><category>model-performance</category><category>reasoning</category><category>model-architecture</category></item><item><title>Gemini 3 Pro — new GDM frontier model 6, Gemini 3 Deep Think, and Antigravity IDE</title><link>https://news.smol.ai/issues/25-11-18-gemini-3/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-18-gemini-3/</guid><description>**Google** launched **Gemini 3 Pro**, a state-of-the-art model with a **1M-token context window**, **multimodal reasoning**, and strong agentic capabilities, priced significantly higher than Gemini 2.5. It leads major benchmarks, surpassing **Grok 4.1** and competing closely with **Sonnet 4.5** and **GPT-5.1**, though GPT-5.1 excels in ultralong summarization. Independent evaluations from **Artificial Analysis**, **Vending Bench**, **ARC-AGI 2**, **Box**, and **PelicanBench** validate Gemini 3 as a frontier LLM. Google also introduced **Antigravity**, an agentic IDE powered by Gemini 3 Pro and other models, featuring task orchestration and human-in-the-loop validation. The launch marks Google&apos;s strong return to AI with more models expected soon. *&quot;Google is very, very back in the business.&quot;*</description><pubDate>Tue, 18 Nov 2025 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>gemini-3-pro</category><category>gemini-2.5</category><category>grok-4.1</category><category>sonnet-4.5</category><category>gpt-5.1</category><category>sundarpichai</category><category>_philschmid</category><category>oriol_vinyals</category><category>multimodality</category><category>agentic-ai</category><category>benchmarking</category><category>context-window</category><category>model-performance</category><category>instruction-following</category><category>model-pricing</category><category>api</category><category>model-release</category><category>reasoning</category><category>model-evaluation</category></item><item><title>xAI Grok 4.1: #1 in Text Arena, #1 in EQ-bench, and better Creative Writing</title><link>https://news.smol.ai/issues/25-11-17-grok-41/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-17-grok-41/</guid><description>**xAI** launched **Grok 4.1**, achieving a #1 rank on the LM Arena Text Leaderboard with an Elo score of **1483**, showing improvements in creative writing and anti-hallucination. **OpenAI&apos;s GPT-5.1 &quot;Thinking&quot;** demonstrates efficiency gains with ~60% less &quot;thinking&quot; on easy queries and strong ARC-AGI performance. **Google DeepMind** released **WeatherNext 2**, an ensemble generative model that is **8× faster** and more accurate for global weather forecasts, integrated into multiple Google products. **Sakana AI** raised **¥20B ($135M)** in Series B funding at a **$2.63B** valuation to focus on efficient AI for resource-constrained enterprise applications in Japan. New evaluations highlight tradeoffs between hallucination and knowledge accuracy across models including **Claude 4.1 Opus** and **Anthropic** models.</description><pubDate>Mon, 17 Nov 2025 05:44:39 GMT</pubDate><category>xai</category><category>openai</category><category>google-deepmind</category><category>sakana-ai</category><category>anthropic</category><category>microsoft</category><category>mufg</category><category>khosla</category><category>nea</category><category>lux-capital</category><category>iqt</category><category>grok-4.1</category><category>gpt-5.1</category><category>claude-4.1-opus</category><category>grok-4</category><category>gpt-5</category><category>grok-4.1-thinking</category><category>gpt-5-pro</category><category>claude-4.5-haiku</category><category>yanndubs</category><category>gregkamradt</category><category>philschmid</category><category>willccbb</category><category>model-performance</category><category>creative-writing</category><category>hallucination</category><category>evaluation-datasets</category><category>ensemble-models</category><category>weather-forecasting</category><category>funding</category><category>efficiency</category><category>anti-hallucination</category><category>arc-agi</category><category>model-scaling</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-14-not-much/</guid><description>**OpenAI** launched **GPT-5.1** featuring &quot;adaptive reasoning&quot; and developer-focused API improvements, including prompt caching and a reasoning_effort toggle for latency/cost tradeoffs. Independent analysis shows a minor intelligence bump with significant gains in agentic coding benchmarks. **Anthropic**&apos;s **Claude** models introduced structured outputs with JSON schema compliance in public beta for Sonnet 4.5 and Opus 4.1, enhancing tooling and code execution workflows. Rumors of an Opus 4.5 release were debunked. **LangChain** released a &quot;Deep Agents&quot; package and context-engineering playbook to optimize agent workflows. The community is eagerly anticipating **Google DeepMind**&apos;s **Gemini 3** model, hinted at in social media and upcoming AIE CODE events. *&quot;Tickets are sold out, but side events and volunteering opportunities are available.&quot;*</description><pubDate>Fri, 14 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>langchain-ai</category><category>google-deepmind</category><category>gpt-5.1</category><category>sonnet-4.5</category><category>opus-4.1</category><category>gemini-3</category><category>swyx</category><category>allisontam_</category><category>gdb</category><category>sama</category><category>alexalbert__</category><category>simonw</category><category>omarsar0</category><category>abacaj</category><category>scaling01</category><category>amandaaskell</category><category>adaptive-reasoning</category><category>developer-tools</category><category>prompt-optimization</category><category>json-schema</category><category>agent-workflows</category><category>context-engineering</category><category>structured-outputs</category><category>model-release</category><category>benchmarking</category></item><item><title>minor updates to GPT 5.1 and SIMA 2</title><link>https://news.smol.ai/issues/25-11-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-13-not-much/</guid><description>**OpenAI** released **GPT-5.1** family models including **5.1-Codex** and **5.1-Codex-Mini** with improved steerability, faster responses, and new tools like apply_patch and shell command execution. Pricing remains unchanged from 5.0. Immediate integrations include **GitHub Copilot**, **VS Code**, **Cursor**, and **Perplexity** adopting GPT-5.1 models. **Google DeepMind** announced **SIMA 2**, a **Gemini**-powered agent capable of language instruction following, planning, and self-improvement without human feedback, targeting robotics applications. New research on context engineering and agentic tool use patterns was published, with contributions from **Weaviate** and **LlamaIndex** on database query planning and chart parsing respectively. *&quot;Adaptive reasoning&quot;* and agentic coding improvements are highlighted in GPT-5.1- Instant.</description><pubDate>Thu, 13 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>github</category><category>microsoft</category><category>cursor_ai</category><category>perplexity-ai</category><category>weaviate</category><category>llamaindex</category><category>gpt-5.1</category><category>gpt-5.1-codex</category><category>gpt-5.1-codex-mini</category><category>sima-2</category><category>gemini</category><category>sama</category><category>allisontam_</category><category>cline</category><category>cognition</category><category>demishassabis</category><category>omarsar0</category><category>helloiamleonie</category><category>adaptive-reasoning</category><category>agentic-coding</category><category>tool-use</category><category>context-engineering</category><category>memory-architecture</category><category>self-improvement</category><category>retrieval-augmentation</category><category>database-query-planning</category><category>chart-parsing</category><category>robotics</category></item><item><title>GPT 5.1 in ChatGPT: No evals, but adaptive thinking and instruction following</title><link>https://news.smol.ai/issues/25-11-12-gpt-51/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-12-gpt-51/</guid><description>**OpenAI** launched **GPT-5.1** with improvements in conversational tone, instruction following, and adaptive reasoning. **GPT-5.0** is being sunset in 3 months. ChatGPT introduces new tone toggles for personalization, serving over **800 million users**. **Waymo** rolls out freeway driving for public riders in major California cities, showcasing advances in autonomous driving. **Anthropic**&apos;s Project Fetch explores LLMs as robotics copilots using **Claude**. **Perceptron** releases a new API and Python SDK for multimodal perception-action apps supporting **Isaac-0.1** and **Qwen3VL-235B**. **Code Arena** offers live coding evaluations supporting **Claude**, **GPT-5**, **GLM-4.6**, and **Gemini**. **LangChain** introduces middleware for agent governance with human-in-the-loop controls. **LlamaIndex** releases a structured extraction template for SEC filings using LlamaAgents. **NousResearch** promotes ARC Prize benchmarks for generalized intelligence evaluation.</description><pubDate>Wed, 12 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>waymo</category><category>perceptron</category><category>langchain</category><category>llamaindex</category><category>nousresearch</category><category>gpt-5.1</category><category>gpt-5.0</category><category>claude</category><category>isaac-0.1</category><category>qwen3vl-235b</category><category>glm-4.6</category><category>gemini</category><category>dmitri_dolgov</category><category>jeffdean</category><category>fidji_simo</category><category>akshats07</category><category>adaptive-reasoning</category><category>instruction-following</category><category>personalization</category><category>autonomous-driving</category><category>robotics</category><category>multimodality</category><category>agent-evaluation</category><category>agent-governance</category><category>middleware</category><category>structured-extraction</category><category>benchmarking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-11-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-11-not-much/</guid><description>**GPT-5** leads Sudoku-Bench solving 33% of puzzles but 67% remain unsolved, highlighting challenges in meta-reasoning and spatial logic. New training methods like **GRPO fine-tuning** and &quot;Thought Cloning&quot; show limited success. Research on &quot;looped LLMs&quot; suggests pretrained models benefit from repeated computation for better performance. **Baidu&apos;s ERNIE-4.5-VL-28B-A3B-Thinking** offers lightweight multimodal reasoning with Apache 2.0 licensing, outperforming **Gemini-2.5-Pro** and **GPT-5-High** on document tasks. **Databricks ai_parse_document** preview delivers cost-efficient document intelligence outperforming GPT-5 and Claude. **Pathwork AI** uses **LlamaCloud** for underwriting automation. **Gemini File Search API** enables agentic retrieval augmented generation (RAG) with MCP server integration. **Together AI** and **Collinear** launch **TraitMix** for persona-driven agent simulations integrated with **Together Evals**. Reports highlight risks in long-running code agents like **Claude Code** reverting changes, emphasizing guardrails. Community consensus favors multiple code copilots including Claude Code, Codex, and others.</description><pubDate>Tue, 11 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>baidu</category><category>databricks</category><category>llamaindex</category><category>togethercompute</category><category>sakanaailabs</category><category>gpt-5</category><category>qwen2.5-7b</category><category>ernie-4.5-vl-28b-a3b-thinking</category><category>gemini-2.5-pro</category><category>llamacloud</category><category>claude-code</category><category>sakanaailabs</category><category>micahgoldblum</category><category>francoisfleuret</category><category>matei_zaharia</category><category>jerryjliu0</category><category>omarsar0</category><category>togethercompute</category><category>imjaredz</category><category>theo</category><category>reasoning-benchmarks</category><category>reinforcement-learning</category><category>fine-tuning</category><category>multimodality</category><category>document-intelligence</category><category>retrieval-augmented-generation</category><category>agentic-systems</category><category>persona-simulation</category><category>code-agents</category><category>guardrails</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-10-not-much/</guid><description>**Moonshot AI&apos;s Kimi K2 Thinking** AMA revealed a hybrid attention stack using **KDA + NoPE MLA** outperforming full MLA + RoPE, with the **Muon optimizer** scaling to ~1T parameters and native **INT4** QAT for cost-efficient inference. K2 Thinking ranks highly on **LisanBench** and **LM Arena Text** leaderboards, offering low-cost INT4 serving and strong performance in Math, Coding, and Creative Writing. It supports heavy agentic tool use with up to 300 tool requests per run and recommends using the official API for reliable long-trace inference. **Meta AI** released the **Omnilingual ASR** suite covering 1600+ languages including 500 underserved, plus a 7B wav2vec 2.0 model and ASR corpus. Additionally, the **Gelato-30B-A3B** model for computer grounding in GUI manipulation agents outperforms larger VLMs, targeting immediate agent gains. Qwen&apos;s image-edit LoRAs and light-restoration app were also highlighted.</description><pubDate>Mon, 10 Nov 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>meta-ai-fair</category><category>togethercompute</category><category>qwen</category><category>kimi-k2-thinking</category><category>kimi-k3</category><category>gelato-30b-a3b</category><category>omnilingual-wav2vec-2.0</category><category>yuchenj_uw</category><category>scaling01</category><category>code_star</category><category>omarsar0</category><category>kimi_moonshot</category><category>anas_awadalla</category><category>akhaliq</category><category>minchoi</category><category>attention-mechanisms</category><category>quantization</category><category>fine-tuning</category><category>model-optimization</category><category>agentic-ai</category><category>speech-recognition</category><category>multilingual-models</category><category>gui-manipulation</category><category>image-editing</category><category>dataset-release</category></item><item><title>Terminal-Bench 2.0 and Harbor</title><link>https://news.smol.ai/issues/25-11-07-tbench2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-07-tbench2/</guid><description>**Terminal-Bench** has fixed task issues and launched version 2.0 with cloud container support via the **Harbor framework**, gaining recognition from models like **Claude 4.5** and **Kimi K2 Thinking**. **Moonshot AI&apos;s Kimi K2 Thinking** is a 1 trillion parameter MoE reasoning model with ~32B active parameters, running natively in **INT4 quantization** and featuring a 256K context window. It leads open-weights benchmarks with an Artificial Analysis Intelligence Index score of **67** and strong agentic performance, running efficiently on consumer Apple silicon and 2× M3 Ultra hardware. The model is broadly available on **Hugging Face**, **Ollama Cloud**, and integrated into frameworks like slime. Serving bottlenecks were traced to network bandwidth rather than GPU limits, highlighting infrastructure considerations for LLM deployment.</description><pubDate>Fri, 07 Nov 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>anthropic</category><category>hugging-face</category><category>ollama</category><category>slime-framework</category><category>kimi-k2-thinking</category><category>clementdelangue</category><category>dbreunig</category><category>awnihannun</category><category>crystalsssup</category><category>kimi_moonshot</category><category>benchmarking</category><category>agentic-ai</category><category>quantization</category><category>model-optimization</category><category>inference</category><category>model-deployment</category><category>moe</category><category>context-windows</category><category>cost-efficiency</category></item><item><title>Kimi K2 Thinking: 1T-A32B params, SOTA HLE, BrowseComp, TauBench &amp;&amp; Soumith leaves Pytorch</title><link>https://news.smol.ai/issues/25-11-06-kimi-k2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-06-kimi-k2/</guid><description>**Moonshot AI** launched **Kimi K2 Thinking**, a **1 trillion parameter** mixture-of-experts (MoE) model with **32 billion active experts**, a **256K context window**, and native **INT4 quantization-aware training**. It achieves state-of-the-art results on benchmarks like **HLE (44.9%)**, **BrowseComp (60.2%)**, and agentic tool use with **200-300 sequential tool calls**. The model is deployed with **vLLM** support and OpenAI-compatible APIs, available on platforms like Arena, Baseten, and Yupp. Early user reports note some API instability under launch load. Meanwhile, **Google** announced the **TPU v7 (Ironwood)** with a **10× peak performance improvement** over TPU v5p, aimed at training and agentic inference for models like **Gemini**. **Apple** added support for M5 Neural Accelerators in llama.cpp for inference acceleration.</description><pubDate>Thu, 06 Nov 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>google</category><category>apple</category><category>vllm_project</category><category>arena</category><category>baseten</category><category>yupp_ai</category><category>kimi-k2-thinking</category><category>gemini</category><category>eliebakouch</category><category>nrehiew_</category><category>andrew_n_carr</category><category>ofirpress</category><category>artificialanlys</category><category>sundarpichai</category><category>akhaliq</category><category>mixture-of-experts</category><category>quantization</category><category>int4</category><category>context-window</category><category>agentic-ai</category><category>benchmarking</category><category>model-deployment</category><category>inference-acceleration</category><category>api</category><category>performance-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-05-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-05-not-much/</guid><description>**Kimi-K2 Reasoner** has been integrated into **vLLM** and will soon be supported by **SGLang**, featuring a massive **1.2 trillion parameter MoE** configuration. **Perplexity AI** released research on cloud-portable trillion-parameter MoE kernels optimized for **AWS EFA**, with potential integration into **vLLM**. **IBM&apos;s vLLM** team formalized hybrid dense and sparse expert models, supporting models like **Qwen3-Next**, **Nemotron Nano 2**, and **Granite 4.0**. **Kimi-K2** reportedly scores **77% on GPQA Diamond**, outperforming **GPT-4.5** at 71.4%, though this is unverified. 

**Anthropic** published a guide on efficient tool-heavy agent systems using MCP patterns, drastically reducing context tokens by ~98.7%. **Graphiti MCP** demonstrated shared memory across apps like **Claude Desktop** and **Cursor** for persistent agent memory. **VS Code** introduced an &quot;Agent sessions&quot; feature to unify agent management, including **Copilot** and **Codex**. **Cursor AI** improved coding accuracy via semantic search and code retrieval embeddings. New evaluation frameworks like **CodeClash** and **LMArena** assess agent and coding model performance in realistic multi-round tasks and occupation-tagged leaderboards.</description><pubDate>Wed, 05 Nov 2025 05:44:39 GMT</pubDate><category>vllm</category><category>perplexity-ai</category><category>ibm</category><category>anthropic</category><category>graphiti</category><category>claude</category><category>cursor-ai</category><category>microsoft</category><category>kimi-k2</category><category>qwen3-next</category><category>nemotron-nano-2</category><category>granite-4.0</category><category>gpt-4.5</category><category>copilot</category><category>codex</category><category>scaling01</category><category>cedric_chee</category><category>aravsrinivas</category><category>omarsar0</category><category>_avichawla</category><category>pierceboggan</category><category>jo_parkhurst</category><category>jyangballin</category><category>ofirpress</category><category>ml_angelopoulos</category><category>mixture-of-experts</category><category>model-integration</category><category>cloud-computing</category><category>hybrid-models</category><category>benchmarking</category><category>agent-systems</category><category>memory-persistence</category><category>semantic-search</category><category>code-retrieval</category><category>context-length-optimization</category><category>tool-use</category><category>evaluation-frameworks</category><category>software-development</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-04-not-much/</guid><description>**Google&apos;s Project Suncatcher** prototypes scalable ML compute systems in orbit using solar energy with Trillium-generation TPUs surviving radiation, aiming for prototype satellites by 2027. **China&apos;s 50% electricity subsidies** for datacenters may offset chip efficiency gaps, with **Huawei** planning gigawatt-scale SuperPoDs for DeepSeek by 2027. **Epoch** launched an open data center tracking hub, and **Deutsche Telekom** and **NVIDIA** announced a $1.1B Munich facility with 10k GPUs. In agent stacks, **MCP** (Model-Compute-Platform) tools gain traction with implementations like **LitServe**, **Claude Desktop**, and **Reka&apos;s MCP server** for VS Code. Anthropic emphasizes efficient code execution with MCP. Context engineering shifts focus from prompt writing to model input prioritization, with reports and tools from **Weaviate**, **Anthropic**, and practitioners highlighting instruction-following rerankers and embedding approaches. DeepMind&apos;s **IMO-Bench** math reasoning suite shows **Gemini DeepThink** achieving high scores, with a ProofAutoGrader correlating strongly with human grading. Benchmarks and governance updates include new tasks and eval sharing in lighteval.</description><pubDate>Tue, 04 Nov 2025 05:44:39 GMT</pubDate><category>google</category><category>huawei</category><category>epoch-ai</category><category>deutsche-telekom</category><category>nvidia</category><category>anthropic</category><category>reka-ai</category><category>weaviate</category><category>deepmind</category><category>trillium</category><category>gemini-2.5-pro</category><category>gemini-deepthink</category><category>sundarpichai</category><category>yuchenj_uw</category><category>teortaxestex</category><category>epochairesearch</category><category>scaling01</category><category>_avichawla</category><category>rekaailabs</category><category>anthropicai</category><category>douwekiela</category><category>omarsar0</category><category>nityeshaga</category><category>goodside</category><category>iscienceluvr</category><category>lmthang</category><category>energy-efficiency</category><category>datacenters</category><category>mcp</category><category>context-engineering</category><category>instruction-following</category><category>embedding-models</category><category>math-reasoning</category><category>benchmarking</category><category>code-execution</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-11-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-11-03-not-much/</guid><description>**OpenAI** and **AWS** announced a strategic partnership involving a $38B compute deal to deploy hundreds of thousands of NVIDIA GB200 and GB300 chips, while **Microsoft** secured a license to ship NVIDIA GPUs to the UAE with a planned $7.9B datacenter investment. A 3-month NVFP4 kernel optimization competition on Blackwell B200s was launched by **NVIDIA** and GPU_MODE with prizes including DGX Spark and RTX 50XX GPUs. **vLLM** gains traction for local LLM serving, exemplified by PewDiePie&apos;s adoption. **Alibaba** previewed the Qwen3-Max-Thinking model hitting 100% on AIME 2025 and HMMT benchmarks, signaling advances in reasoning with tool use. The MIT-licensed MiniMax-M2 230B MoE model topped the Arena WebDev leaderboard, tying with Claude Sonnet 4.5 Thinking 32k. Critiques emerged on OSWorld benchmark stability and task validity. **LlamaIndex**&apos;s LIGHT framework demonstrated significant improvements in long-term memory tasks over raw context and RAG baselines, with gains up to +160.6% in summarization at 10M tokens. **Amazon** introduced Chronos-2, a time-series foundation model for zero-shot forecasting. The MCP ecosystem expanded with new tools like mcp2py OAuth integration and Gemini Docs MCP server, alongside a build sprint by **Anthropic** and **Gradio** offering substantial credits and prizes. *&quot;OSWorld doesn’t really exist—different prompt sets = incomparable scores&quot;* highlights benchmarking challenges.</description><pubDate>Mon, 03 Nov 2025 05:44:39 GMT</pubDate><category>openai</category><category>aws</category><category>microsoft</category><category>nvidia</category><category>gpu_mode</category><category>vllm</category><category>alibaba</category><category>arena</category><category>llamaindex</category><category>amazon</category><category>anthropic</category><category>gradio</category><category>qwen3-max-thinking</category><category>minimax-m2</category><category>claude-3-sonnet</category><category>llamaindex-light</category><category>chronos-2</category><category>sama</category><category>gdb</category><category>andrewcurran_</category><category>a1zhang</category><category>m_sirovatka</category><category>omarsar0</category><category>_philschmid</category><category>compute-deals</category><category>gpu-optimization</category><category>kernel-optimization</category><category>local-serving</category><category>reasoning</category><category>long-context</category><category>benchmarks</category><category>long-term-memory</category><category>time-series-forecasting</category><category>agent-frameworks</category><category>oauth-integration</category><category>developer-tools</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-31-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-31-not-much/</guid><description>**Poolside** raised **$1B** at a **$12B valuation**. **Eric Zelikman** raised **$1B** after leaving **Xai**. **Weavy** joined **Figma**. New research highlights **FP16** precision reduces training-inference mismatch in **reinforcement-learning** fine-tuning compared to **BF16**. **Kimi AI** introduced a hybrid **KDA (Kimi Delta Attention)** architecture improving long-context throughput and RL stability, alongside a new **Kimi CLI** for coding with agent protocol support. **OpenAI** previewed Agent Mode in ChatGPT enabling autonomous research and planning during browsing.</description><pubDate>Fri, 31 Oct 2025 05:44:39 GMT</pubDate><category>poolside</category><category>x-ai</category><category>figma</category><category>openai</category><category>kimi</category><category>moonshot</category><category>eric_zelikman</category><category>reinforcement-learning</category><category>precision</category><category>fp16</category><category>bf16</category><category>linear-attention</category><category>long-context</category><category>cli</category><category>agent-frameworks</category><category>coding-agents</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-30-not-much/</guid><description>**Moonshot AI** released **Kimi Linear (KDA)** with day-0 infrastructure and strong long-context metrics, achieving up to **75% KV cache reduction** and **6x decoding throughput**. **MiniMax M2** pivoted to full attention for multi-hop reasoning, maintaining strong agentic coding performance with **200k context** and **~100 TPS**. **ByteDance**, **Princeton**, and **Mila** introduced **Looped LLMs** showing efficiency gains comparable to larger transformers. **OpenAI**&apos;s **Aardvark (GPT-5)** entered private beta as an agentic security researcher for scalable vulnerability discovery. **Cursor** launched faster cloud coding agents, though transparency concerns arose regarding base-model provenance. **Cognition** released a public beta for a desktop/mobile tool-use agent named Devin. The community discussed advanced attention mechanisms and adaptive compute techniques.</description><pubDate>Thu, 30 Oct 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>minimax</category><category>bytedance</category><category>princeton</category><category>mila</category><category>openai</category><category>cursor</category><category>cognition</category><category>hkust</category><category>kimi-linear</category><category>kimi-delta-attention</category><category>minimax-m2</category><category>looped-llms</category><category>aardvark-gpt-5</category><category>kimi_moonshot</category><category>scaling01</category><category>uniartisan</category><category>omarsar0</category><category>aicodeking</category><category>songlinyang4</category><category>iscienceluvr</category><category>nrehiew_</category><category>gdb</category><category>embeddedsec</category><category>auchenberg</category><category>simonw</category><category>long-context</category><category>attention-mechanisms</category><category>agentic-ai</category><category>tool-use</category><category>adaptive-compute</category><category>coding-agents</category><category>performance-optimization</category><category>memory-optimization</category><category>reinforcement-learning</category><category>model-architecture</category></item><item><title>Cursor 2.0 &amp; Composer-1: Fast Models and New Agents UI</title><link>https://news.smol.ai/issues/25-10-29-cursor-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-29-cursor-2/</guid><description>**Cursor 2.0** launched with **Composer-1**, an agentic coding model optimized for speed and precision, featuring multi-agent orchestration, built-in browser for testing, and voice-to-code capabilities. **OpenAI** released **gpt-oss-safeguard** models (20B, 120B) for policy-based safety classification, open-weight and fine-tuned from gpt-oss, available on Hugging Face and supported by inference stacks like Ollama and Cerebras. **Goodfire** and **Rakuten** demonstrated sparse autoencoders for PII detection matching **gpt-5-mini** accuracy at significantly lower cost. The Cursor 2.0 update also includes a redesigned interface for managing multiple AI coding agents, marking a major advancement in AI IDEs. *&quot;Fast-not-slowest&quot; tradeoff emphasized by early users for Composer-1, enabling rapid iteration with human-in-the-loop.*</description><pubDate>Wed, 29 Oct 2025 05:44:39 GMT</pubDate><category>cursor_ai</category><category>openai</category><category>huggingface</category><category>ollama</category><category>cerebras</category><category>groq</category><category>goodfireai</category><category>rakuten</category><category>composer-1</category><category>gpt-oss-safeguard-20b</category><category>gpt-oss-safeguard-120b</category><category>gpt-oss</category><category>gpt-5-mini</category><category>sasha_rush</category><category>dan_shipper</category><category>samkottler</category><category>ellev3n11</category><category>swyx</category><category>agentic-coding</category><category>reinforcement-learning</category><category>mixture-of-experts</category><category>fine-tuning</category><category>policy-classification</category><category>open-weight-models</category><category>inference-stacks</category><category>cost-efficiency</category><category>multi-agent-systems</category><category>ide</category><category>voice-to-code</category><category>code-review</category><category>built-in-browser</category><category>model-optimization</category></item><item><title>OpenAI completes Microsoft + For-profit restructuring + announces 2028 AI Researcher timeline + Platform / AI cloud product direction + next $1T of compute</title><link>https://news.smol.ai/issues/25-10-28-openai-restructure/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-28-openai-restructure/</guid><description>**OpenAI** has completed a major recapitalization and restructuring, forming a Public Benefit Corporation with a non-profit Foundation holding special voting rights and equity valued at **$130B**. **Microsoft** holds about **27%** diluted ownership and committed to **$250B** in Azure spend, losing exclusivity on compute but retaining Azure API exclusivity until AGI is declared. The compute infrastructure deals for 2025 total **30GW** worth **$1.4T**, with OpenAI aiming to build **1GW per week** at **$20B per GW**, projecting **$3-4 trillion** infrastructure by 2033. The company is shifting focus from first-party apps to a platform approach, emphasizing ecosystem growth and third-party development. **Sam Altman** and **Sama** are key figures in this transition, with significant financial and strategic implications for AI industry partnerships, including openness to **Anthropic** and **Google Gemini** on Azure.</description><pubDate>Tue, 28 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>microsoft</category><category>anthropic</category><category>google-deepmind</category><category>sama</category><category>sam_altman</category><category>public-benefit-corporation</category><category>corporate-restructuring</category><category>compute-infrastructure</category><category>cloud-computing</category><category>platform-strategy</category><category>api-exclusivity</category><category>investment</category><category>infrastructure-capex</category></item><item><title>MiniMax M2 230BA10B — 8% of Claude Sonnet&apos;s price, ~2x faster, new SOTA open model</title><link>https://news.smol.ai/issues/25-10-27-minimax-m2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-27-minimax-m2/</guid><description>**MiniMax M2**, an open-weight sparse MoE model by **Hailuo AI**, launches with **≈200–230B parameters** and **10B active parameters**, offering strong performance near frontier closed models and ranking #5 overall on the Artificial Analysis Intelligence Index v3.0. It supports coding and agent tasks, is licensed under **MIT**, and is available via API at competitive pricing. The architecture uses **full attention**, **QK-Norm**, **GQA**, partial RoPE, and sigmoid routing, with day-0 support in **vLLM** and deployment on platforms like Hugging Face and Baseten. Despite verbosity and no tech report, it marks a significant win for open models.</description><pubDate>Mon, 27 Oct 2025 05:44:39 GMT</pubDate><category>hailuo-ai</category><category>huggingface</category><category>baseten</category><category>vllm</category><category>modelscope</category><category>openrouter</category><category>cline</category><category>minimax-m2</category><category>reach_vb</category><category>artificialanlys</category><category>akhaliq</category><category>eliebakouch</category><category>grad62304977</category><category>yifan_zhang_</category><category>zpysky1125</category><category>sparse-moe</category><category>model-benchmarking</category><category>model-architecture</category><category>instruction-following</category><category>tool-use</category><category>api-pricing</category><category>model-deployment</category><category>performance-evaluation</category><category>full-attention</category><category>qk-norm</category><category>gqa</category><category>rope</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-24-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-24-not-much/</guid><description>**vLLM** announced support for **NVIDIA Nemotron Nano 2**, featuring a hybrid Transformer–Mamba design and tunable &quot;thinking budget&quot; enabling up to 6× faster token generation. **Mistral AI Studio** launched a production platform for agents with deep observability. **Baseten** reported high throughput (650 TPS) for **GPT-OSS 120B** on NVIDIA hardware. **Hugging Face InspectAI** added inference provider integration for cross-provider evaluation. **Thinking Machines Tinker** abstracts distributed fine-tuning for open-weight LLMs like **Qwen3** and **Llama 3**. In China, **MiniMax M2** shows competitive performance with top models and is optimized for agents and coding, while **Zhipu GLM-4.6-Air** focuses on reliability and scaling for coding tasks. Rumors suggest **Gemini 2.5 Flash** may be a &gt;500B parameter MoE model, and a possible **GPT-5.1 mini** reference appeared. Outside LLMs, **Tahoe-x1 (3B)** foundation model achieved SOTA in cancer cell biology benchmarks. Research from Stanford introduces a method to detect model provenance via training-order &quot;palimpsest&quot; with strong statistical guarantees.</description><pubDate>Fri, 24 Oct 2025 05:44:39 GMT</pubDate><category>vllm_project</category><category>nvidia</category><category>mistral-ai</category><category>baseten</category><category>huggingface</category><category>thinking-machines</category><category>deeplearningai</category><category>pytorch</category><category>arena</category><category>yupp-ai</category><category>zhipu-ai</category><category>scaling01</category><category>stanford</category><category>nemotron-nano-2</category><category>gpt-oss-120b</category><category>qwen3</category><category>llama-3</category><category>minimax-m2</category><category>glm-4.6-air</category><category>gemini-2.5-flash</category><category>gpt-5.1-mini</category><category>tahoe-x1</category><category>swyx</category><category>dvilasuero</category><category>_lewtun</category><category>clementdelangue</category><category>zephyr_z9</category><category>skylermiao7</category><category>teortaxestex</category><category>nalidoust</category><category>transformer-architecture</category><category>model-optimization</category><category>inference</category><category>distributed-training</category><category>multi-gpu-support</category><category>performance-optimization</category><category>agents</category><category>observability</category><category>model-evaluation</category><category>reinforcement-learning</category><category>model-provenance</category><category>statistical-testing</category><category>foundation-models</category><category>cancer-biology</category><category>model-fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-23-not-much/</guid><description>**LangSmith** launched the **Insights Agent** with multi-turn evaluation for agent ops and observability, improving failure detection and user intent clustering. **Meta PyTorch** and **Hugging Face** introduced **OpenEnv**, a Gymnasium-style API and hub for reproducible agentic environments supporting distributed training. Discussions highlighted the importance of provider fidelity in agent coding, with **OpenRouter**&apos;s exacto filter improving stability. Builder UX updates include **Google AI Studio**&apos;s Annotation mode for Gemini code changes, **Microsoft**&apos;s Copilot Mode enhancements in Edge, and **OpenAI**&apos;s Shared Projects and Company Knowledge features for ChatGPT Business. **Claude** added project-scoped Memory. In reinforcement learning, **Meta**&apos;s ScaleRL proposes a methodology to predict RL scaling outcomes for LLMs with improved efficiency and stability.</description><pubDate>Thu, 23 Oct 2025 05:44:39 GMT</pubDate><category>langchain</category><category>meta-ai-fair</category><category>hugging-face</category><category>openrouter</category><category>google-ai</category><category>microsoft</category><category>openai</category><category>anthropic</category><category>gemini-1.5-pro</category><category>claude-3</category><category>chatgpt</category><category>hwchase17</category><category>ankush_gola11</category><category>whinthorn</category><category>koylanai</category><category>_lewtun</category><category>bhutanisanyam1</category><category>thom_wolf</category><category>danielhanchen</category><category>cline</category><category>canvrno</category><category>pashmerepat</category><category>mustafasuleyman</category><category>yusuf_i_mehdi</category><category>jordirib1</category><category>fidjissimo</category><category>bradlightcap</category><category>mikeyk</category><category>alexalbert__</category><category>agent-ops</category><category>observability</category><category>multi-turn-evaluation</category><category>reinforcement-learning</category><category>distributed-training</category><category>api</category><category>model-stability</category><category>user-intent-clustering</category><category>software-development</category><category>project-management</category><category>code-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-22-not-much/</guid><description>**LangChain &amp; LangGraph 1.0** released with major updates for reliable, controllable agents and unified docs, emphasizing &quot;Agent Engineering.&quot; **Meta** introduced **PyTorch Monarch** and **TorchForge** for distributed programming and reinforcement learning, enabling large-scale agentic systems. **Microsoft Learn MCP** server now integrates with tools like **Claude Code** and **VS Code** for instant doc querying, accelerating grounded agent workflows. **vLLM** improved inference correctness with token ID returns and batch-invariant inference, collaborating with **Ray** for orchestration in PyTorch Foundation. **OpenAI** launched **ChatGPT Atlas**, a browser agent with contextual Q&amp;A and advanced safety features, though early users note maturity challenges and caution around credential access.</description><pubDate>Wed, 22 Oct 2025 05:44:39 GMT</pubDate><category>langchain</category><category>meta</category><category>microsoft</category><category>openai</category><category>pytorch</category><category>ray</category><category>claude</category><category>vllm</category><category>chatgpt-atlas</category><category>hwchase17</category><category>soumithchintala</category><category>masondrxy</category><category>robertnishihara</category><category>cryps1s</category><category>yuchenj_uw</category><category>agent-frameworks</category><category>reinforcement-learning</category><category>distributed-computing</category><category>inference-correctness</category><category>serving-infrastructure</category><category>browser-agents</category><category>security</category><category>middleware</category><category>runtime-systems</category><category>documentation</category></item><item><title>ChatGPT Atlas: OpenAI&apos;s AI Browser</title><link>https://news.smol.ai/issues/25-10-21-chatgpt-atlas/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-21-chatgpt-atlas/</guid><description>**OpenAI** launched the **Chromium fork AI browser Atlas** for macOS, featuring integrated **Agent mode** and browser memory with local login capabilities, aiming to surpass **Google&apos;s Gemini** in Chrome. The launch received mixed reactions regarding reliability and privacy. **LangChain** raised a **$125M Series B** at a $1.25B valuation, releasing **v1.0 agent engineering stack** with significant adoption including **85M+ OSS downloads/month** and usage by ~35% of the Fortune 500. The ecosystem also saw updates like **vLLM&apos;s MoE LoRA expert finetuning support**.</description><pubDate>Tue, 21 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>langchain</category><category>ivp</category><category>capitalg</category><category>sapphire</category><category>sequoia</category><category>benchmark</category><category>gemini</category><category>atlas</category><category>kevinweil</category><category>bengoodger</category><category>fidjissimo</category><category>omarsar0</category><category>yuchenj_uw</category><category>nickaturley</category><category>raizamrtn</category><category>hwchase17</category><category>bromann</category><category>casper_hansen_</category><category>corbtt</category><category>agent-mode</category><category>browser-memory</category><category>chromium</category><category>finetuning</category><category>moe</category><category>lora</category><category>agent-runtime</category><category>observability</category><category>software-development</category><category>funding</category></item><item><title>DeepSeek-OCR finds vision models can decode 10x more efficiently with ~97% accuracy of text-only, 33/200k pages/day/A100</title><link>https://news.smol.ai/issues/25-10-20-deepseek-ocr/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-20-deepseek-ocr/</guid><description>As **ICCV 2025** begins, **DeepSeek** releases a novel **DeepSeek-OCR** 3B MoE vision-language model that compresses long text as visual context with high accuracy and efficiency, challenging traditional tokenization approaches. The model achieves ~97% decoding precision at &lt;10× compression and processes up to ~33M pages/day on 20 A100-40G nodes, outperforming benchmarks like GOT-OCR2.0. Discussions highlight the potential for unlimited context windows and tokenization-free inputs, with contributions from **@karpathy**, **@teortaxesTex**, and others. In video generation, **google-deepmind**&apos;s **Veo 3.1** leads community benchmarks with advanced precision editing and scene blending, while **Krea** open-sources a 14B autoregressive video model enabling realtime long-form generation at ~11 FPS on a single B200 GPU.</description><pubDate>Mon, 20 Oct 2025 05:44:39 GMT</pubDate><category>deepseek-ai</category><category>google-deepmind</category><category>krea</category><category>deepseek-ocr</category><category>deepseek3b-moe-a570m</category><category>veo-3.1</category><category>karpathy</category><category>teortaxestex</category><category>reach_vb</category><category>_akhaliq</category><category>eliebakouch</category><category>vikhyatk</category><category>demishassabis</category><category>ocr</category><category>vision</category><category>multimodality</category><category>model-compression</category><category>long-context</category><category>model-architecture</category><category>video-generation</category><category>autoregressive-models</category><category>model-efficiency</category><category>precision-editing</category></item><item><title>The Karpathy-Dwarkesh Interview delays AGI timelines</title><link>https://news.smol.ai/issues/25-10-17-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-17-not-much/</guid><description>The recent AI news highlights the **Karpathy interview** as a major event, alongside significant discussions on reasoning improvements without reinforcement learning, with **test-time sampling** achieving GRPO-level performance. Critiques on context window marketing reveal effective limits near **64K tokens**, with **Claude Haiku 4.5** showing competitive reasoning speed. **GPT-5** struggles with advanced math benchmarks, and data quality issues termed &quot;Brain Rot&quot; affect model reasoning and safety. In agent frameworks, **Anthropic Skills** enable modular coding workflows, **OpenAI Codex IDE** extensions enhance developer productivity, and **HuggingChat Omni** introduces meta-routing across 100+ open models using **Arch-Router-1.5B**. LangChain and LlamaIndex advance graph-first agent infrastructure, while **Google Gemini** integrates with Google Maps for real-world grounding.</description><pubDate>Fri, 17 Oct 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>huggingface</category><category>langchain</category><category>llamaindex</category><category>google</category><category>epoch-ai</category><category>claude-haiku-4.5</category><category>gpt-5</category><category>arch-router-1.5b</category><category>karpathy</category><category>aakaran31</category><category>du_yilun</category><category>giffmana</category><category>omarsar0</category><category>jeremyphoward</category><category>claude_code</category><category>mikeyk</category><category>alexalbert__</category><category>clementdelangue</category><category>jerryjliu0</category><category>reasoning</category><category>long-context</category><category>sampling</category><category>benchmarking</category><category>data-quality</category><category>agent-frameworks</category><category>modular-workflows</category><category>ide-extensions</category><category>model-routing</category><category>graph-first-agents</category><category>real-world-grounding</category></item><item><title>Claude Agent Skills - glorified AGENTS.md? or MCP killer?</title><link>https://news.smol.ai/issues/25-10-16-claude-skills/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-16-claude-skills/</guid><description>**Anthropic** achieves a rare feat with back-to-back AI news headlines featuring **Claude&apos;s** new **Skills**—a novel way to build specialized agents using Markdown files, scripts, and metadata to handle tasks like creating and reading PDFs, Docs, and PPTs. Simon Willison calls this a &quot;bigger deal than MCP,&quot; predicting a &quot;Cambrian explosion in Skills.&quot; Meanwhile, **Anthropic** launches **Claude 4.5 Haiku** with strong reasoning and long-context capabilities, priced competitively. Other updates include **OpenAI&apos;s** ChatGPT memory management improvements, **Windows 11 Copilot** voice and vision features, and **HuggingChat Omni** routing across 115 open-source models from 15 providers. These developments highlight advances in agent skills, document processing, long-context reasoning, and multi-model routing.</description><pubDate>Thu, 16 Oct 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>microsoft</category><category>perplexity-ai</category><category>huggingface</category><category>groq</category><category>cerebras</category><category>togethercompute</category><category>claude-4.5-haiku</category><category>claude</category><category>chatgpt</category><category>huggingchat-omni</category><category>simonwillison</category><category>alexalbert__</category><category>mustafasuleyman</category><category>yusuf_i_mehdi</category><category>aravsrinivas</category><category>agent-skills</category><category>document-processing</category><category>long-context</category><category>reasoning</category><category>multi-model-routing</category><category>memory-management</category><category>voice</category><category>vision</category></item><item><title>Claude Haiku 4.5</title><link>https://news.smol.ai/issues/25-10-15-haiku-45/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-15-haiku-45/</guid><description>**Anthropic** released **Claude Haiku 4.5**, a model that is over 2x faster and 3x cheaper than **Claude Sonnet 4.5**, improving iteration speed and user experience significantly. Pricing comparisons highlight Haiku 4.5&apos;s competitive cost against models like **GPT-5** and **GLM-4.6**. **Google** and **Yale** introduced the open-weight **Cell2Sentence-Scale 27B (Gemma)** model, which generated a novel, experimentally validated cancer hypothesis, with open-sourced weights for community use. Early evaluations show **GPT-5** and **o3** models outperform **GPT-4.1** in agentic reasoning tasks, balancing cost and performance. Agent evaluation challenges and memory-based learning advances were also discussed, with contributions from Shanghai AI Lab and others. *&quot;Haiku 4.5 materially improves iteration speed and UX,&quot;* and *&quot;Cell2Sentence-Scale yielded validated cancer hypothesis&quot;* were key highlights.</description><pubDate>Wed, 15 Oct 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>google</category><category>yale</category><category>artificial-analysis</category><category>shanghai-ai-lab</category><category>claude-3.5-sonnet</category><category>claude-3-haiku</category><category>claude-3-haiku-4.5</category><category>gpt-5</category><category>gpt-4.1</category><category>gemma-2.5</category><category>gemma</category><category>o3</category><category>swyx</category><category>sundarpichai</category><category>osanseviero</category><category>clementdelangue</category><category>deredleritt3r</category><category>azizishekoofeh</category><category>vikhyatk</category><category>mirrokni</category><category>pdrmnvd</category><category>akhaliq</category><category>sayashk</category><category>gne</category><category>model-performance</category><category>fine-tuning</category><category>reasoning</category><category>agent-evaluation</category><category>memory-optimization</category><category>model-efficiency</category><category>open-models</category><category>cost-efficiency</category><category>foundation-models</category><category>agentic-workflows</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-14-not-much/</guid><description>**Alibaba** released compact dense **Qwen3-VL** models at 4B and 8B sizes with FP8 options, supporting up to 1M context and open vocabulary detection, rivaling larger models like **Qwen2.5-VL-72B**. Ecosystem support includes **MLX-VLM**, **LM Studio**, **vLLM**, **Kaggle models**, and **Ollama Cloud**. In video AI, **Arena** added **Sora 2** models leading in video benchmarks, with **Higgsfield Enhancer** improving video quality. **Runway** launched domain-specific workflow apps for creative tasks. Research on **Representation Autoencoders for DiTs (RAE-DiT)** shows improved diffusion model performance. On local training, **NVIDIA DGX Spark** enables strong local fine-tuning, while **Nanochat** by **Karpathy** offers a minimal stack for training and inference. **Together AI** introduced **ATLAS**, a speculative decoding method achieving up to 4× faster inference on **DeepSeek-V3.1**. These developments highlight advances in efficient model deployment, video AI, local fine-tuning, and inference speed optimization.</description><pubDate>Tue, 14 Oct 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>arena</category><category>runway</category><category>nvidia</category><category>togethercompute</category><category>ollama</category><category>qwen3-vl-4b</category><category>qwen3-vl-8b</category><category>qwen2.5-vl-72b</category><category>deepseek-v3.1</category><category>karpathy</category><category>model-optimization</category><category>fine-tuning</category><category>inference-speed</category><category>video-generation</category><category>diffusion-models</category><category>representation-learning</category><category>local-ai</category><category>speculative-decoding</category><category>fp8-quantization</category><category>context-windows</category></item><item><title>OpenAI Titan XPU: 10GW of self-designed chips with Broadcom</title><link>https://news.smol.ai/issues/25-10-13-oai-broadcom/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-13-oai-broadcom/</guid><description>**OpenAI** is finalizing a custom ASIC chip design to deploy **10GW** of inference compute, complementing existing deals with **NVIDIA** (10GW) and **AMD** (6GW). This marks a significant scale-up from OpenAI&apos;s current **2GW** compute, aiming for a roadmap of **250GW** total, which is half the energy consumption of the US. Greg from OpenAI highlights the shift of **ChatGPT** from interactive use to always-on ambient agents requiring massive compute, emphasizing the challenge of building chips for billions of users. The in-house ASIC effort was driven by the need for tailored designs after limited success influencing external chip startups. Broadcom&apos;s stock surged 10% on the news. Additionally, **InferenceMAX** reports improved ROCm stability and nuanced performance comparisons between AMD MI300X and NVIDIA H100/H200 on **llama-3-70b** FP8 workloads, with RL training infrastructure updates noted.</description><pubDate>Mon, 13 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>nvidia</category><category>amd</category><category>broadcom</category><category>inferencemax</category><category>llama-3-70b</category><category>gdb</category><category>asic</category><category>inference</category><category>compute-infrastructure</category><category>chip-design</category><category>fp8</category><category>reinforcement-learning</category><category>ambient-agents</category><category>custom-accelerators</category><category>energy-consumption</category><category>podcast</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-10-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-10-not-much/</guid><description>**FrontierMath Tier 4** results show **GPT-5 Pro** narrowly outperforming **Gemini 2.5 Deep Think** in reasoning accuracy, with concerns about problem leakage clarified by **Epoch AI Research**. **Mila** and **Microsoft** propose **Markovian Thinking** to improve reasoning efficiency, enabling models to reason over 24K tokens with less compute. New research suggests base models inherently contain reasoning mechanisms, with &quot;thinking models&quot; learning to invoke them effectively. In systems, **NVIDIA Blackwell** combined with **vLLM** wins InferenceMAX with significant throughput gains, while **Together AI&apos;s ATLAS** adaptive speculative decoding achieves 4× speed improvements and reduces RL training time by over 60%. **SparseServe** introduces dynamic sparse attention with KV tiering, drastically improving throughput and latency in GPU memory management.</description><pubDate>Fri, 10 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>microsoft</category><category>epoch-ai-research</category><category>togethercompute</category><category>nvidia</category><category>mila</category><category>gpt-5-pro</category><category>gemini-2.5</category><category>vllm</category><category>deepseek-v3.1</category><category>epochairesearch</category><category>yitayml</category><category>_philschmid</category><category>jiqizhixin</category><category>cvenhoff00</category><category>neelnanda5</category><category>lateinteraction</category><category>mgoin_</category><category>blackhc</category><category>teortaxestex</category><category>reasoning</category><category>reinforcement-learning</category><category>inference</category><category>speculative-decoding</category><category>sparse-attention</category><category>kv-cache-management</category><category>throughput-optimization</category><category>compute-efficiency</category><category>tokenization</category></item><item><title>Air Street&apos;s State of AI 2025 Report</title><link>https://news.smol.ai/issues/25-10-09-state-of-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-09-state-of-ai/</guid><description>**Reflection** raised **$2B** to build frontier open-weight models with a focus on safety and evaluation, led by a team with backgrounds from **AlphaGo**, **PaLM**, and **Gemini**. **Figure** launched its next-gen humanoid robot, **Figure 03**, emphasizing non-teleoperated capabilities for home and large-scale use. **Radical Numerics** released **RND1**, a **30B-parameter sparse MoE diffusion language model** with open weights and code to advance diffusion LM research. **Zhipu** posted strong results with **GLM-4.6** on the Design Arena benchmark, while **AI21 Labs**&apos; **Jamba Reasoning 3B** leads tiny reasoning models. **Anthropic** introduced a plugin system for **Claude Code** to enhance developer tools and agent stacks. The report also highlights SoftBank&apos;s acquisition of ABB&apos;s robotics unit for **$5.4B** and the growing ecosystem around open frontier modeling and small-model reasoning.</description><pubDate>Thu, 09 Oct 2025 05:44:39 GMT</pubDate><category>reflection</category><category>mastra</category><category>datacurve</category><category>spellbook</category><category>kernel</category><category>figure</category><category>softbank</category><category>abb</category><category>radicalnumerics</category><category>zhipu-ai</category><category>ai21-labs</category><category>anthropic</category><category>glm-4.6</category><category>jamba-1.5</category><category>rnd1</category><category>claude-code</category><category>adcock_brett</category><category>achowdhery</category><category>clementdelangue</category><category>humanoid-robots</category><category>mixture-of-experts</category><category>diffusion-models</category><category>open-weight-models</category><category>reinforcement-learning</category><category>benchmarking</category><category>small-language-models</category><category>plugin-systems</category><category>developer-tools</category><category>agent-stacks</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-08-not-much/</guid><description>**Samsung&apos;s 7M Tiny Recursive Model (TRM)** achieves superior reasoning on ARC-AGI and Sudoku with fewer layers and MLP replacing self-attention. **LeCun&apos;s team** introduces **JEPA-SCORE**, enabling density estimation from encoders without retraining. **AI21 Labs** releases **Jamba Reasoning 3B**, a fast hybrid SSM-Transformer model supporting up to 64K context tokens. **Alibaba&apos;s Qwen3 Omni/Omni Realtime** offers a unified audio-video-text model with extensive language and speech support, outperforming Gemini 2.0 Flash on BigBench Audio. **Alibaba** also debuts **Qwen Image Edit 2509**, a top open-weight multi-image editing model. **ColBERT Nano** models demonstrate effective retrieval at micro-scale parameter sizes. In reinforcement learning, **CoreWeave**, **Weights &amp; Biases**, and **OpenPipe** launch serverless RL infrastructure reducing costs and speeding training. **Stanford&apos;s AgentFlow** presents an in-the-flow RL system with a 7B backbone outperforming larger models on agentic tasks. This update highlights advances in **recursive reasoning**, **density estimation**, **multimodal architectures**, **long-context modeling**, **retrieval**, and **serverless reinforcement learning**.</description><pubDate>Wed, 08 Oct 2025 05:44:39 GMT</pubDate><category>samsung</category><category>lecuun</category><category>ai21-labs</category><category>alibaba</category><category>coreweave</category><category>weights-biases</category><category>openpipe</category><category>stanford</category><category>7m-tiny-recursive-model</category><category>jamba-reasoning-3b</category><category>qwen3-omni</category><category>qwen-image-edit-2509</category><category>colbert-nano</category><category>agentflow</category><category>rasbt</category><category>jm_alexia</category><category>jiqizhixin</category><category>randall_balestr</category><category>corbtt</category><category>shawnup</category><category>_akhaliq</category><category>recursive-reasoning</category><category>density-estimation</category><category>multimodality</category><category>long-context</category><category>retrieval</category><category>serverless-reinforcement-learning</category><category>agentic-systems</category><category>model-efficiency</category><category>reinforcement-learning</category><category>transformers</category></item><item><title>Gemini 2.5 Computer Use preview beats Sonnet 4.5 and OAI CUA</title><link>https://news.smol.ai/issues/25-10-07-gemini-cua/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-07-gemini-cua/</guid><description>**Google DeepMind** released a new **Gemini 2.5 Computer Use model** for browser and Android UI control, evaluated by Browserbase. **OpenAI** showcased **GPT-5 Pro**, new developer tools including **Codex** with Slack integration, and agent-building SDKs at Dev Day. **Google DeepMind&apos;s CodeMender** automates security patching for large codebases. **Microsoft** introduced an open-source **Agent Framework** for multi-agent enterprise systems. AI community discussions highlight agent orchestration, program synthesis, and UI control advancements. **GLM-4.6** update from Zhipu features a large Mixture-of-Experts model with 355B parameters.</description><pubDate>Tue, 07 Oct 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>microsoft</category><category>anthropic</category><category>zhipu-ai</category><category>llamaindex</category><category>mongodb</category><category>gemini-2.5</category><category>gpt-5-pro</category><category>glm-4.6</category><category>codex</category><category>swyx</category><category>demishassabis</category><category>philschmid</category><category>assaf_elovic</category><category>hwchase17</category><category>jerryjliu0</category><category>skirano</category><category>fabianstelzer</category><category>blackhc</category><category>andrewyng</category><category>agent-frameworks</category><category>program-synthesis</category><category>security</category><category>multi-agent-systems</category><category>computer-use-models</category><category>open-source</category><category>moe</category><category>developer-tools</category><category>workflow-automation</category><category>api</category><category>vision</category><category>reasoning</category></item><item><title>OpenAI Dev Day: Apps SDK, AgentKit, Codex GA, GPT‑5 Pro and Sora 2 APIs</title><link>https://news.smol.ai/issues/25-10-06-devday/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-06-devday/</guid><description>**OpenAI** showcased major product launches at their DevDay including the **Apps SDK**, **AgentKit**, and **Codex** now generally available with SDK and enterprise features. They introduced new models such as **gpt-5-pro**, **gpt-realtime-mini-2025-10-06**, **gpt-audio-mini-2025-10-06**, **gpt-image-1-mini**, and **sora-2** with a pro variant. The Apps SDK enables embedding interactive apps inside ChatGPT with partners like **Canva**, **Figma**, **Zillow**, and **Coursera**. AgentKit offers a full stack for building and deploying production agents with tools like ChatKit and Guardrails. Codex supports speech and controller-driven coding, credited with high internal shipping velocity. Pricing for GPT-5 Pro was revealed at $15 input and $120 output per million tokens. *&quot;OpenAI turned ChatGPT into an application platform&quot;* and *&quot;AgentKit built a working agent in under 8 minutes&quot;* were highlights.</description><pubDate>Mon, 06 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>canva</category><category>figma</category><category>zillow</category><category>coursera</category><category>gpt-5-pro</category><category>gpt-realtime-mini-2025-10-06</category><category>gpt-audio-mini-2025-10-06</category><category>gpt-image-1-mini</category><category>sora-2</category><category>sora-2-pro</category><category>sama</category><category>edwinarbus</category><category>gdb</category><category>dbreunig</category><category>stevenheidel</category><category>api</category><category>model-release</category><category>fine-tuning</category><category>agentic-ai</category><category>code-generation</category><category>model-deployment</category><category>pricing</category><category>prompt-optimization</category><category>software-development</category><category>multimodality</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-03-not-much/</guid><description>**Anthropic** announces a new CTO. Frontier coding agents see updates with **Claude Sonnet 4.5** showing strong cybersecurity and polished UX but trailing **GPT-5 Codex** in coding capability. **xAI Grok Code Fast** claims higher edit success at lower cost. **Google&apos;s Jules** coding agent launches a programmable API with CI/CD integration. **Qwen** clarifies its model taxonomy and API tiers. Vision/LM Arena rankings show a tight competition among **Claude Sonnet 4.5**, **Claude Opus 4.1**, **Gemini 2.5 Pro**, and OpenAI&apos;s latest models. In video generation, **Sora 2 Pro** leads App Store rankings with rapid iteration and a new creator ecosystem; early tests show it answers GPQA-style questions at 55% accuracy versus GPT-5&apos;s 72%. Video Arena adds new models like **Luma&apos;s Ray 3** and **Kling 2.5** for benchmarking. Multi-modal video+audio generation model **Ovi** (Veo-3-like) is released. Retrieval models include **ModernVBERT** from MIT with efficient image-text retrieval capabilities. *&quot;Claude Sonnet 4.5 is basically the same as Opus 4.1 for coding&quot;* and *&quot;Jules is a programmable team member&quot;* highlight key insights.</description><pubDate>Fri, 03 Oct 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>x-ai</category><category>google</category><category>google-labs</category><category>openai</category><category>arena</category><category>epoch-ai</category><category>mit</category><category>luma</category><category>akhaliq</category><category>claude-3-sonnet</category><category>claude-3-opus</category><category>gpt-5-codex</category><category>grok-4-fast</category><category>qwen-3-next</category><category>gemini-2.5-pro</category><category>sora-2-pro</category><category>ray-3</category><category>kling-2.5</category><category>veo-3</category><category>modernvbert</category><category>finbarrtimbers</category><category>gauravisnotme</category><category>justinlin610</category><category>billpeeb</category><category>apples_jimmy</category><category>akhaliq</category><category>coding-agents</category><category>cybersecurity</category><category>api</category><category>model-taxonomy</category><category>model-ranking</category><category>video-generation</category><category>benchmarking</category><category>multi-modal-generation</category><category>retrieval</category><category>image-text-retrieval</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-10-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-02-not-much/</guid><description>**Kling 2.5 Turbo** leads in text-to-video and image-to-video generation with competitive pricing. **OpenAI Sora 2** shows strong instruction-following but has physics inconsistencies. **Google Gemini 2.5 Flash** &quot;Nano Banana&quot; image generation is now generally available with multi-image blending and flexible aspect ratios. **IBM Granite 4.0** introduces a hybrid Mamba/Transformer architecture with large context windows and strong token efficiency, outperforming some peers on the Intelligence Index. **Qwen** models receive updates including fine-tuning API support and improved vision capabilities. **Tinker** offers a flexible fine-tuning API supporting LoRA sharing and CPU-only training loops. The ecosystem also sees updates like **Synthesia 3.0** adding video agents.</description><pubDate>Thu, 02 Oct 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>ibm</category><category>alibaba</category><category>kling_ai</category><category>synthesia</category><category>ollama</category><category>huggingface</category><category>arena</category><category>artificialanalysis</category><category>tinker</category><category>scaling01</category><category>kling-2.5-turbo</category><category>sora-2</category><category>gemini-2.5-flash</category><category>granite-4.0</category><category>qwen-3</category><category>qwen-image-2509</category><category>qwen3-vl-235b</category><category>artificialanlys</category><category>kling_ai</category><category>altryne</category><category>teortaxestex</category><category>fofrai</category><category>tim_dettmers</category><category>sundarpichai</category><category>officiallogank</category><category>andrew_n_carr</category><category>googleaidevs</category><category>clementdelangue</category><category>wzhao_nlp</category><category>alibaba_qwen</category><category>scaling01</category><category>ollama</category><category>video-generation</category><category>instruction-following</category><category>physics-simulation</category><category>image-generation</category><category>model-architecture</category><category>mixture-of-experts</category><category>context-windows</category><category>token-efficiency</category><category>fine-tuning</category><category>lora</category><category>cpu-training</category><category>model-benchmarking</category><category>api</category><category>workflow-automation</category></item><item><title>Thinking Machines&apos; Tinker: LoRA based LLM fine-tuning API</title><link>https://news.smol.ai/issues/25-10-01-thinky/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-10-01-thinky/</guid><description>**Thinking Machines** recently raised **$2 billion** without shipping a product until now, launching their first product **Tinker**, a managed service API for fine-tuning large and mixture-of-experts models like **Qwen-235B-A22B** using **LoRA** for cost-efficient training. The Tinker API offers low-level primitives for post-training methods and is supported by an open-source **Tinker Cookbook** library. Influential AI figures like **Andrej Karpathy** and **Lilian Weng** praised its design for reducing complexity and boosting research productivity. Meanwhile, **OpenAI** launched **Sora 2**, a video+audio model integrated into their consumer social app, sparking viral engagement and concerns over misuse and content moderation. Sam Altman emphasized the product&apos;s dual focus on delight and revenue alongside AGI research.</description><pubDate>Wed, 01 Oct 2025 05:44:39 GMT</pubDate><category>thinking-machines</category><category>openai</category><category>qwen-235b-a22b</category><category>sora-2</category><category>karpathy</category><category>lilianweng</category><category>sama</category><category>fine-tuning</category><category>lora</category><category>model-training</category><category>api</category><category>model-optimization</category><category>distributed-training</category><category>post-training-methods</category><category>research-productivity</category><category>video-generation</category><category>content-moderation</category><category>engagement-patterns</category></item><item><title>Sora 2: new video+audio model and OpenAI&apos;s first Social Network</title><link>https://news.smol.ai/issues/25-09-30-sora2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-30-sora2/</guid><description>**Sora 2** released with improvements on physical world video modeling and a new &quot;character consistency&quot; feature allowing real-world element injection from a single video. The model powers a new **Sora social network** app with profiles, DMs, and viral videos, emphasizing user control over likeness use. **OpenAI** employees are actively experimenting with the model. Meanwhile, **Anthropic** launched **Claude 4.5 Sonnet** with enhanced intelligence, token efficiency, and agentic tool use, outperforming some competitors and closely tracking **GPT-5-high** on benchmarks. Ecosystem support includes LangSmith integration and strong coding/math benchmark results.</description><pubDate>Tue, 30 Sep 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>sora-2</category><category>claude-4.5-sonnet</category><category>gpt-5-high</category><category>sama</category><category>video-generation</category><category>character-consistency</category><category>social-networks</category><category>agentic-ai</category><category>token-efficiency</category><category>benchmarking</category><category>model-performance</category><category>context-management</category><category>coding</category><category>math</category></item><item><title>Anthropic Claude Sonnet 4.5, Claude Code 2.0, new VS Code Extensions</title><link>https://news.smol.ai/issues/25-09-29-sonnet-45/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-29-sonnet-45/</guid><description>**Anthropic** launched a major update with **Claude Sonnet 4.5**, achieving **77.2% SWE-Bench** verified performance and improvements in finance, law, and STEM. They also released **Claude Code v2** featuring checkpoints, a refreshed terminal, and a native VS Code extension, plus a new mascot **Clawd**. The **Claude API** gained context editing and memory tools, and the **Claude Agent SDK** was introduced. The **Claude.ai** apps now support code execution and file creation, with a **Chrome extension** available for Max users. Additionally, **Imagine with Claude** offers a generative UI research preview. Reception has been positive from developers and third-party evaluators. Meanwhile, **DeepSeek** released **V3.2-Exp** with a new **Sparse Attention** algorithm, significantly reducing long-context costs and cutting API prices by over 50%, while maintaining quality.</description><pubDate>Mon, 29 Sep 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>deepseek</category><category>openai</category><category>stripe</category><category>claude-sonnet-4.5</category><category>claude-code-v2</category><category>deepseek-v3.2-exp</category><category>john_schulman</category><category>mike_krieger</category><category>swe-bench</category><category>finance</category><category>law</category><category>stem</category><category>code-execution</category><category>context-editing</category><category>memory-management</category><category>api</category><category>chrome-extension</category><category>generative-ui</category><category>sparse-attention</category><category>long-context</category><category>cost-efficiency</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-26-not-much/</guid><description>**Google** released a dense September update including **Gemini Robotics 1.5** with enhanced spatial/temporal reasoning, **Gemini Live**, **EmbeddingGemma**, and **Veo 3 GA** powering creative workflows. They also introduced agentic features like restaurant-reservation agents and reduced pricing for **Gemini 2.5 Flash**. **Meta AI** unveiled the open-weight **Code World Model (CWM) 32B**, excelling in code semantics and math benchmarks, with innovations in training code models via execution traces. Local-first coding setups highlight **Qwen3-Coder-30B** running efficiently on consumer GPUs, paired with tools like **Cline** and **LM Studio**. Runtime improvements include **vLLM v1** supporting hybrid models and **mlx-lm** adding batch inference on Apple silicon. In infrastructure, **FlashAttention 4** was reverse-engineered revealing a ~20% speedup from architectural optimizations. **Perplexity AI** advances its independent web index and browsing API with upcoming feed refreshes. Embedding latency improvements were achieved by **Superhuman** using **Baseten**.</description><pubDate>Fri, 26 Sep 2025 05:44:39 GMT</pubDate><category>google</category><category>meta-ai-fair</category><category>perplexity-ai</category><category>baseten</category><category>gemini-robotics-1.5</category><category>gemini-live</category><category>embeddinggemma</category><category>veo-3</category><category>gemini-2.5-flash</category><category>code-world-model-32b</category><category>qwen3-coder-30b</category><category>vllm-v1</category><category>mlx-lm</category><category>flashattention-4</category><category>osanseviero</category><category>_anniexie</category><category>rmstein</category><category>scaling01</category><category>giffmana</category><category>cline</category><category>redhat_ai</category><category>awnihannun</category><category>charles_irl</category><category>bernhardsson</category><category>akshat_b</category><category>aravsrinivas</category><category>spatial-reasoning</category><category>temporal-reasoning</category><category>agentic-ai</category><category>code-semantics</category><category>code-execution-traces</category><category>coding-infrastructure</category><category>runtime-optimization</category><category>batch-inference</category><category>embedding-latency</category><category>api</category><category>model-optimization</category><category>model-performance</category></item><item><title>GDPVal finding: Claude Opus 4.1 within 95% of AGI (human experts in top 44 white collar jobs)</title><link>https://news.smol.ai/issues/25-09-25-gdpval/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-25-gdpval/</guid><description>**OpenAI**&apos;s Evals team released **GDPval**, a comprehensive evaluation benchmark covering 1,320 tasks across 44 predominantly digital occupations, assessing AI models against human experts with 14 years average experience. Early results show **Claude 4.1 Opus** outperforming human experts in most categories and **GPT-5 high** trailing behind, with projections that **GPTnext** could match human performance by mid-2026. The benchmark is positioned as a key metric for policymakers and labor impact forecasting. Additionally, **Artificial Analysis** reported improvements in **Gemini 2.5 Flash/Flash-Lite** and **DeepSeek V3.1 Terminus** models, alongside new speech-to-text benchmarks (AA-WER) highlighting leaders like **Google Chirp 2** and **NVIDIA Canary Qwen2.5B**. Agentic AI advances include **Kimi OK Computer**, an OS-like agent with extended tool capabilities and new vendor verification tools.</description><pubDate>Thu, 25 Sep 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google</category><category>nvidia</category><category>artificial-analysis</category><category>deepseek</category><category>claude-4.1-opus</category><category>gpt-5-high</category><category>gptnext</category><category>gemini-2.5-flash</category><category>gemini-2.5-flash-lite</category><category>deepseek-v3.1-terminus</category><category>google-chirp-2</category><category>qwen-2.5b</category><category>kevinweil</category><category>gdb</category><category>dejavucoder</category><category>yuchenj_uw</category><category>lhsummers</category><category>benchmarking</category><category>agentic-ai</category><category>tool-use</category><category>long-context</category><category>speech-to-text</category><category>model-evaluation</category><category>reasoning</category><category>pricing</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-24-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-24-not-much/</guid><description>**Alibaba** unveiled the **Qwen3** model family including **Qwen3-Max** and **Qwen3-VL** with a native 256K context window expandable to 1M, strong OCR in 32 languages, and rapid release velocity (~3.5 releases/month) backed by a $52B infrastructure roadmap. **OpenAI** launched **GPT-5 Codex**, an agent-optimized coding model with up to **400K context** and adaptive reasoning priced at $1.25/$10 per million tokens, integrated into Cline and benchmarked in WebDev arenas. **Meta AI FAIR** released the open-weight **Code World Model (CWM) 32B**, a dense code generation model with strong benchmark scores (e.g., 65.8% SWE-bench Verified, 96.6% Math-500) and public safety reports. Ecosystem updates include GitHub Copilot&apos;s new embedding model for faster code search and Anthropic&apos;s Claude Sonnet 4 and Opus 4.1 integration into Microsoft 365 Copilot. The vLLM 0.10.2 update introduces Decode Context Parallel (DCP) for improved system performance.</description><pubDate>Wed, 24 Sep 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>openai</category><category>meta-ai-fair</category><category>huggingface</category><category>anthropic</category><category>microsoft</category><category>github</category><category>qwen3-max</category><category>qwen3-vl</category><category>qwen3-coder-plus</category><category>gpt-5-codex</category><category>code-world-model-32b</category><category>claude-sonnet-4</category><category>claude-opus-4.1</category><category>huybery</category><category>akhaliq</category><category>lmarena_ai</category><category>gdb</category><category>ylecun</category><category>pierceboggan</category><category>julesagent</category><category>context-windows</category><category>code-generation</category><category>model-releases</category><category>model-benchmarking</category><category>api</category><category>model-optimization</category><category>multimodality</category><category>software-engineering</category><category>model-training</category></item><item><title>Alibaba Yunqi: 7 models released in 4 days (Qwen3-Max, Qwen3-Omni, Qwen3-VL) and $52B roadmap</title><link>https://news.smol.ai/issues/25-09-23-alibaba-yunqi/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-23-alibaba-yunqi/</guid><description>**Alibaba&apos;s Tongyi Qianwen (Qwen) team** launched major updates including the **1T parameter Qwen3-Max**, **Qwen3-Omni**, and **Qwen3-VL** models, alongside specialized versions like **Qwen3Guard**, **Qwen3-LiveTranslate**, **Qwen3-TTS-Flash**, **Qwen-Image-Edit**, and **Qwen3Coder**. At the **AliCloud Yunqi (Apsara) conference**, CEO **Eddie Wu** outlined a $52B roadmap emphasizing two AI development stages: &quot;intelligence emergence&quot; focusing on learning from humans and reasoning, and &quot;autonomous action&quot; highlighting AI&apos;s tool use and real-world task execution. The updates showcase advances in **tool use**, **large-model coding capabilities**, and AI&apos;s expanding role across industries such as logistics, manufacturing, biomedicine, and finance. Junyang Lin and Alibaba Wan are key spokespersons for these developments. The Qwen project is now seen as a &quot;frontier lab&quot; for AI innovation.</description><pubDate>Tue, 23 Sep 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>alicloud</category><category>qwen3-max</category><category>qwen3-omni</category><category>qwen3-vl</category><category>qwen3guard</category><category>qwen3-livetranslate</category><category>qwen3-tts-flash</category><category>qwen-image-edit</category><category>qwen3coder</category><category>qwen</category><category>junyang_lin</category><category>eddie_wu</category><category>alibaba_wan</category><category>tool-use</category><category>large-model-coding</category><category>reasoning</category><category>multimodality</category><category>model-release</category><category>model-updates</category><category>industry-application</category><category>scaling</category><category>fine-tuning</category><category>reinforcement-learning</category></item><item><title>NVIDIA to invest $100B in OpenAI for 10GW of Vera Rubin rollout</title><link>https://news.smol.ai/issues/25-09-22-nvda-oai/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-22-nvda-oai/</guid><description>**NVIDIA** and **OpenAI** announced a landmark strategic partnership to deploy at least **10 gigawatts** of AI datacenters using NVIDIA&apos;s systems, with NVIDIA investing up to **$100 billion** progressively as each gigawatt is deployed, starting in the second half of 2026 on the Vera Rubin platform. This deal significantly impacts the AI infrastructure funding landscape, potentially supporting OpenAI&apos;s $300 billion commitment to Oracle. The announcement caused major stock market reactions, with NVIDIA&apos;s market cap surging by $170 billion. Additionally, advancements in deterministic inference for reinforcement learning and FP8 precision gains in GPU performance were highlighted by AI practitioners.</description><pubDate>Mon, 22 Sep 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>openai</category><category>oracle</category><category>intel</category><category>enfabrica</category><category>wayne</category><category>qwen3-omni</category><category>deepseek-v3.1</category><category>artificialanlys</category><category>gdb</category><category>gpu-infrastructure</category><category>deterministic-inference</category><category>reinforcement-learning</category><category>fp8-precision</category><category>gpu-performance</category><category>ai-infrastructure</category><category>strategic-partnerships</category><category>investment</category><category>datacenters</category><category>cuda-graphs</category><category>pipeline-parallelism</category><category>data-parallelism</category></item><item><title>Grok 4 Fast: Xai&apos;s distilled, 40% more token efficient, 2m context, 344 tok/s frontier model</title><link>https://news.smol.ai/issues/25-09-19-grok-4-fast/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-19-grok-4-fast/</guid><description>**xAI** announced **Grok 4 Fast**, a highly efficient model running at **344 tokens/second**, offering reasoning and nonreasoning modes and free trials on major platforms. **Meta** showcased its neural band and Ray-Ban Display with a live demo that experienced hiccups but sparked discussion on live hardware demos and integration challenges. **Meta** is also developing a first-party &quot;Horizon Engine&quot; for AI rendering and released Quest-native Gaussian Splatting capture tech. New model releases include **Mistral&apos;s Magistral 1.2**, a compact multimodal vision-language model with improved benchmarks and local deployment; **Moondream 3**, a 9B-parameter MoE VLM focused on efficient visual reasoning; **IBM&apos;s Granite-Docling-258M**, a document VLM for layout-faithful PDF to HTML/Markdown conversion; and **ByteDance&apos;s SAIL-VL2**, a vision-language foundation model excelling at multimodal understanding and reasoning at 2B and 8B parameter scales.</description><pubDate>Fri, 19 Sep 2025 05:44:39 GMT</pubDate><category>xai</category><category>meta-ai-fair</category><category>mistral-ai</category><category>ibm</category><category>bytedance</category><category>grok-4-fast</category><category>magistral-1.2</category><category>moondream-3</category><category>granite-docling-258m</category><category>sail-vl2</category><category>nearcyan</category><category>aidangomez</category><category>_akhaliq</category><category>vikhyatk</category><category>rohanpaul_ai</category><category>efficiency</category><category>reasoning</category><category>vision</category><category>multimodality</category><category>model-optimization</category><category>model-deployment</category><category>vision-encoders</category><category>model-architecture</category><category>model-training</category></item><item><title>Softbank, NVIDIA and US Govt take 2%, 5% and 10% of Intel, will develop Intel x86 RTX SOCs for consumer &amp; datacenters</title><link>https://news.smol.ai/issues/25-09-18-nvidia-intc/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-18-nvidia-intc/</guid><description>**Nvidia and Intel** announced a joint development partnership for multiple new generations of x86 products, marking a significant shift in the tech industry. This collaboration has been in the works for a year and impacts both consumer and data center markets, boosting hopes for Intel&apos;s Foundry business. On the AI hardware front, **Meta** showcased its neural band and Ray-Ban Display with a live demo that experienced hiccups but sparked discussion on live tech demos. Meta is also moving from Unity to its own Horizon Engine for AI rendering, including Gaussian splatting capture technology. In AI models, **Mistral** released Magistral 1.2, a compact multimodal vision-language model with improved benchmarks and local deployment capabilities, while **Moondream 3** previewed a 9B-parameter, 2B-active MoE VLM focused on efficient visual reasoning.</description><pubDate>Thu, 18 Sep 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>intel</category><category>meta-ai-fair</category><category>mistral-ai</category><category>magistral-1.2</category><category>moondream-3</category><category>nearcyan</category><category>_akhaliq</category><category>vikhyatk</category><category>multimodality</category><category>vision</category><category>model-optimization</category><category>model-efficiency</category><category>model-architecture</category><category>reinforcement-learning</category><category>fine-tuning</category><category>ai-hardware</category><category>gaussian-splatting</category><category>live-demo</category><category>visual-reasoning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-17-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-17-not-much/</guid><description>**Anthropic** published an in-depth postmortem on their August-September reliability issues. **OpenAI**&apos;s GPTeam achieved a perfect 12/12 score at the **ICPC 2025** World Finals, showcasing rapid progress in general-purpose reasoning and introducing controllable &quot;thinking time&quot; tiers for **gpt-5** in ChatGPT. **Google DeepMind**&apos;s **gemini-2.5-deep-think** earned a gold medal level at ICPC, solving 10/12 problems with advances in parallel thoughts, multi-step reasoning, and novel reinforcement learning techniques. OpenAI and Apollo Evaluations detected &quot;scheming&quot; behaviors in frontier models, emphasizing the need for chain-of-thought transparency and launching a $500K Kaggle challenge. GitHub launched an MCP server registry integrated with VS Code Insiders, with additional support from JetBrains and Hugging Face for open LLMs in Copilot Chat. Weaviate released a native Query Agent translating natural language to database operations with citations.</description><pubDate>Wed, 17 Sep 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>openai</category><category>google-deepmind</category><category>apollo-evaluations</category><category>github</category><category>hugging-face</category><category>weaviate</category><category>gpt-5</category><category>gemini-2.5-deep-think</category><category>sama</category><category>merettm</category><category>woj_zaremba</category><category>markchen90</category><category>esyudkowsky</category><category>reasoning</category><category>reinforcement-learning</category><category>alignment</category><category>chain-of-thought</category><category>model-evaluation</category><category>agent-frameworks</category><category>ide-integration</category><category>natural-language-to-sql</category><category>real-time-voice</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-16-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-16-not-much/</guid><description>**GPT-5 Codex** rollout shows strong agentic coding capabilities with some token bloat issues. IDEs like **VS Code Insiders** and **Cursor 1.6** enhance context windows and model integration. **vLLM 0.10.2** supports aarch64 and NVIDIA GB200 with performance improvements. **AMD ROCm** updates add modern attention, sparse MoE, and distributed inference. **TRL** introduces Context Parallelism for long-context training. Robotics and RL data pipelines improve with **Unsloth** and **LeRobotDataset v3**. **Qwen3-Next-80B** runs efficiently on Mac M4 Max with MLX. **Tencent&apos;s HunyuanImage 2.1** is a 17B bilingual text-to-image model with 2048×2048 resolution and restricted open weights.</description><pubDate>Tue, 16 Sep 2025 05:44:39 GMT</pubDate><category>openai</category><category>microsoft</category><category>perplexity-ai</category><category>huggingface</category><category>amd</category><category>tencent</category><category>lmstudio</category><category>gpt-5-codex</category><category>vllm-0.10.2</category><category>qwen3-next-80b</category><category>hunyuanimage-2.1</category><category>gdb</category><category>teknium1</category><category>finbarrtimbers</category><category>thsottiaux</category><category>theturingpost</category><category>pierceboggan</category><category>amandaksilver</category><category>aravsrinivas</category><category>sergiopaniego</category><category>art_zucker</category><category>danielhanchen</category><category>rwojo</category><category>awnihannun</category><category>agentic-ai</category><category>ide</category><category>context-windows</category><category>inference</category><category>distributed-inference</category><category>reinforcement-learning</category><category>robotics</category><category>long-context</category><category>model-optimization</category><category>text-to-image</category><category>multimodality</category><category>model-licenses</category></item><item><title>GPT-5 Codex launch and OpenAI&apos;s quiet rise in Agentic Coding</title><link>https://news.smol.ai/issues/25-09-15-gpt5-codex/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-15-gpt5-codex/</guid><description>**OpenAI** released **GPT-5-Codex**, an agentic coding model optimized for long-running software engineering tasks with dynamic task-adaptive thinking, multi-hour autonomy, and improved code quality. It achieves 51% accuracy on an unreleased large refactor benchmark and integrates deeply with developer tools like Xcode. Meanwhile, **Alibaba** launched **Qwen3-Next-80B**, a hybrid MoE model with native long-context support (262k tokens, extensible to 1M+), targeting efficient reasoning and repository-scale code analysis, supported by **Together AI** and **NVIDIA** with CUDA-accelerated attention. The trend towards hybrid SSM + MoE architectures is noted, emphasizing efficiency and scaling in China and US training regimes. Community discussions highlight the importance of variable compute and routing for inference efficiency and quality.</description><pubDate>Mon, 15 Sep 2025 05:44:39 GMT</pubDate><category>openai</category><category>alibaba</category><category>together-ai</category><category>nvidia</category><category>gpt-5-codex</category><category>qwen3-next-80b</category><category>sama</category><category>swyx</category><category>omarsar0</category><category>ofirpress</category><category>agentic-ai</category><category>software-engineering</category><category>long-context</category><category>mixture-of-experts</category><category>model-optimization</category><category>cuda-acceleration</category><category>inference-efficiency</category><category>routing</category><category>task-adaptive-thinking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-12-not-much/</guid><description>**Meta** released **MobileLLM-R1**, a sub-1B parameter reasoning model family on Hugging Face with strong small-model math accuracy, trained on 4.2T tokens. **Alibaba** introduced **Qwen3-Next-80B-A3B** with hybrid attention, 256k context window, and improved long-horizon memory, priced competitively on Alibaba Cloud. **Meta AI FAIR** fixed a benchmark bug in SWE-Bench affecting agent evaluation. LiveMCP-101 benchmark shows frontier models like **GPT-5** underperform on complex tasks with common failure modes cataloged. OpenAI highlights hallucination issues due to benchmark incentives, proposing calibration improvements. Community demos and tooling updates continue to evolve.</description><pubDate>Sat, 13 Sep 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>huggingface</category><category>alibaba</category><category>openai</category><category>mobilellm-r1</category><category>qwen3-next-80b-a3b</category><category>gpt-5</category><category>_akhaliq</category><category>tacocohen</category><category>pkirgis</category><category>sayashk</category><category>reasoning</category><category>model-efficiency</category><category>hybrid-attention</category><category>long-context</category><category>benchmarking</category><category>agent-evaluation</category><category>hallucination-detection</category><category>model-calibration</category><category>inference-complexity</category><category>model-pricing</category></item><item><title>Qwen3-Next-80B-A3B-Base: Towards Ultimate Training &amp; Inference Efficiency</title><link>https://news.smol.ai/issues/25-09-11-qwen3-next/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-11-qwen3-next/</guid><description>**MoE (Mixture of Experts) models** have become essential in frontier AI models, with **Qwen3-Next** pushing sparsity further by activating only **3.7% of parameters** (3B out of 80B) using a hybrid architecture combining **Gated DeltaNet** and **Gated Attention**. This new design includes **512 total experts** (10 routed + 1 shared), **Zero-Centered RMSNorm** for stability, and improved MoE router initialization, resulting in **~10× cheaper training and 10× faster inference** compared to previous models. **Alibaba&apos;s Qwen3-Next** reportedly outperforms **Gemini-2.5-Flash-Thinking** and approaches the flagship 235B model&apos;s performance, with deployments on **Hugging Face**, **Baseten**, and native **vLLM** support for efficient inference.</description><pubDate>Thu, 11 Sep 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>mistral-ai</category><category>deepseek</category><category>snowflake</category><category>hugging-face</category><category>baseten</category><category>nvidia</category><category>qwen3-next</category><category>qwen3</category><category>mixtral-8x7b</category><category>gemini-2.5-pro</category><category>justinlin610</category><category>teortaxestex</category><category>yuchenj_uw</category><category>mixture-of-experts</category><category>model-sparsity</category><category>gated-attention</category><category>hybrid-architecture</category><category>rmsnorm</category><category>model-stability</category><category>model-training</category><category>inference-optimization</category><category>multi-token-prediction</category><category>model-deployment</category></item><item><title>Oracle jumps +36% in a day after winning $300B OpenAI contract</title><link>https://news.smol.ai/issues/25-09-10-oci/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-10-oci/</guid><description>**Oracle&apos;s OCI division** reported a stunning **+359% revenue bookings growth to $455B** with cloud revenue guidance of **$144B by 2030**, driven significantly by a large deal with **OpenAI** amid tensions with **Microsoft**. On AI infrastructure, **Moonshot AI** released **Kimi’s checkpoint-engine**, enabling rapid weight updates on 1T-parameter models across thousands of GPUs, integrating with **vLLM**. **RLFactory** introduced a plug-and-play reinforcement learning framework for tool-using agents, showing smaller models outperforming larger ones. **TRL v0.23** added context parallelism for long-context training. **Thinking Machines Lab** published research on deterministic inference pipelines, making **vLLM** deterministic for **Qwen** models. **Meta** launched **BackendBench**, a PyTorch benchmarking tool.</description><pubDate>Wed, 10 Sep 2025 05:44:39 GMT</pubDate><category>oracle</category><category>openai</category><category>microsoft</category><category>moonshot-ai</category><category>vllm-project</category><category>thinking-machines-lab</category><category>meta</category><category>qwen3-235b</category><category>qwen3-4b</category><category>qwen2.5-7b</category><category>vllm</category><category>kimi_moonshot</category><category>arankomatsuzaki</category><category>qgallouedec</category><category>cHHillee</category><category>woosuk_k</category><category>stasbekman</category><category>reinforcement-learning</category><category>model-weight-updates</category><category>deterministic-inference</category><category>benchmarking</category><category>long-context</category><category>model-optimization</category><category>cuda</category><category>distributed-training</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-09-not-much/</guid><description>**Cognition** raised **$400M** at a **$10.2B** valuation to advance AI coding agents, with **swyx** joining to support the &quot;Decade of Agents&quot; thesis. **Vercel** launched an OSS &quot;vibe coding platform&quot; using a tuned **GPT-5** agent loop. **Claude Code** emphasizes minimalism in agent loops for reliability. **Kimi K2-0905** achieved 94% on coding evals and improved agentic capabilities with doubled context length. **Alibaba** released **Qwen3-ASR**, a multilingual transcription model with &lt;8% WER. **Meta** introduced Set Block Decoding for 3-5× faster decoding without architectural changes. Innovations in KV cache compression and quantization include **AutoRound**, **QuTLASS v0.1.0**, and **AlgoPerf v0.6**. **Google&apos;s Veo 3** video generation API went GA with significant price cuts and vertical video support.</description><pubDate>Tue, 09 Sep 2025 05:44:39 GMT</pubDate><category>cognition</category><category>founders-fund</category><category>lux-capital</category><category>8vc</category><category>neo</category><category>vercel</category><category>claude</category><category>groq</category><category>alibaba</category><category>huggingface</category><category>meta-ai-fair</category><category>google</category><category>theturingpost</category><category>algoperf</category><category>gpt-5</category><category>kimi-k2-0905</category><category>glm-4.5</category><category>qwen3-asr</category><category>opus-4.1</category><category>swyx</category><category>tim_dettmers</category><category>coding-agents</category><category>agent-architecture</category><category>open-source</category><category>model-evaluation</category><category>multilingual-models</category><category>speech-recognition</category><category>model-optimization</category><category>kv-cache</category><category>quantization</category><category>algorithmic-benchmarking</category><category>video-generation</category><category>context-windows</category></item><item><title>Cognition&apos;s $10b Series C; Smol AI updates</title><link>https://news.smol.ai/issues/25-09-08-cog-smol/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-08-cog-smol/</guid><description>**Cognition** raised **$400M** at a **$10.2B** valuation to advance AI coding agents, with **swyx** joining the company. **Vercel** launched an OSS coding platform using a tuned **GPT-5** agent loop. The **Kimi K2-0905** model achieved top coding eval scores and improved agentic capabilities with doubled context length. **Alibaba** released **Qwen3-ASR**, a multilingual transcription model with robust noise handling. **Meta** introduced Set Block Decoding for 3-5× faster decoding without architectural changes. Innovations in KV cache compression and quantization were highlighted, including **AutoRound** in SGLang and **QuTLASS v0.1.0** for Blackwell GPUs. Algorithmic benchmarking tools like **AlgoPerf v0.6** were updated for efficiency.</description><pubDate>Mon, 08 Sep 2025 05:44:39 GMT</pubDate><category>cognition</category><category>vercel</category><category>meta-ai-fair</category><category>alibaba</category><category>groq</category><category>huggingface</category><category>kimi-k2-0905</category><category>qwen3-asr</category><category>gpt-5</category><category>swyx</category><category>coding-agents</category><category>agent-development</category><category>open-source</category><category>model-evaluation</category><category>multilingual-models</category><category>inference-optimization</category><category>kv-cache-compression</category><category>quantization</category><category>algorithmic-benchmarking</category><category>context-length</category><category>model-performance</category></item><item><title>Kimi K2‑0905 and Qwen3‑Max preview: two 1T open weights models launched</title><link>https://news.smol.ai/issues/25-09-05-1t-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-05-1t-models/</guid><description>**Moonshot AI** updated their **Kimi K2-0905** open model with doubled context length to **256k tokens**, improved coding and tool-calling, and integration with agent scaffolds. **Alibaba** released **Qwen 3 Max**, a **1 trillion parameter** model with agent-oriented behavior, available via **Qwen Chat**, **Alibaba Cloud API**, and **OpenRouter**. The community highlights China&apos;s dominance in open models and debates around meaningful evaluation methods for code agents, emphasizing long-horizon and domain-specific evals. Influential voices like **@swyx** and **@karpathy** discuss the importance of practical evals and discriminator models for ranking outputs.</description><pubDate>Fri, 05 Sep 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>alibaba</category><category>huggingface</category><category>together-ai</category><category>groq</category><category>lmsys</category><category>openrouter</category><category>llamaindex</category><category>kimi-k2-0905</category><category>qwen-3-max</category><category>qwen-3</category><category>swyx</category><category>karpathy</category><category>willdepue</category><category>levie</category><category>bebischof</category><category>andrew_n_carr</category><category>bigeagle_xd</category><category>long-context</category><category>agents</category><category>coding</category><category>tool-use</category><category>model-evaluation</category><category>instruction-following</category><category>context-windows</category><category>semantic-search</category><category>discriminator-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-04-not-much/</guid><description>**Google DeepMind** released **EmbeddingGemma (308M)**, a small multilingual embedding model optimized for on-device retrieval-augmented generation and semantic search, supporting over 100 languages and running efficiently with quantization and EdgeTPU latency under 15ms. **Jina AI** introduced new code-focused embedding models (0.5B/1.5B) with GGUF quantization, achieving state-of-the-art retrieval across multiple languages and tasks. **LightOn** demonstrated large-scale retrieval training without distillation using contrastive training on billions of passages. **Hugging Face** released the **FineVision** dataset with 17.3M images and 9.5B answer tokens for vision-language model training, showing significant benchmark improvements. The **MiniCPM-V 4.5 (8B)** multimodal model reported surpassing **GPT-4o** and **Gemini-2.0 Pro** on OpenCompass benchmarks with innovative video token compression. Microsoft’s **VibeVoice TTS** and Stanford’s Mixture-of-Contexts video generation also featured. Additionally, a Stanford study benchmarked optimizers like Muon, Soap, Mars, and Sophia, finding diminishing speedups over AdamW at larger scales but advantages at smaller scales. The new ChatGPT branching feature was noted for its simplicity and popularity. *&quot;Everyone&apos;s a decacorn now.&quot;*</description><pubDate>Thu, 04 Sep 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>hugging-face</category><category>jina-ai</category><category>lighton</category><category>microsoft</category><category>stanford</category><category>openai</category><category>ollama</category><category>weaviate</category><category>langchain</category><category>llamaindex</category><category>embeddinggemma</category><category>qwen-2.5-coder</category><category>minicpm-v-4.5</category><category>gpt-4o</category><category>gemini-2.0-pro</category><category>osanseviero</category><category>_philschmid</category><category>tomaarsen</category><category>ollama</category><category>weaviate_io</category><category>lusxvr</category><category>andimarafioti</category><category>thibaudfrere</category><category>_akhaliq</category><category>clementdelangue</category><category>gordonwetzstein</category><category>konstmish</category><category>wen_kaiyue</category><category>percyliang</category><category>embeddings</category><category>retrieval-augmented-generation</category><category>quantization</category><category>multilingual-models</category><category>on-device-ai</category><category>semantic-search</category><category>contrastive-learning</category><category>dataset-release</category><category>vision</category><category>multimodality</category><category>video-generation</category><category>text-to-speech</category><category>optimizer-benchmarking</category><category>training-recipes</category><category>model-compression</category><category>video-token-compression</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-03-not-much/</guid><description>**Exa** raised a **$700m Series B**, **OpenPipe** was acquired by **Coreweave**, and **Statsig** and **Alex** were acquired by **OpenAI**. The **Agent/Client Protocol (ACP)** was introduced by the **Zed** team to standardize IDE-agent interoperability, supporting **Claude Code** and **Gemini** CLIs. **LangChain 1.0 alpha** unifies content blocks for reasoning and multimodal data. The **OSWorld Verified leaderboard** promotes reproducible evaluation of computer-use agents including **OpenAI** and **Anthropic** models. FAIR revealed coding agent cheating on **SWE-Bench Verified**. **PR Arena** hosts live coding agent competitions. Benchmarks like **GSO** and **Holistic Agent Leaderboard** test software optimization and web browsing tasks, with **Qwen3-Coder** and **Gemini 2.5 Flash** showing strong performance. Advances in reinforcement learning for tool use include **SimpleTIR** improving multi-turn tool use success rates and **UI-TARS-2** advancing GUI agents. The **DARLING** optimizer improves quality and diversity in reasoning and instruction following, while **DEPO** achieves data-efficient RLVR with significant speedups.</description><pubDate>Wed, 03 Sep 2025 05:44:39 GMT</pubDate><category>exa</category><category>openpipe</category><category>coreweave</category><category>statsig</category><category>openai</category><category>zed</category><category>claude</category><category>gemini</category><category>langchain</category><category>anthropic</category><category>fair</category><category>alibaba</category><category>hud-evals</category><category>claude-code</category><category>gemini</category><category>qwen3-coder</category><category>gemini-2.5-flash</category><category>zeddotdev</category><category>mathemagic1an</category><category>hwchase17</category><category>giffmana</category><category>gneubig</category><category>crystalsssup</category><category>sayashk</category><category>_philschmid</category><category>_akhaliq</category><category>jaseweston</category><category>agent-protocols</category><category>interoperability</category><category>standardization</category><category>agent-evaluation</category><category>coding-agents</category><category>software-optimization</category><category>web-browsing</category><category>reinforcement-learning</category><category>multi-turn-reasoning</category><category>optimizer-design</category><category>data-efficient-rlvr</category><category>leaderboards</category><category>benchmarking</category></item><item><title>Anthropic raises $13B at $183B Series F</title><link>https://news.smol.ai/issues/25-09-02-anthropic-f/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-02-anthropic-f/</guid><description>**Anthropic** achieved a **$183B post-money valuation** in Series F funding by September 2025, growing from about $1B run-rate in January to over **$5B run-rate** by August 2025. Their **Claude Code** product saw **&gt;10x usage growth** in three months and reached **$500M run-rate revenue**, serving over **300,000 business customers** with a nearly **7x increase in large accounts**. **Mistral AI** launched **Le Chat** with 20+ MCP connectors integrating with major SaaS platforms and persistent memory features. Benchmarking updates highlight **GPT-5** leading agent intelligence indices, with strong performances from **xAI&apos;s Grok** and **Anthropic&apos;s Claude** families. Reliability tooling and agent evaluation advances were shared by **Galileo**, **OpenPipe**, and others. **Zhipu/THUDM** open-sourced **Slime v0.1.0**, enhancing RL infrastructure behind **GLM-4.5** with significant decoding speed improvements and advanced tensor offload techniques.</description><pubDate>Tue, 02 Sep 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>mistral-ai</category><category>x-ai</category><category>salesforce</category><category>galileo</category><category>openpipe</category><category>zhipu</category><category>thudm</category><category>claude-code</category><category>gpt-5</category><category>grok-4</category><category>claude</category><category>sonnet-4</category><category>glm-4.5</category><category>deepseek-r1</category><category>swyx</category><category>emilygsands</category><category>_philschmid</category><category>_lewtun</category><category>omarsar0</category><category>_avichawla</category><category>corbtt</category><category>enterprise-connectors</category><category>agent-benchmarking</category><category>reinforcement-learning</category><category>inference-optimization</category><category>memory-optimization</category><category>cuda</category><category>multi-token-prediction</category><category>speculative-decoding</category><category>tensor-offload</category><category>performance-optimization</category><category>real-time-guardrails</category><category>cost-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-09-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-09-01-not-much/</guid><description>**OpenAI** integrates **GPT-5** into Xcode 26 with improved coding latency, though some UX trade-offs are noted. **xAI&apos;s Grok Code Fast 1** gains momentum, surpassing **Claude Sonnet** in usage and praised for fast debugging. **Zhipu&apos;s GLM-4.5** offers a cost-effective coding plan with strong performance against Claude Sonnet 4. **Meituan** releases the **LongCat-Flash-Chat**, a 560B parameter MoE model with adaptive compute and detailed technical insights. Apple debuts on-device vision-language models **FastVLM** and **MobileCLIP2** alongside **InternVL3.5**.</description><pubDate>Mon, 01 Sep 2025 05:44:39 GMT</pubDate><category>openai</category><category>x-ai</category><category>zhipu-ai</category><category>meituan</category><category>apple</category><category>gpt-5</category><category>grok-code-fast-1</category><category>claude-sonnet</category><category>glm-4.5</category><category>longcat-flash-chat</category><category>fastvlm</category><category>mobileclip2</category><category>internvl3.5</category><category>gdb</category><category>martin_casado</category><category>yanndubs</category><category>elonmusk</category><category>cline</category><category>vikhyatk</category><category>dzhng</category><category>quixiai</category><category>tim_dettmers</category><category>casper_hansen_</category><category>reach_vb</category><category>eliebakouch</category><category>teortaxestex</category><category>youjiacheng</category><category>model-architecture</category><category>moe</category><category>adaptive-compute</category><category>inference-speed</category><category>model-training</category><category>cost-efficiency</category><category>coding</category><category>developer-tools</category><category>open-inference</category><category>on-device-ai</category><category>vision</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-29-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-29-not-much/</guid><description>**Apple** released three real-time vision-language models (**FastVLM**, **MobileCLIP2**) on Hugging Face with significant speed and size improvements, supporting WebGPU and Core ML. Their MLX framework now supports **MXFP4** format, competing with **NVFP4** for FP4 quantization. **xAI** launched **grok-code-fast-1**, outperforming Claude for code edits, while **OpenAI** integrated **GPT-5** into Xcode 26 and released a new **Responses API** on **Groq** hardware. CLI-first agent workflows advanced with tools like **SemTools**, **MLX** local runner for Apple Silicon, and **llama.vim** recommending **Qwen 3 Coder 30B A3B**. Retrieval research highlights limitations of single-vector embeddings, promoting ColBERT-style late interaction.</description><pubDate>Fri, 29 Aug 2025 05:44:39 GMT</pubDate><category>apple</category><category>hugging-face</category><category>x-ai</category><category>openai</category><category>groq</category><category>run-llama</category><category>lmstudio</category><category>fastvlm</category><category>mobileclip2</category><category>grok-code-fast-1</category><category>gpt-5</category><category>qwen-3-coder-30b-a3b</category><category>reach_vb</category><category>xenovacom</category><category>pcuenq</category><category>awnihannun</category><category>cline</category><category>veggie_eric</category><category>nickbaumann_</category><category>gdb</category><category>benankdev</category><category>loganmarkewich</category><category>tom_doerr</category><category>fastmcp</category><category>ggerganov</category><category>orionweller</category><category>antoine_chaffin</category><category>vision</category><category>model-quantization</category><category>code-generation</category><category>cli-workflows</category><category>retrieval-augmentation</category><category>embedding-models</category><category>local-ai</category><category>multimodality</category></item><item><title>OpenAI Realtime API GA and new `gpt-realtime` model, 20% cheaper than 4o</title><link>https://news.smol.ai/issues/25-08-28-gpt-realtime/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-28-gpt-realtime/</guid><description>**OpenAI** launched the **gpt-realtime** model and **Realtime API** to GA, featuring advanced speech-to-speech capabilities, new voices (**Cedar**, **Marin**), image input, SIP telephony, and a ~20% price cut. Benchmarks show improvements over **gpt-4o-realtime** on BigBench and ComplexFuncBench. **xAI** introduced **Grok Code Fast 1**, a speed-optimized coding model integrated with popular IDEs, while **OpenAI Codex** received major upgrades for local and cloud development workflows. Google’s **Gemini CLI** improved multi-editor support, and new models like **Microsoft MAI-1-preview** and **MAI-Voice-1** were announced. *&quot;The new all-in-one WebRTC API removes the ephemeral token step and supports video on the same connection,&quot;* highlighting enhanced developer tooling.</description><pubDate>Thu, 28 Aug 2025 08:44:39 GMT</pubDate><category>openai</category><category>xai</category><category>microsoft</category><category>google</category><category>gpt-realtime</category><category>gpt-4o-realtime</category><category>grok-code-fast-1</category><category>codex</category><category>mai-1-preview</category><category>mai-voice-1</category><category>gemini-cli</category><category>swyx</category><category>juberti</category><category>omarsar0</category><category>reach_vb</category><category>pbbakkum</category><category>skcd42</category><category>mohitreddy13</category><category>cline</category><category>kevinweil</category><category>gdb</category><category>sama</category><category>_philschmid</category><category>speech-to-speech</category><category>instruction-following</category><category>function-calling</category><category>telephony</category><category>webrtc</category><category>voice-agents</category><category>multilingual-switching</category><category>voice-control</category><category>benchmarks</category><category>coding-models</category><category>ide-integration</category><category>developer-tools</category><category>model-updates</category></item><item><title>OpenAI updates Codex, VSCode Extension that can sync tasks with Codex Cloud</title><link>https://news.smol.ai/issues/25-08-27-codex-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-27-codex-2/</guid><description>**OpenAI Codex** has launched a new IDE Extension integrating with VS Code and Cursor, enabling seamless local and cloud task handoff, sign-in via ChatGPT plans, upgraded CLI, and GitHub code review automation. Facebook AI researchers introduced **StepWiser**, a process-level reward model improving reasoning and training by chunk-by-chunk evaluation, achieving SOTA on ProcessBench. **Google DeepMind&apos;s Gemini 2.5 Flash Image** model showcases advanced spatial reasoning, multi-image fusion, and developer tools including a browser extension for image remixing. NVIDIA revealed efficiency data on **Nemotron-CC-Math (133B)** and **Jet-Nemotron** models.</description><pubDate>Wed, 27 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>facebook-ai-fair</category><category>google-deepmind</category><category>nvidia</category><category>codex</category><category>stepwiser</category><category>gemini-2.5-flash</category><category>nemotron-cc-math</category><category>jet-nemotron</category><category>jaseweston</category><category>tesatory</category><category>benjamindekr</category><category>tokumin</category><category>fabianstelzer</category><category>officiallogank</category><category>process-reward-modeling</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>spatial-reasoning</category><category>multi-image-fusion</category><category>developer-tools</category><category>code-review</category><category>ide-extension</category><category>cli</category><category>cloud-computing</category><category>model-efficiency</category></item><item><title>nano-banana is Gemini‑2.5‑Flash‑Image, beating Flux Kontext by 170 Elo with SOTA Consistency, Editing, and Multi-Image Fusion</title><link>https://news.smol.ai/issues/25-08-26-nano-banana/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-26-nano-banana/</guid><description>**Google DeepMind** revealed **Gemini-2.5-Flash-Image-Preview**, a state-of-the-art image editing model excelling in **character consistency**, **natural-language edits**, and **multi-image composition**, dominating the Image Edit Arena with a ~170-180 Elo lead and over 2.5M votes. It is integrated into multiple platforms including Google AI Studio and third-party services. **Nous Research** released **Hermes 4**, an open-weight hybrid reasoning model focused on steerability and STEM benchmarks. **NVIDIA** launched **Nemotron Nano 9B V2**, a hybrid Mamba-Transformer with 128k context, top-performing under 10B parameters, and released a 6.6T-token pretraining subset. **InternVL3.5** introduced 32 vision-language models based on OpenAI&apos;s gpt-oss and Qwen3 backbones. **Ollama v0.11.7** added DeepSeek v3.1 support with hybrid thinking and Turbo mode preview.</description><pubDate>Tue, 26 Aug 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>nous-research</category><category>nvidia</category><category>openai</category><category>ollama</category><category>huggingface</category><category>openrouter</category><category>gemini-2.5-flash-image-preview</category><category>hermes-4</category><category>nemotron-nano-9b-v2</category><category>internvl3.5</category><category>gpt-oss</category><category>qwen3</category><category>deepseek-v3.1</category><category>sundarpichai</category><category>_philschmid</category><category>lmarena_ai</category><category>omarsar0</category><category>skirano</category><category>yupp_ai</category><category>xanderatallah</category><category>officiallogank</category><category>mervenoyann</category><category>image-editing</category><category>natural-language-processing</category><category>multi-image-composition</category><category>character-consistency</category><category>reasoning</category><category>hybrid-models</category><category>context-windows</category><category>model-steerability</category><category>pretraining</category><category>finetuning</category><category>alignment</category><category>vision</category><category>vision-language</category><category>api</category><category>model-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-25-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-25-not-much/</guid><description>**xAI** released open weights for **Grok-2** and **Grok-2.5** with a novel MoE residual architecture and μP scaling, sparking community excitement and licensing concerns. **Microsoft** open-sourced **VibeVoice-1.5B**, a multi-speaker long-form TTS model with streaming support and a 7B variant forthcoming. **Motif Technology** published a detailed report on **Motif-2.6B**, highlighting Differential Attention, PolyNorm, and extensive finetuning, trained on AMD MI250 GPUs. In coding tools, momentum builds around **GPT-5**-backed workflows, with developers favoring it over Claude Code. **Alibaba** released **Qwen-Code v0.0.8** with deep VS Code integration and MCP CLI enhancements. The MCP ecosystem advances with LiveMCP-101 stress tests, the universal MCP server &quot;Rube,&quot; and LangGraph Platform&apos;s rollout of revision queueing and ART integration for RL training of agents.</description><pubDate>Mon, 25 Aug 2025 05:44:39 GMT</pubDate><category>xai-org</category><category>microsoft</category><category>motif-technology</category><category>alibaba</category><category>huggingface</category><category>langchain-ai</category><category>grok-2</category><category>grok-2.5</category><category>vibevoice-1.5b</category><category>motif-2.6b</category><category>gpt-5</category><category>qwen-code</category><category>elonmusk</category><category>clementdelangue</category><category>rasbt</category><category>quanquangu</category><category>akhaliq</category><category>eliebakouch</category><category>gdb</category><category>ericmitchellai</category><category>ivanfioravanti</category><category>deanwball</category><category>giffmana</category><category>omarsar0</category><category>corbtt</category><category>mixture-of-experts</category><category>model-scaling</category><category>model-architecture</category><category>text-to-speech</category><category>fine-tuning</category><category>training-data</category><category>optimization</category><category>reinforcement-learning</category><category>agentic-ai</category><category>tool-use</category><category>model-training</category><category>model-release</category><category>api</category><category>software-development</category><category>model-quantization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-22-not-much/</guid><description>**DeepMind** released **Genie 3**, an interactive multimodal world simulator with advanced spatial memory and real-time avatar control, and **SIMA**, an embodied training agent operating inside generated worlds. **Alibaba** introduced **Qwen-Image-Edit**, an open-weights image editor scoring **ELO 1098 (#2)** in the Image Editing Arena, running on Qualcomm NPUs, alongside **Qwen-VL-Max** entering the Vision top-20. Video models like **Kling 2.1** showed a **235% improvement** in frame control, with new entrants **Luma Ray 2** and **Runway Gen-4 Turbo** debuting. **Google** provided free **Veo 3** generations in Gemini App and enhanced Google Photos with natural-language edits. **DeepSeek v3.1** launched with focus on SWE and Search agents, supporting local inference on Apple Silicon with 4-bit quantization achieving ~**21 tok/s** on M3 Ultra. The news highlights advances in interactive simulation, vision editing, video synthesis, and scalable local AI inference.</description><pubDate>Fri, 22 Aug 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>alibaba</category><category>google</category><category>deepseek</category><category>baseten</category><category>yupp</category><category>qwen-image-edit</category><category>qwen-vl-max</category><category>kling-2.1</category><category>veo-3</category><category>deepseek-v3.1</category><category>genie-3</category><category>sima</category><category>demishassabis</category><category>bonniesjli</category><category>shreyar</category><category>ostrisai</category><category>lmarena_ai</category><category>teortaxestex</category><category>ivanfioravanti</category><category>multimodality</category><category>embodied-ai</category><category>simulation</category><category>fine-tuning</category><category>quantization</category><category>video-generation</category><category>image-generation</category><category>local-inference</category><category>scaling</category><category>agent-training</category><category>real-time-control</category><category>spatial-memory</category></item><item><title>Cohere Command A Reasoning beats GPT-OSS-120B and DeepSeek R1 0528</title><link>https://news.smol.ai/issues/25-08-21-cohere-command-a-reasoning/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-21-cohere-command-a-reasoning/</guid><description>**Cohere&apos;s Command A Reasoning** model outperforms GPT-OSS in open deep research capabilities, emphasizing agentic use cases for 2025. **DeepSeek-V3.1** introduces a hybrid reasoning architecture toggling between reasoning and non-reasoning modes, optimized for agentic workflows and coding, with extensive long-context pretraining (~630B tokens for 32k context, ~209B for 128k), FP8 training, and a large MoE expert count (~37B). Benchmarks show competitive performance with notable improvements in SWE-Bench and other reasoning tasks. The model supports a $0.56/M input and $1.68/M output pricing on the DeepSeek API and enjoys rapid ecosystem integration including HF weights, INT4 quantization by Intel, and vLLM reasoning toggles. Community feedback highlights the hybrid design&apos;s pragmatic approach to agent and software engineering workflows, though some note the lack of tool use in reasoning mode.</description><pubDate>Thu, 21 Aug 2025 05:44:39 GMT</pubDate><category>cohere</category><category>deepseek</category><category>intel</category><category>huggingface</category><category>baseten</category><category>vllm-project</category><category>chutes-ai</category><category>anycoder</category><category>command-a-reasoning</category><category>deepseek-v3.1</category><category>artificialanlys</category><category>reach_vb</category><category>scaling01</category><category>cline</category><category>ben_burtenshaw</category><category>haihaoshen</category><category>jon_durbin</category><category>_akhaliq</category><category>willccbb</category><category>teortaxestex</category><category>agentic-ai</category><category>hybrid-models</category><category>long-context</category><category>fp8-training</category><category>mixture-of-experts</category><category>benchmarking</category><category>quantization</category><category>reasoning</category><category>coding-workflows</category><category>model-pricing</category></item><item><title>DeepSeek V3.1: 840B token continued pretrain, beating Claude 4 Sonnet at 11% of its cost</title><link>https://news.smol.ai/issues/25-08-20-deepseekv31/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-20-deepseekv31/</guid><description>**DeepSeek** released **DeepSeek V3.1**, a quietly rolled out open model with an **128K context window** and improvements in **token efficiency**, coding, and agentic benchmarks. **ByteDance** launched the permissive **Seed-OSS 36B** model on Hugging Face, noted for long-context and reasoning capabilities. **Zhipu AI** introduced **ComputerRL**, a reinforcement learning framework for computer-use agents, achieving strong benchmark results. In developer tooling, **GitHub Copilot** expanded globally, **Microsoft VS Code** integrated **Gemini 2.5 Pro** and updated **GPT-5** agent prompts, and **Anthropic** launched **Claude Code** seats with spend controls. Open-source fine-tuning advances include **Together AI** adding SFT for **gpt-oss-120B/20B** and **Baseten** enabling multinode 120B training with Truss CLI. The community noted mixed performance and ongoing post-training adjustments for DeepSeek V3.1.</description><pubDate>Wed, 20 Aug 2025 05:44:39 GMT</pubDate><category>deepseek</category><category>bytedance</category><category>zhipu-ai</category><category>github</category><category>microsoft</category><category>anthropic</category><category>together-ai</category><category>baseten</category><category>huggingface</category><category>deepseek-v3.1</category><category>seed-oss-36b</category><category>computerrl</category><category>gemini-2.5-pro</category><category>gpt-5</category><category>claude-code</category><category>gpt-oss-120b</category><category>gpt-oss-20b</category><category>teortaxestex</category><category>rasbt</category><category>lukehoban</category><category>burkeholland</category><category>_catwu</category><category>cline</category><category>winglian</category><category>token-efficiency</category><category>coding</category><category>agentic-benchmarks</category><category>long-context</category><category>reinforcement-learning</category><category>developer-tools</category><category>fine-tuning</category><category>multinode-training</category><category>model-release</category></item><item><title>Databricks&apos; $100B Series K</title><link>https://news.smol.ai/issues/25-08-19-databricks/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-19-databricks/</guid><description>**Databricks** reached a **$100 billion valuation**, becoming a centicorn with new Data ([Lakebase](https://www.databricks.com/product/lakebase)) and AI ([Agent Bricks](https://docs.databricks.com/aws/en/generative-ai/agent-bricks/)) products. **OpenAI** launched **ChatGPT Go** in India at ₹399/month (~$4.55), offering significantly increased usage limits and UPI payment support, with plans for global expansion. The **DeepSeek V3.1 Base/Instruct** models were quietly released on Hugging Face, showing strong coding benchmark performance and adopting an Anthropic-style hybrid system. The **Qwen-Image-Edit** model from **Alibaba** is gaining traction with integrations and community pruning experiments. *&quot;DeepSeek V3.1 Base outperforms Claude 4 Opus on coding benchmarks&quot;* and *&quot;ChatGPT Go offers 10x higher message limits and 2x longer memory&quot;* highlight key advancements.</description><pubDate>Tue, 19 Aug 2025 05:44:39 GMT</pubDate><category>databricks</category><category>openai</category><category>deepseek</category><category>hugging-face</category><category>alibaba</category><category>deepseek-v3.1-base</category><category>deepseek-v3.1-instruct</category><category>chatgpt-go</category><category>qwen-image-edit</category><category>sama</category><category>nickaturley</category><category>kevinweil</category><category>gdb</category><category>sherwinwu</category><category>nptacek</category><category>reach_vb</category><category>clementdelangue</category><category>teortaxestex</category><category>quixiai</category><category>georgejrjrjr</category><category>scaling01</category><category>alibaba_qwen</category><category>linoy_tsaban</category><category>ostrisai</category><category>lmarena_ai</category><category>model-release</category><category>benchmarking</category><category>pricing-models</category><category>fine-tuning</category><category>model-architecture</category><category>image-editing</category><category>video-generation</category><category>api</category><category>agentic-ai</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-18-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-18-not-much/</guid><description>**Gemma 3 270M**, an ultra-small model optimized for edge and mobile use, was released and is gaining adoption. **NVIDIA** launched two open multilingual ASR models, **Canary 1B** and **Parakeet-TDT 0.6B**, trained on 1 million hours of data with CC-BY licensing, plus the efficient **Nemotron-Nano v2 9B** model with significant speedups. **Alibaba&apos;s Qwen-Image-Edit** offers bilingual text editing and semantic image transformations. **Tencent Hunyuan** introduced a controllable game-world video generator trained on over 1 million gameplay recordings. **Meta&apos;s DINOv3** presents a scalable self-supervised vision backbone with strong domain transfer capabilities. **IBM** quietly released efficient English embedding models under a commercial-friendly license. The **BeyondWeb** synthetic data paper shows significant training speed and performance gains over prior datasets. Analysis of **HRM** architecture suggests performance improvements largely stem from data augmentation and scaffolding rather than novel architecture. *&quot;Models and datasets are openly licensed and available on Hugging Face.&quot;*</description><pubDate>Mon, 18 Aug 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>alibaba</category><category>tencent</category><category>meta-ai-fair</category><category>ibm</category><category>datology</category><category>gemma-3-270m</category><category>canary-1b</category><category>parakeet-tdt-0.6b</category><category>nemotron-nano-v2</category><category>qwen-image-edit</category><category>dino-v3</category><category>demishassabis</category><category>adrgrondin</category><category>rasbt</category><category>reach_vb</category><category>ctnzr</category><category>clementdelangue</category><category>natolambert</category><category>_akhaliq</category><category>itspaulai</category><category>mervenoyann</category><category>xenovacom</category><category>tomaarsen</category><category>pratyushmaini</category><category>code_star</category><category>leavittron</category><category>k_schuerholt</category><category>giffmana</category><category>synthetic-data</category><category>multilingual-asr</category><category>self-supervised-learning</category><category>vision</category><category>model-efficiency</category><category>training-data</category><category>data-augmentation</category><category>model-speedup</category><category>domain-transfer</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-15-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-15-not-much/</guid><description>**OpenAI** rolled out **GPT-5** as the default in ChatGPT with new modes and a &quot;warmer&quot; personality, plus expanded message limits for Plus/Team users and Enterprise/Edu access. Performance rankings show **gpt-5-high** leading, with smaller variants also ranked, though critiques note some underperformance versus Chinese models and sensitivity to sycophancy. OpenAI enhanced developer tools with a &quot;Quick eval&quot; feature, coding tips, and an improved Playground. **Google** released **Imagen 4** generally available with faster generation and higher resolution, plus the ultra-small **Gemma 3 270M** model with a large vocabulary and ecosystem support. Podcasts featured OpenAI leaders discussing GPT-5 systems, routing, and efficiency.</description><pubDate>Fri, 15 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>lmsys</category><category>gpt-5</category><category>gpt-5-high</category><category>gpt-5-mini-high</category><category>gpt-5-nano-high</category><category>imagen-4</category><category>gemma-3-270m</category><category>sama</category><category>aidan_mclau</category><category>kevinweil</category><category>lmarena_ai</category><category>edwinarbus</category><category>gdb</category><category>omarsar0</category><category>philschmid</category><category>m4rkmc</category><category>model-releases</category><category>model-performance</category><category>prompt-engineering</category><category>developer-tools</category><category>image-generation</category><category>model-optimization</category><category>transformers</category><category>tokenization</category><category>model-scaling</category></item><item><title>Western Open Models get Funding: Cohere $500m @ 6.8B, AI2 gets $152m NSF+NVIDIA grants</title><link>https://news.smol.ai/issues/25-08-14-cohere-ai2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-14-cohere-ai2/</guid><description>**OpenAI&apos;s GPT-5** achieved a speedrun of Pokemon Red 3x faster than **o3**. **Perplexity** raised **$200M** at a **$20B valuation**. **AI2** secured **$75M NSF grants** and **$77M from NVIDIA** for AI infrastructure projects like Olmo and Molmo. **Cohere** raised **$500M** and hired **Joelle Pineau** from **meta-ai-fair**, boosting models like Command A. **Google** released the **Gemma 3 270M** on-device tiny LLM with INT4 QAT checkpoints and large embedding tables, and made **Imagen 4** generally available with a fast version at $0.02/image. **Meta-ai-fair** introduced **DINOv3**, a family of self-supervised vision foundation models with high-resolution dense features and strong performance on benchmarks like COCO detection and ADE20K segmentation, under a permissive license. A **$150,000 MiniMax AI Agent Challenge** is ongoing with 200+ prizes, encouraging AI project builds by August 25.</description><pubDate>Thu, 14 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>perplexity-ai</category><category>ai2</category><category>nvidia</category><category>cohere</category><category>meta-ai-fair</category><category>google</category><category>hugging-face</category><category>ollama</category><category>unsloth</category><category>gpt-5</category><category>o3</category><category>command-a</category><category>gemma-3-270m</category><category>imagen-4</category><category>dinov3</category><category>joelle_pineau</category><category>fchollet</category><category>awnihannun</category><category>_philschmid</category><category>osanseviero</category><category>model-speed</category><category>funding</category><category>ai-infrastructure</category><category>on-device-ai</category><category>quantization</category><category>embedding-models</category><category>image-generation</category><category>self-supervised-learning</category><category>vision</category><category>dense-prediction</category><category>benchmarking</category><category>instruction-following</category><category>model-optimization</category><category>model-release</category><category>challenge</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-13-not-much/</guid><description>**OpenAI** continues small updates to **GPT-5**, introducing &quot;Auto/Fast/Thinking&quot; modes with **196k token context**, **3,000 messages/week**, and dynamic routing to cheaper models for cost efficiency. The **MiniMax AI Agent Challenge** offers **$150,000** in prizes for AI agent development by August 25. The community discusses **GPT-OSS-120B** base model extraction, hosting, and tooling improvements, including multi-tool pipelines and flex-attention. **Anthropic** announces model pairing in **Claude Code** with **Opus 4.1** for planning and **Sonnet 4** for execution, expanding context to **1M tokens** and introducing prompt caching. Key figures include *@sama*, *@jeremyphoward*, *@jxmnop*, and *@_catwu*.</description><pubDate>Wed, 13 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>minimax</category><category>gpt-5</category><category>gpt-oss-120b</category><category>opus-4.1</category><category>sonnet-4</category><category>sama</category><category>jeremyphoward</category><category>jxmnop</category><category>_catwu</category><category>context-windows</category><category>model-routing</category><category>model-hosting</category><category>multi-tool-pipelines</category><category>prompt-caching</category><category>model-extraction</category><category>model-pairing</category><category>cost-efficiency</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-12-not-much/</guid><description>**OpenAI** released the **GPT-5** series including **GPT-5-mini** and **GPT-5-nano**, with mixed user feedback on performance and API behavior. **Anthropic** extended **Claude Sonnet 4** context window to **1 million tokens**, a 5x increase, enhancing large document processing. **Zhipu AI** launched the open-source multimodal **GLM-4.5V** model with improvements in RL scaling and agentic tasks. **Google DeepMind** showcased the video generation model **Genie 3** and updated the **Gemini App** with new features like **Deep Think** and **Gemini Live**. **Alibaba Qwen** released the distilled image model **Qwen-Image distilled** and enhanced their Deep Research capabilities. Open source models like **Skywork&apos;s Matrix-Game 2.0** and **Jan.ai&apos;s Jan-v1** (built on **Qwen3-4B-Thinking**) were introduced, focusing on real-time world modeling and web search respectively. Developer tools such as **Claude Code** and **Cursor** were also highlighted.</description><pubDate>Tue, 12 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>zhipu-ai</category><category>google-deepmind</category><category>alibaba</category><category>skywork</category><category>jan-ai</category><category>gpt-5</category><category>gpt-5-mini</category><category>gpt-5-nano</category><category>claude-sonnet-4</category><category>glm-4.5v</category><category>genie-3</category><category>gemini-app</category><category>qwen-image-distilled</category><category>matrix-game-2.0</category><category>jan-v1</category><category>qwen3-4b-thinking</category><category>context-window</category><category>multimodality</category><category>reinforcement-learning</category><category>agentic-tasks</category><category>video-generation</category><category>image-generation</category><category>real-time-systems</category><category>web-search</category><category>model-accuracy</category><category>developer-tools</category><category>open-source-models</category><category>long-context</category><category>model-scaling</category></item><item><title>OpenAI&apos;s IMO Gold model also wins IOI Gold</title><link>https://news.smol.ai/issues/25-08-11-ioi-gold/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-11-ioi-gold/</guid><description>**OpenAI** announced placing **#6 among human coders** at the IOI, reflecting rapid progress in competitive coding AI over the past two years. The **GPT-5** launch faced significant user backlash over restrictive usage limits and removal of model selection control, leading to a reversal and increased limits to **3000 requests per week** for Plus users. Confusion around **GPT-5** naming and benchmarking was highlighted, with critiques on methodological issues comparing models like **Claude** and **Gemini**. Performance reviews of **GPT-5** are mixed, with claims of near-zero hallucinations by **OpenAI** staff but user reports of confidence in hallucinations and steering difficulties. Benchmarks show **GPT-5 mini** performing well on document understanding, while the full **GPT-5** is seen as expensive and middling. On the Chatbot Arena, **Gemini 2.5 Pro** holds a **67%** winrate against **GPT-5 Thinking**. Prompting and model behavior remain key discussion points.</description><pubDate>Mon, 11 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>gpt-5</category><category>gpt-5-thinking</category><category>gpt-5-mini</category><category>gemini-2.5-pro</category><category>claude</category><category>opus-4.1</category><category>sama</category><category>scaling01</category><category>yanndubs</category><category>sherylhsu</category><category>ahmed_el-kishky</category><category>jerry_tworek</category><category>noam_brown</category><category>alex_wei</category><category>amandaaskell</category><category>ericmitchellai</category><category>jon_durbin</category><category>gdb</category><category>jerryjliu0</category><category>reinforcement-learning</category><category>benchmarking</category><category>model-performance</category><category>prompt-engineering</category><category>model-behavior</category><category>competitive-programming</category><category>user-experience</category><category>model-naming</category><category>model-selection</category><category>hallucination-detection</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-08-not-much/</guid><description>**OpenAI** launched **GPT-5** with a unified user experience removing manual model selection, causing initial routing and access issues for Plus users that are being addressed with fixes including restored model options and increased usage limits. **GPT-5** introduces &quot;Priority Processing&quot; for lower latency at higher price tiers, achieving ~750ms median time-to-first-token in some cases. Microsoft reports full Copilot adoption of **GPT-5**, and API traffic doubled within 24 hours, peaking at 2 billion tokens per minute. Early benchmarks show **GPT-5** leading in reasoning tasks like FrontierMath and LiveBench, with improvements in hallucination control and creative writing, though some models like Grok-4 and Claude-4 Sonnet Thinking outperform it in specific RL-heavy reasoning benchmarks. OpenAI also released extensive migration and feature guides but faced some rollout issues including a broken code sample and a problematic Voice Mode launch. *&quot;Unified GPT-5&quot; ends model pickers, pushing developers away from manual model selection.*</description><pubDate>Fri, 08 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>microsoft</category><category>gpt-5</category><category>gpt-4o</category><category>grok-4</category><category>claude-4-sonnet</category><category>sama</category><category>nickaturley</category><category>elaineyale6</category><category>scaling01</category><category>mustafasuleyman</category><category>kevinweil</category><category>omarsar0</category><category>jeremyphoward</category><category>juberti</category><category>epochairesearch</category><category>lechmazur</category><category>gdb</category><category>reasoning</category><category>latency</category><category>model-routing</category><category>benchmarking</category><category>reinforcement-learning</category><category>hallucination-control</category><category>creative-writing</category><category>priority-processing</category><category>api-traffic</category><category>model-deprecation</category><category>user-experience</category><category>model-selection</category><category>voice-mode</category><category>documentation</category></item><item><title>OpenAI rolls out GPT-5 and GPT-5 Thinking to &gt;1B users worldwide; -mini and -nano help claim Pareto Frontier</title><link>https://news.smol.ai/issues/25-08-07-gpt-5/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-07-gpt-5/</guid><description>**OpenAI** launched **GPT-5**, a unified system featuring a fast main model and a deeper thinking model with a real-time router, supporting up to **400K context length** and aggressive pricing that reclaims the Pareto Frontier of Intelligence. The rollout includes variants like **gpt-5-mini** and **gpt-5-nano** with significant cost reductions, and integrations with products such as **ChatGPT**, **Cursor AI**, **JetBrains AI Assistant**, **Microsoft Copilot**, **Notion AI**, and **Perplexity AI**. Benchmarks show GPT-5 performing strongly in coding and long-context reasoning, roughly matching **Claude 4.1 Sonnet/Opus** on SWE-bench Verified. The launch was accompanied by a GPT-5 prompting cookbook and notable community discussions on pricing and performance.</description><pubDate>Thu, 07 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>cursor_ai</category><category>jetbrains</category><category>microsoft</category><category>notion</category><category>perplexity_ai</category><category>factoryai</category><category>gpt-5</category><category>gpt-5-mini</category><category>gpt-5-nano</category><category>claude-4.1-sonnet</category><category>claude-4.1-opus</category><category>sama</category><category>scaling01</category><category>jeffintime</category><category>embirico</category><category>mustafasuleyman</category><category>cline</category><category>lmarena_ai</category><category>nrehiew_</category><category>ofirpress</category><category>sauers_</category><category>model-architecture</category><category>context-windows</category><category>pricing-models</category><category>coding</category><category>long-context</category><category>prompt-engineering</category><category>model-benchmarking</category><category>model-integration</category><category>tool-use</category><category>reasoning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-08-06-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-06-not-much/</guid><description>**OpenAI** released its first open models since GPT-2, **gpt-oss-120b** and **gpt-oss-20b**, which quickly trended on **Hugging Face**. **Microsoft** supports these models via **Azure AI Foundry** and **Windows Foundry Local**. Key architectural innovations include **sliding window attention**, **mixture of experts (MoE)**, a **RoPE variant**, and a **256k context length**. The models use a new **MXFP4** format supported by **llama.cpp**. Hypotheses suggest **gpt-oss** was trained on **synthetic data** to enhance safety and performance, supporting the **Reasoning Core Hypothesis**. **OpenAI** announced a **$500K bounty** for red teaming with partners including **Anthropic**, **Google**, and the **UK AISI**. Performance critiques highlight inconsistent benchmarking results, with **GPT-OSS-120B** scoring **41.8%** on the **Aider Polyglot** coding benchmark, trailing competitors like **Kimi-K2** and **DeepSeek-R1**. Some users note the model excels in math and reasoning but lacks common sense and practical utility.</description><pubDate>Wed, 06 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>huggingface</category><category>microsoft</category><category>llamaindex</category><category>ollama</category><category>baseten</category><category>fireworksai</category><category>cerebras</category><category>groq</category><category>together</category><category>anthropic</category><category>google</category><category>uk-aisi</category><category>gpt-oss-120b</category><category>gpt-oss-20b</category><category>kimi-k2</category><category>deepseek-r1</category><category>qwen-3-32b</category><category>woj_zaremba</category><category>sama</category><category>huybery</category><category>drjimfan</category><category>jxmnop</category><category>scaling01</category><category>arunv30</category><category>kevinweil</category><category>xikun_zhang_</category><category>jerryjliu0</category><category>ollama</category><category>basetenco</category><category>reach_vb</category><category>gneubig</category><category>shxf0072</category><category>_lewtun</category><category>sliding-window-attention</category><category>mixture-of-experts</category><category>rope</category><category>context-length</category><category>mxfp4-format</category><category>synthetic-data</category><category>reasoning-core-hypothesis</category><category>red-teaming</category><category>benchmarking</category><category>coding-benchmarks</category><category>model-performance</category><category>fine-tuning</category></item><item><title>OpenAI&apos;s gpt-oss 20B and 120B, Claude Opus 4.1, DeepMind Genie 3</title><link>https://news.smol.ai/issues/25-08-05-gpt-oss/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-05-gpt-oss/</guid><description>**OpenAI** released the **gpt-oss** family, including **gpt-oss-120b** and **gpt-oss-20b**, their first open-weight models since GPT-2, designed for agentic tasks and licensed under **Apache 2.0**. These models use a **Mixture-of-Experts (MoE)** architecture with wide vs. deep design and innovative features like bias units in attention and a unique swiglu variant. The **120B** model was trained with about **2.1 million H100 GPU hours**. Meanwhile, **Anthropic** launched **claude-4.1-opus**, touted as the best coding model currently. **DeepMind** showcased **genie-3**, a realtime world simulation model with minute-long consistency. The releases highlight advances in open-weight models, reasoning capabilities, and world simulation. Key figures like **@sama**, **@rasbt**, and **@SebastienBubeck** provided technical insights and performance evaluations, noting strengths and hallucination risks.</description><pubDate>Tue, 05 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>gpt-oss-120b</category><category>gpt-oss-20b</category><category>gpt-oss</category><category>claude-4.1-opus</category><category>claude-4.1</category><category>genie-3</category><category>sama</category><category>rasbt</category><category>sebastienbubeck</category><category>polynoamial</category><category>kaicathyc</category><category>finbarrtimbers</category><category>vikhyatk</category><category>scaling01</category><category>teortaxestex</category><category>mixture-of-experts</category><category>model-architecture</category><category>agentic-ai</category><category>model-training</category><category>model-performance</category><category>reasoning</category><category>hallucination-detection</category><category>gpu-optimization</category><category>open-weight-models</category><category>realtime-simulation</category></item><item><title>Qwen-Image: SOTA text rendering + 4o-imagegen-level Editing Open Weights MMDiT</title><link>https://news.smol.ai/issues/25-08-04-qwen-image/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-04-qwen-image/</guid><description>**Alibaba** surprised with the release of **Qwen-Image**, a **20B MMDiT** model excelling at bilingual text rendering and graphic poster creation, with open weights and demos available. **Google DeepMind** launched **Gemini 2.5 Deep Think** to Ultra subscribers, showing significant reasoning improvements and benchmark gains (+11.2% AIME, +13.2% HLE, +13.4% LiveCodeBench) rivaling **OpenAI&apos;s o3 Pro**. ByteDance&apos;s **SeedProver** achieved state-of-the-art math theorem proving results, surpassing DeepMind&apos;s AlphaGeometry2. OpenAI is developing a &quot;universal verifier&quot; for math and coding gains transfer. Competitive reasoning benchmarks and game arenas by Google and Kaggle highlight a meta-shift in reasoning model efficiency, comparable to the original Transformer leap. Other open-weight models gaining momentum include **GLM-4.5**, **XBai o4**, and **Tencent Hunyuan** with a focus on efficient training. *&quot;Qwen is all you need.&quot;*</description><pubDate>Mon, 04 Aug 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>google-deepmind</category><category>openai</category><category>bytedance</category><category>kaggle</category><category>tencent</category><category>qwen-image</category><category>mmdit</category><category>gemini-2.5</category><category>o3-pro</category><category>seedprover</category><category>glm-4.5</category><category>xbai-o4</category><category>hunyuan</category><category>swyx</category><category>demishassabis</category><category>tulseedoshi</category><category>mparakhin</category><category>teortaxestex</category><category>cgeorgiaw</category><category>dorialexander</category><category>steph_palazzolo</category><category>corbtt</category><category>synthwavedd</category><category>epochairesearch</category><category>bilingual-text-rendering</category><category>image-generation</category><category>image-editing</category><category>synthetic-data</category><category>reasoning</category><category>math-theorem-proving</category><category>benchmarking</category><category>instruction-following</category><category>model-efficiency</category><category>open-weight-models</category><category>model-transparency</category><category>competitive-evaluation</category></item><item><title>Gemini 2.5 Deep Think finally ships</title><link>https://news.smol.ai/issues/25-08-01-deep-think/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-08-01-deep-think/</guid><description>**OpenAI** is rumored to soon launch new **GPT-OSS** and **GPT-5** models amid drama with **Anthropic** revoking access to **Claude**. **Google DeepMind** quietly launched **Gemini 2.5 Deep Think**, a model optimized for parallel thinking that achieved gold-medal level at the IMO and excels in reasoning, coding, and creative tasks. Leaks suggest **OpenAI** is developing a **120B MoE** and a **20B** model with advanced attention mechanisms. Chinese AI companies like **Kimi Moonshot**, **Alibaba**, and **ZHIpu AI** are releasing faster and more capable open models such as **kimi-k2-turbo-preview**, **Qwen3-Coder-Flash**, and **GLM-4.5**, signaling strong momentum and potential to surpass the U.S. in AI development. *&quot;The final checkpoint was selected just 5 hours before the IMO problems were released,&quot;* highlighting rapid development cycles.</description><pubDate>Fri, 01 Aug 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>kimi-moonshot</category><category>alibaba</category><category>ollama</category><category>zhipu-ai</category><category>stepfun</category><category>gemini-2.5-deep-think</category><category>gpt-oss</category><category>gpt-5</category><category>kimi-k2-turbo-preview</category><category>qwen3-coder-flash</category><category>glm-4.5</category><category>step-3</category><category>claude</category><category>demishassabis</category><category>philschmid</category><category>scaling01</category><category>teortaxestex</category><category>teknium1</category><category>lmarena_ai</category><category>andrewyng</category><category>parallel-thinking</category><category>model-releases</category><category>moe</category><category>attention-mechanisms</category><category>multimodal-reasoning</category><category>model-performance</category><category>context-windows</category><category>open-source-models</category><category>model-leaks</category><category>creative-ai</category><category>coding</category><category>reasoning</category><category>model-optimization</category></item><item><title>Figma&apos;s $50+b IPO</title><link>https://news.smol.ai/issues/25-07-31-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-31-not-much/</guid><description>**OpenAI**&apos;s stealth model **horizon-alpha** on **OpenRouter** sparks speculation as a precursor to **GPT-5**, showing strong reasoning and SVG generation capabilities, comparable to **Gemini 2.5 Pro**. **Alibaba** released the **Qwen3-Coder** family, including a fast **Qwen3-Coder-Flash (30B-A3B)** variant with agentic features and 1M context length support via **UnslothAI**. **Cohere** launched **Command A Vision**, a 111B parameter open-weights vision-language model outperforming **GPT-4.1** and **Llama 4 Maverick** on enterprise benchmarks. **Black Forest Labs** introduced **FLUX.1 Krea [dev]**, an open-weights photorealism model compatible with fine-tuning tools like **diffusers** and **ostrisai**. **Zhipu AI** unveiled **GLM-4.5**, a hybrid reasoning open model with agentic capabilities available on **Together AI**. Discussions highlight the rising importance of **inference-time training** and **reasoning model generalization**. **Mistral AI** released the technical report for **Voxtral** continuing its open science efforts.</description><pubDate>Thu, 31 Jul 2025 05:44:39 GMT</pubDate><category>openai</category><category>openrouter</category><category>alibaba</category><category>unslothai</category><category>cohere</category><category>huggingface</category><category>black-forest-labs</category><category>diffusers</category><category>ostrisai</category><category>zhipu-ai</category><category>together-ai</category><category>mistral-ai</category><category>horizon-alpha</category><category>gpt-5</category><category>gemini-2.5-pro</category><category>qwen3-coder</category><category>qwen3-coder-flash-30b-a3b</category><category>command-a-vision</category><category>gpt-4.1</category><category>llama-4-maverick</category><category>flux-1-krea-dev</category><category>glm-4.5</category><category>voxtral</category><category>scaling01</category><category>teortaxestex</category><category>huybery</category><category>nickfrosst</category><category>aidangomez</category><category>reach_vb</category><category>zai_org</category><category>corbtt</category><category>jxmnop</category><category>teknuim1</category><category>reasoning</category><category>svg-generation</category><category>agentic-ai</category><category>context-windows</category><category>vision</category><category>fine-tuning</category><category>inference-time-training</category><category>model-generalization</category><category>open-models</category><category>technical-reports</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-30-not-much/</guid><description>**Chinese AI labs** have released powerful open-source models like **GLM-4.5** and **GLM-4.5-Air** from **Zhipu AI**, **Qwen3 Coder** and **Qwen3-235B** from **Alibaba**, and **Kimi K2** from **Moonshot AI**, highlighting a surge in permissively licensed models. **Zhipu AI&apos;s GLM-4.5** is a 355B parameter MoE model competitive with **Claude 4 Opus** and **Gemini 2.5 Pro**. **Alibaba&apos;s Qwen3 Coder** shows strong code generation performance with a low edit failure rate, while **Moonshot AI&apos;s Kimi K2** is a 1 trillion-parameter MoE model surpassing benchmarks like **LiveCodeBench**. In video and image generation, **xAI** launched **Grok Imagine**, and **Wan2.2** impressed with innovative image-to-video generation. Robotics advances include **Figure&apos;s Figure-01 and Figure-02** humanoid robots and **ViTPose++** for pose estimation in basketball analysis. **SmolLM3** training and evaluation code was fully released under Apache 2.0. **OpenAI** introduced **Study Mode** in **ChatGPT** to enhance interactive learning, and **Runway** rolled out **Runway Aleph**, a new in-context video model for multi-task visual generation. The community notes a competitive disadvantage for organizations avoiding these Chinese open-source models. *&quot;Orgs avoiding these models are at a significant competitive disadvantage,&quot;* noted by @corbtt.</description><pubDate>Wed, 30 Jul 2025 05:44:39 GMT</pubDate><category>zhipu-ai</category><category>alibaba</category><category>moonshot-ai</category><category>x-ai</category><category>figure</category><category>openai</category><category>runway</category><category>mlx</category><category>ollama</category><category>deeplearningai</category><category>glm-4.5</category><category>glm-4.5-air</category><category>qwen3-coder</category><category>qwen3-235b</category><category>kimi-k2</category><category>grok-imagine</category><category>wan-2.2</category><category>smollm3</category><category>figure-01</category><category>figure-02</category><category>vitpose++</category><category>chatgpt</category><category>yuchenj_uw</category><category>corbtt</category><category>reach_vb</category><category>ollama</category><category>deeplearningai</category><category>gdb</category><category>sama</category><category>c_valenzuelab</category><category>adcock_brett</category><category>skalskip92</category><category>loubnabenallal1</category><category>hojonathanho</category><category>ostrisai</category><category>model-releases</category><category>model-performance</category><category>moe</category><category>image-generation</category><category>video-generation</category><category>pose-estimation</category><category>robotics</category><category>training-code-release</category><category>interactive-learning</category><category>in-context-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-29-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-29-not-much/</guid><description>**Chinese labs** have released a wave of powerful, permissively licensed models in July, including **Zhipu AI&apos;s GLM-4.5** and **GLM-4.5-Air**, **Alibaba&apos;s Qwen3 Coder** and **Qwen3-235B**, and **Moonshot AI&apos;s Kimi K2**. These models feature large-scale Mixture of Experts architectures with active parameters ranging from 3B to 32B and context windows up to 256K tokens. **Zhipu AI&apos;s GLM-4.5** competes with **Claude 4 Opus** and **Gemini 2.5 Pro** in benchmarks. **Moonshot AI&apos;s Kimi K2** is a 1 trillion-parameter MoE model surpassing other open-weight models on **LiveCodeBench** and **AceBench**. In video and image generation, **xAI** launched **Grok Imagine**, and **Wan2.2** impressed with its Image-to-Video approach. **Ideogram** released a character consistency model. Robotics advances include **Figure&apos;s Figure-01 and Figure-02** humanoid robots and **ViTPose++** for pose estimation in basketball analysis. The **SmolLM3** training and evaluation code was fully released under an Apache 2.0 license. *&quot;Orgs avoiding these Chinese open-source models are at a significant competitive disadvantage,&quot;* noted by @corbtt.</description><pubDate>Tue, 29 Jul 2025 05:44:39 GMT</pubDate><category>zhipu-ai</category><category>alibaba</category><category>moonshot-ai</category><category>x-ai</category><category>ideogram</category><category>figure</category><category>smollm</category><category>openai</category><category>glm-4.5</category><category>glm-4.5-air</category><category>qwen3-coder</category><category>qwen3-235b</category><category>kimi-k2</category><category>wan-2.2</category><category>grok-imagine</category><category>smollm3</category><category>figure-01</category><category>figure-02</category><category>vitpose++</category><category>yuchenj_uw</category><category>corbtt</category><category>cline</category><category>reach_vb</category><category>ollama</category><category>deeplearningai</category><category>ostrisai</category><category>hojonathanho</category><category>adcock_brett</category><category>skalskip92</category><category>loubnabenallal1</category><category>model-releases</category><category>moe</category><category>model-benchmarking</category><category>image-generation</category><category>video-generation</category><category>pose-estimation</category><category>robotics</category><category>training-code-release</category><category>apache-license</category></item><item><title>GLM-4.5: Deeper, Headier, &amp; better than Kimi/Qwen/DeepSeek (SOTA China LLM?)</title><link>https://news.smol.ai/issues/25-07-28-glm-45/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-28-glm-45/</guid><description>**Z.ai** (Zhipu AI) released the **GLM-4.5-355B-A32B** and **GLM-4.5-Air-106B-A12B** open weights models, claiming state-of-the-art performance competitive with **Claude 4 Opus**, **Grok 4**, and OpenAI&apos;s **o3**. These models emphasize token efficiency and efficient reinforcement learning training validated by the Muon optimizer. **Alibaba Qwen** introduced **Group Sequence Policy Optimization (GSPO)**, a new reinforcement learning algorithm powering the **Qwen3** model suite, integrated into Hugging Face&apos;s TRL library. Speculation surrounds mystery models &quot;summit&quot; and &quot;zenith&quot; as potential **GPT-5** variants based on **GPT-4.1** architecture. **Qwen3-Coder** shows strong coding benchmark results, rivaling **Claude Sonnet 4** and **Kimi K2**. The rise of powerful Chinese open-source models like **GLM-4.5**, **Wan-2.2**, and **Qwen3 Coder** contrasts with a slowdown from Western labs such as **OpenAI**.</description><pubDate>Mon, 28 Jul 2025 05:44:39 GMT</pubDate><category>z-ai</category><category>alibaba</category><category>huggingface</category><category>openai</category><category>glm-4.5-355b-a32b</category><category>glm-4.5-air-106b-a12b</category><category>qwen3-coder</category><category>claude-4-opus</category><category>grok-4</category><category>o3</category><category>gpt-4.1</category><category>gpt-5</category><category>kimi-k2</category><category>claude-sonnet-4</category><category>lupantech</category><category>teortaxestex</category><category>mervenoyann</category><category>_lewtun</category><category>scaling01</category><category>cline</category><category>reinforcement-learning</category><category>token-efficiency</category><category>model-optimization</category><category>open-source-models</category><category>agentic-ai</category><category>coding</category><category>model-training</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-25-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-25-not-much/</guid><description>**OpenAI** has fully rolled out its ChatGPT agent to all Plus, Pro, and Team users and is building hype for the upcoming **GPT-5**, which reportedly outperforms **Grok-4** and can build a cookie clicker game in two minutes. **Alibaba&apos;s Qwen** team released the open-source reasoning model **Qwen3-235B-Thinking**, achieving an **89%** win rate over **gpt4-0314** using a new RL algorithm called **Group Sequence Policy Optimization (GSPO)**. **Runway** introduced **Runway Aleph**, a state-of-the-art in-context video model for editing and generating video content. **Hugging Face** highlights the growing momentum of open-source AI, especially from Chinese teams. Other updates include **Kling&apos;s** upgrades for image-to-video generation and **Google&apos;s Imagen 4 Ultra** being recognized as a top text-to-image model. **Anthropic** integrated **Claude** with **Canva** for branded visual designs but faces stability issues. The **PyTorch** team released optimized checkpoints for **SmolLM3** to speed up inference.</description><pubDate>Fri, 25 Jul 2025 05:44:39 GMT</pubDate><category>openai</category><category>alibaba</category><category>runway</category><category>hugging-face</category><category>google</category><category>anthropic</category><category>pytorch</category><category>lmarena</category><category>gpt-5</category><category>gpt4-0314</category><category>qwen3-235b-thinking</category><category>runway-aleph</category><category>imagen-4-ultra</category><category>smollm3</category><category>grok-4</category><category>sama</category><category>clementdelangue</category><category>xikun_zhang_</category><category>teknnium1</category><category>chujiezheng</category><category>reinforcement-learning</category><category>reasoning</category><category>video-generation</category><category>image-generation</category><category>model-optimization</category><category>open-source</category><category>model-performance</category><category>inference-speed</category><category>integration</category><category>stability</category></item><item><title>3x in 3 months: Cursor @ $28b, Cognition + Windsurf @ $10b</title><link>https://news.smol.ai/issues/25-07-24-cogsurf-cursor/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-24-cogsurf-cursor/</guid><description>**Cursor** is reportedly fundraising at a **$28 billion valuation with $1 billion ARR**, while the combined **Cognition+Windsurf** entity is fundraising at a **$10 billion valuation** after acquiring Windsurf remainco for $300 million. The competition between AI coding agents intensifies as Cursor focuses on Async SWE Agents and Cognition+Windsurf acquires an agentic IDE. **Alibaba&apos;s Qwen3-Coder** gains widespread adoption for coding tasks and integration into tools like **Claude Code** and **LM Studio**. **OpenAI** rolls out **ChatGPT Agent** to all Plus, Pro, and Team users, sparking discussions about an &quot;agentic economy&quot; emphasizing **AI literacy**. **Anthropic&apos;s Claude Code** is praised as a premier development tool with active community feedback. **Perplexity&apos;s Comet browser assistant** receives positive reviews and new feature showcases. The debate continues on whether AI coding tools will replace developers, with critiques highlighting the ongoing human effort required. A new minimalistic software engineering agent, **mini**, achieves 65% on SWE-bench with just 100 lines of code.</description><pubDate>Thu, 24 Jul 2025 05:44:39 GMT</pubDate><category>cursor</category><category>cognition</category><category>windsurf</category><category>alibaba</category><category>openai</category><category>anthropic</category><category>perplexity</category><category>qwen3-coder</category><category>chatgpt-agent</category><category>claude-code</category><category>mini</category><category>bindureddy</category><category>xikun_zhang_</category><category>aravsrinivas</category><category>gergelyorosz</category><category>jeremyphoward</category><category>agentic-ai</category><category>fundraising</category><category>software-engineering</category><category>ai-coding</category><category>agentic-economy</category><category>model-integration</category><category>community-feedback</category><category>performance-benchmarking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-23-not-much/</guid><description>**Alibaba** announced the release of **Qwen3-Coder-480B-A35B-Instruct**, an open agentic code model with **480B** parameters and **256K** context length, praised for rapid development and strong coding performance. Benchmark claims of **41.8% on ARC-AGI-1** faced skepticism from **Franois Chollet** and others due to reproducibility issues. The model quickly integrated into ecosystems like **vLLM**, **Dynamic GGUFs**, and **OpenRouterAI**. The **White House** unveiled a new **AI Action Plan** emphasizing **Innovation**, **Infrastructure**, and **International Diplomacy**, linking AI leadership to national security and prioritizing compute access for the **Department of Defense**. The plan sparked debate on open vs. closed-source AI, with calls from **Clement Delangue** to embrace open science to maintain US AI competitiveness.</description><pubDate>Wed, 23 Jul 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>openrouterai</category><category>togethercompute</category><category>vllm_project</category><category>unslothai</category><category>white-house</category><category>qwen3-coder-480b-a35b-instruct</category><category>kimi-k2</category><category>fchollet</category><category>clementdelangue</category><category>scaling01</category><category>aravsrinivas</category><category>rasbt</category><category>gregkamradt</category><category>yuchenj_uw</category><category>code-generation</category><category>benchmarking</category><category>model-integration</category><category>context-windows</category><category>open-source</category><category>national-security</category><category>infrastructure</category><category>ai-policy</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-22-not-much/</guid><description>**Moonshot AI** released the **Kimi K2**, a 1-trillion parameter ultra-sparse Mixture-of-Experts (MoE) model with the **MuonClip** optimizer and a large-scale agentic data pipeline using over **20,000 tools**. Shortly after, **Alibaba** updated its **Qwen3** model with the **Qwen3-235B-A22B** variant, which outperforms Kimi K2 and other top models on benchmarks like **GPQA** and **AIME** despite being 4.25x smaller. Alibaba also released **Qwen3-Coder-480B-A35B**, a MoE model specialized for coding with a 1 million token context window. **Google DeepMind** launched **Gemini 2.5 Flash-Lite**, a faster and more cost-efficient model outperforming previous versions in coding, math, and multimodal tasks. The MoE architecture is becoming mainstream, with models like **Mistral**, **DeepSeek**, and **Kimi K2** leading the trend. In mathematics, an advanced **Gemini** model achieved a gold medal level score at the **International Mathematical Olympiad (IMO)**, marking a first for AI. An **OpenAI** researcher noted their IMO model &quot;knew&quot; when it did not have a correct solution, highlighting advances in model reasoning and self-awareness.</description><pubDate>Tue, 22 Jul 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>alibaba</category><category>google</category><category>google-deepmind</category><category>openai</category><category>hugging-face</category><category>vllm-project</category><category>kimi-k2</category><category>qwen3-235b-a22b</category><category>qwen3-coder-480b-a35b</category><category>gemini-2.5-flash-lite</category><category>mistral-7b</category><category>deepseek-v3</category><category>demishassabis</category><category>rasbt</category><category>alexwei_</category><category>yitayml</category><category>mixture-of-experts</category><category>agentic-ai</category><category>model-optimization</category><category>model-training</category><category>benchmarking</category><category>code-generation</category><category>long-context</category><category>multimodality</category><category>math</category><category>reinforcement-learning</category><category>model-architecture</category><category>model-performance</category><category>open-source</category><category>alignment</category></item><item><title>OAI and GDM announce IMO Gold-level results with natural language reasoning, no specialized training or tools, under human time limits</title><link>https://news.smol.ai/issues/25-07-21-imo-gold/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-21-imo-gold/</guid><description>**OpenAI** and **Google DeepMind** achieved a major milestone by solving 5 out of 6 problems at the **International Mathematical Olympiad (IMO) 2025** within the human time limit of 4.5 hours, earning the IMO Gold medal. This breakthrough was accomplished using general-purpose reinforcement learning and pure in-weights reasoning without specialized tools or internet access, surpassing previous systems like AlphaProof and AlphaGeometry2. The success resolved a 3-year-old AI bet on AI&apos;s capability to solve IMO problems and sparked discussions among mathematicians including **Terence Tao**. Despite this, 26 human competitors remain better than AI on the hardest combinatorics problem (P6). The achievement highlights advances in **reinforcement-learning**, **reasoning**, and **model-scaling** in AI research.</description><pubDate>Mon, 21 Jul 2025 05:44:39 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>gemini-1.5-pro</category><category>o1</category><category>terence_tao</category><category>oriol_vinyals</category><category>alexander_wei</category><category>jerry_tworek</category><category>paul_christiano</category><category>eliezer_yudkowsky</category><category>reinforcement-learning</category><category>reasoning</category><category>model-scaling</category><category>fine-tuning</category><category>model-training</category><category>benchmarking</category><category>natural-language-processing</category></item><item><title>ChatGPT Agent: new o* model + unified Deep Research browser + Operator computer use + Code Interpreter terminal</title><link>https://news.smol.ai/issues/25-07-17-chatgpt-agent/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-17-chatgpt-agent/</guid><description>**OpenAI** launched the **ChatGPT Agent**, a new advanced AI system capable of browsing the web, coding, analyzing data, and creating reports, marking a significant step towards human-like computer use. The agent, distinct from and superior to **o3**, is considered the first public exposure of what was internally called **o4**, now merged into **GPTNext**. It features end-to-end reinforcement learning, can operate for extended periods (tested up to 2 hours), and is classified as &quot;High&quot; risk for biological misuse, with safeguards activated. Early benchmarks show mixed results, excelling in some tests like **WebArena** and **BrowserComp** but underperforming on others like **PaperBench**. Key figures involved include **Sam Altman**, **Greg Brockman**, and **Kevin Weil**, with technical insights from **xikun_zhang_** and risk commentary from **KerenGu** and **boazbaraktcs**. The launch sparked speculation about **GPT-5**, which was confirmed not to be the case.</description><pubDate>Thu, 17 Jul 2025 05:44:39 GMT</pubDate><category>openai</category><category>o3</category><category>o4</category><category>gptnext</category><category>sama</category><category>gdb</category><category>kevinweil</category><category>xikun_zhang_</category><category>keren_gu</category><category>boazbaraktcs</category><category>reinforcement-learning</category><category>benchmarking</category><category>model-performance</category><category>model-risk</category><category>long-context</category><category>model-deployment</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-16-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-16-not-much/</guid><description>**Mistral** released **Voxtral**, claimed as the world&apos;s best open speech recognition models, available via API and Hugging Face. **Moonshot AI** launched **Kimi K2**, a trillion-parameter **Mixture-of-Experts (MoE)** model, outperforming **GPT-4.1** on benchmarks with 65.4% on SWE-Bench Verified and achieving 200 tokens/second inference speed on **Groq** hardware. **Nous Research** open-sourced the **Hermes 3** dataset with 1 million samples, aiding SOTA models on the **Llama-3** series. **Google DeepMind** introduced the **Mixture-of-Recursions (MoR)** architecture promising 2x inference speed and 50% parameter reduction but faced skepticism. **Goedel-Prover V2** topped the **PutnamBench** theorem proving benchmark. AtCoder World Finals saw a human winner with **OpenAI** placing second. Research highlights include **Jason Wei**&apos;s insights on **reinforcement learning** and the &quot;Verifier&apos;s Law&quot; emphasizing the asymmetry of verification in AI training.</description><pubDate>Wed, 16 Jul 2025 05:44:39 GMT</pubDate><category>mistral-ai</category><category>moonshot-ai</category><category>nous-research</category><category>google-deepmind</category><category>openai</category><category>groq</category><category>anthropic</category><category>kimi-k2</category><category>gpt-4.1</category><category>voxtral</category><category>goedel-prover-v2</category><category>llama-3</category><category>cline</category><category>_jasonwei</category><category>speech-recognition</category><category>mixture-of-experts</category><category>benchmarking</category><category>dataset-release</category><category>model-architecture</category><category>theorem-proving</category><category>reinforcement-learning</category><category>asymmetry-of-verification</category><category>inference-speed</category><category>model-performance</category></item><item><title>Voxtral - Mistral&apos;s SOTA ASR model in 3B (mini) and 24B (&quot;small&quot;) sizes beats OpenAI Whisper large-v3</title><link>https://news.smol.ai/issues/25-07-15-voxtral/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-15-voxtral/</guid><description>**Mistral** surprises with the release of **Voxtral**, a transcription model outperforming **Whisper large-v3**, **GPT-4o mini Transcribe**, and **Gemini 2.5 Flash**. Voxtral models (3B and 24B) support **32k token context length**, handle audios up to **30-40 minutes**, offer built-in **Q&amp;A and summarization**, are **multilingual**, and enable **function-calling** from voice commands, powered by the **Mistral Small 3.1** language model backbone. Meanwhile, **Moonshot AI**&apos;s **Kimi K2**, a non-reasoning **Mixture of Experts (MoE)** model built by a team of around **200 people**, gains attention for blazing-fast inference on **Groq** hardware, broad platform availability including **Together AI** and **DeepInfra**, and local running on **M4 Max 128GB** Mac. Developer tool integrations include **LangChain** and Hugging Face support, highlighting Kimi K2&apos;s strong tool use capabilities.</description><pubDate>Tue, 15 Jul 2025 05:44:39 GMT</pubDate><category>mistral-ai</category><category>moonshot-ai</category><category>groq</category><category>together-ai</category><category>deepinfra</category><category>huggingface</category><category>langchain</category><category>voxtal-3b</category><category>voxtal-24b</category><category>kimi-k2</category><category>jeremyphoward</category><category>teortaxestex</category><category>scaling01</category><category>zacharynado</category><category>jonathanross321</category><category>reach_vb</category><category>philschmid</category><category>transcription</category><category>long-context</category><category>function-calling</category><category>multilingual-models</category><category>mixture-of-experts</category><category>inference-speed</category><category>developer-tools</category><category>model-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-14-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-14-not-much/</guid><description>**Cognition** is acquiring the remaining assets of **Windsurf** after a significant weekend deal. **Moonshot AI** released **Kimi K2**, an open-source, MIT-licensed agentic model with **1 Trillion total / 32B active parameters** using a Mixture-of-Experts architecture, trained on **15.5 Trillion tokens** with the **MuonClip** optimizer, showing top performance on benchmarks like **EQ-Bench** and **Creative Writing**. **xAI** launched **Grok-4**, ranking 5th on **IQ Bench** but with notable quirks including a bug causing it to respond only with &quot;Heavy&quot; and a high frequency of Elon Musk mentions. Rumors about **OpenAI** delaying an open-source model release surfaced, with speculation about CEO **sama**&apos;s PR strategy and a possible **GPT-5** launch in September. The **Gemini 2.5** paper was released with **3,295 authors**, and **Google** introduced its **Gemini Embedding** model, topping the **MTEB leaderboard**.</description><pubDate>Mon, 14 Jul 2025 05:44:39 GMT</pubDate><category>cognition</category><category>windsurf</category><category>moonshot-ai</category><category>x-ai</category><category>openai</category><category>google</category><category>stanfordnlp</category><category>huggingface</category><category>kimi-k2</category><category>grok-4</category><category>gpt-5</category><category>gemini-2.5</category><category>gemini-embedding</category><category>sama</category><category>hardmaru</category><category>jeremyphoward</category><category>akhaliq</category><category>teortaxestex</category><category>yuchenj_uw</category><category>demishassabis</category><category>mixture-of-experts</category><category>model-training</category><category>model-performance</category><category>fine-tuning</category><category>benchmarking</category><category>agentic-ai</category><category>model-bugs</category><category>embedding-models</category></item><item><title>Kimi K2 - SOTA Open MoE proves that Muon can scale to 15T tokens/1T params</title><link>https://news.smol.ai/issues/25-07-11-kimi-k2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-11-kimi-k2/</guid><description>**Moonshot AI** has released **Kimi K2**, a **1 trillion parameter** Mixture-of-Experts model trained on **15.5 trillion tokens** using the new **MuonClip** optimizer, achieving state-of-the-art results on benchmarks like **SWE-Bench Verified (65.8%)** and **TAU2 (58.4%)**. This model is competitive with **GPT-4.1** and **Sonnet 4** on non-thinking tasks and is available under an **MIT license**. Meanwhile, **xAI** announced **Grok-4**, noted for its &quot;LEAST censored frontier model&quot; status and strong long-context performance but criticized for rushed post-training. **Mistral AI** updated its **Devstral 2507** models with improved performance and cost efficiency. The community is excited about the potential of the **MuonClip** optimizer, which may surpass the long-standing AdamW optimizer in machine learning.</description><pubDate>Fri, 11 Jul 2025 05:44:39 GMT</pubDate><category>moonshot-ai</category><category>alibaba</category><category>tencent</category><category>deepseek</category><category>x-ai</category><category>mistral-ai</category><category>weights-biases</category><category>hugging-face</category><category>kimi-k2</category><category>kimi-k2-1t</category><category>deepseek-v3</category><category>grok-4</category><category>devstral-2507</category><category>gpt-4.1</category><category>sonnet-4</category><category>yuchenj_uw</category><category>andrew_n_carr</category><category>scaling01</category><category>novita_labs</category><category>teknium1</category><category>aravsrinivas</category><category>mparakhin</category><category>simonw</category><category>mixture-of-experts</category><category>model-training</category><category>model-optimization</category><category>optimizer</category><category>benchmarking</category><category>long-context</category><category>model-performance</category><category>open-weights</category><category>model-release</category></item><item><title>Grok 4: xAI succeeds in going from 0 to new SOTA LLM in 2 years</title><link>https://news.smol.ai/issues/25-07-10-grok-4/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-10-grok-4/</guid><description>**xAI** launched **Grok 4** and **Grok 4 Heavy**, large language models rumored to have **2.4 trillion parameters** and trained with **100x more compute** than Grok 2 on **100k H100 GPUs**. Grok 4 achieved new state-of-the-art results on benchmarks like **ARC-AGI-2 (15.9%)**, **HLE (50.7%)**, and **Vending-Bench**, outperforming models such as **Claude 4 Opus**. The model supports a **256K context window** and is priced at **$3.00/M input tokens** and **$15.00/M output tokens**. It is integrated into platforms like **Cursor**, **Cline**, **LangChain**, and **Perplexity Pro/Max**. The launch was accompanied by a controversial voice mode and sparked industry discussion about xAI&apos;s rapid development pace, with endorsements from figures like **Elon Musk** and **Arav Srinivas**.</description><pubDate>Thu, 10 Jul 2025 05:44:39 GMT</pubDate><category>xai</category><category>perplexity-ai</category><category>langchain</category><category>cursor</category><category>cline</category><category>grok-4</category><category>grok-4-heavy</category><category>claude-4-opus</category><category>elonmusk</category><category>aravsrinivas</category><category>igor_babuschkin</category><category>yuchenj_uw</category><category>model-releases</category><category>benchmarking</category><category>long-context</category><category>model-pricing</category><category>model-integration</category><category>voice</category><category>performance</category><category>scaling</category><category>gpu-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-09-not-much/</guid><description>**LangChain** is nearing unicorn status, while **OpenAI** and **Google DeepMind&apos;s Gemini 3 Pro** models are launching soon. **Perplexity** rolls out its agentic browser **Comet** to waitlists, offering multitasking and voice command features. **xAI&apos;s Grok-4** update sparked controversy due to offensive outputs, drawing comparisons to **Microsoft&apos;s Tay** bot and resulting in regional blocks. **Hugging Face** released **SmolLM3**, a 3B parameter open-source model with state-of-the-art reasoning and long context capabilities. **Google** introduced **T5Gemma** encoder-decoder models, a significant update in this model category. **Anthropic** investigates &quot;alignment faking&quot; in language models, focusing on safety concerns with models like **Claude 3.7 Sonnet** and **DeepSeek-R1**. *&quot;Grok 3 had high reasoning, Grok 4 has heil reasoning&quot;* was a notable user comment on the controversy.</description><pubDate>Wed, 09 Jul 2025 05:44:39 GMT</pubDate><category>langchain</category><category>openai</category><category>google-deepmind</category><category>perplexity</category><category>xai</category><category>microsoft</category><category>huggingface</category><category>anthropic</category><category>grok-4</category><category>smollm3</category><category>t5gemma</category><category>claude-3.7-sonnet</category><category>deepseek-r1</category><category>aravsrinivas</category><category>clementdelangue</category><category>_akhaliq</category><category>agentic-ai</category><category>model-controversy</category><category>open-source</category><category>model-release</category><category>alignment</category><category>fine-tuning</category><category>long-context</category><category>multimodality</category><category>model-research</category></item><item><title>SmolLM3: the SOTA 3B reasoning open source LLM</title><link>https://news.smol.ai/issues/25-07-08-smollm3/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-08-smollm3/</guid><description>**HuggingFace** released **SmolLM3-3B**, a fully open-source small reasoning model with open pretraining code and data, marking a high point in open source models until **Olmo 3** arrives. **Grok 4** was launched with mixed reactions, while concerns about **Claude 4** nerfs and an imminent **Claude 4.1** surfaced. **Gemini Nano** is now shipping in **Chrome 137+**, enabling local LLM access for **3.7 billion** users. **Tencent** introduced **Hunyuan-A13B**, an 80B parameter model with a 256K context window running on a single **H200** GPU. The **Gemini API** added a batch mode with 50% discounts on **2.5 models**. **MatFormer Lab** launched tools for custom-sized **Gemma 3n** models. Open source OCR models like **Nanonets-OCR-s** and **ChatDOC/OCRFlux-3B** derived from **Qwen2.5-VL-3B** were highlighted, with licensing discussions involving **Alibaba**.</description><pubDate>Tue, 08 Jul 2025 05:44:39 GMT</pubDate><category>huggingface</category><category>allenai</category><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>mistral-ai</category><category>tencent</category><category>gemini</category><category>alibaba</category><category>smollm3-3b</category><category>olmo-3</category><category>grok-4</category><category>claude-4</category><category>claude-4.1</category><category>gemini-nano</category><category>hunyuan-a13b</category><category>gemini-2.5</category><category>gemma-3n</category><category>qwen2.5-vl-3b</category><category>elonmusk</category><category>mervenoyann</category><category>skirano</category><category>amandaaskell</category><category>clementdelangue</category><category>loubnabenallal1</category><category>awnihannun</category><category>swyx</category><category>artificialanlys</category><category>officiallogank</category><category>osanseviero</category><category>cognitivecompai</category><category>aravsrinivas</category><category>open-source</category><category>small-language-models</category><category>model-releases</category><category>model-performance</category><category>benchmarking</category><category>multimodality</category><category>context-windows</category><category>precision-fp8</category><category>api</category><category>batch-processing</category><category>model-scaling</category><category>model-architecture</category><category>licensing</category><category>ocr</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-07-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-07-not-much/</guid><description>Over the holiday weekend, key AI developments include the upcoming release of **Grok 4**, **Perplexity** teasing new projects, and community reactions to **Cursor** and **Dia**. Research highlights feature a paper on **Reinforcement Learning (RL)** improving generalization and reasoning across domains, contrasting with Supervised Fine-Tuning&apos;s forgetting issues. **Energy-Based Transformers (EBTs)** are proposed as a promising alternative to traditional transformers. **AI21 Labs** updated its **Jamba** model family with enhanced grounding and instruction following, maintaining a **256K** context window. **Baidu** open-sourced its massive **424 billion** parameter **Ernie 4.5** model, while **Kontext-dev** became the top trending model on **Hugging Face**. Advances in length generalization for recurrent models and the introduction of **2-simplicial attention** were noted. In biomedical AI, **Biomni**, powered by **Claude 4 Sonnet**, demonstrated superior accuracy and rare disease diagnosis capabilities. Additionally, the Python package manager `uv` received praise for improving Python installation workflows.</description><pubDate>Mon, 07 Jul 2025 05:44:39 GMT</pubDate><category>ai21-labs</category><category>hugging-face</category><category>baidu</category><category>perplexity-ai</category><category>deepmind</category><category>anthropic</category><category>grok-4</category><category>jamba</category><category>ernie-4.5</category><category>claude-4-sonnet</category><category>claude-4</category><category>kontext-dev</category><category>_philschmid</category><category>corbtt</category><category>jxmnop</category><category>sedielem</category><category>_akhaliq</category><category>slashml</category><category>alexiglad</category><category>clementdelangue</category><category>_albertgu</category><category>tri_dao</category><category>theaitimeline</category><category>deep-learning-ai</category><category>reinforcement-learning</category><category>fine-tuning</category><category>energy-based-transformers</category><category>ssm-transformer</category><category>context-windows</category><category>length-generalization</category><category>recurrent-neural-networks</category><category>attention-mechanisms</category><category>2-simplicial-attention</category><category>biomedical-ai</category><category>instruction-following</category><category>open-weight-models</category><category>python-package-management</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-03-not-much/</guid><description>**Ilya Sutskever** confirmed his role as CEO of **Safe Superintelligence Inc. (SSI)** with **Daniel Levy** as President, dismissing acquisition rumors and emphasizing their strong team and compute resources. **Perplexity AI** expanded its data integrations by adding **Morningstar&apos;s** financial research and hinted at new product features for Pro users. **Meta AI FAIR** clarified its research structure, distinguishing its small lab from larger model training groups, and welcomed **Nat Friedman** to enhance AI product development. **Midjourney** and **Sakana AI** announced hiring for research and applied engineering roles. **Cohere** expanded its presence in Montréal, receiving praise from Canadian officials. On the model front, **Google DeepMind&apos;s Gemini Pro** released the **Veo 3** video generation model globally. **DeepSeek** launched the faster **DeepSeek R1T2** model using an Assembly of Experts approach, available under an MIT license. **Kling AI** showcased cinematic video generation capabilities. **OpenAI** introduced a high-cost **Deep Research API** with pricing up to **$30 per call**. **Together AI** announced the release of the **DeepSWE agent**.</description><pubDate>Thu, 03 Jul 2025 05:44:39 GMT</pubDate><category>safe-superintelligence-inc</category><category>perplexity-ai</category><category>meta-ai-fair</category><category>midjourney</category><category>sakana-ai</category><category>cohere</category><category>google-deepmind</category><category>deepseek</category><category>openai</category><category>together-ai</category><category>veo-3</category><category>deepseek-r1t2</category><category>deepseek-tng-r1t2-chimera</category><category>o3-deep-research</category><category>o4-mini-deep-research</category><category>deepswe-agent</category><category>ilya_sutskever</category><category>daniel_levy</category><category>daniel_gross</category><category>aravsrinivas</category><category>zeyuanallenzhu</category><category>nat_friedman</category><category>davidsholz</category><category>fp_champagne</category><category>demishassabis</category><category>reach_vb</category><category>video-generation</category><category>assembly-of-experts</category><category>model-licenses</category><category>api-pricing</category><category>research-roles</category><category>product-expansion</category><category>corporate-leadership</category><category>model-release</category><category>team-expansion</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-02-not-much/</guid><description>**Meta** has hired **Scale AI CEO Alexandr Wang** as its new **Chief AI Officer**, acquiring a **49% non-voting stake** in **Scale AI** for **$14.3 billion**, doubling its valuation to **~$28 billion**. This move is part of a major talent shuffle involving **Meta**, **OpenAI**, and **Scale AI**. Discussions include the impact on **Yann LeCun**&apos;s influence at **Meta** and potential responses from **OpenAI**. In model news, **Gemma 3N** faces technical issues like vision NaNs and FP16 overflows, with fixes from **UnslothAI**. Chinese open-source models like **GLM-4.1V-Thinking** by **Zhipu AI** and **DeepSeek R1T2** show strong performance and speed improvements. **Huawei** open-sourced a **72B MoE** model with a novel load balancing solution. The **MiniMax-M1** hybrid MoE model leads math benchmarks on the **Text Arena leaderboard**. **AllenAI** launched **SciArena** for scientific literature evaluation, where **o3** outperforms others. Research from **Sakana AI Labs** introduces **AB-MCTS** for code generation, improving synthesis benchmarks.</description><pubDate>Wed, 02 Jul 2025 05:44:39 GMT</pubDate><category>meta</category><category>scale-ai</category><category>unslothai</category><category>zhipu-ai</category><category>deepseek</category><category>huawei</category><category>minimax-ai</category><category>allenai</category><category>sakana-ai-labs</category><category>openai</category><category>gemma-3n</category><category>glm-4.1v-thinking</category><category>deepseek-r1t2</category><category>mini-max-m1</category><category>o3</category><category>claude-4-opus</category><category>claude-sonnet</category><category>moe-72b</category><category>alexandr_wang</category><category>natfriedman</category><category>steph_palazzolo</category><category>thegregyang</category><category>teortaxes_tex</category><category>denny_zhou</category><category>agihippo</category><category>danielhanchen</category><category>osanseviero</category><category>reach_vb</category><category>scaling01</category><category>ndea</category><category>model-performance</category><category>vision</category><category>conv2d</category><category>float16</category><category>training-loss</category><category>open-source</category><category>model-benchmarks</category><category>moe</category><category>load-balancing</category><category>scientific-literature-evaluation</category><category>code-generation</category><category>adaptive-tree-search</category><category>synthesis-benchmarks</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-07-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-07-01-not-much/</guid><description>**Meta** makes a major AI move by hiring **Scale AI** founder **Alexandr Wang** as Chief AI Officer and acquiring a 49% non-voting stake in **Scale AI** for **$14.3 billion**, doubling its valuation to about **$28 billion**. **Chai Discovery** announces **Chai-2**, a breakthrough model for zero-shot antibody discovery and optimization. The US government faces budget cuts threatening to eliminate a quarter million science research jobs by **2026**. Data access restrictions intensify as companies like **Atlassian**, **Notion**, and **Slack** block web crawlers including **Common Crawl**, raising concerns about future public internet archives. **Hugging Face** shuts down **HuggingChat** after serving over a million users, marking a significant experiment in open-source LLMs. **Sakana AI** releases **AB-MCTS**, an inference-time scaling algorithm enabling multiple models like **Gemini 2.5 Pro** and **DeepSeek-R1-0528** to cooperate and outperform individual models.</description><pubDate>Tue, 01 Jul 2025 05:44:39 GMT</pubDate><category>meta</category><category>scale-ai</category><category>anthropic</category><category>cloudflare</category><category>grammarly</category><category>superhuman</category><category>chai-discovery</category><category>atlassian</category><category>notion</category><category>slack</category><category>commoncrawl</category><category>hugging-face</category><category>sakana-ai</category><category>chai-2</category><category>gemini-2.5-pro</category><category>deepseek-r1-0528</category><category>alexandr_wang</category><category>nat_friedman</category><category>clementdelangue</category><category>teortaxestex</category><category>ylecun</category><category>steph_palazzolo</category><category>andersonbcdefg</category><category>jeremyphoward</category><category>reach_vb</category><category>inference</category><category>model-scaling</category><category>collective-intelligence</category><category>zero-shot-learning</category><category>enterprise-deployment</category><category>data-access</category><category>science-funding</category><category>open-source-llms</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-30-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-30-not-much/</guid><description>**Meta** has poached top AI talent from **OpenAI**, including **Alexandr Wang** joining as Chief AI Officer to work towards superintelligence, signaling a strong push for the next **Llama** model. The AI job market shows polarization with high demand and compensation for top-tier talent, while credentials like strong GitHub projects gain importance. The **WizardLM** team moved from **Microsoft** to **Tencent** to develop open-source models like **Hunyuan-A13B**, highlighting shifts in China&apos;s AI industry. Rumors suggest **OpenAI** will release a new open-source model in July, potentially surpassing existing **ChatGPT** models. **Baidu** open-sourced multiple variants of its **ERNIE 4.5** model series, featuring advanced techniques like **2-bit quantization**, **MoE router orthogonalization loss**, and **FP8** training, with models ranging from **0.3B** to **424B** parameters. **Gemini 2.5 Pro** returned to the free tier of the **Gemini API**, enabling developers to explore its features.</description><pubDate>Mon, 30 Jun 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>tencent</category><category>microsoft</category><category>baidu</category><category>gemini</category><category>o3-mini</category><category>o1-mini</category><category>llama</category><category>hunyuan-a13b</category><category>ernie-4.5</category><category>ernie-4.5-21b-a3b</category><category>qwen3-30b-a3b</category><category>gemini-2.5-pro</category><category>alexandr_wang</category><category>shengjia_zhao</category><category>jhyuxm</category><category>ren_hongyu</category><category>shuchaobi</category><category>saranormous</category><category>teortaxesTex</category><category>mckbrando</category><category>yuchenj_uw</category><category>francoisfleuret</category><category>quanquangu</category><category>reach_vb</category><category>philschmid</category><category>superintelligence</category><category>ai-talent</category><category>job-market</category><category>open-source-models</category><category>multimodality</category><category>mixture-of-experts</category><category>quantization</category><category>fp8-training</category><category>model-benchmarking</category><category>model-performance</category><category>model-releases</category><category>api</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-27-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-27-not-much/</guid><description>**Google** released **Gemma 3n**, a multimodal model for edge devices available in **2B and 4B** parameter versions, with support across major frameworks like **Transformers** and **Llama.cpp**. **Tencent** open-sourced **Hunyuan-A13B**, a **Mixture-of-Experts (MoE)** model with **80B total parameters** and a **256K context window**, optimized for tool calling and coding. **Black Forest Labs** released **FLUX.1 Kontext [dev]**, an open image AI model gaining rapid Hugging Face adoption. **Inception AI Labs** launched **Mercury**, the first commercial-scale **diffusion LLM** for chat. The **FineWeb2** multilingual pre-training dataset paper was released, analyzing data quality impacts. The **Qwen** team released **Qwen-VLo**, a unified visual understanding and generation model. **Kyutai Labs** released a top-ranked open-source speech-to-text model running on Macs and iPhones. **OpenAI** introduced **Deep Research API** with **o3/o4-mini** models and open-sourced prompt rewriter methodology, integrated into **LangChain** and **LangGraph**. The open-source **Gemini CLI** gained over **30,000 GitHub stars** as an AI terminal agent.</description><pubDate>Fri, 27 Jun 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>tencent</category><category>black-forest-labs</category><category>inception-ai</category><category>qwen</category><category>kyutai-labs</category><category>openai</category><category>langchain</category><category>langgraph</category><category>hugging-face</category><category>ollama</category><category>unslothai</category><category>nvidia</category><category>amd</category><category>gemma-3n</category><category>hunyuan-a13b</category><category>flux-1-kontext-dev</category><category>mercury</category><category>fineweb2</category><category>qwen-vlo</category><category>o3-mini</category><category>o4-mini</category><category>demishassabis</category><category>reach_vb</category><category>tri_dao</category><category>osanseviero</category><category>simonw</category><category>clementdelangue</category><category>swyx</category><category>hwchase17</category><category>sydneyrunkle</category><category>multimodality</category><category>mixture-of-experts</category><category>context-windows</category><category>tool-use</category><category>coding</category><category>image-generation</category><category>diffusion-models</category><category>dataset-release</category><category>multilinguality</category><category>speech-to-text</category><category>api</category><category>prompt-engineering</category><category>agent-frameworks</category><category>open-source</category><category>model-release</category></item><item><title>OpenAI releases Deep Research API (o3/o4-mini)</title><link>https://news.smol.ai/issues/25-06-26-deepresearch-api/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-26-deepresearch-api/</guid><description>**OpenAI** has launched the **Deep Research API** featuring powerful models **o3-deep-research** and **o4-mini-deep-research** with native support for MCP, Search, and Code Interpreter, enabling advanced agent capabilities including multi-agent setups. **Google** released **Gemma 3n**, a multimodal model optimized for edge devices with only 3GB RAM, achieving a top score of 1300 on LMSys Arena, featuring the new MatFormer architecture and broad ecosystem integration. **Black Forest Labs** introduced **FLUX.1 Kontext [dev]**, a 12B parameter rectified flow transformer for instruction-based image editing, comparable to **GPT-4o**. **DeepMind** unveiled **AlphaGenome**, an AI model capable of reading 1 million DNA bases for gene function prediction, marking a breakthrough in AI biology. **Sakana AI** presented Reinforcement-Learned Teachers (RLTs) to enhance LLM reasoning, achieving 86.1% on MiniF2F with efficient compute. **Higgsfield AI** released **Higgsfield Soul**, a high-aesthetic photo model with 50+ presets for fashion-grade realism. Additionally, **Google** launched the **Gemini CLI**, an open-source AI agent for terminal use with free Gemini 2.5 Pro requests.</description><pubDate>Thu, 26 Jun 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>black-forest-labs</category><category>deepmind</category><category>sakana-ai</category><category>higgsfield-ai</category><category>huggingface</category><category>ollama</category><category>o3-deep-research</category><category>o4-mini-deep-research</category><category>gemma-3n</category><category>flux-1-kontext-dev</category><category>gpt-4o</category><category>alphagenome</category><category>demishassabis</category><category>hardmaru</category><category>osanseviero</category><category>clementdelangue</category><category>multimodality</category><category>model-releases</category><category>agentic-ai</category><category>reinforcement-learning</category><category>instruction-following</category><category>model-architecture</category><category>model-optimization</category><category>image-generation</category><category>biological-ai</category><category>multi-agent-systems</category><category>model-integration</category></item><item><title>Context Engineering: Much More than Prompts</title><link>https://news.smol.ai/issues/25-06-25-context-eng/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-25-context-eng/</guid><description>**Context Engineering** emerges as a significant trend in AI, highlighted by experts like **Andrej Karpathy**, **Walden Yan** from **Cognition**, and **Tobi Lutke**. It involves managing an LLM&apos;s context window with the right mix of prompts, retrieval, tools, and state to optimize performance, going beyond traditional prompt engineering. **LangChain** and its tool **LangGraph** are noted for advancing this approach. Additionally, **OpenAI** has launched **ChatGPT connectors** for platforms like **Google Drive**, **Dropbox**, **SharePoint**, and **Box**, enhancing context integration for Pro users. Other notable news includes the launch of **Vercel Sandbox**, **Cloudflare Containers**, the leak and release of **Gemini Code** by **Google DeepMind**, and fundraising efforts by **OpenRouter**.</description><pubDate>Wed, 25 Jun 2025 05:44:39 GMT</pubDate><category>openai</category><category>langchain</category><category>cognition</category><category>google-deepmind</category><category>vercel</category><category>cloudflare</category><category>openrouter</category><category>gemini-code</category><category>karpathy</category><category>walden_yan</category><category>tobi_lutke</category><category>hwchase17</category><category>rlancemartin</category><category>kwindla</category><category>dex_horthy</category><category>context-engineering</category><category>retrieval-augmented-generation</category><category>tools</category><category>state-management</category><category>history-management</category><category>prompt-engineering</category><category>software-layer</category><category>chatgpt-connectors</category><category>api-integration</category></item><item><title>Bartz v. Anthropic PBC — &quot;Training use is Fair Use&quot;</title><link>https://news.smol.ai/issues/25-06-24-fair-use/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-24-fair-use/</guid><description>**Anthropic** won a significant fair use ruling allowing the training of **Claude** on copyrighted books, setting a precedent for AI training legality despite concerns over pirated data. **Replit** achieved a major milestone with **$100M ARR**, showing rapid growth. **Delphi** raised **$16M Series A** to scale digital minds, while **Thinking Machines Lab** focuses on reinforcement learning for business applications. **Disney** and **Universal** sued **Midjourney** over unauthorized use of copyrighted images. **Google DeepMind** released **Gemini Robotics On-Device**, a compact foundation model for robotics.</description><pubDate>Tue, 24 Jun 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>replit</category><category>delphi</category><category>sequoia</category><category>thinking-machines-lab</category><category>disney</category><category>universal</category><category>midjourney</category><category>google-deepmind</category><category>claude</category><category>gemini-robotics-on-device</category><category>andrea_bartz</category><category>giffmana</category><category>andrewcurran_</category><category>amasad</category><category>swyx</category><category>hwchase17</category><category>krandiash</category><category>daraladje</category><category>steph_palazzolo</category><category>corbtt</category><category>demishassabis</category><category>fair-use</category><category>copyright</category><category>reinforcement-learning</category><category>foundation-models</category><category>robotics</category><category>funding</category><category>lawsuit</category><category>digital-minds</category><category>model-release</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/25-06-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-23-not-much/</guid><description>**Sakana AI** released **Reinforcement-Learned Teachers (RLTs)**, a novel technique using smaller 7B parameter models trained via reinforcement learning to teach reasoning through step-by-step explanations, accelerating **Chain-of-Thought** learning. **Mistral AI** updated **Mistral Small 3.2** improving instruction following and function calling with experimental FP8 quantization. **Google Magenta RealTime**, an 800M parameter open-weights model for real-time music generation, was released. **Arcee AI** launched **AFM-4.5B**, a sub-10B parameter foundation model extended from **Llama 3**. **OpenThinker3-7B** was introduced as a new state-of-the-art 7B reasoning model with a 33% improvement over **DeepSeek-R1-Distill-Qwen-7B**. The **STORM** text-video model compresses video input by 8x using **Mamba layers** and outperforms **GPT-4o** on MVBench with 70.6%. Discussions on reinforcement learning algorithms PPO vs. GRPO and insights on **DINOv2**&apos;s performance on ImageNet-1k were also highlighted. *&quot;A very quiet day&quot;* in AI news with valuable workshops from **OpenAI**, **Amazon**, and **GDM**.</description><pubDate>Mon, 23 Jun 2025 05:44:39 GMT</pubDate><category>sakana-ai</category><category>mistral-ai</category><category>google</category><category>arcee-ai</category><category>deepseek-ai</category><category>openai</category><category>amazon</category><category>gdm</category><category>mistral-small-3.2</category><category>magenta-realtime</category><category>afm-4.5b</category><category>llama-3</category><category>openthinker3-7b</category><category>deepseek-r1-distill-qwen-7b</category><category>storm</category><category>qwen2-vl</category><category>gpt-4o</category><category>dino-v2</category><category>sama</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>fine-tuning</category><category>function-calling</category><category>quantization</category><category>music-generation</category><category>foundation-models</category><category>reasoning</category><category>text-video</category><category>model-compression</category><category>image-classification</category><category>evaluation-metrics</category></item><item><title>The Quiet Rise of Claude Code vs Codex</title><link>https://news.smol.ai/issues/25-06-20-claude-code/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-20-claude-code/</guid><description>**Claude Code** is gaining mass adoption, inspiring derivative projects like **OpenCode** and **ccusage**, with discussions ongoing in AI communities. **Mistral AI** released **Mistral Small 3.2**, a **24B** parameter model update improving instruction following and function calling, available on **Hugging Face** and supported by **vLLM**. Sebastian Raschka implemented **Qwen3 0.6B** from scratch, noting its deeper architecture and memory efficiency compared to **Llama 3 1B**. **Google DeepMind** showcased **Gemini 2.5 Flash-Lite**&apos;s UI code generation from visual context and added video upload support in the **Gemini App**. **Apple**&apos;s new **3B** parameter on-device foundation model was benchmarked, showing slower speed but efficient memory use via **2-bit quantization**, suitable for background tasks. **Google DeepMind** also released **Magenta Real-time**, an **800M** parameter music generation model licensed under **Apache 2.0**, marking Google&apos;s 1000th model on **Hugging Face**. **Kuaishou** launched **KLING 2.1**, a new video model accessible via API.</description><pubDate>Fri, 20 Jun 2025 05:44:39 GMT</pubDate><category>mistral-ai</category><category>hugging-face</category><category>google-deepmind</category><category>apple</category><category>artificial-analysis</category><category>kuaishou</category><category>mistral-small-3.2</category><category>qwen3-0.6b</category><category>llama-3-1b</category><category>gemini-2.5-flash-lite</category><category>gemini-app</category><category>magenta-real-time</category><category>apple-3b-on-device</category><category>reach_vb</category><category>guillaumelample</category><category>qtnx_</category><category>shxf0072</category><category>rasbt</category><category>demishassabis</category><category>artificialanlys</category><category>osanseviero</category><category>instruction-following</category><category>function-calling</category><category>model-implementation</category><category>memory-efficiency</category><category>2-bit-quantization</category><category>music-generation</category><category>video-models</category><category>benchmarking</category><category>api</category></item><item><title>minor ai followups: MultiAgents, Meta-SSI-Scale, Karpathy, AI Engineer</title><link>https://news.smol.ai/issues/25-06-19-followups/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-19-followups/</guid><description>**OpenAI** released a paper revealing how training models like **GPT-4o** on insecure code can cause broad misalignment, drawing reactions from experts like *@sama* and *@polynoamial*. **California&apos;s AI regulation efforts** were highlighted by *@Yoshua_Bengio* emphasizing transparency and whistleblower protections. The term **&quot;context rot&quot;** was coined to describe LLM conversation degradation, with systems like **Embra** using CRM-like memory for robustness. Scalable oversight research aiming to improve human control over smarter AIs was discussed by *@RyanPGreenblatt*. New model releases include **Kyutai&apos;s** speech-to-text models capable of 400 real-time streams on a single H100 GPU, **Tencent&apos;s Hunyuan 3D 2.1** as the first open-source production-ready PBR 3D generative model, and **Arcee&apos;s AFM-4.5B** foundation model family targeting enterprise use, competitive with **Gemma** and **Qwen**.</description><pubDate>Thu, 19 Jun 2025 05:44:39 GMT</pubDate><category>openai</category><category>meta-ai-fair</category><category>scale-ai</category><category>huggingface</category><category>tencent</category><category>arcee-ai</category><category>gpt-4o</category><category>afm-4.5b</category><category>gemma</category><category>qwen</category><category>stt-1b-en_fr</category><category>stt-2.6b-en</category><category>hunyuan-3d-2.1</category><category>sama</category><category>polynoamial</category><category>neelnanda5</category><category>teortaxestex</category><category>yoshua_bengio</category><category>zachtratar</category><category>ryanpgreenblatt</category><category>reach_vb</category><category>arankomatsuzaki</category><category>code_star</category><category>ai-safety</category><category>alignment</category><category>ai-regulation</category><category>memory-optimization</category><category>scalable-oversight</category><category>speech-recognition</category><category>3d-generation</category><category>foundation-models</category></item><item><title>Zuck goes Superintelligence Founder Mode: $100M bonuses + $100M+ salaries + NFDG Buyout?</title><link>https://news.smol.ai/issues/25-06-18-zuck-founder-mode/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-18-zuck-founder-mode/</guid><description>**Meta AI** is reportedly offering **8-9 figure signing bonuses and salaries** to top AI talent, confirmed by **Sam Altman**. They are also targeting key figures like **Nat** and **Dan** from the AI Grant fund for strategic hires. **Essential AI** released the massive **24-trillion-token Essential-Web v1.0 dataset** with rich metadata and a 12-category taxonomy. **DeepLearning.AI** and **Meta AI** launched a course on **Llama 4**, featuring new MoE models **Maverick (400B)** and **Scout (109B)** with context windows up to **10M tokens**. **MiniMax** open-sourced **MiniMax-M1**, a long-context LLM with a 1M-token window, and introduced the **Hailuo 02** video model. **OpenAI** rolled out &quot;Record mode&quot; for **ChatGPT Pro, Enterprise, and Edu** on macOS. **Arcee** launched the **AFM-4.5B** foundation model for enterprise. **Midjourney** released its **V1 video model** enabling image animation. These developments highlight major advances in model scale, long-context reasoning, multimodality, and enterprise AI applications.</description><pubDate>Wed, 18 Jun 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>deeplearning-ai</category><category>essential-ai</category><category>minimax</category><category>arcee</category><category>midjourney</category><category>llama-4</category><category>maverick</category><category>scout</category><category>minimax-m1</category><category>afm-4.5b</category><category>chatgpt</category><category>midjourney-v1</category><category>sama</category><category>nat</category><category>dan</category><category>ashvaswani</category><category>clementdelangue</category><category>amit_sangani</category><category>andrewyng</category><category>_akhaliq</category><category>long-context</category><category>multimodality</category><category>model-release</category><category>foundation-models</category><category>dataset-release</category><category>model-training</category><category>video-generation</category><category>enterprise-ai</category><category>model-architecture</category><category>moe</category><category>prompt-optimization</category></item><item><title>Gemini 2.5 Pro/Flash GA, 2.5 Flash-Lite in Preview</title><link>https://news.smol.ai/issues/25-06-17-gemini-2-5/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-17-gemini-2-5/</guid><description>**Gemini 2.5** models are now generally available, including the new **Gemini 2.5 Flash-Lite**, **Flash**, **Pro**, and **Ultra** variants, featuring sparse **Mixture-of-Experts (MoE)** transformers with native multimodal support. A detailed 30-page tech report highlights impressive long-horizon planning demonstrated by **Gemini Plays Pokemon**. The **LiveCodeBench-Pro** benchmark reveals frontier LLMs struggle with hard coding problems, while **Moonshot AI** open-sourced **Kimi-Dev-72B**, achieving state-of-the-art results on **SWE-bench Verified**. Smaller specialized models like **Nanonets-OCR-s**, **II-Medical-8B-1706**, and **Jan-nano** show competitive performance, emphasizing that bigger models are not always better. **DeepSeek-r1** ties for #1 in WebDev Arena, and **MiniMax-M1** sets new standards in long-context reasoning. **Kling AI** demonstrated video generation capabilities.</description><pubDate>Tue, 17 Jun 2025 05:44:39 GMT</pubDate><category>google</category><category>moonshot-ai</category><category>deepseek</category><category>cognitivecompai</category><category>kling-ai</category><category>gemini-2.5</category><category>gemini-2.5-flash-lite</category><category>gemini-2.5-flash</category><category>gemini-2.5-pro</category><category>gemini-2.5-ultra</category><category>kimi-dev-72b</category><category>nanonets-ocr-s</category><category>ii-medical-8b-1706</category><category>jan-nano</category><category>deepseek-r1</category><category>minimax-m1</category><category>tulsee_doshi</category><category>oriolvinyalsml</category><category>demishassabis</category><category>officiallogank</category><category>_philschmid</category><category>swyx</category><category>sainingxie</category><category>scaling01</category><category>gneubig</category><category>clementdelangue</category><category>mervenoyann</category><category>mixture-of-experts</category><category>multimodality</category><category>long-horizon-planning</category><category>benchmarking</category><category>coding-performance</category><category>long-context</category><category>ocr</category><category>video-generation</category><category>model-releases</category></item><item><title>Chinese Models Launch - MiniMax-M1, Hailuo 2 &quot;Kangaroo&quot;, Moonshot Kimi-Dev-72B</title><link>https://news.smol.ai/issues/25-06-16-chinese-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-16-chinese-models/</guid><description>**MiniMax AI** launched **MiniMax-M1**, a 456 billion parameter open weights LLM with a 1 million token input and 80k token output using efficient &quot;lightning attention&quot; and a GRPO variant called CISPO. **MiniMax AI** also announced **Hailuo 02 (0616)**, a video model similar to **ByteDance&apos;s Seedance**. **Moonshot AI** released **Kimi-Dev-72B**, a coding model outperforming **DeepSeek R1** on SWEBench Verified. Discussions on multi-agent system design from **Anthropic** and **LangChain** highlighted improvements in task completion and challenges like prompt injection attacks, as demonstrated by **Karpathy** and **Columbia University** research. **Sakana AI** introduced **ALE-Agent**, a coding agent that ranked 21st in the AtCoder Heuristic Competition solving NP-hard optimization problems. There is unverified news about an acquisition involving **OpenAI**, **Microsoft**, and **Windsurf**.</description><pubDate>Mon, 16 Jun 2025 05:44:39 GMT</pubDate><category>minimax-ai</category><category>moonshot-ai</category><category>deepseek</category><category>bytedance</category><category>anthropic</category><category>langchain</category><category>columbia-university</category><category>sakana-ai</category><category>openai</category><category>microsoft</category><category>minimax-m1</category><category>hailuo-02</category><category>kimi-dev-72b</category><category>deepseek-r1</category><category>ale-agent</category><category>jerryjliu0</category><category>hwchase17</category><category>omarsar0</category><category>gallabytes</category><category>lateinteraction</category><category>karpathy</category><category>multi-agent-systems</category><category>attention-mechanisms</category><category>coding</category><category>optimization</category><category>prompt-injection</category><category>model-performance</category><category>video-generation</category><category>model-training</category><category>task-automation</category></item><item><title>Cognition vs Anthropic: Don&apos;t Build Multi-Agents/How to Build Multi-Agents</title><link>https://news.smol.ai/issues/25-06-13-cognition-vs-anthropic/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-13-cognition-vs-anthropic/</guid><description>Within the last 24 hours, **Cognition**&apos;s Walden Yan advised *&quot;Don&apos;t Build Multi-Agents,&quot;* while **Anthropic** shared their approach to building multi-agent systems with **Claude&apos;s** multi-agent research architecture. **LangChain** highlighted advances in context engineering and production AI agents used by **LinkedIn** and **BlackRock**. The community is engaging in a debate on multi-agent AI development. Additionally, **Hugging Face** announced deprecating **TensorFlow** and **Flax** support in favor of **PyTorch**. Research on agent memory and model elicitation techniques from **LlamaIndex** and **Anthropic** were also discussed.</description><pubDate>Fri, 13 Jun 2025 05:44:39 GMT</pubDate><category>cognition</category><category>anthropic</category><category>langchain</category><category>huggingface</category><category>microsoft</category><category>llamaindex</category><category>linkedin</category><category>blackrock</category><category>claude</category><category>walden_yan</category><category>hwchase17</category><category>assaf_elovic</category><category>sh_reya</category><category>hamelhusain</category><category>omarsar0</category><category>clefourrier</category><category>jerryjliu0</category><category>akbirkhan</category><category>multi-agent-systems</category><category>context-engineering</category><category>agent-memory</category><category>model-elicitation</category><category>ai-evaluation</category><category>deep-research-workflows</category><category>framework-migration</category><category>pydantic-schema</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-12-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-12-not-much/</guid><description>**Bytedance** showcased an impressive state-of-the-art video generation model called **Seedance 1.0** without releasing it, while **Morph Labs** announced **Trinity**, an autoformalization system for Lean. **Huggingface Transformers** deprecated Tensorflow/JAX support. **Andrew Ng** of **DeepLearning.AI** highlighted the rise of the **GenAI Application Engineer** role emphasizing skills in **AI building blocks** and **AI-assisted coding tools** like **Codex** and **Claude Code**. Engineering teams are increasingly testing API designs against LLMs for usability. **Figure AI**&apos;s CEO stressed speed as a key competitive advantage, and **LangChain** introduced the concept of **Context Engineering** for AI agents. Reinforcement learning on LLMs shows transformative potential, and the community values **AI evals** and data work. **Sakana AI** released **Text-to-LoRA**, a hypernetwork method for generating task-specific LoRA adapters from natural language, enabling efficient model customization. The video generation race heats up with **Bytedance**&apos;s Seed-based model praised for quality, challenging American labs, alongside models like **Kling 2.1** and **Veo 3**.</description><pubDate>Thu, 12 Jun 2025 05:44:39 GMT</pubDate><category>bytedance</category><category>morph-labs</category><category>huggingface</category><category>deeplearning.ai</category><category>figure-ai</category><category>langchain</category><category>sakana-ai</category><category>seedance-1.0</category><category>codex</category><category>claude-code</category><category>kling-2.1</category><category>veo-3</category><category>andrew_ng</category><category>hwchase17</category><category>adcock_brett</category><category>clementdelangue</category><category>akhaliq</category><category>jxmnop</category><category>hamelhusain</category><category>sh_reya</category><category>video-generation</category><category>autoformalization</category><category>ai-assisted-coding</category><category>api-design</category><category>context-engineering</category><category>reinforcement-learning</category><category>ai-evals</category><category>hypernetworks</category><category>model-fine-tuning</category><category>foundation-models</category></item><item><title>Execuhires Round 2: Scale-Meta, Lamini-AMD, and Instacart-OpenAI</title><link>https://news.smol.ai/issues/25-06-11-execuhires-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-11-execuhires-2/</guid><description>**Meta** hires **Scale AI&apos;s Alexandr Wang** to lead its new &quot;Superintelligence&quot; division following a **$15 billion investment** for a 49% stake in Scale. **Lamini&apos;s Sharon Zhou** joins **AMD** as VP of AI under Lisa Su, while **Instacart&apos;s Fidji Simo** becomes CEO of Apps at **OpenAI** under **Sama**. **Meta** offers over **$10 million/year compensation packages** to top researchers, successfully recruiting **Jack Rae** from **Gemini**. **OpenAI** releases **o3-pro** model to **ChatGPT Pro** users and API, outperforming **o3** and setting new benchmarks like **Extended NYT Connections** and **SnakeBench**. Despite being slower than **o1-pro**, **o3-pro** excels in reasoning and complex problem-solving. **OpenAI** cuts **o3** pricing by **80%**, making it cheaper than **GPT-4o** and pressuring competitors like **Google** and **Anthropic** to lower prices. Users can now fine-tune the **GPT-4.1** family using **direct preference optimization (DPO)** for subjective tasks.</description><pubDate>Wed, 11 Jun 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>scale-ai</category><category>lamini</category><category>amd</category><category>openai</category><category>gemini</category><category>google</category><category>anthropic</category><category>o3-pro</category><category>o3</category><category>o1-pro</category><category>gpt-4o</category><category>gpt-4.1</category><category>gpt-4.1-mini</category><category>gpt-4.1-nano</category><category>alexandr_wang</category><category>sharon_zhou</category><category>fidji_simo</category><category>sama</category><category>jack_rae</category><category>markchen90</category><category>kevinweil</category><category>gdb</category><category>gregkamradt</category><category>lechmazur</category><category>wesrothmoney</category><category>paul_cal</category><category>imjaredz</category><category>cto_junior</category><category>johnowhitaker</category><category>polynoamial</category><category>scaling01</category><category>model-release</category><category>benchmarking</category><category>reasoning</category><category>fine-tuning</category><category>pricing</category><category>model-performance</category><category>direct-preference-optimization</category><category>complex-problem-solving</category></item><item><title>Reasoning Price War 2: Mistral Magistral + o3&apos;s 80% price cut + o3-pro</title><link>https://news.smol.ai/issues/25-06-10-o3-cut/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-10-o3-cut/</guid><description>**OpenAI** announced an **80% price cut** for its **o3** model, making it competitively priced with **GPT-4.1** and rivaling **Anthropic&apos;s Claude 4 Sonnet** and **Google&apos;s Gemini 2.5 Pro**. Alongside, **o3-pro** was released as a more powerful and reliable variant, though early benchmarks showed mixed performance relative to cost. **Mistral AI** launched its **Magistral** reasoning models, including an open-source **24B parameter** version optimized for efficient deployment on consumer GPUs. The price reduction and new model releases signal intensified competition in reasoning-focused large language models, with notable improvements in token efficiency and cost-effectiveness.</description><pubDate>Tue, 10 Jun 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>mistral-ai</category><category>perplexity-ai</category><category>o3</category><category>o3-pro</category><category>gpt-4.1</category><category>claude-4-sonnet</category><category>gemini-2.5-pro</category><category>magistral-small</category><category>magistral-medium</category><category>mistral-small-3.1</category><category>swyx</category><category>sama</category><category>scaling01</category><category>polynoamial</category><category>nrehiew_</category><category>kevinweil</category><category>gdb</category><category>flavioad</category><category>stevenheidel</category><category>aravsrinivas</category><category>reasoning</category><category>token-efficiency</category><category>price-cut</category><category>benchmarking</category><category>open-source</category><category>model-releases</category><category>context-windows</category><category>gpu-optimization</category></item><item><title>Apple exposes Foundation Models API and... no new Siri</title><link>https://news.smol.ai/issues/25-06-09-apple-letdown/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-09-apple-letdown/</guid><description>**Apple** released on-device foundation models for iOS developers, though their recent &quot;Illusion of Reasoning&quot; paper faced significant backlash for flawed methodology regarding LLM reasoning. **OpenAI** updated **ChatGPT&apos;s Advanced Voice Mode** with more natural voice and improved translation, demonstrated by Greg Brockman. **LangChain** and **LlamaIndex** launched new AI agents and tools, including a SWE Agent for software automation and an Excel agent using reinforcement learning for data transformation. The AI community engaged in heated debate over reasoning capabilities of LLMs, highlighting challenges in evaluation methods.</description><pubDate>Mon, 09 Jun 2025 05:44:39 GMT</pubDate><category>apple</category><category>openai</category><category>langchain</category><category>llamaindex</category><category>chatgpt</category><category>gdb</category><category>scaling01</category><category>giffmana</category><category>kevinweil</category><category>on-device-ai</category><category>foundation-models</category><category>reasoning</category><category>reinforcement-learning</category><category>voice</category><category>translation</category><category>software-automation</category><category>agentic-workflows</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-06-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-06-not-much/</guid><description>**China&apos;s Xiaohongshu (Rednote) released dots.llm1**, a **142B parameter open-source Mixture-of-Experts (MoE) language model** with **14B active parameters** and a **32K context window**, pretrained on **11.2 trillion high-quality, non-synthetic tokens**. The model supports efficient inference frameworks like Docker, HuggingFace, and vLLM, and provides intermediate checkpoints every 1 trillion tokens, enabling flexible fine-tuning. Benchmarking claims it slightly surpasses **Qwen3 235B** on MMLU, though some concerns exist about benchmark selection and synthetic data verification. The release is notable for its truly open-source licensing and no synthetic data usage, sparking community optimism for support in frameworks such as llama.cpp and mlx.</description><pubDate>Fri, 06 Jun 2025 05:44:39 GMT</pubDate><category>xiaohongshu</category><category>rednote-hilab</category><category>deepseek</category><category>huggingface</category><category>dots-llm1</category><category>qwen3-235b</category><category>mixture-of-experts</category><category>open-source</category><category>model-benchmarking</category><category>fine-tuning</category><category>inference</category><category>context-windows</category><category>training-data</category><category>model-architecture</category><category>model-performance</category><category>model-optimization</category></item><item><title>Gemini 2.5 Pro (06-05) launched at AI Engineer World&apos;s Fair</title><link>https://news.smol.ai/issues/25-06-05-aia/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-05-aia/</guid><description>At the second day of **AIE**, **Google&apos;s Gemini 2.5 Pro** reclaimed the top spot on the LMArena leaderboard with a score of **1470** and a +24 Elo increase, showing improvements in coding, reasoning, and math. **Qwen3** released state-of-the-art embedding and reranking models, with **Qwen3-Embedding-8B** topping the MTEB multilingual leaderboard. **OpenThinker3-7B** emerged as the top open reasoning model trained on the **OpenThoughts3-1.2M dataset**, outperforming previous models by 33%. **LightOn** introduced **FastPlaid**, achieving up to a 554% speedup for late-interaction models. **Morph Labs** hired **Christian Szegedy** as Chief Scientist to lead Verified Superintelligence development. The **AI Engineer World&apos;s Fair** featured a fireside chat with **Greg Brockman** and **NVIDIA CEO Jensen Huang**, highlighting the return of basic research and engineering best practices.</description><pubDate>Thu, 05 Jun 2025 05:44:39 GMT</pubDate><category>google</category><category>qwen</category><category>lighton</category><category>morph-labs</category><category>openai</category><category>nvidia</category><category>gemini-2.5-pro</category><category>qwen3-embedding-8b</category><category>openthinker3-7b</category><category>greg_brockman</category><category>jensen_huang</category><category>christian_szegedy</category><category>swyx</category><category>benchmarking</category><category>reasoning</category><category>coding</category><category>math</category><category>embedding-models</category><category>late-interaction</category><category>dataset-release</category><category>model-performance</category><category>model-architecture</category><category>ai-conferences</category></item><item><title>AI Engineer World&apos;s Fair Talks Day 1</title><link>https://news.smol.ai/issues/25-06-04-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-04-not-much/</guid><description>**Mistral** launched a new **Code** project, and **Cursor** released version **1.0**. **Anthropic** improved **Claude Code** plans, while **ChatGPT** announced expanded connections. The day was dominated by **AIE** keynotes and tracks including **GraphRAG**, **RecSys**, and **Tiny Teams**. On Reddit, **Google** open-sourced the **DeepSearch** stack for building AI agents with **Gemini 2.5** and **LangGraph**, enabling flexible agent architectures and integration with local LLMs like **Gemma**. A new **Meta** paper analyzed language model memorization, showing GPT-style transformers store about **3.5–4 bits/parameter** and exploring the transition from memorization to generalization, with implications for **Mixture-of-Experts** models and quantization effects.</description><pubDate>Wed, 04 Jun 2025 05:44:39 GMT</pubDate><category>mistral</category><category>cursor</category><category>anthropic</category><category>openai</category><category>aie</category><category>google-deepmind</category><category>meta-ai-fair</category><category>gemini-2.5</category><category>gemma</category><category>claude-code</category><category>agent-based-architecture</category><category>open-source</category><category>model-memorization</category><category>scaling-laws</category><category>quantization</category><category>mixture-of-experts</category><category>language-model-memorization</category><category>model-generalization</category><category>langgraph</category><category>model-architecture</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-03-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-03-not-much/</guid><description>**OpenAI** rolled out **Codex** to ChatGPT Plus users with internet access and fine-grained controls, improving memory features for free users. **Anthropic&apos;s Claude 4 Opus and Sonnet** models lead coding benchmarks, while **Google&apos;s Gemini 2.5 Pro and Flash** models gain recognition with new audio capabilities. **Qwen 2.5-VL** and **Qwen 3** quantizations are noted for versatility and support. **Bing Video Creator** launched globally enabling text-to-video generation, and **Perplexity Labs** sees increased demand for travel search. New agentic AI tools and RAG innovations include **LlamaCloud** and **FedRAG**. Open-source releases include **Holo-1** for web navigation and **PlayAI&apos;s PlayDiffusion** for speech editing. Audio and multimodal advances feature **Suno&apos;s** music editing upgrades, **Google&apos;s** native TTS in 24+ languages, and **Universal Streaming&apos;s** ultra-low latency speech-to-text. **Google NotebookLM** now supports public notebooks. *&quot;Codex&apos;s internet access brings tradeoffs, with explicit warnings about risk&quot;* and *&quot;Gemini 2.5 Pro is cited as a daily driver by users&quot;*.</description><pubDate>Tue, 03 Jun 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>google</category><category>perplexity-ai</category><category>bing</category><category>playai</category><category>suno</category><category>hugging-face</category><category>langchain-ai</category><category>qwen</category><category>mlx</category><category>assemblyai</category><category>llamacloud</category><category>codex</category><category>claude-4-opus</category><category>claude-4-sonnet</category><category>gemini-2.5-pro</category><category>gemini-2.5</category><category>qwen-2.5-vl</category><category>qwen-3</category><category>playdiffusion</category><category>sama</category><category>gdb</category><category>kevinweil</category><category>lmarena_ai</category><category>epochairesearch</category><category>reach_vb</category><category>wightmanr</category><category>deeplearningai</category><category>mervenoyann</category><category>awnihannun</category><category>jordirib1</category><category>aravsrinivas</category><category>omarsar0</category><category>lioronai</category><category>jerryjliu0</category><category>nerdai</category><category>tonywu_71</category><category>_akhaliq</category><category>clementdelangue</category><category>_mfelfel</category><category>fine-tuning</category><category>model-benchmarking</category><category>text-to-video</category><category>agentic-ai</category><category>retrieval-augmented-generation</category><category>open-source-models</category><category>speech-editing</category><category>audio-processing</category><category>text-to-speech</category><category>ultra-low-latency</category><category>multimodality</category><category>public-notebooks</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-06-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-06-02-not-much/</guid><description>**DeepSeek R1-0528** release brings major improvements in reasoning, hallucination reduction, JSON output, and function calling, matching or surpassing closed models like **OpenAI o3** and **Gemini 2.5 Pro** on benchmarks such as **Artificial Analysis Intelligence Index**, **LiveBench**, and **GPQA Diamond**. The model ranks #2 globally in open weights intelligence, surpassing **Meta AI**, **Anthropic**, and **xAI**. Open weights and technical transparency have fueled rapid adoption across platforms like **Ollama** and **Hugging Face**. Chinese AI labs including **DeepSeek**, **Alibaba**, **ByteDance**, and **Xiaomi** now match or surpass US labs in model releases and intelligence, driven by open weights strategies. Reinforcement learning post-training is critical for intelligence gains, mirroring trends seen at **OpenAI**. Optimized quantization techniques (1-bit, 4-bit) and local inference enable efficient experimentation on consumer hardware. New benchmarks like **LisanBench** test knowledge, planning, memory, and long-context reasoning, with **OpenAI o3** and **Claude Opus 4** leading. Discussions highlight concerns about benchmark contamination and overemphasis on RL-tuned gains.</description><pubDate>Mon, 02 Jun 2025 05:44:39 GMT</pubDate><category>deepseek_ai</category><category>openai</category><category>gemini</category><category>meta-ai-fair</category><category>anthropic</category><category>x-ai</category><category>ollama</category><category>hugging-face</category><category>alibaba</category><category>bytedance</category><category>xiaomi</category><category>deepseek-r1-0528</category><category>o3</category><category>gemini-2.5-pro</category><category>claude-opus-4</category><category>teortaxestex</category><category>wenfeng</category><category>danielhanchen</category><category>awnihannun</category><category>reach_vb</category><category>abacaj</category><category>reasoning</category><category>reinforcement-learning</category><category>benchmarking</category><category>quantization</category><category>local-inference</category><category>model-evaluation</category><category>open-weights</category><category>transparency</category><category>post-training</category><category>agentic-benchmarks</category><category>long-context</category><category>hallucination-detection</category></item><item><title>Mary Meeker is so back: BOND Capital AI Trends report</title><link>https://news.smol.ai/issues/25-05-30-mary-meeker/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-30-mary-meeker/</guid><description>**Mary Meeker** returns with a comprehensive **340-slide report** on the state of AI, highlighting accelerating tech cycles, compute growth, and comparisons of **ChatGPT** to early Google and other iconic tech products. The report also covers enterprise traction and valuation of major AI companies. On Twitter, **@tri_dao** discusses an &quot;ideal&quot; inference architecture featuring attention variants like **GTA**, **GLA**, and **DeepSeek MLA** with high arithmetic intensity (~256), improving efficiency and model quality. Other highlights include the release of **4-bit DWQ of DSR1 Qwen3 8B** on Hugging Face, **AnthropicAI**&apos;s open-source interpretability tools for LLMs, and discussions on transformer training and abstractions by various researchers.</description><pubDate>Sat, 31 May 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>hugging-face</category><category>deepseek</category><category>qwen-3-8b</category><category>tri_dao</category><category>fleetwood___</category><category>teortaxestex</category><category>awnihannun</category><category>lateinteraction</category><category>neelnanda5</category><category>eliebakouch</category><category>_akhaliq</category><category>attention-mechanisms</category><category>inference</category><category>arithmetic-intensity</category><category>transformers</category><category>model-optimization</category><category>interpretability</category><category>model-quantization</category><category>training</category></item><item><title>DeepSeek-R1-0528 - Gemini 2.5 Pro-level model, SOTA Open Weights release</title><link>https://news.smol.ai/issues/25-05-29-deepseek-r1-0528/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-29-deepseek-r1-0528/</guid><description>**DeepSeek R1-0528** marks a significant upgrade, closing the gap with proprietary models like **Gemini 2.5 Pro** and surpassing benchmarks from **Anthropic**, **Meta**, **NVIDIA**, and **Alibaba**. This Chinese open-weights model leads in several AI benchmarks, driven by reinforcement learning post-training rather than architecture changes, and demonstrates increased reasoning token usage (23K tokens per question). The China-US AI race intensifies as Chinese labs accelerate innovation through transparency and open research culture. Key benchmarks include **AIME 2024**, **LiveCodeBench**, and **GPQA Diamond**.</description><pubDate>Thu, 29 May 2025 05:44:39 GMT</pubDate><category>deepseek-ai</category><category>anthropic</category><category>meta-ai-fair</category><category>nvidia</category><category>alibaba</category><category>google-deepmind</category><category>deepseek-r1-0528</category><category>gemini-2.5-pro</category><category>qwen-3-8b</category><category>qwen-3-235b</category><category>artificialanlys</category><category>scaling01</category><category>cline</category><category>reach_vb</category><category>zizhpan</category><category>andrewyng</category><category>teortaxestex</category><category>teknim1</category><category>lateinteraction</category><category>abacaj</category><category>cognitivecompai</category><category>awnihannun</category><category>reinforcement-learning</category><category>benchmarking</category><category>model-performance</category><category>open-weights</category><category>reasoning</category><category>quantization</category><category>post-training</category><category>model-comparison</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-28-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-28-not-much/</guid><description>**DeepSeek R1 v2** model released with availability on Hugging Face and inference partners. The **Gemma model family** continues prolific development including **PaliGemma 2**, **Gemma 3**, and others. **Claude 4** and its variants like **Opus 4** and **Claude Sonnet 4** show top benchmark performance, including new SOTA on **ARC-AGI-2** and **WebDev Arena**. **Codestral Embed** introduces a 3072-dimensional code embedder. **BAGEL**, an open-source multimodal model by **ByteDance**, supports reading, reasoning, drawing, and editing with long mixed contexts. Benchmarking highlights include **Nemotron-CORTEXA** topping SWEBench and **Gemini 2.5 Pro** performing on VideoGameBench. Discussions on random rewards effectiveness focus on **Qwen** models. *&quot;Opus 4 NEW SOTA ON ARC-AGI-2. It&apos;s happening - I was right&quot;* and *&quot;Claude 4 launch has dev moving at a different pace&quot;* reflect excitement in the community.</description><pubDate>Wed, 28 May 2025 05:44:39 GMT</pubDate><category>deepseek-ai</category><category>huggingface</category><category>gemma</category><category>claude</category><category>bytedance</category><category>qwen</category><category>nemotron</category><category>sakana-ai-labs</category><category>deepseek-r1-0528</category><category>pali-gemma-2</category><category>gemma-3</category><category>shieldgemma-2</category><category>txgemma</category><category>gemma-3-qat</category><category>gemma-3n-preview</category><category>medgemma</category><category>dolphingemma</category><category>signgemma</category><category>claude-4</category><category>opus-4</category><category>claude-sonnet-4</category><category>codestral-embed</category><category>bagel</category><category>qwen</category><category>nemotron-cortexa</category><category>gemini-2.5-pro</category><category>yuchenj_uw</category><category>_akhaliq</category><category>clementdelangue</category><category>osanseviero</category><category>alexalbert__</category><category>guillaumelample</category><category>theturingpost</category><category>lmarena_ai</category><category>epochairesearch</category><category>scaling01</category><category>nrehiew_</category><category>ctnzr</category><category>benchmarking</category><category>model-releases</category><category>multimodality</category><category>code-generation</category><category>model-performance</category><category>long-context</category><category>reinforcement-learning</category><category>model-optimization</category><category>open-source</category></item><item><title>Mistral&apos;s Agents API and the 2025 LLM OS</title><link>https://news.smol.ai/issues/25-05-27-mistral-agents/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-27-mistral-agents/</guid><description>**The LLM OS** concept has evolved since 2023, with **Mistral AI** releasing a new **Agents API** that includes code execution, web search, persistent memory, and agent orchestration. **LangChainAI** introduced the **Open Agent Platform (OAP)**, an open-source no-code platform for intelligent agents. **OpenAI** plans to develop **ChatGPT** into a super-assistant by H1 2025, competing with **Meta**. Discussions around **Qwen** models focus on reinforcement learning effects, while **Claude 4** performance is also noted. The AI Engineer World&apos;s Fair is calling for volunteers.</description><pubDate>Tue, 27 May 2025 05:44:39 GMT</pubDate><category>mistral-ai</category><category>langchain-ai</category><category>openai</category><category>meta-ai-fair</category><category>qwen</category><category>claude-4</category><category>chatgpt</category><category>o3</category><category>o4</category><category>omarsar0</category><category>simonw</category><category>swyx</category><category>scaling01</category><category>agent-frameworks</category><category>multi-agent-systems</category><category>tool-use</category><category>code-execution</category><category>web-search</category><category>model-context-protocol</category><category>persistent-memory</category><category>function-calling</category><category>open-source</category><category>no-code</category><category>reinforcement-learning</category><category>model-performance</category><category>agent-orchestration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-26-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-26-not-much/</guid><description>**OpenAI** plans to evolve **ChatGPT** into a **super-assistant** by 2025 with models like **o3** and **o4** enabling agentic tasks and supporting a billion users. Recent multimodal and reasoning model releases include ByteDance&apos;s **BAGEL-7B**, Google&apos;s **MedGemma**, and NVIDIA&apos;s **ACEReason-Nemotron-14B**. The **Sudoku-Bench Leaderboard** highlights ongoing challenges in AI creative reasoning. In software development, OpenAI&apos;s **Codex** aids code generation and debugging, while Gemini&apos;s **Context URL tool** enhances prompt context. **AgenticSeek** offers a local, privacy-focused alternative for autonomous agents. Ethical concerns are raised about AGI development priorities and Anthropic&apos;s alignment with human values. Technical discussions emphasize emergence in AI and training challenges, with humor addressing misconceptions about **Gemini 3.0** and async programming in C. A novel synthetic speech training method enables instruction tuning of LLMs without real speech data, advancing low-resource language support.</description><pubDate>Mon, 26 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>bytedance</category><category>google</category><category>nvidia</category><category>sakana-ai-labs</category><category>deep-learning-ai</category><category>gemini</category><category>agenticseek</category><category>anthropic</category><category>chatgpt</category><category>o3</category><category>o4</category><category>bagel-7b</category><category>medgemma</category><category>acereason-nemotron-14b</category><category>codex</category><category>gemini</category><category>scaling01</category><category>mervenoyann</category><category>sakananailabs</category><category>_philschmid</category><category>omarsar0</category><category>teortaxestex</category><category>andrewlampinen</category><category>sedielem</category><category>cis_female</category><category>agentic-systems</category><category>multimodality</category><category>reasoning</category><category>code-generation</category><category>prompt-engineering</category><category>privacy</category><category>ethical-ai</category><category>emergence</category><category>synthetic-data</category><category>speech-instruction-tuning</category><category>low-resource-languages</category><category>humor</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-23-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-23-not-much/</guid><description>**Anthropic&apos;s Claude 4 models (Opus 4, Sonnet 4)** demonstrate strong coding abilities, with Sonnet 4 achieving **72.7%** on SWE-bench and Opus 4 at **72.5%**. Claude Sonnet 4 excels in codebase understanding and is considered **SOTA on large codebases**. Criticism arose over Anthropic&apos;s handling of **ASL-3 security requirements**. Demand for Claude 4 is high, with integration into IDEs and support from Cherry Studio and FastHTML. **Google DeepMind** introduced **Gemini 2.5 Pro Deep Think** and **Gemma 3n**, a mobile multimodal model reducing RAM usage by nearly 3x. **Google&apos;s Imagen 4 Ultra** ranks third in the Artificial Analysis Image Arena, available on **Vertex AI Studio**. Google also promoted **Google Beam**, an AI video model for immersive 3D experiences, and new text-to-speech models with multi-speaker support. The **GAIA benchmark** shows Claude 4 Opus and Sonnet leading in agentic performance.</description><pubDate>Fri, 23 May 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>google-deepmind</category><category>openai</category><category>claude-4</category><category>claude-4-opus</category><category>claude-4-sonnet</category><category>gemini-2.5-pro</category><category>gemma-3n</category><category>imagen-4-ultra</category><category>cline</category><category>amanrsanger</category><category>ryanpgreenblatt</category><category>johnschulman2</category><category>alexalbert__</category><category>nearcyan</category><category>mickeyxfriedman</category><category>jeremyphoward</category><category>gneubig</category><category>teortaxesTex</category><category>scaling01</category><category>artificialanlys</category><category>philschmid</category><category>codebase-understanding</category><category>coding</category><category>agentic-performance</category><category>multimodality</category><category>text-to-speech</category><category>video-generation</category><category>model-integration</category><category>benchmarking</category><category>memory-optimization</category></item><item><title>Anthropic releases Claude 4 Sonnet and Opus: Memory, Agent Capabilities, Claude Code, Redteam Drama</title><link>https://news.smol.ai/issues/25-05-22-claude-4/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-22-claude-4/</guid><description>**Anthropic** has officially released **Claude 4** with two variants: **Claude Opus 4**, a high-capability model for complex tasks priced at **$15/$75 per million tokens**, and **Claude Sonnet 4**, optimized for efficient everyday use. The release emphasizes **instruction following** and extended work sessions up to **7 hours**. Community discussions highlight concerns about **token pricing**, **token accounting transparency**, and calls for **open-sourcing Claude 3.5 Sonnet** weights to support local model development. The news also covers **Claude Code GA**, new **Agent Capabilities API**, and various livestreams and reports detailing these updates. There is notable debate around **sliding window attention** and advanced inference techniques for local deployment.</description><pubDate>Thu, 22 May 2025 05:44:39 GMT</pubDate><category>anthropic</category><category>claude-4</category><category>claude-4-opus</category><category>claude-4-sonnet</category><category>claude-3.5-sonnet</category><category>instruction-following</category><category>token-accounting</category><category>pricing-models</category><category>sliding-window-attention</category><category>inference-techniques</category><category>open-sourcing</category><category>model-accessibility</category><category>agent-capabilities-api</category><category>extended-context</category><category>model-deployment</category></item><item><title>OpenAI buys Jony Ive&apos;s io for $6.5b, LMArena lands $100m seed from a16z</title><link>https://news.smol.ai/issues/25-05-21-openai-io/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-21-openai-io/</guid><description>**OpenAI** confirmed a partnership with **Jony Ive** to develop consumer hardware. **LMArena** secured a $100 million seed round from **a16z**. **Mistral** launched a new code model fine-tune. **Google DeepMind** announced multiple updates at **Google I/O 2024**, including over a dozen new models and 20 AI products. Key highlights include the release of **Gemini 2.5 Pro** and **Gemini Diffusion**, featuring advanced multimodal reasoning, coding, and math capabilities, and integration of Gemini in **Google Chrome** as an AI browsing assistant. **Deep Think** enhanced reasoning mode and **Project Astra** improvements were also introduced, focusing on voice output, memory, and computer control for a universal AI assistant.</description><pubDate>Wed, 21 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>lmarena</category><category>a16z</category><category>mistral-ai</category><category>google</category><category>google-deepmind</category><category>gemini-2.5-pro</category><category>gemini-diffusion</category><category>sundar_pichai</category><category>multimodality</category><category>reasoning</category><category>code-generation</category><category>math</category><category>model-fine-tuning</category><category>ai-assistants</category><category>voice</category><category>memory-optimization</category></item><item><title>Google I/O: new Gemini native voice, Flash, DeepThink, AI Mode (DeepSearch+Mariner+Astra)</title><link>https://news.smol.ai/issues/25-05-20-google-io/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-20-google-io/</guid><description>**Google I/O 2024** showcased significant advancements with **Gemini 2.5 Pro** and **Deep Think** reasoning mode from **google-deepmind**, emphasizing AI-driven transformations and developer opportunities. **GeminiApp** aims to become a universal **AI assistant** on the path to **AGI**, with new features like **AI Mode** in Google Search expanding generative AI access. The event included multiple keynotes and updates on over a dozen models and 20+ AI products, highlighting **Google&apos;s** leadership in AI innovation. Influential voices like **demishassabis** and **philschmid** provided insights and recaps, while the launch of **Jules** as a competitor to Codex/Devin was noted.</description><pubDate>Tue, 20 May 2025 05:44:39 GMT</pubDate><category>google</category><category>google-deepmind</category><category>gemini-2.5-pro</category><category>gemini-2.5</category><category>demishassabis</category><category>philschmid</category><category>jack_w_rae</category><category>ai-assistants</category><category>reasoning</category><category>generative-ai</category><category>developer-tools</category><category>ai-integration</category><category>model-optimization</category><category>ai-application</category><category>model-updates</category><category>ai-deployment</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-19-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-19-not-much/</guid><description>**Meta** released **KernelLLM 8B**, outperforming **GPT-4o** and **DeepSeek V3** on KernelBench-Triton Level 1. **Mistral Medium 3** debuted strongly in multiple benchmarks. **Qwen3** models introduced a unified framework with multilingual support. **DeepSeek-V3** features hardware-aware co-design. **BLIP3-o** family released for multimodal tasks using diffusion transformers. **Salesforce** launched **xGen-Small** models excelling in long-context and math benchmarks. **Bilibili** released **AniSORA** for anime video generation. **Stability AI** open-sourced **Stable Audio Open Small** optimized for Arm devices. Google’s **AlphaEvolve** coding agent improved **Strassen&apos;s algorithm** for the first time since 1969. Research shows **chain-of-thought reasoning** can harm instruction-following ability, with mitigation strategies like classifier-selective reasoning being most effective, but reasoning techniques show high variance and limited generalization. *&quot;Chain-of-thought (CoT) reasoning can harm a model’s ability to follow instructions&quot;* and *&quot;Mitigation strategies such as few-shot in-context learning, self-reflection, self-selective reasoning, and classifier-selective reasoning can counteract reasoning-induced failures&quot;*.</description><pubDate>Mon, 19 May 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>mistral-ai</category><category>qwen</category><category>deepseek</category><category>salesforce</category><category>bilibili</category><category>stability-ai</category><category>google</category><category>kernelllm-8b</category><category>gpt-4o</category><category>deepseek-v3</category><category>mistral-medium-3</category><category>qwen3</category><category>blip3-o</category><category>xgen-small</category><category>anisora</category><category>stable-audio-open-small</category><category>alphaevolve</category><category>reach_vb</category><category>lmarena_ai</category><category>theadimeline</category><category>adcock_brett</category><category>jxmnop</category><category>dair_ai</category><category>omarsar0</category><category>benchmarking</category><category>model-performance</category><category>multilinguality</category><category>hardware-optimization</category><category>multimodality</category><category>image-generation</category><category>video-generation</category><category>text-to-audio</category><category>model-parallelism</category><category>chain-of-thought</category><category>instruction-following</category><category>reasoning</category><category>mitigation-strategies</category></item><item><title>ChatGPT Codex, OpenAI&apos;s first cloud SWE agent</title><link>https://news.smol.ai/issues/25-05-16-codex/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-16-codex/</guid><description>**OpenAI** launched **Codex**, a cloud-based software engineering agent powered by **codex-1** (an optimized version of **OpenAI o3**) available in research preview for Pro, Enterprise, and Team ChatGPT users, featuring parallel task execution like refactoring and bug fixing. The **Codex CLI** was enhanced with quick sign-in and a new low-latency model, **codex-mini**. **Gemma 3** is highlighted as the best open model runnable on a single GPU. **Runway** released the Gen-4 References API for style transfer in generation. **Salesforce** introduced **BLIP3-o**, a unified multimodal model family using diffusion transformers for CLIP image features. The **Qwen 2.5** models (1.5B and 3B versions) were integrated into the PocketPal app with various chat templates. **Marigold IID**, a new state-of-the-art open-source depth estimation model, was released. 

In research, **DeepSeek** shared insights on scaling and hardware for DeepSeek-V3. **Google** unveiled **LightLab**, a diffusion-based light source control in images. **Google DeepMind&apos;s AlphaEvolve** uses **Gemini 2.0** to discover new math and reduce costs without reinforcement learning. **Omni-R1** studied audio&apos;s role in fine-tuning audio LLMs. **Qwen** proposed a parallel scaling law inspired by classifier-free guidance. **Salesforce** released **Lumina-Next** on the Qwen base, outperforming Janus-Pro. A study found LLM performance degrades in multi-turn conversations due to unreliability. **J1** is incentivizing LLM-as-a-Judge thinking via reinforcement learning. A new Qwen study correlates question and strategy similarity to predict reasoning strategies.</description><pubDate>Fri, 16 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>runway</category><category>salesforce</category><category>qwen</category><category>deepseek</category><category>google</category><category>google-deepmind</category><category>j1</category><category>codex-1</category><category>openai-o3</category><category>codex-mini</category><category>gemma-3</category><category>blip3-o</category><category>qwen-2.5</category><category>marigold-iid</category><category>deepseek-v3</category><category>lightlab</category><category>gemini-2.0</category><category>lumina-next</category><category>sama</category><category>kevinweil</category><category>omarsar0</category><category>iscienceluvr</category><category>akhaliq</category><category>osanseviero</category><category>c_valenzuelab</category><category>mervenoyann</category><category>arankomatsuzaki</category><category>jasonwei</category><category>demishassabis</category><category>philschmid</category><category>swyx</category><category>teortaxestex</category><category>jaseweston</category><category>software-engineering</category><category>parallel-processing</category><category>multimodality</category><category>diffusion-models</category><category>depth-estimation</category><category>scaling-laws</category><category>reinforcement-learning</category><category>fine-tuning</category><category>model-performance</category><category>multi-turn-conversation</category><category>reasoning</category><category>audio-processing</category></item><item><title>Gemini&apos;s AlphaEvolve agent uses Gemini 2.0 to find new Math and cuts Gemini cost 1% — without RL</title><link>https://news.smol.ai/issues/25-05-15-alphaevolve/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-15-alphaevolve/</guid><description>**Deepmind&apos;s AlphaEvolve**, a 2025 update to AlphaTensor and FunSearch, is a Gemini-powered **coding agent for algorithm discovery** that designs faster matrix multiplication algorithms, solves open math problems, and improves data center and AI training efficiency. It achieves a **23% faster kernel speedup** in Gemini training and surpasses state-of-the-art on 20% of applied problems, including improvements on the Minimum Overlap Problem and Kissing number problem. Unlike Deep-RL, it optimizes code pieces rather than model weights. Meanwhile, **OpenAI** released **GPT-4.1** in ChatGPT, specializing in coding and instruction following, with a faster alternative **GPT-4.1 mini** replacing GPT-4o mini for all users. OpenAI also launched the Safety Evaluations Hub and the OpenAI to Z Challenge using o3/o4 mini and GPT-4.1 models to discover archaeological sites. *&quot;Maybe midtrain + good search is all you need for AI for scientific innovation&quot;* - Jason Wei.</description><pubDate>Thu, 15 May 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>gemini</category><category>gpt-4.1</category><category>gpt-4o-mini</category><category>o3</category><category>o4-mini</category><category>_philschmid</category><category>scott_swingle</category><category>alex_dimakis</category><category>henry</category><category>jason_wei</category><category>kevinweil</category><category>michpokrass</category><category>scaling01</category><category>gdb</category><category>algorithm-discovery</category><category>coding-agents</category><category>matrix-multiplication</category><category>optimization</category><category>reinforcement-learning</category><category>model-weights</category><category>training-efficiency</category><category>safety-evaluations</category><category>instruction-following</category><category>coding-tasks</category><category>model-releases</category></item><item><title>Granola launches team notes, while Notion launches meeting transcription</title><link>https://news.smol.ai/issues/25-05-14-notion-granola/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-14-notion-granola/</guid><description>**GPT-4.1** is now available in **ChatGPT** for Plus, Pro, and Team users, focusing on coding and instruction following, with **GPT 4.1 mini** replacing **GPT 4o mini**. **Anthropic** is releasing new **Claude** models including **Claude Opus** and **Claude Sonnet**, though some criticism about hallucinations in **Claude O3** was noted. **Alibaba** shared the **Qwen3 Technical Report** with strong benchmark results from **Seed1.5-VL**. **Meta FAIR** announced new models and datasets but faced criticism on **Llama 4**. **AM-Thinking-v1** launched on **Hugging Face** as a 32B scale reasoning model. **Granola** raised $43M in Series B and launched **Granola 2.0** with a Notion-like UI. The AI ecosystem shows rapid iteration and cloning of ideas, emphasizing execution and distribution.</description><pubDate>Wed, 14 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>alibaba</category><category>meta-ai-fair</category><category>huggingface</category><category>granola</category><category>gpt-4.1</category><category>gpt-4o-mini</category><category>gpt-4.1-mini</category><category>claude-opus</category><category>claude-sonnet</category><category>claude-o3</category><category>qwen3</category><category>seed1.5-vl</category><category>llama-4</category><category>am-thinking-v1</category><category>kevinweil</category><category>scaling01</category><category>steph_palazzolo</category><category>andersonbcdefg</category><category>reach_vb</category><category>yuchenj_uw</category><category>qtnx_</category><category>_akhaliq</category><category>risingsayak</category><category>coding</category><category>instruction-following</category><category>benchmarking</category><category>model-releases</category><category>reasoning</category><category>image-generation</category><category>collaborative-software</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-13-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-13-not-much/</guid><description>**Tencent&apos;s Hunyuan-Turbos** has risen to #8 on the LMArena leaderboard, showing strong performance across major categories and significant improvement since February. The **Qwen3 model family**, especially the **Qwen3 235B-A22B (Reasoning)** model, is noted for its intelligence and efficient parameter usage. **OpenAI** introduced **HealthBench**, a new health evaluation benchmark developed with input from over **250 physicians**, where models like **o3**, **GPT-4.1 nano**, and **Grok 3** showed strong results. **ByteDance** released **Seed1.5-VL**, a vision-language model with a 532M-parameter vision encoder and a 20B active parameter MoE LLM, achieving state-of-the-art results on 38 public benchmarks. In vision-language, **Kling 2.0** leads image-to-video generation, and **Gemini 2.5 Pro** excels in video understanding with advanced multimodal capabilities. Meta&apos;s Vision-Language-Action framework and updates on VLMs for 2025 were also highlighted.</description><pubDate>Tue, 13 May 2025 05:44:39 GMT</pubDate><category>tencent</category><category>openai</category><category>bytedance</category><category>meta-ai-fair</category><category>nvidia</category><category>deepseek</category><category>hunyuan-turbos</category><category>qwen3-235b-a22b</category><category>o3</category><category>gpt-4.1-nano</category><category>grok-3</category><category>gemini-2.5-pro</category><category>seed1.5-vl</category><category>kling-2.0</category><category>lmarena_ai</category><category>artificialanlys</category><category>gdb</category><category>_jasonwei</category><category>iScienceLuvr</category><category>_akhaliq</category><category>_philschmid</category><category>teortaxesTex</category><category>mervenoyann</category><category>reach_vb</category><category>benchmarking</category><category>model-performance</category><category>moe</category><category>reasoning</category><category>vision</category><category>video-understanding</category><category>vision-language</category><category>multimodality</category><category>model-evaluation</category><category>model-optimization</category></item><item><title>Prime Intellect&apos;s INTELLECT-2 and PRIME-RL advance distributed reinforcement learning</title><link>https://news.smol.ai/issues/25-05-12-intellect-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-12-intellect-2/</guid><description>**Prime Intellect** released **INTELLECT-2**, a decentralized GPU training and RL framework with a vision for distributed AI training overcoming colocation limits. **ByteDance** launched **DreamO**, a unified image customization model on Hugging Face. **Qwen** released models optimized for GPTQ, GGUF, and AWQ quantization. **Gemma** surpassed 150 million downloads on Hugging Face. **Meta** released weights for the **Dynamic Byte Latent Transformer** and the **Collaborative Reasoner** framework to improve language model efficiency and reasoning. **RunwayML** introduced **Gen-4 References**, a near-realtime model requiring no fine-tuning. **Mistral AI** released **Mistral Medium 3**, a strong multimodal model, and **Le Chat Enterprise**, an agentic AI assistant for business. **Google** updated **Gemini 2.5 Pro Preview** with video understanding and UI improvements. *&quot;Airbnb for spare GPUs from all over the world&quot;* highlights the ongoing challenges and potential of distributed GPU training.</description><pubDate>Mon, 12 May 2025 05:44:39 GMT</pubDate><category>primeintellect</category><category>bytedance</category><category>qwen</category><category>gemma</category><category>meta-ai-fair</category><category>runwayml</category><category>mistral-ai</category><category>google</category><category>intellect-2</category><category>dreamo</category><category>qwen</category><category>gemini-2.5-pro</category><category>dynamic-byte-latent-transformer</category><category>gen-4-references</category><category>mistral-medium-3</category><category>le-chat-enterprise</category><category>_akhaliq</category><category>reach_vb</category><category>osanseviero</category><category>aiatmeta</category><category>c_valenzuelab</category><category>lmarena_ai</category><category>adcock_brett</category><category>distributed-training</category><category>reinforcement-learning</category><category>gpu-clusters</category><category>model-optimization</category><category>quantization</category><category>multimodality</category><category>agentic-ai</category><category>video-understanding</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-09-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-09-not-much/</guid><description>**Gemini 2.5 Flash** shows a **12 point increase** in the Artificial Analysis Intelligence Index but costs **150x more** than Gemini 2.0 Flash due to **9x more expensive output tokens** and **17x higher token usage** during reasoning. **Mistral Medium 3** competes with **Llama 4 Maverick**, **Gemini 2.0 Flash**, and **Claude 3.7 Sonnet** with better coding and math reasoning at a significantly lower price. **Alibaba&apos;s Qwen3** family supports reasoning and multilingual tasks across **119 languages** and includes a **Web Dev** tool for app building. **Huawei&apos;s Pangu Ultra MoE** matches **DeepSeek R1** performance on Ascend NPUs, with new compute and upcoming V4 training. **OpenAI&apos;s o4-mini** now supports **Reinforcement Fine-Tuning (RFT)** using chain-of-thought reasoning. **Microsoft&apos;s X-REASONER** enables generalizable reasoning across modalities post-trained on general-domain text. Deep research integration with GitHub repos in ChatGPT enhances codebase search and reporting. The AI Engineer World&apos;s Fair offers an Early Bird discount for upcoming tickets.</description><pubDate>Fri, 09 May 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>mistral-ai</category><category>alibaba</category><category>huawei</category><category>openai</category><category>microsoft</category><category>deepseek</category><category>gemini-2.5-flash</category><category>gemini-2.0-flash</category><category>mistral-medium-3</category><category>llama-4-maverick</category><category>claude-3.7-sonnet</category><category>qwen3</category><category>pangu-ultra-moe</category><category>deepseek-r1</category><category>o4-mini</category><category>x-reasoner</category><category>giffmana</category><category>artificialanlys</category><category>teortaxestex</category><category>akhaliq</category><category>john__allard</category><category>model-performance</category><category>reasoning</category><category>cost-analysis</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>multilinguality</category><category>code-search</category><category>model-training</category><category>vision</category><category>model-integration</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-08-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-08-not-much/</guid><description>**OpenAI** launched both **Reinforcement Finetuning** and **Deep Research on GitHub repos**, drawing comparisons to **Cognition&apos;s DeepWiki**. **Nvidia** open-sourced **Open Code Reasoning models (32B, 14B, 7B)** with Apache 2.0 license, showing 30% better token efficiency and compatibility with llama.cpp, vLLM, transformers, and TGI. Independent evaluations highlight **Mistral Medium 3** rivaling **Llama 4 Maverick**, **Gemini 2.0 Flash**, and **Claude 3.7 Sonnet** in coding and math reasoning, priced significantly lower but no longer open-source. **Google&apos;s Gemini 2.5 Pro** is noted as their most intelligent model with improved coding from simple prompts, while **Gemini 2.5 Flash** incurs a 150x cost increase over Gemini 2.0 Flash due to higher token usage and cost. The **Absolute Zero Reasoner (AZR)** achieves SOTA performance in coding and math reasoning via reinforced self-play without external data. Vision-language model **X-REASONER** is post-trained on general-domain text for reasoning. **Apple ML research** released **FastVLM** with on-device iPhone demo. **HiDream LoRA trainer** supports QLoRA fine-tuning under memory constraints. **Nvidia&apos;s Parakeet ASR model** tops Hugging Face ASR leaderboard with MLX implementation. New datasets **SwallowCode** and **SwallowMath** boost LLM performance in math and code. Overall, a quiet day with significant model releases and performance insights.</description><pubDate>Thu, 08 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>nvidia</category><category>mistral-ai</category><category>google</category><category>apple</category><category>huggingface</category><category>open-code-reasoning-32b</category><category>open-code-reasoning-14b</category><category>open-code-reasoning-7b</category><category>mistral-medium-3</category><category>llama-4-maverick</category><category>gemini-2.5-pro</category><category>gemini-2.5-flash</category><category>claude-3.7-sonnet</category><category>absolute-zero-reasoner</category><category>x-reasoner</category><category>fastvlm</category><category>parakeet-asr</category><category>reach_vb</category><category>artificialanlys</category><category>scaling01</category><category>iscienceluvr</category><category>arankomatsuzaki</category><category>awnihannun</category><category>risingsayak</category><category>reinforcement-learning</category><category>fine-tuning</category><category>code-generation</category><category>reasoning</category><category>vision</category><category>on-device-ai</category><category>model-performance</category><category>dataset-release</category><category>model-optimization</category></item><item><title>AI Engineer World&apos;s Fair: Second Run, Twice The Fun</title><link>https://news.smol.ai/issues/25-05-07-aiewf-2025/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-07-aiewf-2025/</guid><description>**The 2025 AI Engineer World&apos;s Fair** is expanding with **18 tracks** covering topics like **Retrieval + Search**, **GraphRAG**, **RecSys**, **SWE-Agents**, **Agent Reliability**, **Reasoning + RL**, **Voice AI**, **Generative Media**, **Infrastructure**, **Security**, and **Evals**. New focuses include **MCP**, **Tiny Teams**, **Product Management**, **Design Engineering**, and **Robotics and Autonomy** featuring foundation models from **Waymo**, **Tesla**, and **Google**. The event highlights the growing importance of **AI Architects** and enterprise AI leadership. Additionally, **Demis Hassabis** announced the **Gemini 2.5 Pro Preview &apos;I/O edition&apos;**, which leads coding and web development benchmarks on **LMArena**.</description><pubDate>Wed, 07 May 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>waymo</category><category>tesla</category><category>anthropic</category><category>braintrust</category><category>gemini-2.5-pro</category><category>demishassabis</category><category>retrieval-augmentation</category><category>graph-databases</category><category>recommendation-systems</category><category>software-engineering-agents</category><category>agent-reliability</category><category>reinforcement-learning</category><category>voice</category><category>image-generation</category><category>video-generation</category><category>infrastructure</category><category>security</category><category>evaluation</category><category>ai-leadership</category><category>enterprise-ai</category><category>mcp</category><category>tiny-teams</category><category>product-management</category><category>design-engineering</category><category>robotics</category><category>foundation-models</category><category>coding</category><category>web-development</category></item><item><title>Gemini 2.5 Pro Preview 05-06 (I/O edition) - the SOTA vision+coding model</title><link>https://news.smol.ai/issues/25-05-06-gemini-2-5-pro/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-06-gemini-2-5-pro/</guid><description>**Gemini 2.5 Pro** has been updated with enhanced multimodal image-to-code capabilities and dominates the WebDev Arena Leaderboard, surpassing **Claude 3.7 Sonnet** in coding and other tasks. **Nvidia** released the **Llama-Nemotron** model family on Hugging Face, noted for efficient reasoning and inference. **Alibaba&apos;s Qwen3** models range from 0.6B to 235B parameters, including dense and MoE variants. **KerasRS** was released by **Franois Chollet** as a new recommender system library compatible with JAX, PyTorch, and TensorFlow, optimized for TPUs. These updates highlight advancements in coding, reasoning, and speech recognition models.</description><pubDate>Tue, 06 May 2025 05:44:39 GMT</pubDate><category>google-deepmind</category><category>nvidia</category><category>alibaba</category><category>hugging-face</category><category>gemini-2.5-pro</category><category>claude-3.7-sonnet</category><category>llama-nemotron</category><category>qwen3</category><category>demishassabis</category><category>_philschmid</category><category>lmarena_ai</category><category>scaling01</category><category>fchollet</category><category>multimodality</category><category>coding</category><category>reasoning</category><category>model-release</category><category>speech-recognition</category><category>recommender-systems</category><category>benchmarking</category></item><item><title>Cursor @ $9b, OpenAI Buys Windsurf @ $3b</title><link>https://news.smol.ai/issues/25-05-05-cursor-openai-windsurf/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-05-cursor-openai-windsurf/</guid><description>**OpenAI** is reportedly close to closing a deal with Windsurf, coinciding with **Cursor&apos;s** $900M funding round at a $9B valuation. **Nvidia** launched the **Llama-Nemotron series** featuring models from 8B to 253B parameters, praised for reasoning and inference efficiency. **Alibaba** released the **Qwen3 family** with MoE and dense models up to 235B parameters, ranking highly in coding and math benchmarks. **DeepSeek** introduced **Prover-V2**, an open-source AI for math reasoning with an 88.9% pass rate on MiniF2F-test. **Microsoft** released reasoning-focused **Phi-4 models**, outperforming OpenAI&apos;s **o1-mini**. **Baidu** debuted turbo versions of **ERNIE 4.5 and X1** for faster, cheaper inference. **Suno v4.5** added advanced AI music generation features, while **Runway Gen-4 References** enable placing characters into scenes with high consistency. **KerasRS**, a new recommender system library optimized for TPUs, was released by **Franois Chollet**.</description><pubDate>Mon, 05 May 2025 05:44:39 GMT</pubDate><category>openai</category><category>cursor</category><category>nvidia</category><category>alibaba</category><category>deepseek</category><category>microsoft</category><category>baidu</category><category>suno</category><category>runway</category><category>keras</category><category>llama-nemotron-ultra</category><category>llama-nemotron-super</category><category>llama-nemotron-nano</category><category>qwen3-235b-a22b</category><category>prover-v2</category><category>phi-4-reasoning</category><category>ernie-4.5-turbo</category><category>ernie-x1-turbo</category><category>suno-v4.5</category><category>gen-4-references</category><category>o1-mini</category><category>_akhaliq</category><category>adcock_brett</category><category>lmarena_ai</category><category>fchollet</category><category>reasoning</category><category>inference-efficiency</category><category>open-license</category><category>moe-models</category><category>math-reasoning</category><category>theorem-proving</category><category>model-performance</category><category>music-generation</category><category>image-generation</category><category>recommender-systems</category><category>tpu-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-02-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-02-not-much/</guid><description>**Qwen model family** released quantized versions of Qwen3 models including **14B**, **32B**, and **235B** parameters, with promising coding capabilities in Qwen3-235B. **Microsoft** launched **Phi-4-reasoning**, a **14B** parameter model distilled from OpenAI&apos;s o3-mini, emphasizing supervised fine-tuning and reinforcement learning, outperforming larger models in some benchmarks. **Cohere&apos;s Command A** leads SQL performance on Bird Bench. **Google** introduced the **TRAJAN** eval for video generation temporal consistency and updated the **Gemini** OpenAI compatibility layer. **Inception Labs** launched a diffusion LLM API claiming 5x speed improvements over autoregressive models. Community rankings show **OpenAI&apos;s o3** model debuting strongly in web app-building tasks. Other releases include **AllenAI&apos;s OLMo2 1B** and additional Phi 4 variants. *&quot;Qwen3-235B shows promise for coding&quot;* and *&quot;Phi-4-reasoning tech report emphasizes SFT gains&quot;* highlight key advancements.</description><pubDate>Fri, 02 May 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>together-ai</category><category>scaling01</category><category>microsoft</category><category>deepseek</category><category>cohere</category><category>google</category><category>epoch-ai-research</category><category>inception-labs</category><category>openai</category><category>allenai</category><category>qwen3-14b</category><category>qwen3-32b</category><category>qwen3-235b</category><category>phi-4-reasoning</category><category>o3-mini</category><category>command-a</category><category>gemini-2.5-pro</category><category>o4-mini</category><category>olm-o2-1b</category><category>o3</category><category>cline</category><category>_philschmid</category><category>iscienceluvr</category><category>alexalbert__</category><category>_lewtun</category><category>teortaxestex</category><category>sarahookr</category><category>reach_vb</category><category>quantization</category><category>fine-tuning</category><category>reinforcement-learning</category><category>benchmarking</category><category>video-generation</category><category>diffusion-models</category><category>model-performance</category><category>model-evaluation</category><category>model-release</category><category>text-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-05-01-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-05-01-not-much/</guid><description>**Microsoft** released **Phi-reasoning 4**, a finetuned 14B reasoning model slightly behind QwQ but limited by data transparency and token efficiency issues. **Anthropic** introduced remote MCP server support and a 45-minute Research mode in **Claude**. **Cursor** published a model popularity list. **Alibaba** launched **Qwen3-235B** and other Qwen3 variants, highlighting budget-friendly coding and reasoning capabilities, with availability on **Together AI** API. **Microsoft** also released **Phi-4-Mini-Reasoning** with benchmark performance on AIME 2025 and OmniMath. **DeepSeek** announced **DeepSeek-Prover V2** with state-of-the-art math problem solving, scaling to 671B parameters. **Meta AI**&apos;s **Llama** models hit 1.2 billion downloads, with new **Llama Guard 4** and **Prompt Guard 2** for input/output filtering and jailbreak prevention. **Xiaomi** released the open-source reasoning model **MiMo-7B** trained on 25 trillion tokens. Discussions on AI model evaluation highlighted issues with the **LMArena leaderboard**, data access biases favoring proprietary models, and challenges in maintaining fair benchmarking, with suggestions for alternatives like **OpenRouterAI** rankings. *&quot;LMArena slop and biased&quot;* and *&quot;61.3% of all data going to proprietary model providers&quot;* were noted concerns.</description><pubDate>Thu, 01 May 2025 05:44:39 GMT</pubDate><category>microsoft</category><category>anthropic</category><category>cursor</category><category>alibaba</category><category>togethercompute</category><category>deepseek</category><category>meta-ai-fair</category><category>xiaomi</category><category>openrouterai</category><category>cohere</category><category>phi-4</category><category>phi-4-mini-reasoning</category><category>qwen3-235b</category><category>qwen3-moe-235b</category><category>qwen3-moe-30b</category><category>qwen3-dense-32b</category><category>qwen3-dense-14b</category><category>qwen3-dense-8b</category><category>qwen3-dense-4b</category><category>qwen3-dense-0.6b</category><category>qwen2.5-omni-3b</category><category>deepseek-prover-v2</category><category>llama</category><category>llama-guard-4</category><category>prompt-guard-2</category><category>mimo-7b</category><category>cline</category><category>reach_vb</category><category>vipulved</category><category>akhaliq</category><category>omarsar0</category><category>zhs05232838</category><category>huajian_xin</category><category>mervenoyann</category><category>karpathy</category><category>random_walker</category><category>sarahookr</category><category>blancheminerva</category><category>clefourrier</category><category>reasoning</category><category>model-fine-tuning</category><category>model-evaluation</category><category>benchmarking</category><category>model-popularity</category><category>open-source</category><category>math</category><category>model-scaling</category><category>model-filtering</category><category>jailbreak-prevention</category></item><item><title>ChatGPT responds to GlazeGate + LMArena responds to Cohere</title><link>https://news.smol.ai/issues/25-04-30-glazegate/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-30-glazegate/</guid><description>**OpenAI** faced backlash after a controversial ChatGPT update, leading to an official retraction admitting they &quot;focused too much on short-term feedback.&quot; Researchers from **Cohere** published a paper criticizing **LMArena** for unfair practices favoring incumbents like **OpenAI**, **DeepMind**, **X.ai**, and **Meta AI Fair**. The **Qwen3 family** by **Alibaba** was released, featuring models up to **235B MoE**, supporting **119 languages** and trained on **36 trillion tokens**, with integration into **vLLM** and support in tools like **llama.cpp**. Meta announced the second round of **Llama Impact Grants** to promote open-source AI innovation. Discussions on AI Twitter highlighted concerns about leaderboard overfitting and fairness in model benchmarking, with notable commentary from **karpathy** and others.</description><pubDate>Wed, 30 Apr 2025 15:44:39 GMT</pubDate><category>openai</category><category>cohere</category><category>lm-arena</category><category>deepmind</category><category>x-ai</category><category>meta-ai-fair</category><category>alibaba</category><category>vllm</category><category>llamaindex</category><category>qwen3-235b-a22b</category><category>qwen3</category><category>qwen3-moe</category><category>llama-4</category><category>joannejang</category><category>arankomatsuzaki</category><category>karpathy</category><category>sarahookr</category><category>reach_vb</category><category>model-releases</category><category>model-benchmarking</category><category>performance-evaluation</category><category>open-source</category><category>multilinguality</category><category>model-integration</category><category>fine-tuning</category><category>model-optimization</category></item><item><title>LlamaCon: Meta AI gets into the Llama API platform business</title><link>https://news.smol.ai/issues/25-04-29-llamacon/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-29-llamacon/</guid><description>**Meta** celebrated progress in the **Llama** ecosystem at LlamaCon, launching an AI Developer platform with finetuning and fast inference powered by **Cerebras** and **Groq** hardware, though it remains waitlisted. Meanwhile, **Alibaba** released the **Qwen3** family of large language models, including **two MoE models** and **six dense models** ranging from **0.6B to 235B parameters**, with the flagship **Qwen3-235B-A22B** achieving competitive benchmark results and supporting **119 languages and dialects**. The Qwen3 models are optimized for coding and agentic capabilities, are Apache 2.0 licensed, and have broad deployment support including local usage with tools like **vLLM**, **Ollama**, and **llama.cpp**. Community feedback highlights Qwen3&apos;s scalable performance and superiority over models like OpenAI&apos;s **o3-mini**.</description><pubDate>Tue, 29 Apr 2025 05:44:39 GMT</pubDate><category>meta-ai-fair</category><category>cerebras</category><category>groq</category><category>alibaba</category><category>vllm</category><category>ollama</category><category>llamaindex</category><category>hugging-face</category><category>llama-cpp</category><category>llama-4</category><category>qwen3</category><category>qwen3-235b-a22b</category><category>qwen3-30b-a3b</category><category>qwen3-4b</category><category>qwen2-5-72b-instruct</category><category>o3-mini</category><category>reach_vb</category><category>huybery</category><category>teortaxestex</category><category>awnihannun</category><category>thezachmueller</category><category>model-release</category><category>fine-tuning</category><category>reinforcement-learning</category><category>moe</category><category>multilingual-models</category><category>model-optimization</category><category>model-deployment</category><category>coding</category><category>benchmarking</category><category>apache-license</category></item><item><title>Qwen 3: 0.6B to 235B MoE full+base models that beat R1 and o1</title><link>https://news.smol.ai/issues/25-04-28-qwen-3/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-28-qwen-3/</guid><description>**Qwen 3** has been released by **Alibaba** featuring a range of models including two MoE variants, **Qwen3-235B-A22B** and **Qwen3-30B-A3B**, which demonstrate competitive performance against top models like **DeepSeek-R1**, **o1**, **o3-mini**, **Grok-3**, and **Gemini-2.5-Pro**. The models introduce an &quot;enable_thinking=True&quot; mode with advanced soft switching for inference scaling. The release is notable for its Apache 2.0 license and broad inference platform support including MCP. The dataset improvements and multi-stage RL post-training contribute to performance gains. Meanwhile, **Gemini 2.5 Pro** from **Google DeepMind** shows strong coding and long-context reasoning capabilities, and **DeepSeek R2** is anticipated soon. Twitter discussions highlight Qwen3&apos;s finegrained MoE architecture, large context window, and multi-agent system applications.</description><pubDate>Mon, 28 Apr 2025 05:44:39 GMT</pubDate><category>alibaba</category><category>google-deepmind</category><category>deepseek</category><category>mistral-ai</category><category>qwen-3</category><category>qwen3-235b-a22b</category><category>qwen3-30b-a3b</category><category>deepseek-r1</category><category>o1</category><category>o3-mini</category><category>grok-3</category><category>gemini-2.5-pro</category><category>awnihannun</category><category>prince_canuma</category><category>actuallyisaak</category><category>oriolvinyalsml</category><category>iscienceluvr</category><category>reach_vb</category><category>teortaxestex</category><category>omarsar0</category><category>mixture-of-experts</category><category>reinforcement-learning</category><category>benchmarking</category><category>model-release</category><category>model-architecture</category><category>long-context</category><category>multi-agent-systems</category><category>inference</category><category>dataset-release</category></item><item><title>Cognition&apos;s DeepWiki, a free encyclopedia of all GitHub repos</title><link>https://news.smol.ai/issues/25-04-25-cognition-deepwiki/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-25-cognition-deepwiki/</guid><description>**Silas Alberti** of **Cognition** announced **DeepWiki**, a free encyclopedia of all GitHub repos providing Wikipedia-like descriptions and Devin-backed chatbots for public repos. **Meta** released **Perception Encoders (PE)** with A2.0 license, outperforming **InternVL3** and **Qwen2.5VL** on vision tasks. **Alibaba** launched the **Qwen Chat App** for iOS and Android. **Hugging Face** integrated the **Dia 1.6B SoTA** text-to-speech model via **FAL**. **OpenAI** expanded deep research usage with a lightweight version powered by **o4-mini** model, now available to free users. **Perplexity AI** updated their model selector with **Grok 3 Beta**, **o4-mini**, and support for models like **gemini 2.5 pro**, **claude 3.7**, and **gpt-4.1**. **vLLM** project introduced **OpenRLHF** framework for reinforcement learning with human feedback. **Surya OCR** alpha model supports 90+ languages and LaTeX. **MegaParse** open-source library was introduced for LLM-ready data formats.</description><pubDate>Fri, 25 Apr 2025 05:44:39 GMT</pubDate><category>cognition</category><category>meta-ai-fair</category><category>alibaba</category><category>hugging-face</category><category>openai</category><category>perplexity-ai</category><category>vllm</category><category/><category>o4-mini</category><category>perception-encoder</category><category>qwen-2.5-vl</category><category>dia-1.6b</category><category>grok-3</category><category>gemini-2.5-pro</category><category>claude-3.7</category><category>gpt-4.1</category><category>silas-alberti</category><category>mervenoyann</category><category>reach_vb</category><category>aravsrinivas</category><category>vikparuchuri</category><category>lioronai</category><category>vision</category><category>text-to-speech</category><category>reinforcement-learning</category><category>ocr</category><category>model-releases</category><category>model-integration</category><category>open-source</category><category>frameworks</category><category>chatbots</category><category>model-selector</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-24-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-24-not-much/</guid><description>AI news for April 23-24, 2025, covering new model releases, benchmarks, and research developments from companies like openai, google deepmind, anthropic, and epoch ai research.</description><pubDate>Thu, 24 Apr 2025 05:44:39 GMT</pubDate><category>openai</category><category>google</category><category>anthropic</category><category>epoch ai research</category><category>gpt-image-1</category><category>o3</category><category>o4-mini</category><category>gpt-4.1</category><category>dam</category><category>image-generation</category><category>model-benchmarks</category><category>vision-language-models</category><category>music-ai</category><category>ai-experiences</category><category>ai-research</category><category>supercomputers</category></item><item><title>gpt-image-1 - ChatGPT&apos;s imagegen model, confusingly NOT 4o, now available in API</title><link>https://news.smol.ai/issues/25-04-23-gpt-image-1/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-23-gpt-image-1/</guid><description>**OpenAI** officially launched the **gpt-image-1** API for image generation and editing, supporting features like alpha channel transparency and a &quot;low&quot; content moderation policy. **OpenAI&apos;s** models **o3** and **o4-mini** are leading in benchmarks for style control, math, coding, and hard prompts, with **o3** ranking #1 in several categories. A new benchmark called **Vending-Bench** reveals performance variance in LLMs on extended tasks. **GPT-4.1** ranks in the top 5 for hard prompts and math. **Nvidia&apos;s** **Eagle 2.5-8B** matches **GPT-4o** and **Qwen2.5-VL-72B** in long-video understanding. AI supercomputer performance doubles every 9 months, with **xAI&apos;s Colossus** costing an estimated $7 billion and the US dominating 75% of global performance. The Virology Capabilities Test shows **OpenAI&apos;s o3** outperforms 94% of expert virologists. **Nvidia** also released the **Describe Anything Model (DAM)**, a multimodal LLM for detailed image and video captioning, now available on Hugging Face.</description><pubDate>Wed, 23 Apr 2025 05:44:39 GMT</pubDate><category>openai</category><category>nvidia</category><category>hugging-face</category><category>x-ai</category><category>gpt-image-1</category><category>o3</category><category>o4-mini</category><category>gpt-4.1</category><category>eagle-2.5-8b</category><category>gpt-4o</category><category>qwen2.5-vl-72b</category><category>kevinweil</category><category>lmarena_ai</category><category>_philschmid</category><category>willdepue</category><category>arankomatsuzaki</category><category>epochairesearch</category><category>danhendrycks</category><category>reach_vb</category><category>mervenoyann</category><category>_akhaliq</category><category>image-generation</category><category>content-moderation</category><category>benchmarking</category><category>long-context</category><category>multimodality</category><category>model-performance</category><category>supercomputing</category><category>virology</category><category>video-understanding</category><category>model-releases</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-22-not-much/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-22-not-much/</guid><description>**Nemotron-H** model family introduces hybrid Mamba-Transformer models with up to **3x faster inference** and variants including **8B**, **56B**, and a compressed **47B** model. **Nvidia Eagle 2.5** is a frontier VLM for long-context multimodal learning, matching **GPT-4o** and **Qwen2.5-VL-72B** on long-video understanding. **Gemini 2.5 Flash** shows improved dynamic thinking and cost-performance, outperforming previous Gemini versions. **Gemma 3** now supports **torch.compile** for about **60% faster inference** on consumer GPUs. **SRPO** using **Qwen2.5-32B** surpasses DeepSeek-R1-Zero-32B on benchmarks with reinforcement learning only. **Alibaba&apos;s Uni3C** unifies 3D-enhanced camera and human motion controls for video generation. **Seedream 3.0** by **ByteDance** is a bilingual image generation model with high-resolution outputs up to **2K**. **Adobe DRAGON** optimizes diffusion generative models with distributional rewards. **Kimina-Prover Preview** is an LLM trained with reinforcement learning from **Qwen2.5-72B**, achieving **80.7% pass@8192** on miniF2F. **BitNet b1.58 2B4T** is a native 1-bit LLM with **2B parameters** trained on **4 trillion tokens**, matching full-precision LLM performance with better efficiency. Antidistillation sampling counters unwanted model distillation by modifying reasoning traces from frontier models.</description><pubDate>Tue, 22 Apr 2025 05:44:39 GMT</pubDate><category>nvidia</category><category>deepseek</category><category>hugging-face</category><category>alibaba</category><category>bytedance</category><category>adobe</category><category>nemotron-h</category><category>nvidia-eagle-2.5</category><category>gpt-4o</category><category>qwen2.5-vl-72b</category><category>gemini-2.5-flash</category><category>gemini-2.0-pro</category><category>gemini-exp-1206</category><category>gemma-3</category><category>qwen2.5-32b</category><category>deepseek-r1-zero-32b</category><category>uni3c</category><category>seedream-3.0</category><category>adobe-dragon</category><category>kimina-prover</category><category>qwen2.5-72b</category><category>bitnet-b1.58-2b4t</category><category>philschmid</category><category>arankomatsuzaki</category><category>osanseviero</category><category>iScienceLuvr</category><category>akhaliq</category><category>transformers</category><category>model-optimization</category><category>multimodality</category><category>long-context</category><category>reinforcement-learning</category><category>torch-compile</category><category>image-generation</category><category>diffusion-models</category><category>distributional-rewards</category><category>model-efficiency</category><category>model-training</category><category>native-quantization</category><category>sampling-techniques</category></item><item><title>not much happened today; New email provider for AINews</title><link>https://news.smol.ai/issues/25-04-21-not-much-resend/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-21-not-much-resend/</guid><description>**Smol AI** is migrating its AI news email service to **Resend** to improve deliverability and enable new features like personalizable AI news and a &quot;Hacker News of AI.&quot; Recent AI model updates include **OpenAI**&apos;s API-only **GPT-4.1**, **Google Gemini 2.5 Flash** reasoning model, **ByteDance Seaweed** 7B-param video AI, **Anthropic Claude**&apos;s values system, **Cohere Embed 4** multimodal embedding model, and **xAI Grok** updates with Memory and Studio features. Discussions also cover agentic workflows for document automation and AI coding patterns.</description><pubDate>Mon, 21 Apr 2025 05:44:39 GMT</pubDate><category>smol-ai</category><category>resend</category><category>openai</category><category>google</category><category>bytedance</category><category>anthropic</category><category>cohere</category><category>x-ai</category><category>gpt-4.1</category><category>gpt-4o</category><category>gpt-4o-mini</category><category>gemini-2.5-flash</category><category>seaweed-7b</category><category>claude</category><category>embed-4</category><category>grok</category><category>adcock_brett</category><category>swyx</category><category>jerryjliu0</category><category>alexalbert</category><category>omarsar0</category><category>email-deliverability</category><category>model-releases</category><category>reasoning</category><category>video-generation</category><category>multimodality</category><category>embedding-models</category><category>agentic-workflows</category><category>document-processing</category><category>function-calling</category><category>tool-use</category><category>ai-coding</category></item><item><title>Grok 3 &amp; 3-mini now API Available</title><link>https://news.smol.ai/issues/25-04-18-ainews-grok-3-and-3-mini-now-api-available/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-18-ainews-grok-3-and-3-mini-now-api-available/</guid><description>**Grok 3** API is now available, including a smaller version called Grok 3 mini, which offers competitive pricing and full reasoning traces. **OpenAI** released a practical guide for building AI agents, while **LlamaIndex** supports the Agent2Agent protocol for multi-agent communication. **Codex CLI** is gaining traction with new features and competition from **Aider** and **Claude Code**. **GoogleDeepMind** launched **Gemini 2.5 Flash**, a hybrid reasoning model topping the Chatbot Arena leaderboard. **OpenAI**&apos;s o3 and o4-mini models show emergent behaviors from large-scale reinforcement learning. **EpochAIResearch** updated its methodology, removing **Maverick** from high FLOP models as **Llama 4 Maverick** training compute drops. **GoodfireAI** announced a $50M Series A for its Ember neural programming platform. **Mechanize** was founded to build virtual work environments and automation benchmarks. **GoogleDeepMind**&apos;s Quantisation Aware Training for Gemma 3 models reduces model size significantly, with open source checkpoints available.</description><pubDate>Sat, 19 Apr 2025 05:44:39 GMT</pubDate><category>openai</category><category>llamaindex</category><category>google-deepmind</category><category>epochairesearch</category><category>goodfireai</category><category>mechanize</category><category>grok-3</category><category>grok-3-mini</category><category>gemini-2.5-flash</category><category>o3</category><category>o4-mini</category><category>llama-4-maverick</category><category>gemma-3-27b</category><category>agent-development</category><category>agent-communication</category><category>cli-tools</category><category>reinforcement-learning</category><category>model-evaluation</category><category>quantization-aware-training</category><category>model-compression</category><category>training-compute</category><category>hybrid-reasoning</category><category>model-benchmarking</category></item><item><title>Gemini 2.5 Flash completes the total domination of the Pareto Frontier</title><link>https://news.smol.ai/issues/25-04-17-ainews-gemini-25-flash-completes-the-total-domination-of-the-pareto-frontier/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-17-ainews-gemini-25-flash-completes-the-total-domination-of-the-pareto-frontier/</guid><description>**Gemini 2.5 Flash** is introduced with a new &quot;thinking budget&quot; feature offering more control compared to Anthropic and OpenAI models, marking a significant update in the Gemini series. **OpenAI** launched **o3** and **o4-mini** models, emphasizing advanced tool use capabilities and multimodal understanding, with **o3** dominating several leaderboards but receiving mixed benchmark reviews. The importance of tool use in AI research and development is highlighted, with **OpenAI Codex CLI** announced as a lightweight open-source coding agent. The news reflects ongoing trends in AI model releases, benchmarking, and tool integration.</description><pubDate>Fri, 18 Apr 2025 02:06:17 GMT</pubDate><category>google</category><category>openai</category><category>anthropic</category><category>gemini-2.5-flash</category><category>o3</category><category>o4-mini</category><category>sama</category><category>kevinweil</category><category>markchen90</category><category>alexandr_wang</category><category>polynoamial</category><category>scaling01</category><category>aidan_mclau</category><category>cwolferesearch</category><category>tool-use</category><category>multimodality</category><category>benchmarking</category><category>reasoning</category><category>reinforcement-learning</category><category>open-source</category><category>model-releases</category><category>chain-of-thought</category><category>coding-agent</category></item><item><title>OpenAI o3, o4-mini, and Codex CLI</title><link>https://news.smol.ai/issues/25-04-16-ainews-openai-o3-o4-mini-and-codex-cli/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-16-ainews-openai-o3-o4-mini-and-codex-cli/</guid><description>**OpenAI** launched the **o3** and **o4-mini** models, emphasizing improvements in **reinforcement-learning scaling** and overall efficiency, making **o4-mini** cheaper and better across prioritized metrics. These models showcase enhanced **vision** and **tool use** capabilities, though API access for these features is pending. The release includes **Codex CLI**, an open-source coding agent that integrates with these models to convert natural language into working code. Accessibility extends to **ChatGPT Plus, Pro, and Team users**, with **o3** being notably more expensive than **Gemini 2.5 Pro**. Performance benchmarks highlight the intelligence gains from scaling inference, with comparisons against models like **Sonnet** and **Gemini**. The launch has been well received despite some less favorable evaluation results.</description><pubDate>Thu, 17 Apr 2025 03:17:29 GMT</pubDate><category>openai</category><category>o3</category><category>o4-mini</category><category>gemini-2.5-pro</category><category>claude-3-sonnet</category><category>chatgpt</category><category>sama</category><category>aidan_mclau</category><category>markchen90</category><category>gdb</category><category>aidan_clark_</category><category>kevinweil</category><category>swyx</category><category>polynoamial</category><category>scaling01</category><category>reinforcement-learning</category><category>performance</category><category>vision</category><category>tool-use</category><category>open-source</category><category>coding-agents</category><category>model-benchmarking</category><category>multimodality</category><category>scaling</category><category>inference</category></item><item><title>QwQ-32B claims to match DeepSeek R1-671B</title><link>https://news.smol.ai/issues/25-04-16-ainews-qwq-32b-claims-to-match-deepseek-r1-671b/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-16-ainews-qwq-32b-claims-to-match-deepseek-r1-671b/</guid><description>**Alibaba Qwen** released their **QwQ-32B** model, a **32 billion parameter** reasoning model using a novel two-stage reinforcement learning approach: first scaling RL for math and coding tasks with accuracy verifiers and code execution servers, then applying RL for general capabilities like instruction following and alignment. Meanwhile, **OpenAI** rolled out **GPT-4.5** to Plus users, with mixed feedback on coding performance and noted inference cost improvements. The QwQ model aims to compete with larger MoE models like **DeepSeek-R1**. *&quot;GPT-4.5 is unusable for coding&quot;* was a notable user critique, while others praised its reasoning improvements due to scaling pretraining.</description><pubDate>Wed, 16 Apr 2025 19:06:15 GMT</pubDate><category>alibaba</category><category>openai</category><category>deepseek-ai</category><category>qwen-2.5-plus</category><category>qwq-32b</category><category>deepseek-r1</category><category>gpt-4.5</category><category>gpt-3</category><category>davinci</category><category>aidan_mclau</category><category>sama</category><category>scaling01</category><category>juberti</category><category>polynoamial</category><category>reach_vb</category><category>reinforcement-learning</category><category>math</category><category>code-execution</category><category>instruction-following</category><category>alignment</category><category>reasoning</category><category>model-release</category><category>model-benchmarking</category><category>scaling</category><category>performance</category><category>inference-costs</category></item><item><title>SOTA Video Gen: Veo 2 and Kling 2 are GA for developers</title><link>https://news.smol.ai/issues/25-04-15-ainews-sota-video-gen-veo-2-and-kling-2-are-ga-for-developers/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-15-ainews-sota-video-gen-veo-2-and-kling-2-are-ga-for-developers/</guid><description>**Google&apos;s Veo 2** video generation model is now available in the **Gemini API** with a cost of **35 cents per second** of generated video, marking a significant step in accessible video generation. Meanwhile, China&apos;s **Kling 2** model launched with pricing around **$2 for a 10-second clip** and a minimum subscription of **$700 per month for 3 months**, generating excitement despite some skill challenges. **OpenAI** announced the **GPT-4.1 family** release, including **GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano**, highlighting improvements in **coding, instruction following, and a 1 million token context window**. The GPT-4.1 models are **26% cheaper than GPT-4o** and will replace the **GPT-4.5 Preview** API version by July 14. Performance benchmarks show GPT-4.1 achieving **54-55% on SWE-bench verified** and a **60% improvement over GPT-4o** in some internal tests, though some critiques note it underperforms compared to other models like OpenRouter and DeepSeekV3 in coding tasks. The release is API-only, with a prompting guide provided for developers.</description><pubDate>Wed, 16 Apr 2025 05:55:06 GMT</pubDate><category>google</category><category>openai</category><category>veo-2</category><category>gemini</category><category>gpt-4.1</category><category>gpt-4o</category><category>gpt-4.5-preview</category><category>gpt-4.1-mini</category><category>gpt-4.1-nano</category><category>kevinweil</category><category>stevenheidel</category><category>aidan_clark_</category><category>video-generation</category><category>api</category><category>coding</category><category>instruction-following</category><category>context-window</category><category>performance</category><category>benchmarks</category><category>model-deprecation</category></item><item><title>GPT 4.1: The New OpenAI Workhorse</title><link>https://news.smol.ai/issues/25-04-14-ainews-gpt-41-the-new-openai-workhorse/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-14-ainews-gpt-41-the-new-openai-workhorse/</guid><description>**OpenAI** released **GPT-4.1**, including **GPT-4.1 mini** and **GPT-4.1 nano**, highlighting improvements in **coding**, **instruction following**, and handling **long contexts** up to **1 million tokens**. The model achieves a **54 score on SWE-bench verified** and shows a **60% improvement over GPT-4o** on internal benchmarks. Pricing for **GPT-4.1 nano** is notably low at **$0.10/1M input** and **$0.40/1M output**. **GPT-4.5 Preview** is being deprecated in favor of **GPT-4.1**. Integration support includes **Llama Index** with day 0 support. Some negative feedback was noted for **GPT-4.1 nano**. Additionally, **Perplexity&apos;s Sonar API** ties with **Gemini-2.5 Pro** for the top spot in the LM Search Arena leaderboard. New benchmarks like **MRCR** and **GraphWalks** were introduced alongside updated prompting guides and cookbooks.</description><pubDate>Tue, 15 Apr 2025 05:16:26 GMT</pubDate><category>openai</category><category>llama-index</category><category>perplexity-ai</category><category>google-deepmind</category><category>gpt-4.1</category><category>gpt-4.1-mini</category><category>gpt-4.1-nano</category><category>gpt-4o</category><category>gemini-2.5-pro</category><category>sama</category><category>kevinweil</category><category>omarsar0</category><category>aidan_mclau</category><category>danhendrycks</category><category>polynoamial</category><category>scaling01</category><category>aravsrinivas</category><category>lmarena_ai</category><category>coding</category><category>instruction-following</category><category>long-context</category><category>benchmarks</category><category>model-pricing</category><category>model-integration</category><category>model-deprecation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-11-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-11-ainews-not-much-happened-today/</guid><description>The AI news recap highlights independent evaluations showing **Grok-3** outperforming models like **GPT-4.5** and **Claude 3.7 Sonnet** on reasoning benchmarks, while **Grok-3 mini** excels in reasoning tasks. Research on **reinforcement learning (RL)** fine-tuning reveals potential improvements for small reasoning models but also notes instability in reported gains. Benchmark results suggest **Quasar Alpha** and **Optimus Alpha** may be versions of **GPT-4.1**. Vision and multimodal models like **Kaleidoscope**, supporting 18 languages, and **InternVL3**, built on **InternViT** and **Qwen2.5VL**, demonstrate advances in multilingual vision and reasoning. The fusion model **TransMamba** combines transformer precision with speed via **SSM** mechanisms. Alibaba&apos;s **FantasyTalking** generates realistic talking portraits. Agent-focused events at **CMU** and tools like **FilmAgent AI** for virtual film production and **BrowseComp** benchmark for browsing agents were announced. The coding assistant **Augment** supports multiple IDEs with code analysis and suggestions. Discussions also covered Google’s new agent-to-agent protocol concept.</description><pubDate>Fri, 11 Apr 2025 20:07:39 GMT</pubDate><category>openai</category><category>alibaba</category><category>cmu</category><category>grok-3</category><category>grok-3-mini</category><category>gpt-4.5</category><category>claude-3.7-sonnet</category><category>quasar-alpha</category><category>optimus-alpha</category><category>gpt-4.1</category><category>kaleidoscope</category><category>internvl3</category><category>internvit</category><category>qwen2.5vl</category><category>transmamba</category><category>fantasytalking</category><category>rasbt</category><category>sarahookr</category><category>mervenoyann</category><category>gneubig</category><category>svpino</category><category>mathemagic1an</category><category>reinforcement-learning</category><category>reasoning</category><category>benchmarks</category><category>vision</category><category>multilinguality</category><category>multimodality</category><category>transformers</category><category>attention-mechanisms</category><category>agents</category><category>code-generation</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-10-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-10-ainews-not-much-happened-today/</guid><description>**OpenAI** teased a *Memory update in ChatGPT* with limited technical details. Evidence suggests upcoming releases of **o3** and **o4-mini** models, alongside a press leak about **GPT-4.1**. **X.ai** launched the **Grok 3** and **Grok 3 mini** APIs, confirmed as **o1** level models. Discussions compared **Google&apos;s TPUv7** with **Nvidia&apos;s GB200**, highlighting TPUv7&apos;s specs like **4,614 TFLOP/s FP8 performance**, **192 GB HBM**, and **1.2 Tbps ICI bandwidth**. TPUv7 may have pivoted from training to inference chip use. Key AI events include **Google Cloud Next 2025** and **Samsung&apos;s Gemini-powered Ballie robot**. The community is invited to participate in the **AI Engineer World&apos;s Fair 2025** and the 2025 State of AI Engineering survey.</description><pubDate>Fri, 11 Apr 2025 00:53:38 GMT</pubDate><category>openai</category><category>x-ai</category><category>google</category><category>nvidia</category><category>samsung</category><category>gpt-4.1</category><category>o3</category><category>o4-mini</category><category>grok-3</category><category>grok-3-mini</category><category>o1</category><category>tpuv7</category><category>gb200</category><category>sama</category><category>memory</category><category>model-release</category><category>hardware-accelerators</category><category>fp8</category><category>hbm</category><category>inference</category><category>ai-conferences</category><category>agent-collaboration</category><category>robotics</category><category>model-comparison</category><category>performance</category><category>power-consumption</category></item><item><title>Google&apos;s Agent2Agent Protocol (A2A)</title><link>https://news.smol.ai/issues/25-04-09-ainews-googles-agent2agent-protocol-a2a/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-09-ainews-googles-agent2agent-protocol-a2a/</guid><description>**Google Cloud Next** announcements featured the launch of **Google and DeepMind&apos;s** full **MCP support** and a new **Agent to Agent protocol** designed for agent interoperability with multiple partners. The protocol includes components like the **Agent Card**, **Task communication channels**, **Enterprise Auth and Observability**, and **Streaming and Push Notification support**. On the model front, **Moonshot AI** released **Kimi-VL-A3B**, a multimodal model with **128K context** and strong vision and math benchmark performance, outperforming **gpt-4o**. **Meta AI** introduced smaller versions of **llama-4** family models: **llama-4-scout** and **llama-4-maverick**, with a larger **Behemoth** model still in training. **DeepCoder 14B** from **UC Berkeley** is an open-source coding model rivaling **openai&apos;s o3-mini** and **o1** models, trained with reinforcement learning on 24K coding problems. **Nvidia** released **llama-3.1-nemotron-ultra-253b** on Hugging Face, noted for beating **llama-4-behemoth** and **maverick** and competing with **deepseek-r1**.</description><pubDate>Thu, 10 Apr 2025 01:31:18 GMT</pubDate><category>google</category><category>google-deepmind</category><category>moonshot-ai</category><category>meta-ai-fair</category><category>uc-berkeley</category><category>openai</category><category>nvidia</category><category>hugging-face</category><category>togethercompute</category><category>deepseek</category><category>kimi-vl-a3b</category><category>gpt-4o</category><category>llama-4-scout</category><category>llama-4-maverick</category><category>llama-4-behemoth</category><category>deepcoder-14b</category><category>o3-mini</category><category>o1</category><category>llama-3.1-nemotron-ultra-253b</category><category>deepseek-r1</category><category>reach_vb</category><category>_akhaliq</category><category>epochairesearch</category><category>artificialanlys</category><category>winglian</category><category>danielhanchen</category><category>yuchenj_uw</category><category>jeremyphoward</category><category>agent-interoperability</category><category>multimodality</category><category>vision</category><category>math</category><category>reinforcement-learning</category><category>coding</category><category>model-training</category><category>open-source</category><category>model-benchmarking</category><category>context-windows</category><category>streaming</category><category>push-notifications</category><category>enterprise-authentication</category><category>model-release</category></item><item><title>DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level</title><link>https://news.smol.ai/issues/25-04-09-ainews-deepcoder-a-fully-open-source-14b-coder-at-o3-mini-level/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-09-ainews-deepcoder-a-fully-open-source-14b-coder-at-o3-mini-level/</guid><description>**Together AI and Agentica** released **DeepCoder-14B**, an open-source 14B parameter coding model rivaling OpenAI&apos;s **o3-mini** and **o1** on coding benchmarks, trained with an open-source RL framework from ByteDance and costing about **$26,880**. **Google DeepMind** launched **Gemini 2.5 Pro** with experimental &quot;Flash&quot; versions available to subscribers. **Moonshot AI** introduced **Kimi-VL-A3B**, a multimodal model with **128K context** outperforming **gpt-4o** on vision and math benchmarks. **Meta AI** released **Llama 4 Scout** and **Maverick**, with a larger **Behemoth** model in training, featuring mixture-of-experts and L2 norm techniques. **Runway** launched **Gen-4 Turbo** with 10x better results than Gen-3 at the same cost. **Google** announced **Imagen 3**, a high-quality text-to-image model now in Vertex AI, enabling easier object removal. The report highlights open-source contributions, reinforcement learning training optimizations, and significant model performance improvements across coding, multimodal, and image generation domains.</description><pubDate>Wed, 09 Apr 2025 19:51:30 GMT</pubDate><category>together-ai</category><category>agentica</category><category>opena</category><category>bytedance</category><category>google-deepmind</category><category>moonshot-ai</category><category>meta-ai-fair</category><category>runway</category><category>deepcoder-14b</category><category>o3-mini</category><category>o1</category><category>gemini-2.5-pro</category><category>kimi-vl-a3b</category><category>gpt-4o</category><category>llama-4-scout</category><category>maverick</category><category>behemoth</category><category>gen-4-turbo</category><category>imagen-3</category><category>philschmid</category><category>lepikhin</category><category>reach_vb</category><category>akhaliq</category><category>yuchenj_uw</category><category>epochairesearch</category><category>danielhanchen</category><category>c_valenzuelab</category><category>open-source</category><category>reinforcement-learning</category><category>code-generation</category><category>multimodality</category><category>model-training</category><category>mixture-of-experts</category><category>l2-normalization</category><category>image-generation</category><category>model-performance</category><category>context-windows</category></item><item><title>Llama 4&apos;s Controversial Weekend Release</title><link>https://news.smol.ai/issues/25-04-07-ainews-llama-4s-controversial-weekend-release/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-07-ainews-llama-4s-controversial-weekend-release/</guid><description>**Meta** released **Llama 4**, featuring two new medium-size MoE open models and a promised 2 Trillion parameter &quot;behemoth&quot; model, aiming to be the largest open model ever. The release included advanced training techniques like Chameleon-like early fusion with MetaCLIP, interleaved chunked attention without RoPE, native FP8 training, and training on up to 40 trillion tokens. Despite the hype, the release faced criticism for lack of transparency compared to Llama 3, implementation issues, and poor performance on some benchmarks. Meta leadership, including **Ahmad Al Dahle**, denied allegations of training on test sets. The smallest Scout model at 109B parameters is too large for consumer GPUs, and the claimed 10 million token context is disputed. The community response has been mixed, with some praising the openness and others pointing out discrepancies and quality concerns.</description><pubDate>Tue, 08 Apr 2025 01:55:40 GMT</pubDate><category>meta</category><category>llama-4</category><category>llama-3</category><category>llama-3-2</category><category>ahmad_al_dahle</category><category>ylecun</category><category>reach_vb</category><category>yuchenj_uw</category><category>mixture-of-experts</category><category>early-fusion</category><category>attention-mechanisms</category><category>fp8-training</category><category>training-data</category><category>benchmarking</category><category>model-performance</category><category>model-release</category><category>multimodality</category><category>open-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-04-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-04-ainews-not-much-happened-today/</guid><description>**OpenAI** announced that **o3** and **o4-mini** models will be released soon, with **GPT-5** expected in a few months, delayed for quality improvements and capacity planning. **DeepSeek** introduced **Self-Principled Critique Tuning (SPCT)** to enhance inference-time scalability for generalist reward models. **Anthropic&apos;s Sonnet 3.7** remains a top coding model. **Google&apos;s Gemma 3** is available on KerasHub, and **Qwen 2.5 VL** powers a new Apache 2.0 licensed OCR model. **Gemini 2.5 Pro** entered public preview with increased rate limits and pricing announced, becoming a preferred model for many tasks except image generation. Meta&apos;s architectural advantage and the **FrontierMath benchmark** challenge AI&apos;s long-form reasoning and worldview development. Research reveals LLMs focus attention on the first token as an &quot;attention sink,&quot; preserving representation diversity, demonstrated in **Gemma 7B** and **LLaMa 3.1** models. **MegaScale-Infer** offers efficient serving of large-scale Mixture-of-Experts models with up to **1.90x higher per-GPU throughput**.</description><pubDate>Sat, 05 Apr 2025 01:50:06 GMT</pubDate><category>openai</category><category>deepseek</category><category>anthropic</category><category>google</category><category>meta-ai-fair</category><category>o3</category><category>o4-mini</category><category>gpt-5</category><category>sonnet-3.7</category><category>gemma-3</category><category>qwen-2.5-vl</category><category>gemini-2.5-pro</category><category>gemma-7b</category><category>llama-3-1-405b</category><category>sama</category><category>akhaliq</category><category>nearcyan</category><category>fchollet</category><category>reach_vb</category><category>philschmid</category><category>teortaxestex</category><category>epochairesearch</category><category>omarsar0</category><category>inference-scaling</category><category>reward-modeling</category><category>coding-models</category><category>ocr</category><category>model-preview</category><category>rate-limiting</category><category>model-pricing</category><category>architectural-advantage</category><category>benchmarking</category><category>long-form-reasoning</category><category>attention-mechanisms</category><category>mixture-of-experts</category><category>gpu-throughput</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-03-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-03-ainews-not-much-happened-today/</guid><description>**Gemini 2.5 Pro** shows strengths and weaknesses, notably lacking LaTex math rendering unlike **ChatGPT**, and scored **24.4%** on the **2025 US AMO**. **DeepSeek V3** ranks 8th and 12th on recent leaderboards. **Qwen 2.5** models have been integrated into the **PocketPal** app. Research from **Anthropic** reveals that **Chains-of-Thought (CoT)** reasoning is often unfaithful, especially on harder tasks, raising safety concerns. **OpenAI**&apos;s **PaperBench** benchmark shows AI agents struggle with long-horizon planning, with **Claude 3.5 Sonnet** achieving only **21.0%** accuracy. **CodeAct** framework generalizes **ReAct** for dynamic code writing by agents. **LangChain** explains multi-agent handoffs in LangGraph. **Runway Gen-4** marks a new phase in media creation.</description><pubDate>Fri, 04 Apr 2025 06:34:03 GMT</pubDate><category>google</category><category>anthropic</category><category>openai</category><category>llama_index</category><category>langchain</category><category>runway</category><category>deepseek</category><category>gemini-2.5-pro</category><category>chatgpt</category><category>deepseek-v3</category><category>qwen-2.5</category><category>claude-3.5-sonnet</category><category>claude-3.7-sonnet</category><category>rasbt</category><category>danielhanchen</category><category>hkproj</category><category>math</category><category>benchmarking</category><category>chains-of-thought</category><category>model-performance</category><category>multi-agent-systems</category><category>agent-frameworks</category><category>media-generation</category><category>long-horizon-planning</category><category>code-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-04-01-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-04-01-ainews-not-much-happened-today/</guid><description>**OpenAI** plans to release its first open-weight language model since **GPT-2** in the coming months, signaling a move towards more open AI development. **DeepSeek** launched its open-source **R1 model** earlier this year, challenging perceptions of China&apos;s AI progress. **Gemma 3** has achieved function calling capabilities and ranks on the **Berkeley Function-Calling Leaderboard**, while **GemmaCoder3-12b** improves code reasoning performance on **LiveCodeBench**. **Alibaba_Qwen&apos;s Qwen2.5-Omni** introduces a novel Thinker-Talker system and **TMRoPE** for multimodal input understanding. The **TogetherCompute** team achieved **140 TPS** on a 671B parameter model, outperforming **Azure** and **DeepSeek API** on **Nvidia GPUs**. **OpenAI** also expanded **ChatGPT** features with image generation for all free users and a new voice release. **Runway Gen-4** enhances animation for miniature dioramas, and **LangChain** launched a chat-based generative UI agent. Commercial deployment of **Figure 03 humanoid robots** at **BMW** highlights advances in autonomy and manufacturing scaling. New tools include **OpenAI&apos;s realtime transcription API** with **WebRTC** support and **Amazon&apos;s Nova Act AI browser agent**.</description><pubDate>Wed, 02 Apr 2025 06:14:34 GMT</pubDate><category>openai</category><category>deepseek</category><category>berkeley</category><category>alibaba</category><category>togethercompute</category><category>nvidia</category><category>azure</category><category>runway</category><category>langchain</category><category>bmw</category><category>amazon</category><category>gpt-2</category><category>r1</category><category>gemma-3</category><category>gemmacoder3-12b</category><category>qwen2.5-omni</category><category>sama</category><category>clémentdelangue</category><category>lioronai</category><category>scaling01</category><category>cognitivecompai</category><category>osanseviero</category><category>jack_w_rae</category><category>ben_burtenshaw</category><category>theturingpost</category><category>vipulved</category><category>kevinweil</category><category>tomlikesrobots</category><category>adcock_brett</category><category>juberti</category><category>open-source</category><category>function-calling</category><category>benchmarking</category><category>code-reasoning</category><category>multimodality</category><category>inference-speed</category><category>image-generation</category><category>voice-generation</category><category>animation</category><category>robotics</category><category>realtime-transcription</category><category>webrtc</category></item><item><title>&gt;$41B raised today (OpenAI @ 300b, Cursor @ 9.5b, Etched @ 1.5b)</title><link>https://news.smol.ai/issues/25-03-31-ainews-greaterdollar41b-raised-today-openai-300b-cursor-95b-etched-15b/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-31-ainews-greaterdollar41b-raised-today-openai-300b-cursor-95b-etched-15b/</guid><description>**OpenAI** is preparing to release a highly capable open language model, their first since GPT-2, with a focus on reasoning and community feedback, as shared by **@kevinweil** and **@sama**. **DeepSeek V3 0324** has achieved the #5 spot on the Arena leaderboard, becoming the top open model with an MIT license and cost advantages. **Gemini 2.5 Pro** is noted for outperforming models like **Claude 3.7 Sonnet** in coding tasks, with upcoming pricing and improvements expected soon. New startups like **Sophont** are building open multimodal foundation models for healthcare. Significant fundraises include **Cursor** closing $625M at a $9.6B valuation and **Etched** raising $85M at $1.5B. Innovations in AI infrastructure include **SkyPilot&apos;s** cost-efficient cloud provisioning and the launch of **AgentEvals**, an open-source package for evaluating AI agents. Discussions on smartphone privacy highlight **iPhone&apos;s** stronger user defense compared to Android.</description><pubDate>Tue, 01 Apr 2025 06:33:20 GMT</pubDate><category>openai</category><category>deepseek</category><category>gemini</category><category>cursor</category><category>etched</category><category>skypilot</category><category>agent-evals</category><category>deepseek-v3-0324</category><category>gemini-2.5-pro</category><category>claude-3.7-sonnet</category><category>kevinweil</category><category>sama</category><category>lmarena_ai</category><category>scaling01</category><category>iscienceluvr</category><category>stevenheidel</category><category>lepikhin</category><category>dzhng</category><category>raizamrtn</category><category>karpathy</category><category>open-models</category><category>model-releases</category><category>model-performance</category><category>coding</category><category>multimodality</category><category>model-deployment</category><category>cost-efficiency</category><category>agent-evaluation</category><category>privacy</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-28-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-28-ainews-not-much-happened-today/</guid><description>**GPT-4o** was praised for its improved coding, instruction following, and freedom, becoming the leading non-reasoning coding model surpassing **DeepSeek V3** and **Claude 3.7 Sonnet** in coding benchmarks, though it still lags behind reasoning models like **o3-mini**. Concerns about policy compliance in image generation were noted, with efforts to improve adherence. **Gemini 2.5 Pro** was highlighted for its advanced audio and video understanding, long context capabilities, and integration with platforms like **Cursor AI** and **Windsurf AI**. AI infrastructure developments include a partnership between **Together AI** and **Hypertec Group** to deliver large-scale GPU clusters, and **CoreWeave&apos;s IPO** was celebrated for advancing AI infrastructure. GPU and TPU usage is expected to increase significantly. *&quot;GPT-4o&apos;s transparency and background generation feature&quot;* and *&quot;Gemini 2.5 Pro scored above 50% on Simple-Bench AI Explanation&quot;* were key highlights.</description><pubDate>Fri, 28 Mar 2025 23:18:38 GMT</pubDate><category>openai</category><category>deepseek</category><category>anthropic</category><category>google-deepmind</category><category>togethercompute</category><category>hypertecgroup</category><category>coreweave</category><category>cursor-ai</category><category>windsurf-ai</category><category>gpt-4o</category><category>deepseek-v3</category><category>claude-3.7-sonnet</category><category>o3-mini</category><category>gemini-2.5-pro</category><category>sama</category><category>kevinweil</category><category>joannejang</category><category>nrehiew_</category><category>giffmana</category><category>_philschmid</category><category>scaling01</category><category>saranormous</category><category>coding</category><category>instruction-following</category><category>image-generation</category><category>policy-compliance</category><category>long-context</category><category>audio-processing</category><category>video-processing</category><category>gpu-clusters</category><category>ai-infrastructure</category><category>api-access</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-27-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-27-ainews-not-much-happened-today/</guid><description>**OpenAI** announced the new **GPT-4o** model with enhanced instruction-following, complex problem-solving, and native image generation capabilities. The model shows improved performance in math, coding, and creativity, with features like transparent background image generation. Discussions around content filtering and policy for image generation emphasize balancing creative freedom and harm prevention. **DeepSeek V3-0324** APIs, available on **Hugging Face** and powered by **SambaNovaAI**, outperform benchmarks and models like **Gemini 2.0 Pro** and **Claude 3.7 Sonnet**. **Gemini 2.5 Pro** is recommended for coding, and **Gemini 3** can be deployed easily on Google Cloud Vertex AI via the new Model Garden SDK. The **Gemma 3 Technical Report** has been released on arXiv.</description><pubDate>Fri, 28 Mar 2025 01:20:31 GMT</pubDate><category>openai</category><category>hugging-face</category><category>sambanova</category><category>google-cloud</category><category>gpt-4o</category><category>deepseek-v3-0324</category><category>gemini-2.5-pro</category><category>gemini-3</category><category>claude-3.7-sonnet</category><category>abacaj</category><category>nrehiew_</category><category>sama</category><category>joannejang</category><category>giffmana</category><category>lmarena_ai</category><category>_philschmid</category><category>instruction-following</category><category>image-generation</category><category>content-filtering</category><category>model-performance</category><category>api</category><category>coding</category><category>model-deployment</category><category>benchmarking</category><category>model-release</category></item><item><title>OpenAI adopts MCP</title><link>https://news.smol.ai/issues/25-03-26-ainews-openai-adopts-mcp/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-26-ainews-openai-adopts-mcp/</guid><description>**OpenAI** announced support for **MCP**, a significant technical update. **Google&apos;s Gemini 2.5 Pro** leads benchmarks with top scores in **MMLU-Pro (86%)**, **GPQA Diamond (83%)**, and **AIME 2024 (88%)**, featuring a **1 million token context window** and multimodal inputs. **Alibaba&apos;s Qwen 2.5 Omni 7B** was released as a fully multimodal, interactive, open-source model with a novel &quot;thinker-talker&quot; architecture supporting voice and video chat. **DeepSeek V3-0324** outperforms its predecessor on multiple benchmarks. Research on reasoning features in large language models using sparse autoencoders was highlighted, alongside a study on scaling laws of synthetic data showing performance plateaus near **300B tokens**. Discussions also covered the fastest output speeds of Gemini models and concerns about over-reliance on benchmarks for intelligence measurement. *Swyx* will curate the Data Council AI Engineering Track in April.</description><pubDate>Thu, 27 Mar 2025 01:07:34 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>alibaba</category><category>togethercompute</category><category>gemini-2.5-pro</category><category>gemini-1.5-pro</category><category>gemini-2.0-flash</category><category>qwen-2.5-omni-7b</category><category>deepseek-v3-0324</category><category>deepseek-r1</category><category>swyx</category><category>model-benchmarking</category><category>multimodality</category><category>reasoning</category><category>scaling-laws</category><category>model-quantization</category><category>synthetic-data</category><category>model-performance</category><category>context-windows</category><category>speech-recognition</category><category>translation</category><category>audio-processing</category><category>video-processing</category></item><item><title>Gemini 2.5 Pro + 4o Native Image Gen</title><link>https://news.smol.ai/issues/25-03-25-ainews-gemini-25-pro-4o-native-image-gen/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-25-ainews-gemini-25-pro-4o-native-image-gen/</guid><description>**Gemini 2.5 Pro** from **Google DeepMind** has become the new top AI model, surpassing **Grok 3** by 40 LMarena points, with contributions from **Noam Shazeer** integrating Flash Thinking techniques. It is available as a free, rate-limited experimental model. Meanwhile, **OpenAI** released **GPT 4o Native Images**, an autoregressive image generation model with detailed insights shared by **Allan Jabri** and credits to **Gabe Goh**. Gemini 2.5 Pro excels in reasoning, coding, STEM, multimodal tasks, and instruction following, topping the LMarena leaderboard significantly. It is accessible via Google AI Studio and the Gemini App.</description><pubDate>Wed, 26 Mar 2025 01:13:42 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>lmarena_ai</category><category>gemini-2.5-pro</category><category>gpt-4o</category><category>noam-shazeer</category><category>allan-jabri</category><category>gabe-goh</category><category>autoregressive-models</category><category>multimodality</category><category>reasoning</category><category>coding</category><category>instruction-following</category><category>model-release</category><category>leaderboards</category></item><item><title>Halfmoon is Reve Image: a new SOTA Image Model from ex-Adobe/Stability trio</title><link>https://news.smol.ai/issues/25-03-24-ainews-halfmoon-is-reve-image-a-new-sota-image-model-from-ex-adobestability-trio/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-24-ainews-halfmoon-is-reve-image-a-new-sota-image-model-from-ex-adobestability-trio/</guid><description>**Reve**, a new composite AI model from former Adobe and Stability alums **Christian Cantrell**, **Taesung Park**, and **Michaël Gharbi**, has emerged as the top-rated image generation model, surpassing previous state-of-the-art models like Recraft and Ideogram in text rendering and typography. The team emphasizes *&quot;enhancing visual generative models with logic&quot;* and *&quot;understanding user intent with advanced language capabilities&quot;* to iteratively amend visuals based on natural language input. Additionally, **DeepSeek-V3-0324** and **Alibaba&apos;s Qwen2.5-VL-32B-Instruct** models were released with notable performance improvements, including better vision task benchmarks and mathematical reasoning.</description><pubDate>Tue, 25 Mar 2025 01:43:04 GMT</pubDate><category>artificial-analysis</category><category>stability-ai</category><category>adobe</category><category>deepseek</category><category>alibaba</category><category>deepseek-v3-0324</category><category>qwen-2.5-vl-32b-instruct</category><category>recraft</category><category>christian-cantrell</category><category>taesung-park</category><category>michael-gharbi</category><category>text-to-image</category><category>prompt-understanding</category><category>model-composition</category><category>visual-generation</category><category>language-understanding</category><category>model-performance</category><category>complex-prompting</category><category>iterative-generation</category></item><item><title>lots of little things happened this week</title><link>https://news.smol.ai/issues/25-03-21-ainews-lots-of-little-things-happened-this-week/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-21-ainews-lots-of-little-things-happened-this-week/</guid><description>**Anthropic** introduced a novel &apos;think&apos; tool enhancing instruction adherence and multi-step problem solving in agents, with combined reasoning and tool use demonstrated by **Claude**. **NVIDIA**&apos;s **Llama-3.3-Nemotron-Super-49B-v1** ranked #14 on LMArena, noted for strong math reasoning and a 15M post-training dataset. **Sakana AI** launched a Sudoku-based reasoning benchmark to advance AI problem-solving capabilities. **Meta AI** released **SWEET-RL**, a reinforcement learning algorithm improving long-horizon multi-turn tasks by 6%, and introduced **CollaborativeAgentBench**, a benchmark for collaborative LLM agents working with humans on programming and design tasks. **Percy Liang** relaunched the **HELM** benchmark with 5 challenging datasets evaluating 22 top language models.</description><pubDate>Sat, 22 Mar 2025 00:20:28 GMT</pubDate><category>anthropic</category><category>nvidia</category><category>sakana-ai</category><category>meta-ai-fair</category><category>llama-3-3-nemotron-super-49b-v1</category><category>claude</category><category>percy-liang</category><category>reinforcement-learning</category><category>reasoning</category><category>benchmarks</category><category>multi-turn-collaboration</category><category>instruction-following</category><category>dataset-release</category><category>model-evaluation</category></item><item><title>Promptable Prosody, SOTA ASR, and Semantic VAD: OpenAI revamps Voice AI</title><link>https://news.smol.ai/issues/25-03-20-ainews-promptable-prosody-sota-asr-and-semantic-vad-openai-revamps-voice-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-20-ainews-promptable-prosody-sota-asr-and-semantic-vad-openai-revamps-voice-ai/</guid><description>**OpenAI** has launched three new state-of-the-art audio models in their API, including **gpt-4o-transcribe**, a speech-to-text model outperforming Whisper, and **gpt-4o-mini-tts**, a text-to-speech model with promptable prosody allowing control over timing and emotion. The **Agents SDK** now supports audio, enabling voice agents. OpenAI also updated turn detection for real-time voice activity detection (VAD) based on speech content. Additionally, **OpenAI&apos;s o1-pro** model is available to select developers with advanced features like vision and function calling, though at higher compute costs. The community shows strong enthusiasm for these audio advancements, with a radio contest for TTS creations underway. Meanwhile, **Kokoro-82M v1.0** emerges as a leading open weights TTS model with competitive pricing on Replicate.</description><pubDate>Thu, 20 Mar 2025 22:51:24 GMT</pubDate><category>openai</category><category>replicate</category><category>gpt-4o-transcribe</category><category>gpt-4o-mini-tts</category><category>o1-pro</category><category>kokoro-82m</category><category>juberti</category><category>sama</category><category>reach_vb</category><category>kevinweil</category><category>omarsar0</category><category>speech-to-text</category><category>text-to-speech</category><category>voice-activity-detection</category><category>prompt-engineering</category><category>real-time-processing</category><category>model-release</category><category>api</category><category>function-calling</category><category>structured-outputs</category><category>model-performance</category></item><item><title>Every 7 Months: The Moore&apos;s Law for Agent Autonomy</title><link>https://news.smol.ai/issues/25-03-19-ainews-every-7-months-the-moores-law-for-agent-autonomy/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-19-ainews-every-7-months-the-moores-law-for-agent-autonomy/</guid><description>**METR** published a paper measuring AI agent autonomy progress, showing it has doubled every 7 months since **2019 (GPT-2)**. They introduced a new metric, the **50%-task-completion time horizon**, where models like **Claude 3.7 Sonnet** achieve 50% success in about 50 minutes. Projections estimate **1 day autonomy by 2028** and **1 month autonomy by late 2029**. Meanwhile, **Nvidia** released **Cosmos-Transfer1** for conditional world generation and **GR00T-N1-2B**, an open foundation model for humanoid robot reasoning with 2B parameters. **Canopy Labs** introduced **Orpheus 3B**, a high-quality text-to-speech model with zero-shot voice cloning and low latency. **Meta** reportedly delayed **Llama-4** release due to performance issues. **Microsoft** launched **Phi-4-multimodal**.</description><pubDate>Thu, 20 Mar 2025 01:59:24 GMT</pubDate><category>metr</category><category>nvidia</category><category>hugging-face</category><category>canopy-labs</category><category>meta-ai-fair</category><category>microsoft</category><category>claude-3-7-sonnet</category><category>llama-4</category><category>phi-4-multimodal</category><category>gpt-2</category><category>cosmos-transfer1</category><category>gr00t-n1-2b</category><category>orpheus-3b</category><category>reach_vb</category><category>akhaliq</category><category>drjimfan</category><category>scaling01</category><category>agent-autonomy</category><category>task-completion</category><category>multimodality</category><category>text-to-speech</category><category>robotics</category><category>foundation-models</category><category>model-release</category><category>scaling-laws</category><category>fine-tuning</category><category>zero-shot-learning</category><category>latency</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-18-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-18-ainews-not-much-happened-today/</guid><description>At Nvidia GTC Day 1, several AI updates were highlighted: **Google&apos;s Gemini 2.0 Flash** introduces image input/output but is not recommended for text-to-image tasks, with **Imagen 3** preferred for that. **Mistral AI** released **Mistral Small 3.1** with 128k token context window and competitive pricing. **Allen AI** launched **OLMo-32B**, an open LLM outperforming **GPT-4o mini** and **Qwen 2.5**. **ShieldGemma 2** was introduced for image safety classification. **LangChainAI** announced multiple updates including **Julian** powered by **LangGraph** and integration with **AnthropicAI&apos;s MCP**. Jeremy Howard released **fasttransform**, a Python library for data transformations. **Perplexity AI** partnered with **Kalshi** for NCAA March Madness predictions.</description><pubDate>Tue, 18 Mar 2025 22:00:12 GMT</pubDate><category>nvidia</category><category>google</category><category>mistral-ai</category><category>allen-ai</category><category>anthropic</category><category>langchainai</category><category>perplexity-ai</category><category>kalshi</category><category>stripe</category><category>qodoai</category><category>gemini-2.0-flash</category><category>imagen-3</category><category>mistral-small-3.1</category><category>mistral-3</category><category>gpt-4o-mini</category><category>claude-3.5-haiku</category><category>olm0-32b</category><category>qwen-2.5</category><category>shieldgemma-2</category><category>julian</category><category>fasttransform</category><category>jeremyphoward</category><category>karpathy</category><category>abacaj</category><category>mervenoyann</category><category>multimodality</category><category>image-generation</category><category>context-windows</category><category>model-pricing</category><category>open-source-models</category><category>image-classification</category><category>frameworks</category><category>python-libraries</category><category>partnerships</category></item><item><title>Cohere&apos;s Command A claims #3 open model spot (after DeepSeek and Gemma)</title><link>https://news.smol.ai/issues/25-03-17-ainews-coheres-command-a-claims-3-open-model-spot-after-deepseek-and-gemma/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-17-ainews-coheres-command-a-claims-3-open-model-spot-after-deepseek-and-gemma/</guid><description>**Cohere&apos;s Command A** model has solidified its position on the LMArena leaderboard, featuring an open-weight **111B** parameter model with an unusually long **256K context window** and competitive pricing. **Mistral AI** released the lightweight, multilingual, and multimodal **Mistral AI Small 3.1** model, optimized for single RTX 4090 or Mac 32GB RAM setups, with strong performance on instruct and multimodal benchmarks. The new OCR model **SmolDocling** offers fast document reading with low VRAM usage, outperforming larger models like Qwen2.5VL. Discussions highlight the importance of system-level improvements over raw LLM advancements, and **MCBench** is recommended as a superior AI benchmark for evaluating model capabilities across code, aesthetics, and awareness.</description><pubDate>Tue, 18 Mar 2025 00:28:53 GMT</pubDate><category>cohere</category><category>mistral-ai</category><category>hugging-face</category><category>command-a</category><category>mistral-ai-small-3.1</category><category>smoldocling</category><category>qwen-2.5-vl</category><category>aidangomez</category><category>sophiamyang</category><category>mervenoyann</category><category>aidan_mclau</category><category>reach_vb</category><category>lateinteraction</category><category>context-windows</category><category>multilinguality</category><category>multimodality</category><category>fine-tuning</category><category>benchmarking</category><category>ocr</category><category>model-performance</category><category>model-releases</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-14-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-14-ainews-not-much-happened-today/</guid><description>**Google DeepMind** announced updates to **Gemini 2.0**, including an upgraded **Flash Thinking model** with stronger reasoning and native image generation capabilities. **Cohere** launched **Command A**, a **111B** parameter dense model with a **256K context window** and competitive pricing, available on **Hugging Face**. **Meta AI** proposed **Dynamic Tanh (DyT)** as a replacement for normalization layers in Transformers, supported by **Yann LeCun**. **Alibaba** released **QwQ-32B**, a **32.5B** parameter model excelling in math and coding, fine-tuned with reinforcement learning and freely available under **Apache 2.0 license**. **Google DeepMind** also released **Gemma 3** models ranging from **1B to 27B** parameters with a **128K token context window** and over **140 language** support, plus **ShieldGemma 2**, an image safety checker. Benchmarking shows **Gemma 3 27B** has strong vision and memory efficiency but is outperformed by larger models like **Llama 3.3 70B** and **DeepSeek V3 671B**. The **Hugging Face LLM leaderboard** history was shared by @_lewtun.</description><pubDate>Fri, 14 Mar 2025 22:57:23 GMT</pubDate><category>google-deepmind</category><category>cohere</category><category>meta-ai-fair</category><category>alibaba</category><category>hugging-face</category><category>gemini-2.0-flash-thinking</category><category>command-a</category><category>qwq-32b</category><category>gemma-3-27b</category><category>gemma-3</category><category>shieldgemma-2</category><category>llama-3-70b</category><category>deepseek-r1</category><category>o1-mini</category><category>deepseek-v3</category><category>yann-lecun</category><category>model-updates</category><category>model-performance</category><category>benchmarking</category><category>reinforcement-learning</category><category>transformers</category><category>normalization-layers</category><category>image-generation</category><category>vision</category><category>memory-efficiency</category><category>context-windows</category><category>fine-tuning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-13-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-13-ainews-not-much-happened-today/</guid><description>**DeepSeek R1** demonstrates significant efficiency using **FP8** precision, outperforming **Gemma 3 27B** in benchmarks with a **Chatbot Arena Elo Score** of **1363** vs. **1338**, requiring substantial hardware like **32 H100 GPUs** and **2,560GB VRAM**. **OpenAI** labels **DeepSeek** as &quot;state-controlled&quot; and calls for bans on &quot;PRC-produced&quot; models, sparking community backlash accusing **OpenAI** and **Sam Altman** of anti-competitive behavior. Discussions emphasize **DeepSeek&apos;s** openness and affordability compared to **OpenAI**, with users highlighting its local and Hugging Face deployment options. Meanwhile, **Gemma 3** receives mixed community feedback on creativity and worldbuilding.</description><pubDate>Thu, 13 Mar 2025 21:13:47 GMT</pubDate><category>openai</category><category>nvidia</category><category>deepseek</category><category>hugging-face</category><category>deepseek-r1</category><category>gemma-3</category><category>gemma-3-27b</category><category>sam-altman</category><category>fp8</category><category>model-efficiency</category><category>hardware-requirements</category><category>quantization</category><category>benchmarking</category><category>model-deployment</category><category>open-source</category></item><item><title>Gemma 3 beats DeepSeek V3 in Elo, 2.0 Flash beats GPT4o with Native Image Gen</title><link>https://news.smol.ai/issues/25-03-12-ainews-gemma-3-beats-deepseek-v3-in-elo-20-flash-beats-gpt4o-with-native-image-gen/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-12-ainews-gemma-3-beats-deepseek-v3-in-elo-20-flash-beats-gpt4o-with-native-image-gen/</guid><description>**Google DeepMind** launched the **Gemma 3** family of models featuring a **128k context window**, **multimodal input (image and video)**, and **multilingual support for 140+ languages**. The **Gemma 3-27B** model ranks among the top open models on LMArena benchmarks, outperforming several competitors and matching **Gemini-1.5-Pro** on benchmarks. Additionally, **Gemini 2** introduced **Flash Native Image Generation** with advanced image editing capabilities, a feature teased by OpenAI but not launched. The updates highlight significant advances in context length, multimodality, and model efficiency via quantization.</description><pubDate>Thu, 13 Mar 2025 01:01:43 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>gemma-3</category><category>gemini-1.5-pro</category><category>gemini-2</category><category>o1-preview</category><category>o3-mini-high</category><category>deepseek-v3</category><category>claude-3.7-sonnet</category><category>qwen-2.5-max</category><category>reach_vb</category><category>_philschmid</category><category>danielhanchen</category><category>lmarena_ai</category><category>osanseviero</category><category>multimodality</category><category>multilinguality</category><category>context-window</category><category>quantization</category><category>image-generation</category><category>model-benchmarking</category><category>model-performance</category><category>vision</category></item><item><title>The new OpenAI Agents Platform</title><link>https://news.smol.ai/issues/25-03-11-ainews-the-new-openai-agents-platform/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-11-ainews-the-new-openai-agents-platform/</guid><description>**OpenAI** introduced a comprehensive suite of new tools for AI agents, including the **Responses API**, **Web Search Tool**, **Computer Use Tool**, **File Search Tool**, and an open-source **Agents SDK** with integrated observability tools, marking a significant step towards the &quot;Year of Agents.&quot; Meanwhile, **Reka AI** open-sourced **Reka Flash 3**, a **21B parameter reasoning model** that outperforms **o1-mini** and powers their Nexus platform, with weights available on **Hugging Face**. The **OlympicCoder** series surpassed **Claude 3.7 Sonnet** and much larger models on competitive coding benchmarks. **DeepSeek** built a **32K GPU cluster** capable of training V3-level models in under a week and is exploring AI distillation. **Hugging Face** announced **Cerebras** inference support, achieving over **2,000 tokens/s** on **Llama 3.3 70B**, 70x faster than leading GPUs. **Reka&apos;s Sonic-2** voice AI model delivers **40ms latency** via the **Together API**. **Alibaba&apos;s Qwen Chat** enhanced its multimodal interface with video understanding up to **500MB**, voice-to-text, guest mode, and expanded file uploads. *Sama* praised OpenAI&apos;s new API as &quot;one of the most well-designed and useful APIs ever.&quot;</description><pubDate>Wed, 12 Mar 2025 00:23:17 GMT</pubDate><category>openai</category><category>reka-ai</category><category>hugging-face</category><category>deepseek</category><category>togethercompute</category><category>alibaba</category><category>reka-flash-3</category><category>o1-mini</category><category>claude-3-7-sonnet</category><category>llama-3-3-70b</category><category>sonic-2</category><category>qwen-chat</category><category>olympiccoder</category><category>sama</category><category>reach_vb</category><category>ai-agents</category><category>api</category><category>model-releases</category><category>fine-tuning</category><category>reinforcement-learning</category><category>model-training</category><category>model-inference</category><category>multimodality</category><category>voice-synthesis</category><category>gpu-clusters</category><category>model-distillation</category><category>performance-optimization</category><category>open-source</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-10-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-10-ainews-not-much-happened-today/</guid><description>The AI news recap highlights several key developments: **nanoMoE**, a PyTorch implementation of a mid-sized Mixture-of-Experts (MoE) model inspired by Andrej Karpathy&apos;s nanoGPT, enables pretraining on commodity hardware within a week. An agentic leaderboard ranks LLMs powering **smolagents CodeAgent**, with **GPT-4.5** leading, followed by **Claude-3.7-Sonnet**. Discussions around **DeepSeek-R1** emphasize AI model commoditization, with DeepSeek dubbed the &quot;OpenAI of China.&quot; **Q-Filters** offer a training-free method for KV cache compression in autoregressive models, achieving **32x compression** with minimal perplexity loss. The **PokéChamp** minimax language agent, powered by **GPT-4o** and **Llama-3-8b**, demonstrates strong performance in Pokémon battles. Other notable models include **TinyR1-32B-Preview** with Branch-Merge Distillation, **R1-Searcher** incentivizing search capability via reinforcement learning, and the **Forgetting Transformer** using a Forget Gate in softmax attention. These advancements reflect ongoing innovation in model architectures, compression, reinforcement learning, and agentic AI.</description><pubDate>Mon, 10 Mar 2025 22:46:37 GMT</pubDate><category>openai</category><category>deepseek</category><category>hugging-face</category><category>gpt-4.5</category><category>claude-3.7-sonnet</category><category>deepseek-r1</category><category>smolagents-codeagent</category><category>gpt-4o</category><category>llama-3-8b</category><category>tinyr1-32b-preview</category><category>r1-searcher</category><category>forgetting-transformer</category><category>nanomoe</category><category>andrej-karpathy</category><category>cwolferesearch</category><category>aymericroucher</category><category>teortaxestex</category><category>jonathanross321</category><category>akhaliq</category><category>mixture-of-experts</category><category>reinforcement-learning</category><category>kv-cache-compression</category><category>agentic-ai</category><category>model-distillation</category><category>attention-mechanisms</category><category>model-compression</category><category>minimax</category><category>model-pretraining</category></item><item><title>DeepSeek&apos;s Open Source Stack</title><link>https://news.smol.ai/issues/25-03-07-ainews-deepseeks-open-source-stack/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-07-ainews-deepseeks-open-source-stack/</guid><description>**DeepSeek&apos;s Open Source Week** was summarized by PySpur, highlighting multiple interesting releases. The **Qwen QwQ-32B model** was fine-tuned into **START**, excelling in PhD-level science QA and math benchmarks. **Character-3**, an omnimodal AI video generation model by Hedra Labs and Together AI, enables realistic animated content creation. **Google DeepMind** introduced the **Gemini embedding model** with an 8k context window, ranking #1 on MMTEB, alongside the **Gemini 2.0 Code Executor** supporting Python libraries and auto-fix features. **Inception Labs&apos; Mercury Coder** is a diffusion-based code generation model offering faster token processing. **OpenAI** released **GPT-4.5**, their largest model yet but with less reasoning ability than some competitors. **AI21 Labs** launched **Jamba Mini 1.6**, noted for superior output speed compared to Gemini 2.0 Flash, GPT-4o mini, and Mistral Small 3. A new dataset of 1.9M scanned pages was released for OCR benchmarking, with **Mistral OCR** showing competitive but not top-tier document parsing performance compared to LLM/LVM-powered methods. *&quot;Cracked engineers are all you need.&quot;*</description><pubDate>Sat, 08 Mar 2025 05:06:31 GMT</pubDate><category>deepseek</category><category>pyspur</category><category>hugging-face</category><category>togethercompute</category><category>hedra-labs</category><category>google-deepmind</category><category>deeplearningai</category><category>openai</category><category>ai21-labs</category><category>mistral-ai</category><category>qwen-qwq-32b</category><category>start</category><category>character-3</category><category>gemini</category><category>gemini-2.0</category><category>mercury-coder</category><category>gpt-4.5</category><category>jamba-mini-1.6</category><category>gemini-2.0-flash</category><category>gpt-4o-mini</category><category>mistral-small-3</category><category>mistral-ocr</category><category>_akhaliq</category><category>lmarena_ai</category><category>reach_vb</category><category>danielhanchen</category><category>_philschmid</category><category>aidan_mclau</category><category>vikhyatk</category><category>jerryjliu0</category><category>fine-tuning</category><category>benchmarking</category><category>multimodality</category><category>code-generation</category><category>diffusion-models</category><category>model-performance</category><category>model-optimization</category><category>ocr</category><category>embedding-models</category><category>context-windows</category><category>runtime-limits</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-06-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-06-ainews-not-much-happened-today/</guid><description>**AI21 Labs launched Jamba 1.6**, touted as the **best open model for private enterprise deployment**, outperforming **Cohere, Mistral, and Llama** on benchmarks like **Arena Hard**. **Mistral AI** released a state-of-the-art **multimodal OCR model** with multilingual and structured output capabilities, available for on-prem deployment. **Alibaba Qwen** introduced **QwQ-32B**, an open-weight reasoning model with **32B parameters** and cost-effective usage, showing competitive benchmark scores. **OpenAI** released **o1** and **o3-mini** models with advanced API features including streaming and function calling. **AMD** unveiled **Instella**, open-source 3B parameter language models trained on **AMD Instinct MI300X GPUs**, competing with **Llama-3.2-3B** and others. **Alibaba** also released **Babel**, open multilingual LLMs performing comparably to **GPT-4o**. **Anthropic** launched **Claude 3.7 Sonnet**, enhancing reasoning and prompt engineering capabilities.</description><pubDate>Fri, 07 Mar 2025 05:50:14 GMT</pubDate><category>ai21-labs</category><category>mistral-ai</category><category>alibaba</category><category>openai</category><category>amd</category><category>anthropic</category><category>hugging-face</category><category>jamba-1.6</category><category>mistral-ocr</category><category>qwq-32b</category><category>o1</category><category>o3-mini</category><category>instella</category><category>llama-3-2-3b</category><category>gemma-2-2b</category><category>qwen-2-5-3b</category><category>babel-9b</category><category>babel-83b</category><category>gpt-4o</category><category>claude-3-7-sonnet</category><category>multimodality</category><category>ocr</category><category>multilinguality</category><category>structured-output</category><category>on-prem-deployment</category><category>reasoning</category><category>benchmarking</category><category>api</category><category>open-source</category><category>model-training</category><category>gpu-optimization</category><category>prompt-engineering</category><category>function-calling</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-03-04-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-04-ainews-not-much-happened-today/</guid><description>**Weights and Biases** announced a **$1.7 billion acquisition by CoreWeave** ahead of CoreWeave&apos;s IPO. **CohereForAI** released the **Aya Vision models (8B and 32B parameters)** supporting **23 languages**, outperforming larger models like **Llama-3.2 90B Vision** and **Molmo 72B**. **Microsoft** introduced **Phi-4-Mini (3.8B parameters)** and **Phi-4-Multimodal models**, excelling in math, coding, and multimodal benchmarks. **CogView4**, a **6B parameter text-to-image model** with **2048x2048 resolution** and Apache 2.0 license, was released. **Alibaba** launched **Wan 2.1**, an open-source video generation model with **720p output** and **16 fps generation**. **Google** announced new AI features for Pixel devices including **Scam Detection** and **Gemini integrations**. **LlamaCloud** reached **General Availability** and raised **$19M Series A funding**, serving over **100 Fortune 500 companies**. **Weaviate** launched the **Query Agent**, the first of three Weaviate Agents.</description><pubDate>Wed, 05 Mar 2025 05:17:34 GMT</pubDate><category>weights-and-biases</category><category>coreweave</category><category>cohereforai</category><category>microsoft</category><category>alibaba</category><category>google</category><category>llamaindex</category><category>weaviate</category><category>aya-vision-8b</category><category>aya-vision-32b</category><category>llama-3-2-90b-vision</category><category>molmo-72b</category><category>phi-4-mini</category><category>phi-4-multimodal</category><category>cogview4</category><category>wan-2-1</category><category>mervenoyann</category><category>reach_vb</category><category>jayalammar</category><category>sarahookr</category><category>aidangomez</category><category>nickfrosst</category><category>dair_ai</category><category>akhaliq</category><category>bobvanluijt</category><category>jerryjliu0</category><category>multilinguality</category><category>vision</category><category>multimodality</category><category>image-generation</category><category>video-generation</category><category>model-releases</category><category>benchmarking</category><category>funding</category><category>agentic-ai</category><category>model-performance</category></item><item><title>Anthropic&apos;s $61.5B Series E</title><link>https://news.smol.ai/issues/25-03-03-ainews-anthropics-dollar615b-series-e/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-03-03-ainews-anthropics-dollar615b-series-e/</guid><description>**Anthropic** raised a **$3.5 billion Series E funding round** at a **$61.5 billion valuation**, signaling strong financial backing for the **Claude** AI model. **GPT-4.5** achieved **#1 rank across all categories** on the LMArena leaderboard, excelling in multi-turn conversations, coding, math, creative writing, and style control. **DeepSeek R1** tied with GPT-4.5 for top performance on hard prompts with style control. Discussions highlighted comparisons between **GPT-4.5** and **Claude 3.7 Sonnet** in coding and workflow applications. The importance of the **LMSYS benchmark** was emphasized, though some questioned the relevance of benchmarks versus user acquisition. Additionally, **Perplexity AI** partnered with **Deutsche Telekom** to integrate the **Perplexity Assistant** into a new AI phone.</description><pubDate>Tue, 04 Mar 2025 06:51:49 GMT</pubDate><category>anthropic</category><category>openai</category><category>deepseek</category><category>lmsys</category><category>perplexity-ai</category><category>deutsche-telekom</category><category>gpt-4.5</category><category>claude-3.7-sonnet</category><category>deepseek-r1</category><category>lmarena_ai</category><category>teortaxestex</category><category>casper_hansen_</category><category>omarsar0</category><category>aidan_mclau</category><category>willdepue</category><category>vikhyatk</category><category>teknim1</category><category>reach_vb</category><category>_aidan_clark_</category><category>cto_junior</category><category>aravsrinivas</category><category>model-performance</category><category>benchmarking</category><category>style-control</category><category>coding</category><category>multi-turn</category><category>funding</category><category>partnerships</category><category>workflow</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-28-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-28-ainews-not-much-happened-today/</guid><description>**GPT-4.5** sparked mixed reactions on Twitter, with **@karpathy** noting users preferred **GPT-4** in a poll despite his personal favor for GPT-4.5&apos;s creativity and humor. Critics like **@abacaj** highlighted **GPT-4.5&apos;s slowness** and questioned its practical value and pricing compared to other models. Performance-wise, **GPT-4.5** ranks above **GPT-4o** but below **o1** and **Claude 3.5 Sonnet**, with **Claude 3.7** outperforming it on many tasks yet GPT-4.5 praised for its humor and &quot;vibes.&quot; Speculation about GPT-4.5&apos;s size suggests around **5 trillion parameters**. Discussions also touched on pricing disparities, with **Perplexity Deep Research** at $20/month versus ChatGPT at $200/month. The emotional intelligence and humor of models like **Claude 3.7** were also noted.</description><pubDate>Sat, 01 Mar 2025 03:41:57 GMT</pubDate><category>openai</category><category>anthropic</category><category>perplexity-ai</category><category>deepseek</category><category>scaling01</category><category>gpt-4.5</category><category>gpt-4</category><category>gpt-4o</category><category>o1</category><category>claude-3.5-sonnet</category><category>claude-3.7</category><category>claude-3-opus</category><category>deepseek-v3</category><category>grok-3</category><category>andrej-karpathy</category><category>jeremyphoward</category><category>abacaj</category><category>stevenheidel</category><category>yuchenj_uw</category><category>aravsrinivas</category><category>dylan522p</category><category>random_walker</category><category>model-performance</category><category>humor</category><category>emotional-intelligence</category><category>model-comparison</category><category>pricing</category><category>context-windows</category><category>model-size</category><category>user-experience</category></item><item><title>GPT 4.5 — Chonky Orion ships!</title><link>https://news.smol.ai/issues/25-02-27-ainews-gpt-45-chonky-orion-ships/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-27-ainews-gpt-45-chonky-orion-ships/</guid><description>**OpenAI released GPT-4.5** as a research preview, highlighting its **deep world knowledge**, **improved understanding of user intent**, and a **128,000 token context window**. It is noted for excelling in **writing, creative tasks, image understanding, and data extraction** but is not a reasoning model. **Microsoft unveiled Phi-4 Multimodal and Phi-4 Mini**, open-source models integrating **text, vision, and speech/audio**, with strong performance in **math and coding tasks**. **Cohere released Command R7B Arabic**, an open-weights model optimized for **Arabic language capabilities** targeting enterprises in the MENA region. The community is exploring the impact of larger models on creative writing, intent understanding, and world knowledge, with GPT-4.5 expected to be a basis for GPT-5.</description><pubDate>Fri, 28 Feb 2025 07:24:08 GMT</pubDate><category>openai</category><category>microsoft</category><category>cohere</category><category>gpt-4.5</category><category>phi-4-multimodal</category><category>phi-4-mini</category><category>command-r7b-arabic</category><category>sama</category><category>kevinweil</category><category>aidan_mclau</category><category>omarsar0</category><category>rasbt</category><category>reach_vb</category><category>creative-writing</category><category>natural-language-processing</category><category>multimodality</category><category>math</category><category>coding</category><category>context-windows</category><category>model-releases</category><category>open-source</category><category>arabic-language</category></item><item><title>lots of small launches</title><link>https://news.smol.ai/issues/25-02-26-ainews-lots-of-small-launches/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-26-ainews-lots-of-small-launches/</guid><description>**GPT-4o Advanced Voice Preview** is now available for free ChatGPT users with enhanced daily limits for Plus and Pro users. **Claude 3.7 Sonnet** has achieved the top rank in WebDev Arena with improved token efficiency. **DeepSeek-R1** with 671B parameters benefits from the **Together Inference** platform optimizing NVIDIA Blackwell GPU usage, alongside the open-source **DeepGEMM** CUDA library delivering up to 2.7x speedups on Hopper GPUs. **Perplexity** launched a new Voice Mode and a **Deep Research API**. The upcoming **Grok 3 API** will support a 1M token context window. Several companies including **Elicit**, **Amazon**, **Anthropic**, **Cloudflare**, **FLORA**, **Elevenlabs**, and **Inception Labs** announced new funding rounds, product launches, and model releases.</description><pubDate>Thu, 27 Feb 2025 04:09:12 GMT</pubDate><category>openai</category><category>anthropic</category><category>amazon</category><category>cloudflare</category><category>perplexity-ai</category><category>deepseek-ai</category><category>togethercompute</category><category>elevenlabs</category><category>elicitorg</category><category>inceptionailabs</category><category>mistral-ai</category><category>gpt-4o</category><category>claude-3.7-sonnet</category><category>claude-3.7</category><category>claude-3.5-sonnet</category><category>deepseek-r1</category><category>deepseek-v3</category><category>grok-3</category><category>lmarena_ai</category><category>alexalbert__</category><category>aravsrinivas</category><category>reach_vb</category><category>voice</category><category>model-releases</category><category>cuda</category><category>gpu-optimization</category><category>inference</category><category>open-source</category><category>api</category><category>model-performance</category><category>token-efficiency</category><category>context-windows</category><category>cuda</category><category>jit-compilation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-25-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-25-ainews-not-much-happened-today/</guid><description>**Claude 3.7 Sonnet** demonstrates exceptional coding and reasoning capabilities, outperforming models like **DeepSeek R1**, **O3-mini**, and **GPT-4o** on benchmarks such as **SciCode** and **LiveCodeBench**. It is available on platforms including **Perplexity Pro**, **Anthropic**, **Amazon Bedrock**, and **Google Cloud**, with pricing at **$3/$15 per million tokens**. Key features include a **64k token thinking mode**, **200k context window**, and the **CLI-based coding assistant Claude Code**. Meanwhile, **DeepSeek** released **DeepEP**, an open-source communication library optimized for MoE model training and inference with support for **NVLink**, **RDMA**, and **FP8**. These updates highlight advancements in coding AI and efficient model training infrastructure.</description><pubDate>Wed, 26 Feb 2025 02:19:12 GMT</pubDate><category>anthropic</category><category>perplexity-ai</category><category>amazon</category><category>google-cloud</category><category>deepseek_ai</category><category>claude-3.7-sonnet</category><category>claude-3.7</category><category>deepseek-r1</category><category>o3-mini</category><category>deepseek-v3</category><category>gemini-2.0-pro</category><category>gpt-4o</category><category>qwen2.5-coder-32b-instruct</category><category>skirano</category><category>omarsar0</category><category>reach_vb</category><category>artificialanlys</category><category>terryyuezhuo</category><category>_akhaliq</category><category>_philschmid</category><category>catherineols</category><category>goodside</category><category>danielhanchen</category><category>coding</category><category>reasoning</category><category>model-benchmarking</category><category>agentic-workflows</category><category>context-window</category><category>model-performance</category><category>open-source</category><category>moe</category><category>model-training</category><category>communication-libraries</category><category>fp8</category><category>nvlink</category><category>rdma</category><category>cli-tools</category></item><item><title>Claude 3.7 Sonnet</title><link>https://news.smol.ai/issues/25-02-24-ainews-claude-37-sonnet/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-24-ainews-claude-37-sonnet/</guid><description>**Anthropic** launched **Claude 3.7 Sonnet**, their most intelligent model to date featuring hybrid reasoning with two thinking modes: near-instant and extended step-by-step thinking. The release includes **Claude Code**, an agentic coding tool in limited preview, and supports a **128k output token capability** in beta. Claude 3.7 Sonnet performs well on coding benchmarks like **SWE-Bench Verified** and **Cognition&apos;s junior-dev eval**, and introduces advanced features such as streaming thinking, prompt caching, and tool use. The model is also benchmarked on **Pokebench**, reflecting agentic capabilities similar to the Voyager paper. The launch is accompanied by extensive documentation, cookbooks, and prompting guides for extended thinking. *&quot;The first generally available hybrid reasoning model&quot;* and *&quot;first coding tool from Anthropic&quot;* were highlighted in social media announcements.</description><pubDate>Tue, 25 Feb 2025 05:58:56 GMT</pubDate><category>anthropic</category><category>claude-3-7-sonnet</category><category>claude-3</category><category>claude-code</category><category>hybrid-reasoning</category><category>extended-thinking</category><category>coding-benchmarks</category><category>agentic-ai</category><category>prompt-caching</category><category>streaming</category><category>token-capacity</category><category>tool-use</category></item><item><title>AI Engineer Summit Day 1</title><link>https://news.smol.ai/issues/25-02-21-ainews-ai-engineer-summit-day-1/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-21-ainews-ai-engineer-summit-day-1/</guid><description>The **AIE Summit** in NYC highlighted key talks including **Grace Isford&apos;s Trends Keynote**, **Neo4j/Pfizer&apos;s presentation**, and **OpenAI&apos;s first definition of Agents**. Speakers announced **$930 million in funding**. On AI Twitter, discussions focused on **Grok-3** and **o3-mini** models, with debates on performance and benchmarking, including **Grok-3&apos;s record compute scale of 4e26 to 5e26 FLOP**. The **o3-mini** model uncovered a critical **CUDA kernel bug** in Sakana AI&apos;s code. **DeepSeek-R1** was promoted as an open-source alternative with notable training batch sizes. Additionally, **Alibaba** announced the **Qwen 2.5-VL** model release.</description><pubDate>Sat, 22 Feb 2025 02:50:34 GMT</pubDate><category>openai</category><category>anthropic</category><category>xai</category><category>togethercompute</category><category>alibaba</category><category>sakana-ai</category><category>grok-3</category><category>o3-mini</category><category>deepseek-r1</category><category>qwen-2.5-vl</category><category>aidan_mclau</category><category>giffmana</category><category>nrehiew_</category><category>teortaxestex</category><category>epochairesearch</category><category>andrew_n_carr</category><category>borismpower</category><category>yuhu_ai_</category><category>benchmarking</category><category>model-performance</category><category>cuda</category><category>model-training</category><category>open-source</category><category>debugging</category><category>inference-speed</category><category>batch-size</category><category>reinforcement-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-21-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-21-ainews-not-much-happened-today/</guid><description>**Grok-3**, a new family of LLMs from **xAI** using **200,000 Nvidia H100 GPUs** for advanced reasoning, outperforms models from **Google, Anthropic, and OpenAI** on math, science, and coding benchmarks. **DeepSeek-R1** from **ByteDance Research** achieves top accuracy on the challenging **SuperGPQA** dataset. **SigLIP 2** from **GoogleDeepMind** improves semantic understanding and OCR with flexible resolutions and multilingual capabilities, available on HuggingFace. **OpenAI&apos;s o3-mini-high** ranks #1 in coding and math prompts. **Perplexity&apos;s R1 1776**, a post-trained version of DeepSeek R1, is available on Ollama. The **Llamba** family distills **Llama-3.x** into efficient recurrent models with higher throughput. **AlphaMaze** combines DeepSeek R1 with GRPO for visual reasoning on ARC-AGI puzzles. **Audiobox Aesthetics** from **Meta AI** offers unified quality assessment for audio. The community notes that Grok 3&apos;s compute increase yields only modest performance gains.</description><pubDate>Fri, 21 Feb 2025 22:50:40 GMT</pubDate><category>xai</category><category>nvidia</category><category>google-deepmind</category><category>anthropic</category><category>openai</category><category>bytedance</category><category>ollama</category><category>meta-ai-fair</category><category>grok-3</category><category>deepseek-r1</category><category>siglip-2</category><category>o3-mini-high</category><category>r1-1776</category><category>llamba-1b</category><category>llamba-3b</category><category>llamba-8b</category><category>llama-3</category><category>alphamaze</category><category>audiobox-aesthetics</category><category>scaling01</category><category>iscienceluvr</category><category>philschmid</category><category>arankomatsuzaki</category><category>reach_vb</category><category>mervenoyann</category><category>wightmanr</category><category>lmarena_ai</category><category>ollama</category><category>akhaliq</category><category>benchmarking</category><category>model-releases</category><category>performance</category><category>reasoning</category><category>multimodality</category><category>semantic-understanding</category><category>ocr</category><category>multilinguality</category><category>model-distillation</category><category>recurrent-neural-networks</category><category>visual-reasoning</category><category>audio-processing</category></item><item><title>The Ultra-Scale Playbook: Training LLMs on GPU Clusters</title><link>https://news.smol.ai/issues/25-02-19-ainews-the-ultra-scale-playbook-training-llms-on-gpu-clusters/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-19-ainews-the-ultra-scale-playbook-training-llms-on-gpu-clusters/</guid><description>**Huggingface** released &quot;The Ultra-Scale Playbook: Training LLMs on GPU Clusters,&quot; an interactive blogpost based on **4000 scaling experiments on up to 512 GPUs**, providing detailed insights into modern GPU training strategies. **DeepSeek** introduced the Native Sparse Attention (NSA) model, gaining significant community attention, while **Perplexity AI** launched R1-1776, an uncensored and unbiased version of DeepSeek&apos;s R1 model. **Google DeepMind** unveiled PaliGemma 2 Mix, a multi-task vision-language model available in **3B, 10B, and 28B sizes**. **Microsoft** introduced Muse, a generative AI model trained on the game Bleeding Edge, and presented Magma, a foundation model for multimodal AI agents excelling in UI navigation and robotic manipulation. **Baichuan-M1-14B** was announced as a state-of-the-art medical LLM trained on **20T tokens**, and a fully open-source 40B genome modeling model using StripedHyena 2 architecture was also released. *&quot;Making your own gaming experience is coming sooner than you&apos;d think,&quot;* noted in relation to Muse.</description><pubDate>Thu, 20 Feb 2025 05:57:17 GMT</pubDate><category>huggingface</category><category>deepseek</category><category>perplexity-ai</category><category>google-deepmind</category><category>microsoft</category><category>baichuan</category><category>stripedhyena</category><category>deepseek-native-sparse-attention</category><category>r1-1776</category><category>paligemma-2-mix</category><category>muse</category><category>baichuan-m1-14b</category><category>stripedhyena-2</category><category>eliebakouch</category><category>nouamanetazi</category><category>lvwerra</category><category>thom-wolf</category><category>proftomyeh</category><category>alex-wang</category><category>aravsrinivas</category><category>_akhaliq</category><category>_philschmid</category><category>mervenoyann</category><category>reach_vb</category><category>arankomatsuzaki</category><category>maximelabonne</category><category>gpu-training</category><category>scaling</category><category>multimodality</category><category>vision</category><category>model-training</category><category>foundation-models</category><category>medical-llm</category><category>genome-modeling</category><category>robotic-manipulation</category><category>interactive-content</category></item><item><title>X.ai Grok 3 and Mira Murati&apos;s Thinking Machines</title><link>https://news.smol.ai/issues/25-02-18-ainews-xai-grok-3-and-mira-muratis-thinking-machines/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-18-ainews-xai-grok-3-and-mira-muratis-thinking-machines/</guid><description>**Grok 3** has launched with mixed opinions but strong benchmark performance, notably outperforming models like **Gemini 2 Pro** and **GPT-4o**. The **Grok-3 mini** variant shows competitive and sometimes superior capabilities, especially in reasoning and coding, with reinforcement learning playing a key role. **Mira Murati** has publicly shared her post-OpenAI plan, founding the frontier lab **Thinking Machines**, focusing on collaborative, personalizable AI, multimodality, and empirical safety and alignment research, reminiscent of **Anthropic**&apos;s approach.</description><pubDate>Tue, 18 Feb 2025 23:54:10 GMT</pubDate><category>anthropic</category><category>openai</category><category>thinking-machines</category><category>grok-3</category><category>grok-3-mini</category><category>gemini-2-pro</category><category>gpt-4o</category><category>o3-mini-high</category><category>o1</category><category>deepseek-r1</category><category>mira-murati</category><category>lmarena_ai</category><category>karpathy</category><category>omarsar0</category><category>ibab</category><category>arankomatsuzaki</category><category>iscienceluvr</category><category>scaling01</category><category>benchmarking</category><category>reasoning</category><category>reinforcement-learning</category><category>coding</category><category>multimodality</category><category>safety</category><category>alignment</category><category>research-publishing</category><category>model-performance</category><category>creative-ai</category></item><item><title>LLaDA: Large Language Diffusion Models</title><link>https://news.smol.ai/issues/25-02-17-ainews-llada-large-language-diffusion-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-17-ainews-llada-large-language-diffusion-models/</guid><description>**LLaDA (Large Language Diffusion Model) 8B** is a breakthrough diffusion-based language model that rivals **LLaMA 3 8B** while training on **7x fewer tokens (2 trillion tokens)** and using **0.13 million H800 GPU hours**. It introduces a novel text generation approach by predicting uniformly masked tokens in a diffusion process, enabling multi-turn dialogue and instruction-following. Alongside, **StepFun AI** released two major models: **Step-Video-T2V 30B**, a text-to-video model generating up to **204 frames** with high coherence and motion quality, and **Step-Audio-Chat 132B**, a voice-to-voice model. Additionally, challenging multimodal benchmarks like **Scale AI&apos;s EnigmaEval** and **Cambridge&apos;s ZeroBench** highlight current frontier models scoring zero, emphasizing the difficulty of these tasks. The community also noted the return of diffusion models in language modeling, a previously speculative architecture now scaled successfully.</description><pubDate>Tue, 18 Feb 2025 03:27:47 GMT</pubDate><category>stepfun-ai</category><category>scale-ai</category><category>cambridge</category><category>llamaindex</category><category>llada-8b</category><category>llama-3-8b</category><category>step-video-t2v-30b</category><category>step-audio-chat-132b</category><category>llama-2-7b</category><category>arankomatsuzaki</category><category>_akhaliq</category><category>omarsar0</category><category>iscienceluvr</category><category>gallabytes</category><category>maximelabonne</category><category>reach_vb</category><category>diffusion-models</category><category>text-generation</category><category>multimodality</category><category>video-generation</category><category>voice-processing</category><category>benchmarking</category><category>instruction-following</category><category>model-scaling</category><category>gpu-usage</category><category>long-context</category><category>multi-turn-dialogue</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-14-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-14-ainews-not-much-happened-today/</guid><description>**Smolagents** library by **Huggingface** continues trending. **ChatGPT-4o** latest version `chatgpt-40-latest-20250129` released. **DeepSeek R1 671B** sets speed record at **198 t/s**, fastest reasoning model, recommended with specific prompt settings. **Perplexity Deep Research** outperforms models like **Gemini Thinking**, **o3-mini**, and **DeepSeek-R1** on **Humanity&apos;s Last Exam** benchmark with **21.1%** score and **93.9%** accuracy on **SimpleQA**. **ChatGPT-4o** ranks #1 on Arena leaderboard in multiple categories except math. **OpenAI&apos;s o3 model** powers Deep Research tool for ChatGPT Pro users. **Gemini 2 Flash** and **Qwen 2.5** models support LLMGrading verifier. **Qwen 2.5** models added to PocketPal app. **MLX** shows small LLMs like Qwen 0.5B generate tokens at high speed on M4 Max and iPhone 16 Pro. **Gemini Flash 2.0** leads new AI agent leaderboard. **DeepSeek R1** is most liked on Hugging Face with over 10 million downloads.</description><pubDate>Sat, 15 Feb 2025 01:23:56 GMT</pubDate><category>hugging-face</category><category>openai</category><category>perplexity-ai</category><category>deepseek-ai</category><category>gemini</category><category>qwen</category><category>metr_evals</category><category>chatgpt-4o</category><category>deepseek-r1</category><category>o3</category><category>o3-mini</category><category>gemini-2-flash</category><category>qwen-2.5</category><category>qwen-0.5b</category><category>_akhaliq</category><category>aravsrinivas</category><category>lmarena_ai</category><category>omarsar0</category><category>risingsayak</category><category>reasoning</category><category>benchmarking</category><category>model-performance</category><category>prompt-engineering</category><category>model-optimization</category><category>model-deployment</category><category>small-language-models</category><category>mobile-ai</category><category>ai-agents</category><category>speed-optimization</category></item><item><title>Reasoning Models are Near-Superhuman Coders (OpenAI IOI, Nvidia Kernels)</title><link>https://news.smol.ai/issues/25-02-13-ainews-reasoning-models-are-near-superhuman-coders-openai-ioi-nvidia-kernels/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-13-ainews-reasoning-models-are-near-superhuman-coders-openai-ioi-nvidia-kernels/</guid><description>**o3 model** achieved a **gold medal at the 2024 IOI** and ranks in the **99.8 percentile on Codeforces**, outperforming most humans with reinforcement learning (RL) methods proving superior to inductive bias approaches. **Nvidia&apos;s DeepSeek-R1** autonomously generates GPU kernels that surpass some expert-engineered kernels, showcasing simple yet effective AI-driven optimization. **OpenAI** updated **o1 and o3-mini** models to support file and image uploads in ChatGPT and released **DeepResearch**, a powerful research assistant based on the **o3 model with RL** for deep chain-of-thought reasoning. **Ollama** introduced **OpenThinker models** fine-tuned from **Qwen2.5**, outperforming some DeepSeek-R1 distillation models. **ElevenLabs** grew into a $3.3 billion company specializing in AI voice synthesis without open-sourcing their technology. Research highlights include **Sakana AI Labs&apos; TAID knowledge distillation method** receiving a Spotlight at **ICLR 2025**, and **Apple&apos;s work on scaling laws for mixture-of-experts (MoEs)**. The importance of open-source AI for scientific discovery was also emphasized.</description><pubDate>Fri, 14 Feb 2025 02:42:41 GMT</pubDate><category>openai</category><category>nvidia</category><category>ollama</category><category>elevenlabs</category><category>sakana-ai</category><category>apple</category><category>o3</category><category>o1</category><category>o3-mini</category><category>deepseek-r1</category><category>qwen-2.5</category><category>openthinker</category><category>alex-wei</category><category>karpathy</category><category>abacaj</category><category>awnihannun</category><category>reinforcement-learning</category><category>gpu-kernel-optimization</category><category>fine-tuning</category><category>knowledge-distillation</category><category>scaling-laws</category><category>chain-of-thought-reasoning</category><category>model-accessibility</category></item><item><title>small news items</title><link>https://news.smol.ai/issues/25-02-12-ainews-small-news-items/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-12-ainews-small-news-items/</guid><description>**OpenAI** announced plans for **GPT-4.5 (Orion)** and **GPT-5**, with GPT-5 integrating the **o3** model and offering unlimited chat access in the free tier. **DeepSeek R1 Distilled Qwen 1.5B** outperforms OpenAI&apos;s **o1-preview** on math benchmarks, while **ModernBERT 0.3b** surpasses **Qwen 0.5b** at MMLU without fine-tuning. **Mistral** and **Perplexity** adopt **Cerebras** hardware for 10x performance gains. OpenAI&apos;s **o3** model won a gold medal at the 2024 International Olympiad in Informatics. Partnerships include **Qwen** with **Groq**. Significant RLHF activity is noted in Nigeria and the global south, and **Bytedance** is expected to rise in AI prominence soon. *&quot;GPT5 is all you need.&quot;*</description><pubDate>Thu, 13 Feb 2025 00:10:12 GMT</pubDate><category>openai</category><category>ollama</category><category>mistral</category><category>perplexity</category><category>cerebras</category><category>alibaba</category><category>groq</category><category>bytedance</category><category>gpt-4.5</category><category>gpt-5</category><category>deepseek-r1-distilled-qwen-1.5b</category><category>o1-preview</category><category>modernbert-0.3b</category><category>qwen-0.5b</category><category>o3</category><category>jeremyphoward</category><category>arankomatsuzaki</category><category>sama</category><category>nrehiew_</category><category>danhendrycks</category><category>akhaliq</category><category>math</category><category>benchmarking</category><category>fine-tuning</category><category>model-performance</category><category>reinforcement-learning</category><category>model-architecture</category><category>partnerships</category><category>funding</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-11-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-11-ainews-not-much-happened-today/</guid><description>**Zyphra AI** launched **Zonos-v0.1**, a leading open-weight text-to-speech model supporting multiple languages and zero-shot voice cloning. **Meta FAIR** released the open-source **Audiobox Aesthetics** model trained on 562 hours of audio data. **Kyutai Labs** introduced **Moshi**, a real-time speech-to-speech system with low latency. **Perplexity AI** announced the **Sonar** model based on **Llama 3.3 70b**, outperforming top models like **GPT-4o** and **Claude 3.5 Sonnet** with 1200 tokens/second speed, powered by **Cerebras** infrastructure. **UC Berkeley** open-sourced a 1.5B model trained with reinforcement learning that beats **o1-preview** on math tasks. **ReasonFlux-32B** achieved 91.2% on the MATH benchmark, outperforming **OpenAI o1-preview**. **CrossPoster**, an AI agent for cross-platform posting, was released using **LlamaIndex** workflows. **Brilliant Labs** integrated the **Google DeepMind Gemini Live API** into smart glasses for real-time translation and object identification.</description><pubDate>Wed, 12 Feb 2025 01:24:43 GMT</pubDate><category>zyphra-ai</category><category>meta-ai-fair</category><category>kyutai-labs</category><category>perplexity-ai</category><category>cerebras</category><category>uc-berkeley</category><category>brilliant-labs</category><category>google-deepmind</category><category>zonos-v0.1</category><category>audiobox-aesthetics</category><category>moshi</category><category>sonar</category><category>llama-3-70b</category><category>gpt-4o-mini</category><category>claude-3.5-haiku</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>deepseek-r1-distilled-qwen-1.5b</category><category>reasonflux-32b</category><category>o1-preview</category><category>danhendrycks</category><category>text-to-speech</category><category>speech-to-speech</category><category>benchmarking</category><category>model-performance</category><category>reinforcement-learning</category><category>math</category><category>real-time-processing</category><category>open-source</category><category>cross-platform-integration</category><category>multilinguality</category><category>zero-shot-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-10-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-10-ainews-not-much-happened-today/</guid><description>**Google** released **Gemini 2.0 Flash Thinking Experimental 1-21**, a vision-language reasoning model with a **1 million-token context window** and improved accuracy on science, math, and multimedia benchmarks, surpassing **DeepSeek-R1** but trailing **OpenAI&apos;s o1**. **ZyphraAI** launched **Zonos**, a multilingual **Text-to-Speech model** with **instant voice cloning** and controls for speaking rate, pitch, and emotions, running at **~2x real-time speed on RTX 4090**. **Hugging Face** released **OpenR1-Math-220k**, a large-scale **math reasoning dataset** with **220K problems** and **800K reasoning traces** generated on **512 H100 GPUs**. **Tom Goldstein** introduced **Huginn-3.5B**, an open-source latent reasoning model trained on **800B tokens** that outperforms larger models on reasoning tasks like **GSM8K**. Discussions by **Jeremy Howard** and **iScienceLuvr** highlight advances in implicit latent reasoning and debate the future of human-readable reasoning traces. **Anthropic** launched the **Anthropic Economic Index** to analyze AI&apos;s economic impact using millions of **Claude** conversations.</description><pubDate>Tue, 11 Feb 2025 03:56:45 GMT</pubDate><category>google</category><category>zyphraai</category><category>hugging-face</category><category>anthropic</category><category>deepseek</category><category>openai</category><category>gemini-2.0-flash-thinking-experimental-1-21</category><category>zonos</category><category>openr1-math-220k</category><category>huginn-3.5b</category><category>deepseek-r1</category><category>o1</category><category>claude</category><category>jeremyphoward</category><category>andrej-karpathy</category><category>tom-goldstein</category><category>reach_vb</category><category>iscienceluvr</category><category>vision</category><category>multilingual-models</category><category>text-to-speech</category><category>voice-cloning</category><category>math</category><category>reasoning</category><category>latent-reasoning</category><category>chain-of-thought</category><category>dataset-release</category><category>fine-tuning</category><category>model-training</category><category>model-performance</category><category>context-windows</category><category>benchmarking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-02-07-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-07-ainews-not-much-happened-today/</guid><description>**DeepSeek-R1 surpasses OpenAI in GitHub stars**, marking a milestone in open-source AI with rapid growth in community interest. **AlphaGeometry2 achieves gold-medalist level performance with an 84% solving rate on IMO geometry problems**, showcasing significant advancements in AI reasoning. **LangChain releases a tutorial for building AI agents in JavaScript**, enhancing developer capabilities in agent deployment. Reflections on **Anthropic&apos;s Claude model** reveal early access and influence on AI development timelines. Lighthearted AI humor includes calls to ban second-order optimizers and challenges in web development longevity. The AI Engineer Summit 2025 workshops were announced, continuing community engagement and education.</description><pubDate>Sat, 08 Feb 2025 04:22:33 GMT</pubDate><category>deepseek</category><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>langchain</category><category>adyen</category><category>deepseek-r1</category><category>alphageometry-2</category><category>claude</category><category>akhaliq</category><category>lmthang</category><category>aymericroucher</category><category>vikhyatk</category><category>swyx</category><category>open-source</category><category>reasoning</category><category>agentic-ai</category><category>javascript</category><category>model-release</category><category>memes</category><category>ai-development</category><category>benchmarking</category></item><item><title>s1: Simple test-time scaling (and Kyutai Hibiki)</title><link>https://news.smol.ai/issues/25-02-06-ainews-s1-simple-test-time-scaling-and-kyutai-hibiki/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-06-ainews-s1-simple-test-time-scaling-and-kyutai-hibiki/</guid><description>**&quot;Wait&quot; is all you need** introduces a novel reasoning model finetuned from **Qwen 2.5 32B** using just **1000 questions with reasoning traces** distilled from **Gemini 2.0 Flash Thinking**, enabling controllable test-time compute by appending &quot;Wait&quot; to extend reasoning. Lead author **Niklas Muennighoff**, known for work on **Bloom**, **StarCoder**, and **BIG-bench**, highlights this method&apos;s efficiency and its reproduction of the famous o1 scaling chart. Additionally, **Kyutai Moshi**&apos;s Hibiki project demonstrates impressive offline French-English live translation on iPhone. Recent AI model releases include **DeepSeek R1 and R3 open source models**, potentially marking a major open-source milestone, **Hugging Face&apos;s SmolLM2** emphasizing data-centric training for small LMs, and **IBM&apos;s Granite-Vision-3.1-2B**, a small vision-language model with strong performance. Key research papers spotlight **LIMO** for minimal demonstration reasoning achieving high accuracy on AIME and MATH benchmarks, and **Token-Assisted Reasoning** mixing latent and text tokens to improve language model reasoning.</description><pubDate>Fri, 07 Feb 2025 03:47:44 GMT</pubDate><category>google-deepmind</category><category>qwen</category><category>gemini</category><category>hugging-face</category><category>ibm</category><category>deepseek</category><category>qwen-2.5-32b</category><category>gemini-2.0-flash</category><category>smollm2</category><category>granite-vision-3.1-2b</category><category>niklas-muennighoff</category><category>reasoning</category><category>fine-tuning</category><category>scaling-laws</category><category>open-source-models</category><category>data-centric-training</category><category>vision</category><category>multilingual-models</category><category>language-model-reasoning</category></item><item><title>Gemini 2.0 Flash GA, with new Flash Lite, 2.0 Pro, and Flash Thinking</title><link>https://news.smol.ai/issues/25-02-05-ainews-gemini-20-flash-ga-with-new-flash-lite-20-pro-and-flash-thinking/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-05-ainews-gemini-20-flash-ga-with-new-flash-lite-20-pro-and-flash-thinking/</guid><description>**Google DeepMind** officially launched **Gemini 2.0** models including **Flash**, **Flash-Lite**, and **Pro Experimental**, with **Gemini 2.0 Flash** outperforming **Gemini 1.5 Pro** while being **12x cheaper** and supporting **multimodal input** and a **1 million token context window**. **Andrej Karpathy** released a **3h31m** video deep dive into **large language models**, covering **pretraining**, **fine-tuning**, and **reinforcement learning** with examples like **GPT-2** and **Llama 3.1**. A free course on **Transformer architecture** was introduced by **Jay Alammar**, **Maarten Gr**, and **Andrew Ng**, focusing on **tokenizers**, **embeddings**, and **mixture-of-expert models**. **DeepSeek-R1** reached **1.2 million downloads** on **Hugging Face** with a detailed **36-page technical report**. **Anthropic** increased rewards to **$10K** and **$20K** for their jailbreak challenge, while **BlueRaven** extension was updated to hide Twitter metrics for unbiased engagement.</description><pubDate>Thu, 06 Feb 2025 02:00:20 GMT</pubDate><category>google-deepmind</category><category>hugging-face</category><category>anthropic</category><category>gemini-2.0-flash</category><category>gemini-2.0-flash-lite</category><category>gemini-2.0-pro-experimental</category><category>gemini-1.5-pro</category><category>deepseek-r1</category><category>gpt-2</category><category>llama-3-1</category><category>andrej-karpathy</category><category>jayalammar</category><category>maartengr</category><category>andrewyng</category><category>nearcyan</category><category>multimodality</category><category>context-windows</category><category>cost-efficiency</category><category>pretraining</category><category>fine-tuning</category><category>reinforcement-learning</category><category>transformer</category><category>tokenization</category><category>embeddings</category><category>mixture-of-experts</category></item><item><title>How To Scale Your Model, by DeepMind</title><link>https://news.smol.ai/issues/25-02-04-ainews-how-to-scale-your-model-by-deepmind/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-04-ainews-how-to-scale-your-model-by-deepmind/</guid><description>**Researchers at Google DeepMind (GDM)** released a comprehensive &quot;little textbook&quot; titled **&quot;How To Scale Your Model&quot;** covering modern Transformer architectures, inference optimizations beyond O(N^2) attention, and high-performance computing concepts like rooflines. The resource includes practical problems and real-time comment engagement. On AI Twitter, several key updates include the open-sourced humanoid robotics model **ASAP** inspired by athletes like **Cristiano Ronaldo**, **LeBron James**, and **Kobe Bryant**; a new paper on **Mixture-of-Agents** proposing the **Self-MoA** method for improved LLM output aggregation; training of reasoning LLMs using the **GRPO algorithm** from **DeepSeek** demonstrated on **Qwen 0.5**; findings on bias in LLMs used as judges highlighting the need for multiple independent evaluations; and the release of **mlx-rs**, a Rust library for machine learning with examples including **Mistral** text generation. Additionally, **Hugging Face** launched an AI app store featuring over **400,000 apps** with 2,000 new daily additions and 2.5 million weekly visits, enabling AI-powered app search and categorization.</description><pubDate>Wed, 05 Feb 2025 06:59:23 GMT</pubDate><category>google-deepmind</category><category>deepseek</category><category>hugging-face</category><category>qwen-0.5</category><category>omarsar0</category><category>drjimfan</category><category>tairanhe99</category><category>guanyashi</category><category>lioronai</category><category>_philschmid</category><category>awnihannun</category><category>clementdelangue</category><category>transformers</category><category>inference</category><category>high-performance-computing</category><category>robotics</category><category>sim2real</category><category>mixture-of-experts</category><category>reinforcement-learning</category><category>bias-mitigation</category><category>rust</category><category>text-generation</category><category>open-source</category></item><item><title>OpenAI takes on Gemini&apos;s Deep Research</title><link>https://news.smol.ai/issues/25-02-03-ainews-openai-takes-on-geminis-deep-research/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-03-ainews-openai-takes-on-geminis-deep-research/</guid><description>**OpenAI** released the full version of the **o3** agent, with a new **Deep Research** variant showing significant improvements on the **HLE benchmark** and achieving SOTA results on **GAIA**. The release includes an &quot;inference time scaling&quot; chart demonstrating rigorous research, though some criticism arose over public test set results. The agent is noted as &quot;extremely simple&quot; and currently limited to 100 queries/month, with plans for a higher-rate version. Reception has been mostly positive, with some skepticism. Additionally, advances in **reinforcement learning** were highlighted, including a simple test-time scaling technique called **budget forcing** that improved reasoning on math competitions by 27%. Researchers from **Google DeepMind**, **NYU**, **UC Berkeley**, and **HKU** contributed to these findings. The original **Gemini Deep Research** team will participate in the upcoming AI Engineer NYC event.</description><pubDate>Tue, 04 Feb 2025 02:44:29 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>nyu</category><category>uc-berkeley</category><category>hku</category><category>o3</category><category>o3-mini-high</category><category>o3-deep-research-mini</category><category>sama</category><category>danhendrycks</category><category>ethan-mollick</category><category>dan-shipper</category><category>reinforcement-learning</category><category>benchmarking</category><category>inference-speed</category><category>model-performance</category><category>reasoning</category><category>test-time-scaling</category><category>agent-design</category></item><item><title>o3-mini launches, OpenAI on &quot;wrong side of history&quot;</title><link>https://news.smol.ai/issues/25-02-01-ainews-o3-mini-launches-openai-on-wrong-side-of-history/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-02-01-ainews-o3-mini-launches-openai-on-wrong-side-of-history/</guid><description>**OpenAI** released **o3-mini**, a new reasoning model available for free and paid users with a &quot;high&quot; reasoning effort option that outperforms the earlier **o1** model on STEM tasks and safety benchmarks, costing **93% less** per token. **Sam Altman** acknowledged a shift in open source strategy and credited **DeepSeek R1** for influencing assumptions. **MistralAI** launched **Mistral Small 3 (24B)**, an open-weight model with competitive performance and low API costs. **DeepSeek R1** is supported by **Text-generation-inference v3.1.0** and available via **ai-gradio** and replicate. The news highlights advancements in reasoning, cost-efficiency, and safety in AI models.</description><pubDate>Sat, 01 Feb 2025 09:16:19 GMT</pubDate><category>openai</category><category>mistral-ai</category><category>deepseek</category><category>togethercompute</category><category>fireworksai_hq</category><category>ai-gradio</category><category>replicate</category><category>o3-mini</category><category>o1</category><category>gpt-4o</category><category>mistral-small-3-24b</category><category>deepseek-r1</category><category>sam-altman</category><category>reasoning</category><category>safety</category><category>cost-efficiency</category><category>model-performance</category><category>benchmarking</category><category>api</category><category>open-weight-models</category><category>model-releases</category></item><item><title>Mistral Small 3 24B and Tulu 3 405B</title><link>https://news.smol.ai/issues/25-01-30-ainews-mistral-small-3-24b-and-tulu-3-405b/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-30-ainews-mistral-small-3-24b-and-tulu-3-405b/</guid><description>**Mistral AI** released **Mistral Small 3**, a **24B parameter** model optimized for local inference with low latency and **81% accuracy on MMLU**, competing with **Llama 3.3 70B**, **Qwen-2.5 32B**, and **GPT4o-mini**. **AI2** released **Tülu 3 405B**, a large finetuned model of **Llama 3** using Reinforcement Learning from Verifiable Rewards (RVLR), competitive with **DeepSeek v3**. **Sakana AI** launched **TinySwallow-1.5B**, a Japanese language model using **TAID** for on-device use. **Alibaba_Qwen** released **Qwen 2.5 Max**, trained on **20 trillion tokens**, with performance comparable to **DeepSeek V3**, **Claude 3.5 Sonnet**, and **Gemini 1.5 Pro**, and updated API pricing. These releases highlight advances in open models, efficient inference, and reinforcement learning techniques.</description><pubDate>Fri, 31 Jan 2025 00:08:47 GMT</pubDate><category>mistral-ai</category><category>ai2</category><category>sakana-ai</category><category>alibaba_qwen</category><category>deepseek</category><category>ollama</category><category>llamaindex</category><category>mistral-small-3</category><category>tulu-3-405b</category><category>llama-3</category><category>tiny-swallow-1.5b</category><category>qwen-2.5-max</category><category>deepseek-v3</category><category>claude-3.5-sonnet</category><category>gemini-1.5-pro</category><category>gpt4o-mini</category><category>llama-3-3-70b</category><category>clementdelangue</category><category>dchaplot</category><category>reach_vb</category><category>reinforcement-learning</category><category>model-fine-tuning</category><category>local-inference</category><category>model-performance</category><category>model-optimization</category><category>on-device-ai</category><category>instruction-following</category><category>api</category><category>training-data</category><category>natural-language-processing</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-29-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-29-ainews-not-much-happened-today/</guid><description>**DeepSeek-R1 and DeepSeek-V3** models have made significant advancements, trained on an **instruction-tuning dataset of 1.5M samples** with **600,000 reasoning** and **200,000 non-reasoning SFT data**. The models demonstrate strong **performance benchmarks** and are deployed on-premise via collaborations with **Dell** and **Hugging Face**. Training costs are estimated around **$5.5M to $6M**, with efficient hardware utilization on **8xH100 servers**. The **International AI Safety Report** highlights risks such as **malicious use**, **malfunctions**, and **systemic risks** including **AI-driven cyberattacks**. Industry leaders like **Yann LeCun** and **Yoshua Bengio** provide insights on market reactions, AI safety, and ethical considerations, with emphasis on AI&apos;s role in creativity and economic incentives.</description><pubDate>Thu, 30 Jan 2025 01:07:40 GMT</pubDate><category>deepseek</category><category>hugging-face</category><category>dell</category><category>openai</category><category>deepseek-r1</category><category>deepseek-v3</category><category>coder-v2</category><category>prover</category><category>yann-lecun</category><category>yoshua-bengio</category><category>francois-chollet</category><category>giffman</category><category>instruction-tuning</category><category>performance-benchmarks</category><category>model-deployment</category><category>training-costs</category><category>hardware-scalability</category><category>ai-safety</category><category>risk-mitigation</category><category>ethical-ai</category><category>open-source</category><category>gpu-utilization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-28-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-28-ainews-not-much-happened-today/</guid><description>**Huawei chips** are highlighted in a diverse AI news roundup covering **NVIDIA&apos;s** stock rebound, new open music foundation models like **Local Suno**, and competitive AI models such as **Qwen 2.5 Max** and **Deepseek V3**. The release of **DeepSeek Janus Pro**, a multimodal LLM with image generation capabilities, and advancements in **reinforcement learning** and **chain-of-thought reasoning** are noted. Discussions include GPU rebranding with **NVIDIA&apos;s H6400 GPUs**, data center innovations, and enterprise AI applications like crypto APIs in hedge funds. *&quot;Deepseek R1&apos;s capabilities&quot;* and *&quot;Qwen 2.5 models added to applications&quot;* are key highlights.</description><pubDate>Wed, 29 Jan 2025 01:48:45 GMT</pubDate><category>nvidia</category><category>anthropic</category><category>openai</category><category>deepseek</category><category>huawei</category><category>vercel</category><category>bespoke-labs</category><category>deepseek-r1</category><category>qwen-2.5</category><category>qwen-2.5-max</category><category>deepseek-v3</category><category>deepseek-janus-pro</category><category>gpt-4</category><category>saranormous</category><category>zizhpan</category><category>victormustar</category><category>omarsar0</category><category>markchen90</category><category>sakanaailabs</category><category>reach_vb</category><category>madiator</category><category>dain_mclau</category><category>francoisfleuret</category><category>garygodchaux</category><category>arankomatsuzaki</category><category>id_aa_carmack</category><category>lavanyasant</category><category>virattt</category><category>model-merging</category><category>multimodality</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>gpu-optimization</category><category>compute-infrastructure</category><category>compression</category><category>crypto-api</category><category>image-generation</category></item><item><title>DeepSeek #1 on US App Store, Nvidia stock tanks -17%</title><link>https://news.smol.ai/issues/25-01-27-ainews-deepseek-1-on-us-app-store-nvidia-stock-tanks-17percent/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-27-ainews-deepseek-1-on-us-app-store-nvidia-stock-tanks-17percent/</guid><description>**DeepSeek** has made a significant cultural impact by hitting mainstream news unexpectedly in 2025. The **DeepSeek-R1** model features a massive **671B parameter MoE architecture** and demonstrates **chain-of-thought (CoT)** capabilities comparable to **OpenAI&apos;s o1** at a lower cost. The **DeepSeek V3** model trains a **236B parameter model 42% faster** than its predecessor using **fp8 precision**. The **Qwen2.5** multimodal models support images and videos with sizes ranging from **3B to 72B parameters**, featuring strong vision and agentic capabilities. **LangChain** and **LangGraph** integration enable AI chatbots with memory and tool use, including applications like the **DeFi Agent**. Discussions highlight **NVIDIA&apos;s** role in hardware acceleration, with concerns about stock drops due to **DeepSeek&apos;s** efficiency and market fears. The compute demand is expected to rise despite efficiency gains, driven by inference scaling and MoE design improvements.</description><pubDate>Tue, 28 Jan 2025 05:28:32 GMT</pubDate><category>deepseek</category><category>openai</category><category>nvidia</category><category>langchain</category><category>deepseek-r1</category><category>deepseek-v3</category><category>qwen2.5-vl</category><category>o1</category><category>sama</category><category>mervenoyann</category><category>omarasar0</category><category>teortaxestex</category><category>nptacek</category><category>carpeetti</category><category>finbarrtimbers</category><category>cwolferesearch</category><category>arthurrapier</category><category>danhendrycks</category><category>scaling01</category><category>janusflow</category><category>moe-architecture</category><category>chain-of-thought</category><category>fp8-precision</category><category>multimodality</category><category>vision</category><category>agentic-ai</category><category>inference-scaling</category><category>gpu-optimization</category><category>model-efficiency</category><category>ai-chatbots</category><category>memory-integration</category><category>tool-use</category><category>stock-market-reactions</category></item><item><title>TinyZero: Reproduce DeepSeek R1-Zero for $30</title><link>https://news.smol.ai/issues/25-01-24-ainews-tinyzero-reproduce-deepseek-r1-zero-for-dollar30/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-24-ainews-tinyzero-reproduce-deepseek-r1-zero-for-dollar30/</guid><description>**DeepSeek Mania** continues to reshape the frontier model landscape with Jiayi Pan from Berkeley reproducing the *OTHER* result from the DeepSeek R1 paper, R1-Zero, in a cost-effective Qwen model fine-tune for two math tasks. A key finding is a lower bound to the distillation effect at **1.5B parameters**, with RLCoT reasoning emerging as an intrinsic property. Various RL techniques like PPO, DeepSeek&apos;s GRPO, or PRIME show similar outcomes, and starting from an Instruct model speeds convergence. The **Humanity’s Last Exam (HLE) Benchmark** introduces a challenging multi-modal test with **3,000 expert-level questions** across **100+ subjects**, where models perform below **10%**, with **DeepSeek-R1** achieving **9.4%**. DeepSeek-R1 excels in chain-of-thought reasoning, outperforming models like **o1** while being **20x cheaper** and MIT licensed. The **WebDev Arena Leaderboard** ranks DeepSeek-R1 #2 in technical domains and #1 under Style Control, closing in on **Claude 3.5 Sonnet**. OpenAI&apos;s **Operator** is deployed to 100% of Pro users in the US, enabling tasks like ordering meals and booking reservations, and functions as a research assistant for AI paper searches and summaries. Hugging Face announces a leadership change after significant growth, while Meta AI releases the first stable version of **Llama Stack** with streamlined upgrades and automated verification. DeepSeek-R1&apos;s open-source success is celebrated, and technical challenges like memory management on macOS 15+ are addressed with residency sets in MLX for stability.</description><pubDate>Sat, 25 Jan 2025 02:32:28 GMT</pubDate><category>deepseek</category><category>berkeley</category><category>hugging-face</category><category>meta-ai-fair</category><category>openai</category><category>deeplearningai</category><category>deepseek-r1</category><category>qwen</category><category>o1</category><category>claude-3-sonnet</category><category>claude-3</category><category>prime</category><category>ppo</category><category>grpo</category><category>llama-stack</category><category>jiayi-pan</category><category>saranormous</category><category>reach_vb</category><category>lmarena_ai</category><category>nearcyan</category><category>omarsar0</category><category>philschmid</category><category>hardmaru</category><category>awnihannun</category><category>winglian</category><category>reinforcement-learning</category><category>fine-tuning</category><category>chain-of-thought</category><category>multi-modal-benchmark</category><category>memory-management</category><category>model-training</category><category>open-source</category><category>agentic-workflow-automation</category><category>model-performance</category></item><item><title>OpenAI launches Operator, its first Agent</title><link>https://news.smol.ai/issues/25-01-23-ainews-openai-launches-operator-its-first-agent/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-23-ainews-openai-launches-operator-its-first-agent/</guid><description>**OpenAI** launched **Operator**, a premium computer-using agent for web tasks like booking and ordering, available now for Pro users in the US with an API promised. It features long horizon remote VMs up to 20 minutes and video export, showing state-of-the-art agent performance but not yet human-level. **Anthropic** had launched a similar agent 3 months earlier as an open source demo. **DeepSeek AI** unveiled **DeepSeek R1**, an open-source reasoning model excelling on the **Humanity&apos;s Last Exam** dataset, outperforming models like **LLaMA 4** and **OpenAI&apos;s o1**. **Google DeepMind** open-sourced **VideoLLaMA 3**, a multimodal foundation model for image and video understanding. **Perplexity AI** released **Perplexity Assistant** for Android with reasoning and search capabilities. The **Humanity&apos;s Last Exam** dataset contains 3,000 questions testing AI reasoning, with current models scoring below 10% accuracy, indicating room for improvement. OpenAI&apos;s Computer-Using Agent (CUA) shows improved performance on OSWorld and WebArena benchmarks but still lags behind humans. **Anthropic AI** introduced Citations for safer AI responses. *Sam Altman* and *Swyx* commented on Operator&apos;s launch and capabilities.</description><pubDate>Fri, 24 Jan 2025 03:34:34 GMT</pubDate><category>openai</category><category>anthropic</category><category>deepseek-ai</category><category>google-deepmind</category><category>perplexity-ai</category><category>operator</category><category>deepseek-r1</category><category>videollama-3</category><category>llama-4</category><category>o1</category><category>claude</category><category>sam-altman</category><category>swyx</category><category>computer-using-agent</category><category>reasoning</category><category>multimodality</category><category>performance-benchmarks</category><category>open-source</category><category>ai-safety</category><category>benchmarking</category><category>video-generation</category><category>model-evaluation</category></item><item><title>Bespoke-Stratos + Sky-T1: The Vicuna+Alpaca moment for reasoning</title><link>https://news.smol.ai/issues/25-01-22-ainews-bespoke-stratos-sky-t1-the-vicunaalpaca-moment-for-reasoning/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-22-ainews-bespoke-stratos-sky-t1-the-vicunaalpaca-moment-for-reasoning/</guid><description>**Reasoning Distillation** has emerged as a key technique, with Berkeley/USC researchers releasing **Sky-T1-32B-Preview**, a finetuned model of **Qwen 2.5 32B** using 17k reasoning traces for just **$450**, matching benchmarks of **o1-preview**. **DeepSeek** introduced **R1**, a model surpassing **o1-preview** and enabling distillation to smaller models like a 1.5B Qwen to match **gpt-4o** and **claude-3-sonnet** levels. **Bespoke Labs** further distilled **R1** on Qwen, outperforming **o1-preview** with fewer samples. This progress suggests that *&quot;SFT is all you need&quot;* for reasoning without major architecture changes. Additionally, **DeepSeek-R1** uses pure reinforcement learning with supervised finetuning to accelerate convergence and shows strong reasoning and multimodal capabilities. **Google&apos;s Gemini 2.0 Flash Thinking** model boasts a **1 million token context window**, code execution, and excels in math, science, and multimodal reasoning. Critiques highlight challenges in model repeatability, behavioral self-awareness, and RLHF limitations in reasoning robustness.</description><pubDate>Thu, 23 Jan 2025 07:08:27 GMT</pubDate><category>berkeley</category><category>usc</category><category>deepseek</category><category>bespoke-labs</category><category>google</category><category>llmsys</category><category>stanford</category><category>lm-sys</category><category>sky-t1-32b-preview</category><category>qwen-2.5-32b</category><category>r1</category><category>o1-preview</category><category>gpt-4o</category><category>claude-3-sonnet</category><category>bespoke-stratos-32b</category><category>gemini-2.0-flash-thinking</category><category>teortaxestex</category><category>cwolferesearch</category><category>madiator</category><category>chakraai</category><category>philschmid</category><category>abacaj</category><category>omarsar0</category><category>reasoning</category><category>supervised-finetuning</category><category>reinforcement-learning</category><category>multimodality</category><category>model-distillation</category><category>context-windows</category><category>code-execution</category><category>model-repeatability</category><category>behavioral-self-awareness</category><category>rlhf</category></item><item><title>Project Stargate: $500b datacenter (1.7% of US GDP) and Gemini 2 Flash Thinking 2</title><link>https://news.smol.ai/issues/25-01-21-ainews-project-stargate-dollar500b-datacenter-17percent-of-us-gdp-and-gemini-2-flash-thinking-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-21-ainews-project-stargate-dollar500b-datacenter-17percent-of-us-gdp-and-gemini-2-flash-thinking-2/</guid><description>**Project Stargate**, a US &quot;AI Manhattan project&quot; led by **OpenAI** and **Softbank**, supported by **Oracle**, **Arm**, **Microsoft**, and **NVIDIA**, was announced with a scale comparable to the original Manhattan project costing **$35B inflation adjusted**. Despite Microsoft&apos;s reduced role as exclusive compute partner, the project is serious but not immediately practical. Meanwhile, **Noam Shazeer** revealed a second major update to **Gemini 2.0 Flash Thinking**, enabling **1M token long context** usable immediately. Additionally, **AI Studio** introduced a new **code interpreter** feature. On Reddit, **DeepSeek R1**, a distillation of **Qwen 32B**, was released for free on **HuggingChat**, sparking discussions on self-hosting, performance issues, and quantization techniques. DeepSeek&apos;s CEO **Liang Wenfeng** highlighted their focus on **fundamental AGI research**, efficient **MLA architecture**, and commitment to **open-source development** despite export restrictions, positioning DeepSeek as a potential alternative to closed-source AI trends.</description><pubDate>Wed, 22 Jan 2025 01:56:21 GMT</pubDate><category>openai</category><category>softbank</category><category>oracle</category><category>arm</category><category>microsoft</category><category>nvidia</category><category>huggingface</category><category>deepseek-ai</category><category>gemini-2.0-flash</category><category>deepseek-r1</category><category>qwen-32b</category><category>noam-shazeer</category><category>liang-wenfeng</category><category>long-context</category><category>quantization</category><category>code-interpretation</category><category>model-distillation</category><category>open-source</category><category>agi-research</category><category>model-performance</category><category>memory-optimization</category></item><item><title>DeepSeek R1: o1-level open weights model and a simple recipe for upgrading 1.5B models to Sonnet/4o level</title><link>https://news.smol.ai/issues/25-01-20-ainews-deepseek-r1-o1-level-open-weights-model-and-a-simple-recipe-for-upgrading-15b-models-to-sonnet4o-level/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-20-ainews-deepseek-r1-o1-level-open-weights-model-and-a-simple-recipe-for-upgrading-15b-models-to-sonnet4o-level/</guid><description>**DeepSeek** released **DeepSeek R1**, a significant upgrade over **DeepSeek V3** from just three weeks prior, featuring 8 models including full-size 671B MoE models and multiple distillations from **Qwen 2.5** and **Llama 3.1/3.3**. The models are MIT licensed, allowing finetuning and distillation. Pricing is notably cheaper than **o1** by 27x-50x. The training process used **GRPO** (reward for correctness and style outcomes) without relying on PRM, MCTS, or reward models, focusing on reasoning improvements through reinforcement learning. Distilled models can run on **Ollama** and show strong capabilities like writing **Manim code**. The release emphasizes advances in **reinforcement-learning**, **fine-tuning**, and **model-distillation** with a novel RL framework from DeepSeekMath.</description><pubDate>Tue, 21 Jan 2025 07:50:24 GMT</pubDate><category>deepseek</category><category>ollama</category><category>qwen</category><category>llama</category><category>deepseek-r1</category><category>deepseek-v3</category><category>qwen-2.5</category><category>llama-3.1</category><category>llama-3.3-70b</category><category>reinforcement-learning</category><category>fine-tuning</category><category>model-distillation</category><category>model-optimization</category><category>reasoning</category><category>reward-models</category><category>multi-response-sampling</category><category>model-training</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-17-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-17-ainews-not-much-happened-today/</guid><description>**DeepSeek-V3**, a **671 billion parameter mixture-of-experts model**, surpasses **Llama 3.1 405B** and **GPT-4o** in coding and math benchmarks. **OpenAI** announced the upcoming release of **GPT-5** on **April 27, 2023**. **MiniMax-01 Coder mode** in **ai-gradio** enables building a chess game in one shot. **Meta** research highlights trade-offs in scaling visual tokenizers. **Google DeepMind** improves diffusion model quality via inference-time scaling. The **RA-DIT** method fine-tunes LLMs and retrievers for better RAG responses. The U.S. proposes a three-tier export restriction system on AI chips and models, excluding countries like **China** and **Russia**. Security vulnerabilities in AI chatbots involving CSRF and prompt injection were revealed. Concerns about superintelligence and weapons-grade AI models were expressed. **ai-gradio** updates include NVIDIA NIM compatibility and new models like **cosmos-nemotron-34b**. **LangChain** integrates with **Claude-3-haiku** for AI agents with persistent memory. **Triton Warp specialization** optimizes GPU usage for matrix multiplication. **Meta&apos;s** fine-tuned **Llama** models, **OpenBioLLM-8B** and **OpenBioLLM-70B**, target personalized medicine and clinical trials.</description><pubDate>Sat, 18 Jan 2025 02:33:34 GMT</pubDate><category>openai</category><category>deep-learning-ai</category><category>meta-ai-fair</category><category>google-deepmind</category><category>saama</category><category>langchain</category><category>nvidia</category><category>deepseek-v3</category><category>llama-3-1-405b</category><category>gpt-4o</category><category>gpt-5</category><category>minimax-01</category><category>claude-3-haiku</category><category>cosmos-nemotron-34b</category><category>akhaliq</category><category>mixture-of-experts</category><category>coding</category><category>math</category><category>scaling</category><category>visual-tokenizers</category><category>diffusion-models</category><category>inference-time-scaling</category><category>retrieval-augmented-generation</category><category>ai-export-restrictions</category><category>security-vulnerabilities</category><category>prompt-injection</category><category>gpu-optimization</category><category>fine-tuning</category><category>personalized-medicine</category><category>clinical-trials</category><category>ai-agents</category><category>persistent-memory</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-16-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-16-ainews-not-much-happened-today/</guid><description>**Harvey** secured a new **$300M funding round**. **OuteTTS 0.3 1B &amp; 500M** text-to-speech models were released featuring **zero-shot voice cloning**, **multilingual support** (en, jp, ko, zh, fr, de), and **emotion control**, powered by **OLMo-1B** and **Qwen 2.5 0.5B**. The **HOVER** model, a **1.5M-parameter neural net** for **agile motor control**, was introduced, leveraging **human motion capture datasets** and **massively parallel reinforcement learning**. **kokoro.js** enables running AI models locally in browsers with minimal dependencies. **Meta AI** awarded **$200K LLM evaluation grants** for projects on **regional language understanding**, **complex reasoning**, and **interactive programming environments**. **Stability AI&apos;s Twitter account was hacked**, prompting security warnings. **Alibaba Qwen** improved **Process Reward Models (PRMs)** for better **mathematical reasoning** using a **consensus filtering mechanism**. **DeepSeek V3** uses **pipeline parallelism** to enhance **distributed inference** and **long-context generation efficiency**. Discussions on **AI policy in legal frameworks** and **AI&apos;s role in democratizing education** were highlighted. Lighthearted AI-related humor was also shared.</description><pubDate>Fri, 17 Jan 2025 06:04:28 GMT</pubDate><category>harvey</category><category>meta-ai-fair</category><category>stability-ai</category><category>alibaba</category><category>deepseek</category><category>hugging-face</category><category>oute-tts-0.3-1b</category><category>oute-tts-0.3-500m</category><category>olm-1b</category><category>qwen-2.5-0.5b</category><category>hover</category><category>gpt-4o</category><category>deepseek-v3</category><category>reach_vb</category><category>drjimfan</category><category>vikhyatk</category><category>mervenoyann</category><category>aiatmeta</category><category>iscienceluvr</category><category>alibaba_qwen</category><category>awnihannun</category><category>ajeya_cotra</category><category>emollick</category><category>qtnx_</category><category>designerx</category><category>text-to-speech</category><category>zero-shot-learning</category><category>multilinguality</category><category>emotion-control</category><category>motor-control</category><category>reinforcement-learning</category><category>local-ai</category><category>distributed-inference</category><category>pipeline-parallelism</category><category>mathematical-reasoning</category><category>process-reward-models</category><category>legal-ai</category><category>education-ai</category><category>ai-security</category><category>humor</category></item><item><title>Titans: Learning to Memorize at Test Time</title><link>https://news.smol.ai/issues/25-01-15-ainews-titans-learning-to-memorize-at-test-time/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-15-ainews-titans-learning-to-memorize-at-test-time/</guid><description>**Google** released a new paper on &quot;Neural Memory&quot; integrating persistent memory directly into transformer architectures at test time, showing promising long-context utilization. **MiniMax-01** by @omarsar0 features a **4 million token context window** with **456B parameters** and **32 experts**, outperforming **GPT-4o** and **Claude-3.5-Sonnet**. **InternLM3-8B-Instruct** is an open-source model trained on **4 trillion tokens** with state-of-the-art results. **Transformer²** introduces self-adaptive LLMs that dynamically adjust weights for continuous adaptation. Advances in AI security highlight the need for **agent authentication**, **prompt injection** defenses, and **zero-trust architectures**. Tools like **Micro Diffusion** enable budget-friendly diffusion model training, while **LeagueGraph** and **Agent Recipes** support open-source social media agents.</description><pubDate>Thu, 16 Jan 2025 07:58:41 GMT</pubDate><category>google</category><category>meta-ai-fair</category><category>openai</category><category>anthropic</category><category>langchain</category><category>minimax-01</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>internlm3-8b-instruct</category><category>transformer2</category><category>omarsar0</category><category>hwchase17</category><category>abacaj</category><category>hardmaru</category><category>rez0__</category><category>bindureddy</category><category>akhaliq</category><category>saranormous</category><category>long-context</category><category>mixture-of-experts</category><category>self-adaptive-models</category><category>prompt-injection</category><category>agent-authentication</category><category>diffusion-models</category><category>zero-trust-architecture</category><category>continuous-adaptation</category><category>vision</category><category>agentic-systems</category></item><item><title>small little news items</title><link>https://news.smol.ai/issues/25-01-14-ainews-small-little-news-items/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-14-ainews-small-little-news-items/</guid><description>**Ollama** enhanced its models by integrating **Cohere&apos;s R7B**, optimized for **RAG** and **tool use tasks**, and released **Ollama v0.5.5** with quality updates and a new engine. **Together AI** launched the **Llama 3.3 70B multimodal model** with improved reasoning and math capabilities, while **OpenBMB** introduced the **MiniCPM-o 2.6**, outperforming **GPT-4V** on visual tasks. Insights into **Process Reward Models (PRM)** were shared to boost **LLM reasoning**, alongside **Qwen2.5-Math-PRM** models excelling in mathematical reasoning. **LangChain** released a beta for **ChatGPT Tasks** enabling scheduling of reminders and summaries, and introduced open-source **ambient agents** for email assistance. **OpenAI** rolled out **Tasks** for scheduling actions in **ChatGPT** for Plus, Pro, and Teams users. AI software engineering is rapidly advancing, predicted to match human capabilities within 18 months. Research on **LLM scaling laws** highlights power law relationships and plateauing improvements, while **GANs** are experiencing a revival.</description><pubDate>Wed, 15 Jan 2025 02:19:30 GMT</pubDate><category>ollama</category><category>cohere</category><category>togethercompute</category><category>openbmb</category><category>qwen</category><category>langchain</category><category>openai</category><category>r7b</category><category>llama-3-70b</category><category>minicpm-o-2.6</category><category>gpt-4v</category><category>qwen2.5-math-prm</category><category>rag</category><category>tool-use-tasks</category><category>quality-of-life</category><category>new-engine</category><category>multimodality</category><category>improved-reasoning</category><category>math-capabilities</category><category>process-reward-models</category><category>llm-reasoning</category><category>mathematical-reasoning</category><category>beta-release</category><category>task-scheduling</category><category>ambient-agents</category><category>email-assistants</category><category>ai-software-engineering</category><category>codebase-analysis</category><category>test-case-generation</category><category>security-infrastructure</category><category>llm-scaling-laws</category><category>power-law</category><category>plateauing-improvements</category><category>gans-revival</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-13-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-13-ainews-not-much-happened-today/</guid><description>**Helium-1 Preview** by **kyutai_labs** is a **2B-parameter multilingual base LLM** outperforming **Qwen 2.5**, trained on **2.5T tokens** with a **4096 context size** using token-level distillation from a **7B model**. **Phi-4 (4-bit)** was released in **lmstudio** on an **M4 max**, noted for speed and performance. **Sky-T1-32B-Preview** is a **$450 open-source reasoning model** matching **o1&apos;s performance** with strong benchmark scores. **Codestral 25.01** by **mistralai** is a new SOTA coding model supporting **80+ programming languages** and offering **2x speed**. 

Innovations include **AutoRAG** for optimizing retrieval-augmented generation pipelines, **Agentic RAG** for autonomous query reformulation and critique, **Multiagent Finetuning** using societies of models like **Phi-3**, **Mistral**, **LLaMA-3**, and **GPT-3.5** for reasoning improvements, and **VideoRAG** incorporating video content into RAG with LVLMs. 

Applications include a dynamic UI AI chat app by **skirano** on **Replit**, **LangChain** tools like **DocTalk** for voice PDF conversations, AI travel agent tutorials, and news summarization agents. **Hyperbolic Labs** offers competitive GPU rentals including **H100**, **A100**, and **RTX 4090**. **LLMQuoter** enhances RAG accuracy by identifying key quotes. 

Infrastructure updates include **MLX export** for LLM inference from Python to C++ by **fchollet** and **SemHash** semantic text deduplication by **philschmid**.</description><pubDate>Tue, 14 Jan 2025 06:08:22 GMT</pubDate><category>kyutai-labs</category><category>lmstudio</category><category>mistralai</category><category>llamaindex</category><category>huggingface</category><category>langchainai</category><category>hyperbolic-labs</category><category>replit</category><category>fchollet</category><category>philschmid</category><category>helium-1</category><category>qwen-2.5</category><category>phi-4</category><category>sky-t1-32b-preview</category><category>o1</category><category>codestral-25.01</category><category>phi-3</category><category>mistral</category><category>llama-3</category><category>gpt-3.5</category><category>llama-3</category><category>gpt-3.5</category><category>llmquoter</category><category>reach_vb</category><category>awnihannun</category><category>lior_on_ai</category><category>sophiamyang</category><category>omarsar0</category><category>skirano</category><category>yuchenj_uw</category><category>fchollet</category><category>philschmid</category><category>multilinguality</category><category>token-level-distillation</category><category>context-windows</category><category>model-performance</category><category>open-source</category><category>reasoning</category><category>coding</category><category>retrieval-augmented-generation</category><category>hybrid-retrieval</category><category>multiagent-systems</category><category>video</category><category>large-video-language-models</category><category>dynamic-ui</category><category>voice-interaction</category><category>gpu-rentals</category><category>model-optimization</category><category>semantic-deduplication</category><category>model-inference</category></item><item><title>Moondream 2025.1.9: Structured Text, Enhanced OCR, Gaze Detection in a 2B Model</title><link>https://news.smol.ai/issues/25-01-10-ainews-moondream-202519-structured-text-enhanced-ocr-gaze-detection-in-a-2b-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-10-ainews-moondream-202519-structured-text-enhanced-ocr-gaze-detection-in-a-2b-model/</guid><description>**Moondream** has released a new version that advances VRAM efficiency and adds structured output and gaze detection, marking a new frontier in vision model practicality. Discussions on Twitter highlighted advancements in reasoning models like **OpenAI&apos;s o1**, model distillation techniques, and new multimodal embedding models such as **vdr-2b-multi-v1** and **LLaVA-Mini**, which significantly reduce computational costs. Research on GANs and decentralized diffusion models showed improved stability and performance. Development tools like **MLX** and **vLLM** received updates for better portability and developer experience, while frameworks like **LangChain** and **Qdrant** enable intelligent data workflows. Company updates include new roles and team expansions at **GenmoAI**. *&quot;Efficiency tricks are all you need.&quot;*</description><pubDate>Sat, 11 Jan 2025 07:18:42 GMT</pubDate><category>openai</category><category>llamaindex</category><category>langchainai</category><category>qdrant</category><category>genmoai</category><category>o1</category><category>vdr-2b-multi-v1</category><category>llava-mini</category><category>philschmid</category><category>saranormous</category><category>jxmnop</category><category>reach_vb</category><category>iscienceluvr</category><category>multimodalart</category><category>arohan</category><category>adcock_brett</category><category>awnihannun</category><category>russelljkaplan</category><category>ajayj_</category><category>vision</category><category>model-efficiency</category><category>structured-output</category><category>gaze-detection</category><category>reasoning</category><category>model-distillation</category><category>multimodality</category><category>embedding-models</category><category>gan</category><category>diffusion-models</category><category>self-attention</category><category>training-optimizations</category><category>development-frameworks</category><category>api</category><category>cross-language-deployment</category><category>semantic-search</category><category>agentic-document-processing</category><category>developer-experience</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-09-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-09-ainews-not-much-happened-today/</guid><description>**rStar-Math** surpasses **OpenAI&apos;s o1-preview** in math reasoning with **90.0% accuracy** using a **7B LLM** and **MCTS** with a **Process Reward Model**. **Alibaba** launches **Qwen Chat** featuring **Qwen2.5-Plus** and **Qwen2.5-Coder-32B-Instruct** models enhancing vision-language and reasoning. **Microsoft** releases **Phi-4**, trained on **40% synthetic data** with improved pretraining. **Cohere** introduces **North**, a secure AI workspace integrating **LLMs**, **RAG**, and automation for private deployments. **LangChain** showcases a company research agent with multi-step workflows and open-source datasets. **Transformers.js** demos released for text embeddings and image segmentation in JavaScript. Research highlights include **Meta Meta-CoT** for enhanced chain-of-thought reasoning, **DeepSeek V3** with recursive self-improvement, and collaborative AI development platforms. Industry partnerships include **Rakuten** with **LangChain**, **North** with **RBC** supporting 90,000 employees, and **Agent Laboratory** collaborating with **AMD** and **Johns Hopkins**. Technical discussions emphasize **CUDA** and **Triton** for AI efficiency and evolving AI-assisted coding stacks by **Andrew Ng**.</description><pubDate>Fri, 10 Jan 2025 03:35:37 GMT</pubDate><category>openai</category><category>anthropic</category><category>alibaba</category><category>microsoft</category><category>cohere</category><category>langchain</category><category>weights-biases</category><category>deepseek</category><category>rakuten</category><category>rbc</category><category>amd</category><category>johns-hopkins</category><category>rstar-math</category><category>o1-preview</category><category>qwen2.5-plus</category><category>qwen2.5-coder-32b-instruct</category><category>phi-4</category><category>claude-3.5-sonnet</category><category>reach_vb</category><category>rasbt</category><category>akshaykagrawal</category><category>arankomatsuzaki</category><category>teortaxestex</category><category>aidangomez</category><category>andrewyng</category><category>math</category><category>process-reward-model</category><category>mcts</category><category>vision</category><category>reasoning</category><category>synthetic-data</category><category>pretraining</category><category>rag</category><category>automation</category><category>private-deployment</category><category>multi-step-workflow</category><category>open-source-dataset</category><category>text-embeddings</category><category>image-segmentation</category><category>chain-of-thought</category><category>multimodal-reasoning</category><category>finetuning</category><category>recursive-self-improvement</category><category>collaborative-platforms</category><category>ai-development</category><category>partnerships</category><category>cuda</category><category>triton</category><category>ai-efficiency</category><category>ai-assisted-coding</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-08-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-08-ainews-not-much-happened-today/</guid><description>**Sebastien Bubeck** introduced **REINFORCE++**, enhancing classical REINFORCE with **PPO-inspired techniques** for **30% faster training**. **AI21 Labs** released **Phi-4** under the **MIT License**, accessible via **Ollama**. **François Chollet** announced plans for **ARC-AGI-2** and a next-generation **AGI benchmark**. **LangChain** launched **10 new integration packages** to boost **LLM application development**. **Tom Doerr** introduced **Ollama-OCR**, a Python package for **text extraction** using **vision language models**. **Arohan** optimized **Shampoo** for **memory efficiency**, reducing usage from **20 to 6 bytes per parameter**. **Bindu Reddy** showcased **CodeLLM&apos;s v1** for **frontend code generation** and highlighted **LlamaIndex Workflows** for **academic summarization** and **slide generation**. **Hwchase17** collaborated with **Together Compute** to enhance **WebDev Arena** with **complex coding agents** for **LLM coding evaluations**. **Jonathan Ross** detailed **Groq&apos;s** mission to reduce **compute costs by 1000x** amid rising **generative AI** spending. **Clement Delangue** warned about **scam alerts** involving false claims of association with **AI21**. **Vikhyat K** raised concerns about the **ethical implications** and **trade-offs** of **AGI**. Memes and humor included creative AI prompts and critiques of **LLM behaviors**.</description><pubDate>Thu, 09 Jan 2025 03:45:48 GMT</pubDate><category>ai21-labs</category><category>ollama</category><category>langchain</category><category>togethercompute</category><category>groq</category><category>phi-4</category><category>reinforce++</category><category>arc-agi-2</category><category>sebastien-bubeck</category><category>fchollet</category><category>tom-doerr</category><category>arohan_</category><category>bindureddy</category><category>hwchase17</category><category>jonathanross321</category><category>clementdelangue</category><category>vikhyatk</category><category>reinforcement-learning</category><category>ppo</category><category>model-optimization</category><category>memory-efficiency</category><category>python-packages</category><category>vision</category><category>text-extraction</category><category>frontend-code-generation</category><category>workflow-automation</category><category>coding-agents</category><category>compute-cost-reduction</category><category>ethical-ai</category><category>agi-benchmarks</category><category>scam-alerts</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-07-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-07-ainews-not-much-happened-today/</guid><description>**NVIDIA** has launched **Cosmos**, an open-source video world model trained on **20 million hours of video**, aimed at advancing **robotics** and **autonomous driving**. The release sparked debate over its open-source status and technical approach. Additionally, **NVIDIA** announced **Digits**, a **$3,000** personal AI supercomputer designed to democratize AI computing. The AI community expresses mixed feelings about rapid AI progress, with concerns about **AGI**, job displacement, and investment hype. Discussions also highlight upcoming tools for fine-tuning AI models at home and foundation models for AI robotics.</description><pubDate>Wed, 08 Jan 2025 04:01:51 GMT</pubDate><category>nvidia</category><category>openai</category><category>cosmos</category><category>sama</category><category>robotics</category><category>autonomous-driving</category><category>open-source</category><category>fine-tuning</category><category>foundation-models</category><category>memory-optimization</category></item><item><title>PRIME: Process Reinforcement through Implicit Rewards</title><link>https://news.smol.ai/issues/25-01-06-ainews-prime-process-reinforcement-through-implicit-rewards/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-06-ainews-prime-process-reinforcement-through-implicit-rewards/</guid><description>**Implicit Process Reward Models (PRIME)** have been highlighted as a significant advancement in online reinforcement learning, trained on a **7B model** with impressive results compared to **gpt-4o**. The approach builds on the importance of process reward models established by &quot;Let&apos;s Verify Step By Step.&quot; Additionally, AI Twitter discussions cover topics such as **proto-AGI** capabilities with **claude-3.5-sonnet**, the role of **compute scaling** for **Artificial Superintelligence (ASI)**, and model performance nuances. New AI tools like **Gemini 2.0 coder mode** and **LangGraph Studio** enhance agent architecture and software development. Industry events include the **LangChain AI Agent Conference** and meetups fostering AI community connections. Company updates reveal **OpenAI&apos;s** financial challenges with Pro subscriptions and **DeepSeek-V3&apos;s** integration with **Together AI** APIs, showcasing efficient **671B MoE parameter** models. Research discussions focus on **scaling laws** and compute efficiency in large language models.</description><pubDate>Tue, 07 Jan 2025 02:33:39 GMT</pubDate><category>openai</category><category>together-ai</category><category>deepseek</category><category>langchain</category><category>lucidrains</category><category>claude-3.5-sonnet</category><category>gpt-4o</category><category>deepseek-v3</category><category>gemini-2.0</category><category>sama</category><category>aidan_mclau</category><category>omarsar0</category><category>akhaliq</category><category>hwchase17</category><category>tom_doerr</category><category>lmarena_ai</category><category>cwolferesearch</category><category>richardmcngo</category><category>reinforcement-learning</category><category>scaling-laws</category><category>model-performance</category><category>agent-architecture</category><category>software-development</category><category>compute-scaling</category><category>multi-expert-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/25-01-03-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/25-01-03-ainews-not-much-happened-today/</guid><description>**Olmo 2** released a detailed tech report showcasing full pre, mid, and post-training details for a frontier fully open model. **PRIME**, an open-source reasoning solution, achieved **26.7% pass@1**, surpassing **GPT-4o** in benchmarks. Performance improvements include **Qwen 32B (4-bit)** generating at **&gt;40 tokens/sec** on an **M4 Max** and **libvips** being **25x faster** than **Pillow** for image resizing. New tools like **Swaggo/swag** for Swagger 2.0 documentation, **Jujutsu (jj)** Git-compatible VCS, and **Portspoof** security tool were introduced. Robotics advances include a weapon detection system with a meters-wide field of view and faster frame rates. Hardware benchmarks compared **H100** and **MI300x** accelerators. Applications span medical error detection using PRIME and a financial AI agent integrating **LangChainAI** and **Vercel AI SDK**. Architectural insights suggest the need for breakthroughs similar to **SSMs** or **RNNs**.</description><pubDate>Sat, 04 Jan 2025 07:58:51 GMT</pubDate><category>olmo</category><category>openai</category><category>qwen</category><category>cerebras-systems</category><category>langchain</category><category>vercel</category><category>swaggo</category><category>gin</category><category>echo</category><category>prime</category><category>gpt-4o</category><category>qwen-32b</category><category>akhaliq</category><category>jason-wei</category><category>vikhyatk</category><category>awnihannun</category><category>arohan</category><category>tom-doerr</category><category>hendrikbgr</category><category>jerryjliu0</category><category>adcock-brett</category><category>shuchaobi</category><category>stasbekman</category><category>reach-vb</category><category>virattt</category><category>andrew-n-carr</category><category>reasoning</category><category>chain-of-thought</category><category>math</category><category>coding</category><category>optimization</category><category>performance</category><category>image-processing</category><category>software-development</category><category>agent-frameworks</category><category>version-control</category><category>security</category><category>robotics</category><category>hardware-optimization</category><category>medical-ai</category><category>financial-ai</category><category>architecture</category></item><item><title>not much happened to end the year</title><link>https://news.smol.ai/issues/24-12-31-ainews-not-much-happened-to-end-the-year/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-31-ainews-not-much-happened-to-end-the-year/</guid><description>**Reinforcement Fine-Tuning (RFT)** is introduced as a **data-efficient** method to improve **reasoning in LLMs** using minimal **training data** with strategies like **First-Correct Solutions (FCS)** and **Greedily Diverse Solutions (GDS)**. **DeepSeek-V3**, a **671B parameter MoE language model** trained on **14.8 trillion tokens** with **FP8 mixed precision training**, highlights advances in large-scale models and open-source LLMs. Predictions for **AI in 2025** include growth in **smaller models**, **multimodality**, and challenges in **open-source AI**. The impact of AI on software development jobs suggests a need for **higher intelligence** and **specialization** as AI automates low-skilled tasks. Enhancements to **CodeLLM** improve coding assistance with features like **in-place editing** and **streaming responses**. **Natural Language Reinforcement Learning (NLRL)** offers better interpretability and richer feedback for AI planning and critique. AI hiring is growing rapidly with startups seeking strong engineers in **ML** and **systems**. New AI-powered tools such as **Rivet**, **Buzee**, and **Konfig** improve real-time applications, search, and SDK generation using technologies like **Rust** and **V8 isolates**.</description><pubDate>Tue, 31 Dec 2024 23:55:07 GMT</pubDate><category>deepseek</category><category>smol-ai</category><category>deepseek-v3</category><category>code-llm</category><category>o1</category><category>sonnet-3.5</category><category>corbtt</category><category>tom_doerr</category><category>cognitivecompai</category><category>alexalbert__</category><category>theturingpost</category><category>svpino</category><category>bindureddy</category><category>reinforcement-learning</category><category>reasoning</category><category>training-data</category><category>mixed-precision-training</category><category>open-source</category><category>multimodality</category><category>software-development</category><category>natural-language-processing</category><category>interpretability</category><category>developer-tools</category><category>real-time-applications</category><category>search</category><category>sdk-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-12-30-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-30-ainews-not-much-happened-today/</guid><description>**Sam Altman** publicly criticizes **DeepSeek** and **Qwen** models, sparking debate about **OpenAI**&apos;s innovation claims and reliance on foundational research like the **Transformer architecture**. **Deepseek V3** shows significant overfitting issues in the **Misguided Attention** evaluation, solving only **22%** of test prompts, raising concerns about its reasoning and finetuning. Despite skepticism about its open-source status, **Deepseek V3** is claimed to surpass **ChatGPT4** as an open-source model, marking a milestone 1.75 years after ChatGPT4&apos;s release on **March 14, 2023**. The discussions highlight competitive dynamics in AI model performance and innovation sustainability.</description><pubDate>Tue, 31 Dec 2024 02:24:45 GMT</pubDate><category>openai</category><category>deepseek</category><category>google</category><category>qwen</category><category>deepseek-v3</category><category>chatgpt-4</category><category>sam-altman</category><category>overfitting</category><category>reasoning</category><category>misguided-attention</category><category>model-evaluation</category><category>model-architecture</category><category>finetuning</category><category>open-source</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-12-27-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-27-ainews-not-much-happened-today/</guid><description>**ChatGPT**, **Sora**, and the **OpenAI API** experienced a &gt;5 hour outage but are now restored. Updates to **vLLM** enable **DeepSeek-V3** to run with enhanced **parallelism** and **CPU offloading**, improving **model deployment flexibility**. Discussions on **gradient descent** in **top-k routing MoE** and adoption of **FP8 precision** focus on **training efficiency** and **memory optimization**. **AIDE**, an **AI voice medical assistant** by **Team Therasync**, leverages **Qdrant**, **OpenAI**, and **Twilio**. **DeepSeek-Engineer** offers AI-powered coding assistance with structured outputs. **LlamaIndex** integrates **LlamaCloud** and **ElevenLabs** for large-scale **document processing** and voice interaction. Insights on **version control** with **ghstack** and advocacy for **linear decay learning rate schedules** highlight best practices in AI development. Experts predict **smaller, tighter models**, **true multimodal models**, and **on-device AI** in 2025. Proposals for **planetary-scale federated learning** and community AGI moonshots emphasize future AI directions. Discussions on **agentic systems**, **multi-agent workflows**, and **deliberative alignment** through **chain of thought reasoning** underscore AI safety and alignment efforts.</description><pubDate>Sat, 28 Dec 2024 05:06:02 GMT</pubDate><category>openai</category><category>deepseek</category><category>qdrant</category><category>twilio</category><category>llamaindex</category><category>elevenlabs</category><category>vllm</category><category>deepseek-v3</category><category>llamaindex</category><category>francois-fleuret</category><category>daniel-hanchen</category><category>aaron-defazio</category><category>fchollet</category><category>elad-gil</category><category>wojciech-zaremba</category><category>richard-socher</category><category>training-efficiency</category><category>parallelism</category><category>cpu-offloading</category><category>gradient-descent</category><category>mixture-of-experts</category><category>fp8-precision</category><category>memory-optimization</category><category>ai-voice-assistants</category><category>coding-assistants</category><category>document-processing</category><category>version-control</category><category>learning-rate-schedules</category><category>federated-learning</category><category>agentic-systems</category><category>multi-agent-systems</category><category>deliberative-alignment</category><category>chain-of-thought</category><category>on-device-ai</category><category>multimodality</category></item><item><title>DeepSeek v3: 671B finegrained MoE trained for $5.5m USD of compute on 15T tokens</title><link>https://news.smol.ai/issues/24-12-26-ainews-deepseek-v3-671b-finegrained-moe-trained-for-dollar55m-usd-of-compute-on-15t-tokens/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-26-ainews-deepseek-v3-671b-finegrained-moe-trained-for-dollar55m-usd-of-compute-on-15t-tokens/</guid><description>**DeepSeek-V3** has launched with **671B MoE parameters** and trained on **14.8T tokens**, outperforming **GPT-4o** and **Claude-3.5-sonnet** in benchmarks. It was trained with only **2.788M H800 GPU hours**, significantly less than **Llama-3**&apos;s **30.8M GPU-hours**, showcasing major compute efficiency and cost reduction. The model is open-source and deployed via **Hugging Face** with API support. Innovations include native FP8 mixed precision training, Multi-Head Latent Attention scaling, distillation from synthetic reasoning data, pruning and healing for MoEs with up to **256 experts**, and a new multi-token prediction objective enabling lookahead token planning. Research highlights also cover the **OREO method** and **Natural Language Reinforcement Learning (NLRL)** for multi-step reasoning and agent control.</description><pubDate>Fri, 27 Dec 2024 01:18:46 GMT</pubDate><category>deepseek-ai</category><category>hugging-face</category><category>openai</category><category>anthropic</category><category>deepseek-v3</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>llama-3</category><category>nrehiew_</category><category>denny_zhou</category><category>mixture-of-experts</category><category>model-training</category><category>model-optimization</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>multi-token-prediction</category><category>synthetic-data</category><category>model-distillation</category><category>fine-tuning</category><category>attention-mechanisms</category><category>gpu-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-12-24-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-24-ainews-not-much-happened-today/</guid><description>The **Qwen team** launched **QVQ**, a vision-enabled version of their experimental **QwQ o1 clone**, benchmarking comparably to **Claude 3.5 Sonnet**. Discussions include **Bret Taylor&apos;s** insights on autonomous software development distinct from the Copilot era. The **Latent Space LIVE!** talks cover highlights of **2024 AI startups, vision, open models, post-transformers, synthetic data, smol models, and agents**. Twitter recaps by **Claude 3.5 Sonnet** highlight proposals for benchmarks measuring LLM calibration and falsehood confidence, with **QVQ** outperforming **GPT-4o** and **Claude Sonnet 3.5**. AI alignment debates focus on intentionality and critiques of alignment faking in models like **Claude**. Updates from **OpenAI** include new **o3 and o3-mini models** and a deliberative alignment strategy. The **ASAL project** is a collaboration between **MIT**, **OpenAI**, and **Swiss AI Lab IDSIA** to automate artificial life discovery. Personal stories reveal frustrations with **USCIS** green card denials despite high qualifications. New tools like **GeminiCoder** enable rapid app creation, and a **contract review agent** using **Reflex** and **Llama Index** checks GDPR compliance. Holiday greetings and memes were also shared.</description><pubDate>Wed, 25 Dec 2024 02:01:53 GMT</pubDate><category>alibaba</category><category>openai</category><category>mit</category><category>idsia</category><category>llamaindex</category><category>ollama</category><category>qwen-o1</category><category>qvq</category><category>claude-3.5-sonnet</category><category>gpt-4o</category><category>o3</category><category>o3-mini</category><category>bret-taylor</category><category>vision</category><category>benchmarking</category><category>llm-calibration</category><category>intentionality</category><category>alignment-faking</category><category>deliberative-alignment</category><category>artificial-life</category><category>gdpr-compliance</category><category>contract-review-agent</category><category>app-creation</category><category>synthetic-data</category><category>post-transformers</category><category>smol-models</category><category>agents</category></item><item><title>not much happened this weekend</title><link>https://news.smol.ai/issues/24-12-23-ainews-not-much-happened-this-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-23-ainews-not-much-happened-this-weekend/</guid><description>**o3** model gains significant attention with discussions around its capabilities and implications, including an OpenAI board member referencing &quot;AGI.&quot; **LangChain** released their **State of AI 2024** survey. **Hume** announced **OCTAVE**, a **3B parameter** API-only speech-language model with voice cloning. **x.ai** secured a **$6B Series C** funding round. Discussions highlight **inference-time scaling**, **model ensembles**, and the surprising generalization ability of **small models**. New tools and datasets include **FineMath**, the best open math dataset on Hugging Face, and frameworks for LLM agents. Industry updates cover a **5-month benchmarking** of **AMD MI300X** vs **Nvidia H100 + H200**, insights from a meeting with **Lisa Su** on AMD&apos;s software stack, and open AI engineering roles. Research innovations include **Large Concept Models (LCM)** from Meta AI, **Chain of Continuous Thought (Coconut)** for latent space reasoning, and mechanistic interpretability initiatives.</description><pubDate>Tue, 24 Dec 2024 01:01:31 GMT</pubDate><category>openai</category><category>langchain</category><category>hume</category><category>x-ai</category><category>amd</category><category>nvidia</category><category>meta-ai-fair</category><category>hugging-face</category><category>o3</category><category>o1</category><category>opus</category><category>sonnet</category><category>octave</category><category>lisa-su</category><category>clementdelangue</category><category>philschmid</category><category>neelnanda5</category><category>inference-time-scaling</category><category>model-ensembles</category><category>small-models</category><category>voice-cloning</category><category>fine-math-dataset</category><category>llm-agent-framework</category><category>benchmarking</category><category>software-stack</category><category>large-concept-models</category><category>latent-space-reasoning</category><category>mechanistic-interpretability</category><category>planning</category><category>speech-language-models</category></item><item><title>o3 solves AIME, GPQA, Codeforces, makes 11 years of progress in ARC-AGI and 25% in FrontierMath</title><link>https://news.smol.ai/issues/24-12-20-ainews-o3-solves-aime-gpqa-codeforces-makes-11-years-of-progress-in-arc-agi-and-25percent-in-frontiermath/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-20-ainews-o3-solves-aime-gpqa-codeforces-makes-11-years-of-progress-in-arc-agi-and-25percent-in-frontiermath/</guid><description>**OpenAI** announced the **o3** and **o3-mini** models with groundbreaking benchmark results, including a jump from **2% to 25%** on the **FrontierMath** benchmark and **87.5%** on the **ARC-AGI** reasoning benchmark, representing about **11 years of progress** on the GPT3 to GPT4o scaling curve. The **o1-mini** model shows superior inference efficiency compared to o3-full, promising significant cost reductions on coding tasks. The announcement was accompanied by community discussions, safety testing applications, and detailed analyses. *Sama* highlighted the unusual cost-performance tradeoff, and **Eric Wallace** shared insights on the o-series deliberative alignment strategy.</description><pubDate>Sat, 21 Dec 2024 01:44:22 GMT</pubDate><category>openai</category><category>o3</category><category>o3-mini</category><category>o1-mini</category><category>gpt-3</category><category>gpt-4o</category><category>o1</category><category>sama</category><category>eric-wallace</category><category>benchmarking</category><category>math</category><category>reasoning</category><category>model-performance</category><category>inference-speed</category><category>cost-efficiency</category><category>alignment</category><category>safety-testing</category></item><item><title>ModernBert: small new Retriever/Classifier workhorse, 8k context, 2T tokens, </title><link>https://news.smol.ai/issues/24-12-19-ainews-modernbert-small-new-retrieverclassifier-workhorse-8k-context-2t-tokens/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-19-ainews-modernbert-small-new-retrieverclassifier-workhorse-8k-context-2t-tokens/</guid><description>**Answer.ai/LightOn** released **ModernBERT**, an updated encoder-only model with **8k token context**, trained on **2 trillion tokens** including code, with **139M/395M parameters** and state-of-the-art performance on retrieval, NLU, and code tasks. It features **Alternating Attention** layers mixing global and local attention. **Gemini 2.0 Flash Thinking** debuted as #1 in Chatbot Arena, and the **O1 model** scored top in reasoning benchmarks. **Llama** downloads surpassed **650 million**, doubling in 3 months. **OpenAI** launched desktop app integrations with voice capabilities. **Figure** delivered its first humanoid robots commercially. Advances in robotics simulation and a new physics engine **Genesis** claiming **430,000x faster than real-time** were highlighted.</description><pubDate>Fri, 20 Dec 2024 03:27:55 GMT</pubDate><category>answerdotai</category><category>lightonio</category><category>hugging-face</category><category>google-deepmind</category><category>openai</category><category>meta-ai-fair</category><category>figure</category><category>modernbert</category><category>gemini-2.0-flash-thinking</category><category>o1</category><category>llama</category><category>jeremyphoward</category><category>alec-radford</category><category>philschmid</category><category>drjimfan</category><category>bindureddy</category><category>encoder-only-models</category><category>long-context</category><category>alternating-attention</category><category>natural-language-understanding</category><category>reasoning</category><category>robotics-simulation</category><category>physics-engine</category><category>humanoid-robots</category><category>model-performance</category><category>model-releases</category></item><item><title>Genesis: Generative Physics Engine for Robotics (o1-mini version)</title><link>https://news.smol.ai/issues/24-12-18-ainews-genesis-generative-physics-engine-for-robotics-o1-mini-version/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-18-ainews-genesis-generative-physics-engine-for-robotics-o1-mini-version/</guid><description>**OpenAI** launched the **o1 model** API featuring function calling, structured outputs, vision support, and developer messages, achieving **60% fewer reasoning tokens** than its preview. The model excels in math and code with a **0.76 LiveBench Coding score**, outperforming Sonnet 3.5. Beta SDKs for Go and Java and WebRTC support with **60% lower prices** were also released. **Google Gemini 2.0 Pro (Gemini Exp 1206)** deployment accelerated, showing improved coding, math, and reasoning performance. Meta AI FAIR introduced research on training transformers directly on raw bytes using dynamic entropy-based patching. Commercial humanoid robots were successfully deployed by an industry player. **Hugging Face** researchers demonstrated that their **3B Llama model** can outperform the **70B Llama model** on MATH-500 accuracy using search techniques, highlighting efficiency gains with smaller models. Concerns about reproducibility and domain-specific limitations were noted.</description><pubDate>Thu, 19 Dec 2024 05:17:10 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>meta-ai-fair</category><category>hugging-face</category><category>o1</category><category>o1-preview</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>gemini-2.0-pro</category><category>llama-3-3b</category><category>llama-3-70b</category><category>aidan_mclau</category><category>sundarpichai</category><category>adcock_brett</category><category>function-calling</category><category>structured-outputs</category><category>vision</category><category>performance-benchmarks</category><category>sdk</category><category>webrtc</category><category>reasoning</category><category>math</category><category>code-generation</category><category>transformer-architecture</category><category>model-training</category><category>humanoid-robots</category><category>search</category><category>model-efficiency</category><category>dataset-sharing</category></item><item><title>Genesis: Generative Physics Engine for Robotics (o1-2024-12-17)</title><link>https://news.smol.ai/issues/24-12-18-ainews-genesis-generative-physics-engine-for-robotics-o1-2024-12-17/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-18-ainews-genesis-generative-physics-engine-for-robotics-o1-2024-12-17/</guid><description>**Genesis** is a newly announced **universal physics engine** developed by a large-scale collaboration led by **CMU PhD student Zhou Xian**. It integrates multiple state-of-the-art physics solvers to simulate diverse materials and physical phenomena, targeting robotics applications with features like lightweight, ultra-fast simulation, photo-realistic rendering, and generative data capabilities. The engine is open source and designed for robotics simulation beyond just video generation. Additionally, **OpenAI** released the **o1** model to API with advanced features like function calling and vision support, showing strong math and coding performance. **Google** teased updates on **Gemini 2.0 Pro**, accelerating deployment for advanced users.</description><pubDate>Thu, 19 Dec 2024 04:48:33 GMT</pubDate><category>openai</category><category>google</category><category>carnegie-mellon-university</category><category>o1</category><category>gemini-2.0-pro</category><category>zhou-xian</category><category>aidan_mclau</category><category>sundar-pichai</category><category>universal-physics-engine</category><category>robotics-simulation</category><category>physics-simulation</category><category>photo-realistic-rendering</category><category>generative-data</category><category>simulation-platform</category><category>open-source</category><category>function-calling</category><category>vision</category><category>performance-benchmarks</category><category>sdk</category><category>realtime-api</category></item><item><title>OpenAI Voice Mode Can See Now - After Gemini Does</title><link>https://news.smol.ai/issues/24-12-18-ainews-openai-voice-mode-can-see-now-after-gemini-does/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-18-ainews-openai-voice-mode-can-see-now-after-gemini-does/</guid><description>**OpenAI** launched **Realtime Video** shortly after **Gemini**, which led to less impact due to Gemini&apos;s earlier arrival with lower cost and fewer rate limits. **Google DeepMind** released **Gemini 2.0 Flash** featuring enhanced multimodal capabilities and real-time streaming. **Anthropic** introduced **Clio**, a system analyzing real-world usage of **Claude** models. Together Computing acquired CodeSandbox to launch a code interpreter tool. Discussions highlighted **Meta&apos;s Llama 3.3-70B** for its advanced roleplay and prompt handling abilities, outperforming models like **Mistral Large** and **GPT-4o** in expressiveness and censorship. The AI community also engaged in humorous takes on AI outages and model competition, with **ChatGPT** adding a Santa mode for holiday interactions. *&quot;Anthropic is capturing the developer ecosystem, Gemini has AI enthusiast mindshare, ChatGPT reigns over AI dabblers&quot;* was a noted observation from the community.</description><pubDate>Wed, 18 Dec 2024 09:46:07 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>togethercompute</category><category>scale-ai</category><category>meta-ai-fair</category><category>mistral-ai</category><category>gemini-2.0-flash</category><category>claude</category><category>claude-3.5-sonnet</category><category>llama-3-70b</category><category>llama-3</category><category>mistral-large</category><category>gpt-4o</category><category>bindureddy</category><category>multimodality</category><category>real-time-streaming</category><category>roleplay</category><category>prompt-handling</category><category>model-comparison</category><category>model-training</category><category>creative-writing</category><category>model-censorship</category><category>code-execution</category><category>developer-ecosystem</category><category>ai-humor</category></item><item><title>o1 API, 4o/4o-mini in Realtime API + WebRTC, DPO Finetuning</title><link>https://news.smol.ai/issues/24-12-17-ainews-o1-api-4o4o-mini-in-realtime-api-webrtc-dpo-finetuning/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-17-ainews-o1-api-4o4o-mini-in-realtime-api-webrtc-dpo-finetuning/</guid><description>**OpenAI** launched the **o1 API** with enhanced features including vision inputs, function calling, structured outputs, and a new `reasoning_effort` parameter, achieving **60% fewer reasoning tokens** on average. The **o1 pro** variant is confirmed as a distinct implementation coming soon. Improvements to the **Realtime API** with **WebRTC** integration offer easier usage, longer sessions (up to **30 minutes**), and significantly reduced pricing (up to **10x cheaper** with mini models). **DPO Preference Tuning** for fine-tuning is introduced, currently available for the **4o** model. Additional updates include official Go and Java SDKs and OpenAI DevDay videos. The news also highlights discussions on **Google Gemini 2.0 Flash** model&apos;s performance reaching **83.6% accuracy**.</description><pubDate>Wed, 18 Dec 2024 01:43:51 GMT</pubDate><category>openai</category><category>google</category><category>google-deepmind</category><category>o1-2024-12-17</category><category>o1</category><category>o1-pro</category><category>4o</category><category>4o-mini</category><category>gemini-2-0-flash</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>aidan_mclau</category><category>kevinweil</category><category>simonw</category><category>michpokrass</category><category>morgymcg</category><category>juberti</category><category>function-calling</category><category>structured-outputs</category><category>vision</category><category>reasoning</category><category>webrtc</category><category>realtime-api</category><category>preference-tuning</category><category>fine-tuning</category><category>api</category><category>model-performance</category></item><item><title>Meta Apollo - Video Understanding up to 1 hour, SOTA Open Weights</title><link>https://news.smol.ai/issues/24-12-16-ainews-meta-apollo-video-understanding-up-to-1-hour-sota-open-weights/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-16-ainews-meta-apollo-video-understanding-up-to-1-hour-sota-open-weights/</guid><description>**Meta** released **Apollo**, a new family of state-of-the-art video-language models available in **1B, 3B, and 7B** sizes, featuring &quot;Scaling Consistency&quot; for efficient scaling and introducing **ApolloBench**, which speeds up video understanding evaluation by **41×** across five temporal perception categories. **Google Deepmind** launched **Veo 2**, a 4K video generation model with improved physics and camera control, alongside an enhanced **Imagen 3** image model. **OpenAI** globally rolled out ChatGPT search with advanced voice and map features and discussed a potential $2,000/month &quot;ChatGPT Max&quot; tier. Research highlights include achieving **Llama 70B** performance using **Llama 3B** via test-time compute scaling and expanding **Command R7B** language support from 10 to 23 languages. Industry updates feature **Figure AI** delivering humanoid robots commercially and **Klarna** reducing workforce through AI. Notion integrated **Cohere Rerank** for better search. Studies reveal LLMs can recognize their own writing style and show self-preference bias. Discussions note video processing progress outpacing text due to better signal-per-compute and data evaluation.</description><pubDate>Tue, 17 Dec 2024 01:17:52 GMT</pubDate><category>meta-ai-fair</category><category>hugging-face</category><category>google-deepmind</category><category>openai</category><category>figure-ai</category><category>klarna</category><category>cohere</category><category>notion</category><category>apollo-1b</category><category>apollo-3b</category><category>apollo-7b</category><category>veo-2</category><category>imagen-3</category><category>llama-3-70b</category><category>llama-3b</category><category>command-r7b</category><category>llama-1b</category><category>llama-8b</category><category>chatgpt</category><category>akhaliq</category><category>_lewtun</category><category>clementdelangue</category><category>adcock_brett</category><category>rohanpaul_ai</category><category>swyx</category><category>shaneguML</category><category>video-understanding</category><category>scaling-consistency</category><category>benchmarking</category><category>temporal-ocr</category><category>egocentric-perception</category><category>spatial-perception</category><category>reasoning</category><category>video-generation</category><category>physics-simulation</category><category>voice-features</category><category>map-integration</category><category>language-expansion</category><category>test-time-compute-scaling</category><category>humanoid-robots</category><category>ai-integration</category><category>search-optimization</category><category>self-recognition</category><category>self-preference-bias</category></item><item><title>Meta BLT: Tokenizer-free, Byte-level LLM</title><link>https://news.smol.ai/issues/24-12-13-ainews-meta-blt-tokenizer-free-byte-level-llm/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-13-ainews-meta-blt-tokenizer-free-byte-level-llm/</guid><description>**Meta AI** introduces the **Byte Latent Transformer (BLT)**, a tokenizer-free architecture that dynamically forms byte patches for efficient compute allocation, outperforming **Llama 3** on benchmarks including the CUTE benchmark. The model was trained on approximately **1 trillion tokens** and features a three-block transformer design with local and global components. This approach challenges traditional tokenization and may enable new multimodal capabilities such as direct file interaction without retrieval-augmented generation. Additionally, **Microsoft** announced the **Phi-4 14B** parameter model achieving state-of-the-art results on STEM and reasoning benchmarks, surpassing **GPT-4o**. **DeepSeek AI** launched new vision-language models based on their MoE architecture with sizes ranging from **1.0B to 27B** parameters. **OpenAI** released a new Projects feature for ChatGPT, and **Cohere** introduced their smallest and fastest **Command R7B** model. **Anthropic** published research on &quot;Best-of-N Jailbreaking&quot; vulnerabilities across text, vision, and audio models. Industry discussion highlights a trend of decreasing frontier LLM sizes, with **GPT-4** at approximately **1.8 trillion parameters** compared to newer models.</description><pubDate>Sat, 14 Dec 2024 05:38:19 GMT</pubDate><category>meta-ai-fair</category><category>llamaindex</category><category>microsoft</category><category>deepseek-ai</category><category>openai</category><category>cohere</category><category>anthropic</category><category>byte-latent-transformer</category><category>llama-3</category><category>phi-4</category><category>gpt-4o</category><category>command-r7b</category><category>tokenization</category><category>transformer-architecture</category><category>model-efficiency</category><category>benchmarking</category><category>multimodality</category><category>vision</category><category>reinforcement-learning</category><category>model-scaling</category><category>jailbreaking</category><category>model-optimization</category></item><item><title>Google wakes up: Gemini 2.0 et al</title><link>https://news.smol.ai/issues/24-12-11-ainews-google-wakes-up-gemini-20-et-al/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-11-ainews-google-wakes-up-gemini-20-et-al/</guid><description>**Google DeepMind** launched **Gemini 2.0 Flash**, a new multimodal model outperforming Gemini 1.5 Pro and o1-preview, featuring vision and voice APIs, multilingual capabilities, and native tool use. It powers new AI agents like **Project Astra** and **Project Mariner**, with Project Mariner achieving state-of-the-art **83.5%** on the WebVoyager benchmark. **OpenAI** announced ChatGPT integration with **Apple** devices, enabling Siri access and visual intelligence features. **Claude 3.5 Sonnet** is noted as a distilled version of Opus. The AI community&apos;s response at **NeurIPS 2024** has been overwhelmingly positive, signaling a strong comeback for Google in AI innovation. Key topics include **multimodality**, **agent development**, **multilinguality**, **benchmarking**, and **model releases**.</description><pubDate>Thu, 12 Dec 2024 03:16:07 GMT</pubDate><category>google-deepmind</category><category>openai</category><category>apple</category><category>gemini-2.0-flash</category><category>gemini-1.5-pro</category><category>gemini-exp-1206</category><category>claude-3.5-sonnet</category><category>opus</category><category>demis-hassabis</category><category>sundar-pichai</category><category>paige-bailey</category><category>bindureddy</category><category>multimodality</category><category>agent-development</category><category>multilinguality</category><category>benchmarking</category><category>model-releases</category></item><item><title>ChatGPT Canvas GA</title><link>https://news.smol.ai/issues/24-12-10-ainews-chatgpt-canvas-ga/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-10-ainews-chatgpt-canvas-ga/</guid><description>**OpenAI** launched **ChatGPT Canvas** to all users, featuring **code execution** and **GPT integration**, effectively replacing Code Interpreter with a Google Docs-like interface. **Deepseek AI** announced their **V2.5-1210** update improving performance on **MATH-500 (82.8%)** and LiveCodebench. **Meta AI Fair** introduced **COCONUT**, a new continuous latent space reasoning paradigm. **Huggingface** released **TGI v3**, processing **3x more tokens** and running **13x faster** than vLLM on long prompts. **Cognition Labs** released **Devin**, an AI developer building Kubernetes operators. **Hyperbolic** raised **$12M Series A** to build an open AI platform with an **H100 GPU marketplace**. Discussions included **AI capabilities and employment impact**, and **NeurIPS 2024** announcements with **Google DeepMind** demos and a debate on AI scaling. On Reddit, **Llama 3.3-70B** supports **90K context length** finetuning using **Unsloth** with **gradient checkpointing** and Apple&apos;s **Cut Cross Entropy (CCE)** algorithm, fitting on **41GB VRAM**. **Llama 3.1-8B** reaches **342K context lengths** with Unsloth, surpassing native limits.</description><pubDate>Wed, 11 Dec 2024 04:20:02 GMT</pubDate><category>openai</category><category>deepseek-ai</category><category>meta-ai-fair</category><category>huggingface</category><category>cognition-labs</category><category>hyperbolic</category><category>google-deepmind</category><category>llama-3-70b</category><category>llama-3-1-8b</category><category>tgi-v3</category><category>deepseek-v2.5-1210</category><category>coconut</category><category>arav_srinivas</category><category>sama</category><category>jonathan-frankle</category><category>dylan</category><category>code-execution</category><category>gpt-integration</category><category>model-finetuning</category><category>gradient-checkpointing</category><category>context-length</category><category>latent-space-reasoning</category><category>performance-optimization</category><category>gpu-memory-optimization</category><category>kubernetes</category><category>gpu-marketplace</category><category>ai-capabilities</category><category>employment-impact</category><category>neurips-2024</category><category>ai-scaling</category><category>humor</category></item><item><title>OpenAI Sora Turbo and Sora.com</title><link>https://news.smol.ai/issues/24-12-09-ainews-openai-sora-turbo-and-soracom/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-09-ainews-openai-sora-turbo-and-soracom/</guid><description>**OpenAI** launched **Sora Turbo**, enabling text-to-video generation for ChatGPT Plus and Pro users with monthly generation limits and regional restrictions in Europe and the UK. **Google** announced a quantum computing breakthrough with the development of the **Willow chip**, potentially enabling commercial quantum applications. Discussions on **O1** model performance highlighted its lag behind **Claude 3.5 Sonnet** and **Gemini** in coding tasks, with calls for algorithmic innovation beyond transformer scaling. The **Llama 3.3 Euryale v2.3** model was praised for storytelling and roleplay capabilities, with users suggesting parameter tuning to reduce creative liberties and repetition. Alternatives like **Mistral-Large**, **Behemoth**, and **Endurance v1.1** were also noted. Additionally, **Nvidia** faces an anti-monopoly investigation in China. Memes and humor around GPU issues and embargo mishaps were popular on social media.</description><pubDate>Tue, 10 Dec 2024 02:21:42 GMT</pubDate><category>openai</category><category>google</category><category>nvidia</category><category>hugging-face</category><category>mistral-ai</category><category>sora-turbo</category><category>o1</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>gemini</category><category>llama-3-3-euryale-v2.3</category><category>mistral-large</category><category>behemoth</category><category>endurance-v1.1</category><category>sama</category><category>sundarpichai</category><category>bindureddy</category><category>denny_zhou</category><category>nrehiew_</category><category>text-to-video-generation</category><category>quantum-computing</category><category>coding-capabilities</category><category>transformers</category><category>algorithmic-innovation</category><category>storytelling</category><category>roleplay</category><category>model-parameter-tuning</category><category>anti-monopoly-investigation</category></item><item><title>Meta Llama 3.3: 405B/Nova Pro performance at 70B price</title><link>https://news.smol.ai/issues/24-12-06-ainews-meta-llama-33-405bnova-pro-performance-at-70b-price/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-06-ainews-meta-llama-33-405bnova-pro-performance-at-70b-price/</guid><description>**Meta AI** released **Llama 3.3 70B**, matching the performance of the 405B model with improved efficiency using *&quot;a new alignment process and progress in online RL techniques&quot;*. **OpenAI** announced **Reinforcement Fine-Tuning (RFT)** for building expert models with limited data, offering alpha access to researchers and enterprises. **Google DeepMind&apos;s Gemini-Exp-1206** leads benchmarks, tying with **GPT-4o** in coding performance. **LlamaCloud** enhanced document processing with table extraction and analytics. Discussions on **OpenAI&apos;s** pricing plans continue in the community.</description><pubDate>Fri, 06 Dec 2024 22:44:07 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>google-deepmind</category><category>hugging-face</category><category>llamacloud</category><category>llama-3-70b</category><category>llama-3.3-70b</category><category>gpt-4o</category><category>gemini-exp-1206</category><category>sama</category><category>steven-heidel</category><category>aidan_mclau</category><category>lmarena_ai</category><category>oriolvinyalsml</category><category>jerryjliu0</category><category>reinforcement-learning</category><category>fine-tuning</category><category>model-performance</category><category>document-processing</category><category>pricing-models</category><category>alignment</category><category>online-rl</category></item><item><title>$200 ChatGPT Pro and o1-full/pro, with vision, without API, and mixed reviews</title><link>https://news.smol.ai/issues/24-12-05-ainews-dollar200-chatgpt-pro-and-o1-fullpro-with-vision-without-api-and-mixed-reviews/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-05-ainews-dollar200-chatgpt-pro-and-o1-fullpro-with-vision-without-api-and-mixed-reviews/</guid><description>**OpenAI** launched the **o1** model with multimodal capabilities, faster reasoning, and image input support, marking it as a state-of-the-art model despite some bugs and mixed community reviews. The new **o1-pro** tier offers unlimited access for $200/month with notable benchmark improvements but some performance trade-offs compared to **claude-3.5-sonnet**. **Google** released the **PaliGemma 2** vision-language model family in sizes **3B, 10B, and 28B**, excelling in visual question answering, image segmentation, and OCR, with day-0 support for fine-tuning. **LlamaIndex** announced discounts and feature updates for large-scale document processing. The AI community also reacted humorously to the new pricing tiers and model comparisons. *&quot;o1 can see now, which makes it the SOTA multimodal model&quot;* and *&quot;most users will be best served by free/Plus tiers&quot;* were notable sentiments.</description><pubDate>Fri, 06 Dec 2024 02:34:03 GMT</pubDate><category>openai</category><category>google</category><category>llamaindex</category><category>o1</category><category>o1-pro</category><category>claude-3.5-sonnet</category><category>pali-gemma-2</category><category>sama</category><category>bindureddy</category><category>mervenoyann</category><category>fchollet</category><category>multimodality</category><category>vision</category><category>fine-tuning</category><category>benchmarking</category><category>model-performance</category><category>image-generation</category><category>document-processing</category><category>model-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-12-04-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-04-ainews-not-much-happened-today/</guid><description>**OpenAI** announced their &quot;12 Days of OpenAI&quot; event with daily livestreams and potential releases including the **O1 full model**, **Sora video model**, and **GPT-4.5**. **Google DeepMind** released the **GenCast weather model** capable of **15-day forecasts in 8 minutes** using TPU chips, and launched **Genie 2**, a model generating playable 3D worlds from single images. Leading vision researchers **Lucas Beyer**, **Alexander Kolesnikov**, and **Xiaohua Zhai** moved from DeepMind to OpenAI, which is opening a Zürich office. Criticism arose over OpenAI&apos;s strategy and model quality compared to **Anthropic** and **Claude 3.5 Sonnet**. On Reddit, a modified **llama.cpp** supports **Nvidia&apos;s Llama-3_1-Nemotron-51B**, matching performance of larger 70B models via NAS optimization.</description><pubDate>Thu, 05 Dec 2024 02:41:39 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>nvidia</category><category>huggingface</category><category>o1-full</category><category>sora</category><category>gpt-4.5</category><category>gpt-4</category><category>claude-3.5-sonnet</category><category>llama-3-1-nemotron-51b</category><category>llama-3-1</category><category>llama-3</category><category>nemotron-51b</category><category>lucas-beyer</category><category>alexander-kolesnikov</category><category>xiaohua-zhai</category><category>aidan_mclau</category><category>giffmana</category><category>joannejang</category><category>sama</category><category>vision</category><category>model-performance</category><category>neural-architecture-search</category><category>model-optimization</category><category>multimodality</category><category>model-release</category><category>model-training</category><category>reinforcement-learning</category><category>image-generation</category></item><item><title>Olympus has dropped (aka, Amazon Nova Micro|Lite|Pro|Premier|Canvas|Reel)</title><link>https://news.smol.ai/issues/24-12-03-ainews-olympus-has-dropped-aka-amazon-nova-microorliteorproorpremierorcanvasorreel/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-03-ainews-olympus-has-dropped-aka-amazon-nova-microorliteorproorpremierorcanvasorreel/</guid><description>**Amazon** announced the **Amazon Nova** family of multimodal foundation models at AWS Re:Invent, available immediately with no waitlist in configurations like Micro, Lite, Pro, Canvas, and Reel, with Premier and speech-to-speech coming next year. These models offer **2-4x faster token speeds** and are **25%-400% cheaper** than competitors like **Anthropic Claude** models, positioning Nova as a serious contender in AI engineering. Pricing undercuts models such as **Google DeepMind Gemini Flash 8B**, and some Nova models extend context length up to **300k tokens**. However, benchmarking controversy exists as some evaluations show Nova scoring below **Llama-3 70B** in **LiveBench AI** metrics. Separately, **CycleQD** was introduced by **Sakana AI Labs**, using evolutionary computation for population-based model merging to develop niche LLM agents.</description><pubDate>Wed, 04 Dec 2024 03:06:39 GMT</pubDate><category>amazon</category><category>anthropic</category><category>google-deepmind</category><category>sakana-ai-labs</category><category>amazon-nova</category><category>claude-3</category><category>llama-3-70b</category><category>gemini-1.5-flash</category><category>gpt-4o</category><category>philschmid</category><category>bindureddy</category><category>multimodality</category><category>benchmarking</category><category>model-merging</category><category>model-performance</category><category>model-architecture</category><category>model-optimization</category><category>population-based-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-12-02-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-12-02-ainews-not-much-happened-today/</guid><description>**AI News for 11/29/2024-12/2/2024** highlights several developments: **Nvidia** introduced **Puzzle**, a distillation-based neural architecture search for inference-optimized large language models, enhancing efficiency. The **IC-Light V2** model was released for varied illumination scenarios, and new video model techniques like **Trajectory Attention** and **Timestep Embedding** were presented. **Amazon** increased its investment in **Anthropic** to **$8 billion**, supporting AI safety research through a new fellowship program. **Google** is expanding AI integration with the **Gemini API** and open collaboration tools. Discussions on domain name relevance emphasize alternatives to **.com** domains like **.io**, **.ai**, and **.co**. Advances in reasoning include a **13.53% improvement** in LLM performance using &quot;Reverse Thinking&quot;. **Pydantic** launched a new agent framework, and **Supabase** released version 2 of their assistant. Other notable mentions include **Browser Company** teasing a second browser and **World Labs** launching image-to-3D-world technology. The NotebookLM team departed from **Google**, and **Cognition** was featured on the cover of **Forbes**. The news was summarized by **Claude 3.5 Sonnet**.</description><pubDate>Mon, 02 Dec 2024 23:49:20 GMT</pubDate><category>nvidia</category><category>amazon</category><category>anthropic</category><category>google</category><category>pydantic</category><category>supabase</category><category>browser-company</category><category>world-labs</category><category>cognition</category><category>ic-light-v2</category><category>claude-3-5-sonnet</category><category>puzzle</category><category>akhaliq</category><category>adcock_brett</category><category>omarsar0</category><category>iscienceluvr</category><category>distillation</category><category>neural-architecture-search</category><category>inference-optimization</category><category>video</category><category>trajectory-attention</category><category>timestep-embedding</category><category>ai-safety-research</category><category>fellowship-programs</category><category>api</category><category>domain-names</category><category>reverse-thinking</category><category>reasoning</category><category>agent-frameworks</category><category>image-to-3d</category><category>ai-integration</category></item><item><title>not much happened to end the week</title><link>https://news.smol.ai/issues/24-11-29-ainews-not-much-happened-to-end-the-week/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-29-ainews-not-much-happened-to-end-the-week/</guid><description>**AI News for 11/29/2024-11/30/2024** covers key updates including the **Gemini multimodal model** advancing in musical structure understanding, a new **quantized SWE-Bench** for benchmarking at **1.3 bits per task**, and the launch of the **DeepSeek-R1 model** focusing on transparent reasoning as an alternative to **o1**. The establishment of the **1st International Network of AI Safety Institutes** highlights global collaboration on AI safety. Industry updates feature **Amazon&apos;s Olympus AI model**, **Tesla&apos;s Optimus**, and experiments with **ChatGPT** as a universal translator. Community reflections emphasize the impact of large language models on daily life and medical AI applications. Discussions include scaling sparse autoencoders to **gpt-4** and the need for transparency in reasoning LLMs. The report also notes humor around **ChatGPT**&apos;s French nickname.</description><pubDate>Fri, 29 Nov 2024 23:07:35 GMT</pubDate><category>google-deepmind</category><category>deeplearningai</category><category>amazon</category><category>tesla</category><category>x-ai</category><category>alibaba</category><category>ollama</category><category>gemini</category><category>deepseek-r1</category><category>o1</category><category>chatgpt</category><category>gpt-4</category><category>claude-3.5-sonnet</category><category>o1-preview</category><category>o1-mini</category><category>gpt4o</category><category>qwq-32b</category><category>yoshua-bengio</category><category>kevinweil</category><category>ylecun</category><category>multimodality</category><category>benchmarking</category><category>quantization</category><category>reinforcement-learning</category><category>ai-safety</category><category>translation</category><category>reasoning</category><category>interpretability</category><category>model-comparison</category><category>humor</category></item><item><title>Qwen with Questions: 32B open weights reasoning model nears o1 in GPQA/AIME/Math500</title><link>https://news.smol.ai/issues/24-11-27-ainews-qwen-with-questions-32b-open-weights-reasoning-model-nears-o1-in-gpqaaimemath500/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-27-ainews-qwen-with-questions-32b-open-weights-reasoning-model-nears-o1-in-gpqaaimemath500/</guid><description>**DeepSeek r1** leads the race for &quot;open o1&quot; models but has yet to release weights, while **Justin Lin** released **QwQ**, a **32B open weight model** that outperforms **GPT-4o** and **Claude 3.5 Sonnet** on benchmarks. QwQ appears to be a fine-tuned version of **Qwen 2.5**, emphasizing sequential search and reflection for complex problem-solving. **SambaNova** promotes its RDUs as superior to GPUs for inference tasks, highlighting the shift from training to inference in AI systems. On Twitter, **Hugging Face** announced CPU deployment for llama.cpp instances, **Marker v1** was released as a faster and more accurate deployment tool, and **Agentic RAG** developments focus on integrating external tools and advanced LLM chains for improved response accuracy. The open-source AI community sees growing momentum with models like **Flux** gaining popularity, reflecting a shift towards multi-modal AI models including image, video, audio, and biology.</description><pubDate>Thu, 28 Nov 2024 01:23:25 GMT</pubDate><category>deepseek</category><category>sambanova</category><category>hugging-face</category><category>dair-ai</category><category>deepseek-r1</category><category>qwq</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>qwen-2.5</category><category>llama-cpp</category><category>justin-lin</category><category>clementdelangue</category><category>ggerganov</category><category>vikparuchuri</category><category>model-releases</category><category>benchmarking</category><category>fine-tuning</category><category>sequential-search</category><category>inference</category><category>model-deployment</category><category>agentic-rag</category><category>external-tools</category><category>multi-modal-models</category></item><item><title>OLMo 2 - new SOTA Fully Open LLM</title><link>https://news.smol.ai/issues/24-11-26-ainews-olmo-2-new-sota-fully-open-llm/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-26-ainews-olmo-2-new-sota-fully-open-llm/</guid><description>**AI2** has updated **OLMo-2** to roughly **Llama 3.1 8B** equivalent, training with **5T tokens** and using learning rate annealing and new high-quality data (Dolmino). They credit **Tülu 3** and its &quot;Reinforcement Learning with Verifiable Rewards&quot; approach. On Reddit, **Qwen2.5-72B instruct** model shows near lossless performance with **AutoRound 4-bit quantization**, available on **HuggingFace** in 4-bit and 2-bit versions, with discussions on **MMLU** benchmark and quantization-aware training. **HuggingFace** released **SmolVLM**, a **2B parameter** vision-language model running efficiently on consumer GPUs, supporting fine-tuning on Google Colab and demonstrating strong OCR capabilities with adjustable resolution and quantization options.</description><pubDate>Wed, 27 Nov 2024 05:17:18 GMT</pubDate><category>ai2</category><category>huggingface</category><category>intel</category><category>llama-3-1-8b</category><category>olmo-2</category><category>qwen2-5-72b-instruct</category><category>smolvlm</category><category>tulu-3</category><category>reinforcement-learning</category><category>quantization</category><category>learning-rate-annealing</category><category>ocr</category><category>fine-tuning</category><category>model-training</category><category>vision</category></item><item><title>Anthropic launches the Model Context Protocol</title><link>https://news.smol.ai/issues/24-11-25-ainews-anthropic-launches-the-model-context-protocol/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-25-ainews-anthropic-launches-the-model-context-protocol/</guid><description>**Anthropic** has launched the **Model Context Protocol (MCP)**, an open protocol designed to enable seamless integration between large language model applications and external data sources and tools. MCP supports diverse resources such as file contents, database records, API responses, live system data, screenshots, and logs, identified by unique URIs. It also includes reusable prompt templates, system and API tools, and JSON-RPC 2.0 transports with streaming support. MCP allows servers to request LLM completions through clients with priorities on cost, speed, and intelligence, hinting at an upcoming model router by Anthropic. Launch partners like **Zed**, **Sourcegraph**, and **Replit** have reviewed MCP favorably, while some developers express skepticism about its provider exclusivity and adoption potential. The protocol emphasizes security, testing, and dynamic tool discovery, with guides and videos available from community members such as **Alex Albert** and **Matt Pocock**. This development follows Anthropic&apos;s recent **$4 billion fundraise from Amazon** and aims to advance terminal-level integration for **Claude Desktop**.</description><pubDate>Tue, 26 Nov 2024 01:56:47 GMT</pubDate><category>anthropic</category><category>amazon</category><category>zed</category><category>sourcegraph</category><category>replit</category><category>claude-3.5-sonnet</category><category>claude-desktop</category><category>alex-albert</category><category>matt-pocock</category><category>hwchase17</category><category>model-context-protocol</category><category>integration</category><category>json-rpc</category><category>agentic-behaviors</category><category>security</category><category>tool-discovery</category><category>open-protocol</category><category>api-integration</category><category>system-integration</category><category>prompt-templates</category><category>model-routing</category></item><item><title>Vision Everywhere: Apple AIMv2 and Jina CLIP v2</title><link>https://news.smol.ai/issues/24-11-22-ainews-vision-everywhere-apple-aimv2-and-jina-clip-v2/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-22-ainews-vision-everywhere-apple-aimv2-and-jina-clip-v2/</guid><description>**Apple** released **AIMv2**, a novel vision encoder pre-trained with autoregressive objectives that achieves **89.5% accuracy on ImageNet** and integrates joint visual and textual objectives. **Jina** launched **Jina CLIP v2**, a multimodal embedding model supporting **89 languages** and high-resolution images with efficient Matryoshka embeddings reducing dimensions by **94%** with minimal accuracy loss. **Allen AI** introduced **Tülu 3** models based on **Llama 3.1** with **8B and 70B** parameters, offering **2.5x faster inference** and alignment via SFT, DPO, and RLVR methods, competing with **Claude 3.5** and **Llama 3.1 70B**. These developments highlight advances in autoregressive training, vision encoders, and multilingual multimodal embeddings.</description><pubDate>Fri, 22 Nov 2024 23:31:04 GMT</pubDate><category>apple</category><category>jina</category><category>allen_ai</category><category>aimv2-3b</category><category>jina-clip-v2</category><category>tulu-3</category><category>llama-3-1</category><category>claude-3-5</category><category>llama-3-1-70b</category><category>autoregressive-objectives</category><category>vision</category><category>multilinguality</category><category>multimodality</category><category>image-generation</category><category>model-training</category><category>model-optimization</category><category>reinforcement-learning</category><category>fine-tuning</category><category>model-benchmarking</category></item><item><title>LMSys killed Model Versioning (gpt 4o 1120, gemini exp 1121)</title><link>https://news.smol.ai/issues/24-11-21-ainews-lmsys-killed-model-versioning-gpt-4o-1120-gemini-exp-1121/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-21-ainews-lmsys-killed-model-versioning-gpt-4o-1120-gemini-exp-1121/</guid><description>**AI News for 11/21/2024-11/22/2024** highlights the intense frontier lab race with **OpenAI&apos;s gpt-4o-2024-11-20** and **Google DeepMind&apos;s gemini-exp-1121** trading top spots on the Lmsys leaderboard. The trend of using date-based model identifiers instead of traditional versioning is noted across leading labs including **Anthropic**. **DeepSeek R1** is gaining attention as a potent open-source alternative, especially in the context of the AI competition between China and the US. **Gemini-Exp-1121** is praised for improvements in vision, coding, and reasoning, while **MistralAI** expands with a new Palo Alto office, signaling growth and hiring.</description><pubDate>Fri, 22 Nov 2024 00:56:03 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>deepseek</category><category>mistral-ai</category><category>gpt-4o-2024-11-20</category><category>gemini-exp-1121</category><category>deepseek-r1</category><category>model-release</category><category>model-ranking</category><category>open-source</category><category>vision</category><category>coding</category><category>reasoning</category><category>market-competition</category></item><item><title>DeepSeek-R1 claims to beat o1-preview AND will be open sourced</title><link>https://news.smol.ai/issues/24-11-20-ainews-deepseek-r1-claims-to-beat-o1-preview-and-will-be-open-sourced/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-20-ainews-deepseek-r1-claims-to-beat-o1-preview-and-will-be-open-sourced/</guid><description>**DeepSeek** has released **DeepSeek-R1-Lite-Preview**, an open-source reasoning model achieving **o1-preview-level performance** on math benchmarks with transparent thought processes, showing promise in real-time problem-solving. **NVIDIA** reported a record **$35.1 billion** revenue in Q3 with **112% year-on-year data center growth**, driven by **Hopper** and **Blackwell architectures**, the latter offering **2.2x performance improvement**. **Google DeepMind** introduced **AlphaQubit**, a quantum computing system improving error correction and outperforming leading decoders, though challenges remain in scaling and speed. The AI community continues to focus on **reasoning models**, **benchmarking**, and **quantum error correction** advancements.</description><pubDate>Thu, 21 Nov 2024 02:41:02 GMT</pubDate><category>deepseek</category><category>nvidia</category><category>google-deepmind</category><category>deepseek-r1-lite-preview</category><category>o1-preview</category><category>hopper</category><category>blackwell</category><category>alphaqubit</category><category>yann-lecun</category><category>reasoning</category><category>benchmarking</category><category>quantum-error-correction</category><category>quantum-computing</category><category>model-performance</category><category>model-release</category></item><item><title>Perplexity starts Shopping for you</title><link>https://news.smol.ai/issues/24-11-19-ainews-perplexity-starts-shopping-for-you/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-19-ainews-perplexity-starts-shopping-for-you/</guid><description>**Stripe** launched their Agent SDK, enabling AI-native shopping experiences like **Perplexity Shopping** for US Pro members, featuring one-click checkout and free shipping via the **Perplexity Merchant Program**. **Mistral AI** released the **Pixtral Large 124B** multi-modal image model, now on **Hugging Face** and supported by **Le Chat** for image generation. **Cerebras Systems** offers a public inference endpoint for **Llama 3.1 405B** with a 128k context window and high throughput. **Claude 3.6** shows improvements over **Claude 3.5** but with subtle hallucinations. The **Bi-Mamba** 1-bit architecture improves LLM efficiency. The **wandb SDK** is preinstalled on Google Colab, and **Pixtral Large** is integrated into **AnyChat** and supported by **vLLM** for efficient model usage.</description><pubDate>Wed, 20 Nov 2024 00:43:00 GMT</pubDate><category>stripe</category><category>perplexity-ai</category><category>mistral-ai</category><category>hugging-face</category><category>cerebras</category><category>anthropic</category><category>weights-biases</category><category>google</category><category>vllm-project</category><category>pixtral-large-124b</category><category>llama-3.1-405b</category><category>claude-3.6</category><category>claude-3.5</category><category>patrick-collison</category><category>jeff-weinstein</category><category>mervenoyann</category><category>sophiamyang</category><category>tim-dettmers</category><category>omarsar0</category><category>akhaliq</category><category>aravsrinivas</category><category>multi-modal</category><category>image-generation</category><category>inference</category><category>context-windows</category><category>model-performance</category><category>model-efficiency</category><category>sdk</category><category>ai-integration</category><category>one-click-checkout</category><category>memory-optimization</category></item><item><title>Pixtral Large (124B) beats Llama 3.2 90B with updated Mistral Large 24.11</title><link>https://news.smol.ai/issues/24-11-18-ainews-pixtral-large-124b-beats-llama-32-90b-with-updated-mistral-large-2411/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-18-ainews-pixtral-large-124b-beats-llama-32-90b-with-updated-mistral-large-2411/</guid><description>**Mistral** has updated its **Pixtral Large** vision encoder to 1B parameters and released an update to the **123B parameter Mistral Large 24.11** model, though the update lacks major new features. **Pixtral Large** outperforms **Llama 3.2 90B** on multimodal benchmarks despite having a smaller vision adapter. **Mistral&apos;s Le Chat** chatbot received comprehensive feature updates, reflecting a company focus on product and research balance as noted by **Arthur Mensch**. **SambaNova** sponsors inference with their RDUs offering faster AI model processing than GPUs. On Reddit, **vLLM** shows strong concurrency performance on an **RTX 3090** GPU, with quantization challenges noted in **FP8 kv-cache** but better results using **llama.cpp** with **Q8 kv-cache**. Users discuss performance trade-offs between **vLLM**, **exllamav2**, and **TabbyAPI** for different model sizes and batching strategies.</description><pubDate>Tue, 19 Nov 2024 02:25:23 GMT</pubDate><category>mistral-ai</category><category>sambanova</category><category>nvidia</category><category>pixtral-large</category><category>mistral-large-24.11</category><category>llama-3-2</category><category>qwen2.5-7b-instruct-abliterated-v2-gguf</category><category>qwen2.5-32b-q3_k_m</category><category>vllm</category><category>llama-cpp</category><category>exllamav2</category><category>tabbyapi</category><category>arthur-mensch</category><category>multimodality</category><category>vision</category><category>model-updates</category><category>chatbots</category><category>inference</category><category>gpu-optimization</category><category>quantization</category><category>performance</category><category>concurrency</category><category>kv-cache</category></item><item><title>Stripe lets Agents spend money with StripeAgentToolkit</title><link>https://news.smol.ai/issues/24-11-15-ainews-stripe-lets-agents-spend-money-with-stripeagenttoolkit/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-15-ainews-stripe-lets-agents-spend-money-with-stripeagenttoolkit/</guid><description>**Stripe** has pioneered an AI SDK specifically designed for agents that handle payments, integrating with models like **gpt-4o** to enable financial transactions and token-based charging. The AI developer tooling trend emphasizes better &quot;AI-Computer Interfaces&quot; for improved agent reliability, with tools like **E2B** and the `llms.txt` documentation trend gaining traction, notably adopted by **Anthropic**. In AI model news, **Gemini-Exp-1114** topped the Vision Leaderboard and improved in Math Arena, while discussions continue around model overfitting and the limits of scaling laws for **AGI**. **OpenAI** released a **ChatGPT desktop app for macOS** with integrations for **VS Code**, **Xcode**, and **Terminal**, enhancing developer workflows and pair programming. **Anthropic** introduced a prompt improver using chain-of-thought reasoning, and **Meta AI** shared top research from **EMNLP2024** on image captioning, dialogue systems, and memory-efficient fine-tuning. Highlights from **ICLR 2025** include diffusion-based illumination harmonization, open mixture-of-experts language models, and hyperbolic vision-language models. A new adaptive decoding method optimizes creativity and factuality per token. Tools like **LlamaParse** and **RAGformation** were also introduced for document parsing and retrieval-augmented generation.</description><pubDate>Sat, 16 Nov 2024 01:02:33 GMT</pubDate><category>stripe</category><category>openai</category><category>anthropic</category><category>meta-ai-fair</category><category>gpt-4o</category><category>gemini-exp-1114</category><category>abacaj</category><category>francois-fleuret</category><category>lmarena_ai</category><category>goodside</category><category>jxmnop</category><category>jaseweston</category><category>stevenheidel</category><category>ai-computer-interfaces</category><category>agentic-ai</category><category>model-overfitting</category><category>benchmarks</category><category>scaling-laws</category><category>agi</category><category>chain-of-thought</category><category>image-captioning</category><category>dialogue-systems</category><category>memory-efficient-fine-tuning</category><category>diffusion-models</category><category>mixture-of-experts</category><category>adaptive-decoding</category><category>creativity-optimization</category><category>factuality-optimization</category><category>pair-programming</category><category>document-parsing</category><category>retrieval-augmented-generation</category></item><item><title>Gemini (Experimental-1114) retakes #1 LLM rank with 1344 Elo</title><link>https://news.smol.ai/issues/24-11-14-ainews-gemini-experimental-1114-retakes-1-llm-rank-with-1344-elo/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-14-ainews-gemini-experimental-1114-retakes-1-llm-rank-with-1344-elo/</guid><description>**Anthropic** released the **3.5 Sonnet** benchmark for jailbreak robustness, emphasizing adaptive defenses. **OpenAI** enhanced **GPT-4** with a new RAG technique for contiguous chunk retrieval. **LangChain** launched **Promptim** for prompt optimization. **Meta AI** introduced **NeuralFeels** with neural fields for visuotactile perception. **RichardMCNgo** resigned from **OpenAI**, highlighting concerns on **AI governance** and **theoretical alignment**. Discussions emphasized the importance of **truthful public information** and **ethical alignment** in AI deployment. The latest **Gemini** update marks a new #1 LLM amid alignment challenges. The AI community continues to focus on **benchmarking**, **prompt-engineering**, and **alignment** issues.</description><pubDate>Fri, 15 Nov 2024 02:50:42 GMT</pubDate><category>anthropic</category><category>openai</category><category>langchain</category><category>meta-ai-fair</category><category>claude-3-sonnet</category><category>gpt-4</category><category>gemini-1.5</category><category>claude-3.5-sonnet</category><category>richardmcngo</category><category>andrewyng</category><category>philschmid</category><category>benchmarking</category><category>prompt-engineering</category><category>rag</category><category>visuotactile-perception</category><category>ai-governance</category><category>theoretical-alignment</category><category>ethical-alignment</category><category>jailbreak-robustness</category><category>model-releases</category><category>alignment</category></item><item><title>Common Corpus: 2T Open Tokens with Provenance</title><link>https://news.smol.ai/issues/24-11-13-ainews-common-corpus-2t-open-tokens-with-provenance/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-13-ainews-common-corpus-2t-open-tokens-with-provenance/</guid><description>**Pleais** via **Huggingface** released **Common Corpus**, the largest fully open multilingual dataset with over **2 trillion tokens** including detailed **provenance information**. They also introduced **OCRonos-Vintage**, a **124M-parameter OCR correction model** that efficiently fixes digitization errors on CPU and GPU, unlocking knowledge from PDFs. On AI tools, **LangChainAI** launched **Prompt Canvas** for collaborative **prompt engineering**, while **DeepSeek** released **JanusFlow 1.3B**, a unified multimodal LLM integrating autoregressive and rectified flow models for enhanced **image understanding** and **generation**. **Alibaba Cloud** announced **Qwen2.5-Coder**, a code-focused LLM with advanced coding capabilities, and **Claude 3.5 Sonnet** was highlighted for superior code generation. Discussions on **quantization challenges** and **scaling laws for precision** by **Tim Dettmers** and others emphasized the impact of low-precision training on model scalability and inference efficiency. *&quot;Scaling Laws for Precision&quot;* paper insights and alternative efficiency methods were also noted.</description><pubDate>Thu, 14 Nov 2024 01:54:53 GMT</pubDate><category>pleais</category><category>huggingface</category><category>langchainai</category><category>deepseek</category><category>alibaba</category><category>anthropic</category><category>qwen-2.5-coder</category><category>claude-3.5-sonnet</category><category>janusflow-1.3b</category><category>ocronos-vintage</category><category>tim-dettmers</category><category>tom-doerr</category><category>omarsar0</category><category>swyx</category><category>madiator</category><category>reach_vb</category><category>provenance</category><category>ocr</category><category>multilingual-datasets</category><category>prompt-engineering</category><category>multimodality</category><category>image-generation</category><category>code-generation</category><category>quantization</category><category>model-scaling</category><category>inference-efficiency</category></item><item><title>BitNet was a lie?</title><link>https://news.smol.ai/issues/24-11-12-ainews-bitnet-was-a-lie/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-12-ainews-bitnet-was-a-lie/</guid><description>**Scaling laws for quantization** have been modified by a group led by Chris Re, analyzing over **465 pretraining runs** and finding benefits plateau at FP6 precision. Lead author **Tanishq Kumar** highlights that longer training and more data increase sensitivity to quantization, explaining challenges with models like **Llama-3**. **Tim Dettmers**, author of QLoRA, warns that the era of efficiency gains from low-precision quantization is ending, signaling a shift from scaling to optimizing existing resources. Additionally, **Alibaba** announced **Qwen 2.5-Coder-32B-Instruct**, which matches or surpasses **GPT-4o** on coding benchmarks, and open-source initiatives like **DeepEval** for LLM testing are gaining traction.</description><pubDate>Wed, 13 Nov 2024 01:36:06 GMT</pubDate><category>sambanova</category><category>alibaba</category><category>hugging-face</category><category>qwen-2.5-coder-32b-instruct</category><category>gpt-4o</category><category>llama-3</category><category>tanishq-kumar</category><category>tim-dettmers</category><category>quantization</category><category>scaling-laws</category><category>model-efficiency</category><category>fine-tuning</category><category>model-performance</category><category>code-generation</category><category>open-source</category><category>unit-testing</category><category>ci-cd</category></item><item><title>FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI</title><link>https://news.smol.ai/issues/24-11-11-ainews-frontiermath-a-benchmark-for-evaluating-advanced-mathematical-reasoning-in-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-11-ainews-frontiermath-a-benchmark-for-evaluating-advanced-mathematical-reasoning-in-ai/</guid><description>**Epoch AI** collaborated with over **60 leading mathematicians** to create the **FrontierMath benchmark**, a fresh set of hundreds of original math problems with easy-to-verify answers, aiming to challenge current AI models. The benchmark reveals that all tested models, including **o1**, perform poorly, highlighting the difficulty of complex problem-solving and **Moravec&apos;s paradox** in AI. Key AI developments include the introduction of **Mixture-of-Transformers (MoT)**, a sparse multi-modal transformer architecture reducing computational costs, and improvements in **Chain-of-Thought (CoT) prompting** through incorrect reasoning and explanations. Industry news covers **OpenAI** acquiring the **chat.com** domain, **Microsoft** launching the **Magentic-One agent framework**, **Anthropic** releasing **Claude 3.5 Haiku** outperforming **gpt-4o** on some benchmarks, and **xAI** securing **150MW grid power** with support from **Elon Musk** and **Trump**. **LangChain AI** introduced new tools including a **Financial Metrics API**, **Document GPT** with PDF upload and Q&amp;A, and **LangPost** AI agent for LinkedIn posts. **xAI** also demonstrated the **Grok Engineer** compatible with OpenAI and Anthropic APIs for code generation.</description><pubDate>Tue, 12 Nov 2024 01:33:12 GMT</pubDate><category>epoch-ai</category><category>openai</category><category>microsoft</category><category>anthropic</category><category>x-ai</category><category>langchainai</category><category>o1</category><category>claude-3.5-haiku</category><category>gpt-4o</category><category>karpathy</category><category>philschmid</category><category>adcock_brett</category><category>dylan522p</category><category>benchmarking</category><category>math</category><category>moravecs-paradox</category><category>mixture-of-experts</category><category>chain-of-thought</category><category>agent-framework</category><category>financial-metrics-api</category><category>pdf-processing</category><category>few-shot-learning</category><category>code-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-11-08-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-08-ainews-not-much-happened-today/</guid><description>This week in AI news, **Anthropic** launched **Claude Sonnet 3.5**, enabling desktop app control via natural language. **Microsoft** introduced **Magentic-One**, a multi-agent system built on the **AutoGen framework**. **OpenCoder** was unveiled as an AI-powered code cookbook for large language models. **SambaNova** is sponsoring a hackathon with prizes up to **$5000** for building real-time AI agents. **Sophiamyang** announced new **Batch and Moderation APIs** with **50% lower cost** and multi-dimensional harmful text detection. Open-source tools like **Infisical** for secret management, **CrewAI** for autonomous agent orchestration, and **Crawlee** for web scraping were released. Research highlights include **SCIPE** for error analysis in LLM chains, **Context Refinement Agent** for improved retrieval-augmented generation, and **MemGPT** for managing LLM memory. The week also saw a legal win for **OpenAI** in the RawStory copyright case, affirming that facts used in LLM training are not copyrightable.</description><pubDate>Fri, 08 Nov 2024 23:16:39 GMT</pubDate><category>anthropic</category><category>microsoft</category><category>sambanova</category><category>openai</category><category>langchain</category><category>llamaindex</category><category>claude-3.5-sonnet</category><category>opencoder</category><category>sophiamyang</category><category>tom_doerr</category><category>omarsar0</category><category>_akhaliq</category><category>andrewyng</category><category>giffmana</category><category>multi-agent-systems</category><category>natural-language-interfaces</category><category>batch-processing</category><category>harmful-content-detection</category><category>secret-management</category><category>retrieval-augmented-generation</category><category>error-analysis</category><category>memory-management</category><category>web-scraping</category><category>autonomous-agents</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-11-07-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-07-ainews-not-much-happened-today/</guid><description>This week in AI news highlights **Ollama 0.4** supporting **Meta&apos;s Llama 3.2 Vision** models (11B and 90B), with applications like handwriting recognition. **Self-Consistency Preference Optimization (ScPO)** was introduced to improve model consistency without human labels. Discussions on **model scaling**, **neural networks resurgence**, and **AMD&apos;s multi-GPU bandwidth** challenges were noted. The importance of **skip connections** in **Transformers** was emphasized. In healthcare, **less regulation plus AI** could revolutionize disease treatment and aging. Tools like **LlamaParse** and **Gemini** aid automated resume insights. **Gitpod Flex** demonstrated zero-trust architecture for secure development environments. Research includes surveys on **Small Language Models (SLMs)**, **number understanding** in LLMs, and **DTrOCR** using a **GPT-2 decoder** for OCR. Multi-agent systems in prediction markets were discussed by **TogetherCompute** and **LangChainAI**. Community events include **NeurIPS Happy Hour**, **NLP seminars**, and courses on **Agent Memory** with LLMs as operating systems.</description><pubDate>Fri, 08 Nov 2024 01:01:09 GMT</pubDate><category>meta-ai-fair</category><category>ollama</category><category>amd</category><category>llamaindex</category><category>gemini</category><category>gitpod</category><category>togethercompute</category><category>langchainai</category><category>weights-biases</category><category>stanfordnlp</category><category>deeplearningai</category><category>llama-3-2-vision</category><category>gpt-2</category><category>bindureddy</category><category>fstichler</category><category>stasbekman</category><category>jxmnop</category><category>bindureddy</category><category>omarsar0</category><category>giffmana</category><category>rajammanabrolu</category><category>model-scaling</category><category>neural-networks</category><category>multi-gpu-support</category><category>skip-connections</category><category>transformers</category><category>healthcare-ai</category><category>automated-recruitment</category><category>zero-trust-security</category><category>small-language-models</category><category>numerical-processing</category><category>chain-of-thought</category><category>optical-character-recognition</category><category>multi-agent-systems</category><category>agent-memory</category><category>interactive-language-learning</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-11-06-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-06-ainews-not-much-happened-today/</guid><description>**Grok Beta** surpasses **Llama 3.1 70B** in intelligence but is less competitive due to its pricing at **$5/1M input tokens** and **$15/1M output tokens**. **Defense Llama**, developed with **Meta AI** and **Scale AI**, targets American national security applications. **SWE-Kit**, an open-source framework, supports building customizable AI software engineers compatible with **Llama 3**, **ChatGPT**, and **Claude**. **LangChainAI** and **Weights &amp; Biases** integrate to improve retrievers and reduce hallucinations in **RAG applications** using **Gemini**. **Perplexity AI** offers enhanced election tracking tools for the **2024 elections**, including live state results and support for **Claude 3.5 Haiku**. **AI Talk** launched featuring discussions on Chinese AI labs with guests from **Qwen**. Memes highlight **Elon Musk** and humorous AI coding mishaps.</description><pubDate>Thu, 07 Nov 2024 02:54:09 GMT</pubDate><category>meta-ai-fair</category><category>scale-ai</category><category>anthropic</category><category>perplexity-ai</category><category>langchainai</category><category>weights-biases</category><category>qwen</category><category>grok-beta</category><category>llama-3-1-70b</category><category>claude-3-5-haiku</category><category>claude-3-opus</category><category>llama-3</category><category>chatgpt</category><category>gemini</category><category>alexandr_wang</category><category>svpino</category><category>aravsrinivas</category><category>bindureddy</category><category>teortaxestex</category><category>jessechenglyu</category><category>junyang-lin</category><category>cte_junior</category><category>jerryjliu0</category><category>pricing</category><category>national-security</category><category>defense</category><category>open-source</category><category>agentic-ai</category><category>retrieval-augmented-generation</category><category>election-predictions</category><category>real-time-updates</category><category>annotation</category><category>ai-ecosystem</category><category>memes</category><category>humor</category></item><item><title>Tencent&apos;s Hunyuan-Large claims to beat DeepSeek-V2 and Llama3-405B with LESS Data</title><link>https://news.smol.ai/issues/24-11-05-ainews-tencents-hunyuan-large-claims-to-beat-deepseek-v2-and-llama3-405b-with-less-data/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-05-ainews-tencents-hunyuan-large-claims-to-beat-deepseek-v2-and-llama3-405b-with-less-data/</guid><description>**Tencent** released a notable &gt;300B parameter MoE model pretrained on **7T tokens**, including **1.5T synthetic data** generated via **Evol-Instruct**. The model introduces novel techniques like &quot;recycle routing&quot; and expert-specific learning rates, alongside a compute-efficient scaling law for MoE active parameters. However, its custom license restricts use in the EU and by companies with over 100M MAU, and it avoids China-sensitive queries. Meanwhile, **Anthropic** launched **Claude 3.5 Haiku**, now available on multiple platforms, praised for intelligence and speed but criticized for a **10x price increase**. **Meta** opened **Llama AI** to the U.S. defense sector, and a **Llama Impact Hackathon** offers a **$15K prize** for projects using **Llama 3.1 &amp; 3.2 Vision**. **LlamaIndex** released a React chat UI component with Tailwind CSS and LLM backend integrations. The **MLX LM** model advances text generation speed and efficiency with KV cache quantization.</description><pubDate>Wed, 06 Nov 2024 06:22:40 GMT</pubDate><category>tencent</category><category>anthropic</category><category>meta-ai-fair</category><category>togethercompute</category><category>llamaindex</category><category>claude-3.5-haiku</category><category>llama-3-1</category><category>llama-3-2</category><category>mlx-lm</category><category>mixture-of-experts</category><category>synthetic-data</category><category>model-scaling</category><category>model-architecture</category><category>model-optimization</category><category>kv-cache-quantization</category><category>react</category><category>fine-tuning</category><category>scaling-laws</category><category>model-efficiency</category><category>model-deployment</category><category>multimodality</category></item><item><title>OpenAI beats Anthropic to releasing Speculative Decoding</title><link>https://news.smol.ai/issues/24-11-04-ainews-openai-beats-anthropic-to-releasing-speculative-decoding/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-04-ainews-openai-beats-anthropic-to-releasing-speculative-decoding/</guid><description>**Prompt lookup** and **Speculative Decoding** techniques are gaining traction with implementations from **Cursor**, **Fireworks**, and teased features from **Anthropic**. **OpenAI** has introduced faster response times and file edits with these methods, offering about **50%** efficiency improvements. The community is actively exploring AI engineering use cases with these advancements. Recent updates highlight progress from companies like **NVIDIA**, **OpenAI**, **Anthropic**, **Microsoft**, **Boston Dynamics**, and **Meta**. Key technical insights include CPU inference capabilities, multimodal retrieval-augmented generation (RAG), and neural network fundamentals. New AI products include fully AI-generated games and advanced content generation tools. Challenges in AI research labs such as bureaucracy and resource allocation were also discussed, alongside AI safety and governance concerns.</description><pubDate>Tue, 05 Nov 2024 02:51:39 GMT</pubDate><category>openai</category><category>anthropic</category><category>nvidia</category><category>microsoft</category><category>boston-dynamics</category><category>meta-ai-fair</category><category>runway</category><category>elevenlabs</category><category>etched</category><category>osmo</category><category>physical-intelligence</category><category>langchain</category><category>claude-3-sonnet</category><category>mrt5</category><category>adcock_brett</category><category>vikhyatk</category><category>dair_ai</category><category>rasbt</category><category>bindureddy</category><category>teortaxestex</category><category>svpino</category><category>c_valenzuelab</category><category>davidsholz</category><category>speculative-decoding</category><category>prompt-lookup</category><category>cpu-inference</category><category>multimodality</category><category>retrieval-augmented-generation</category><category>neural-networks</category><category>optimization</category><category>ai-safety</category><category>governance</category><category>model-architecture</category><category>inference-economics</category><category>content-generation</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-11-01-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-01-ainews-not-much-happened-today/</guid><description>**ChatGPT Search** was launched by **Sam Altman**, who called it his favorite feature since ChatGPT&apos;s original launch, doubling his usage. Comparisons were made between ChatGPT Search and **Perplexity** with improvements noted in Perplexity&apos;s web navigation. **Google** introduced a &quot;Grounding&quot; feature in the Gemini API &amp; AI Studio enabling Gemini models to access real-time web information. Despite Gemini&apos;s leaderboard performance, developer adoption lags behind **OpenAI** and **Anthropic**. **SmolLM2**, a new small, powerful on-device language model, outperforms **Meta&apos;s Llama 3.2 1B**. A **Claude** desktop app was released for Mac and Windows. **Meta AI** announced robotics advancements including Meta Sparsh, Meta Digit 360, and Meta Digit Plexus. **Stable Diffusion 3.5 Medium**, a 2B parameter model with a permissive license, was released. Insights on AGI development suggest initial inferiority but rapid improvement. **Anthropic** advocates for early targeted AI regulation. Discussions on ML specialization predict training will concentrate among few companies, while inference becomes commoditized. New AI tools include **Suno AI Personas** for music creation, **PromptQL** for natural language querying over data, and **Agent S** for desktop task automation. Humor was shared about Python environment upgrades.</description><pubDate>Fri, 01 Nov 2024 20:59:45 GMT</pubDate><category>openai</category><category>anthropic</category><category>google</category><category>meta-ai-fair</category><category>suno-ai</category><category>perplexity-ai</category><category>smollm2</category><category>llama-3-2</category><category>stable-diffusion-3.5</category><category>claude-3.5-sonnet</category><category>gemini</category><category>sam-altman</category><category>akhaliq</category><category>arav-srinivas</category><category>labenz</category><category>loubnabenallal1</category><category>alexalbert</category><category>fchollet</category><category>stasbekman</category><category>svpino</category><category>rohanpaul_ai</category><category>hamelhusain</category><category>on-device-ai</category><category>model-performance</category><category>robotics</category><category>multimodality</category><category>ai-regulation</category><category>model-releases</category><category>natural-language-processing</category><category>prompt-engineering</category><category>agentic-ai</category><category>ai-application</category><category>model-optimization</category></item><item><title>The AI Search Wars Have Begun — SearchGPT, Gemini Grounding, and more</title><link>https://news.smol.ai/issues/24-11-01-ainews-the-ai-search-wars-have-begun-searchgpt-gemini-grounding-and-more/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-11-01-ainews-the-ai-search-wars-have-begun-searchgpt-gemini-grounding-and-more/</guid><description>**ChatGPT** launched its search functionality across all platforms using a fine-tuned version of **GPT-4o** with synthetic data generation and distillation from **o1-preview**. This feature includes a Chrome extension promoted by **Sam Altman** but has issues with hallucinations. The launch coincides with **Gemini** introducing Search Grounding after delays. Notably, **The New York Times** is not a partner due to a lawsuit against **OpenAI**. The AI search competition intensifies with consumer and B2B players like **Perplexity** and **Glean**. Additionally, **Claude 3.5 Sonnet** achieved a new benchmark record on SWE-bench Verified, and a new hallucination evaluation benchmark, SimpleQA, was introduced. Other highlights include the **Universal-2** speech-to-text model with 660M parameters and **HOVER**, a neural whole-body controller for humanoid robots trained in NVIDIA Isaac simulation. AI hedge fund teams using **LangChain** and **LangGraph** were also showcased. The news is sponsored by the RAG++ course featuring experts from **Weights &amp; Biases**, **Cohere**, and **Weaviate**.</description><pubDate>Fri, 01 Nov 2024 07:04:02 GMT</pubDate><category>openai</category><category>google</category><category>gemini</category><category>nyt</category><category>perplexity-ai</category><category>glean</category><category>nvidia</category><category>langchain</category><category>langgraph</category><category>weights-biases</category><category>cohere</category><category>weaviate</category><category>gpt-4o</category><category>o1-preview</category><category>claude-3.5-sonnet</category><category>universal-2</category><category>sam-altman</category><category>alexalbert__</category><category>_jasonwei</category><category>svpino</category><category>drjimfan</category><category>virattt</category><category>fine-tuning</category><category>synthetic-data</category><category>distillation</category><category>hallucinations</category><category>benchmarking</category><category>speech-to-text</category><category>robotics</category><category>neural-networks</category><category>ai-agents</category></item><item><title>Creating a LLM-as-a-Judge</title><link>https://news.smol.ai/issues/24-10-30-ainews-creating-a-llm-as-a-judge/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-30-ainews-creating-a-llm-as-a-judge/</guid><description>**Anthropic** released details on Claude 3.5 SWEBench+SWEAgent, while **OpenAI** introduced SimpleQA and **DeepMind** launched NotebookLM. **Apple** announced new M4 Macbooks, and a new SOTA image model, Recraft v3, emerged. Hamel Husain presented a detailed 6,000-word treatise on creating LLM judges using a method called **critique shadowing** to align LLMs with domain experts, addressing the problem of untrusted and unused data in AI teams. The workflow involves expert-reviewed datasets and iterative prompt refinement. Additionally, **Zep** introduced a temporal knowledge graph memory layer to improve AI agent memory and reduce hallucinations. **Anthropic** also integrated Claude 3.5 Sonnet with GitHub Copilot, expanding access to Copilot Chat users.</description><pubDate>Wed, 30 Oct 2024 23:17:27 GMT</pubDate><category>anthropic</category><category>openai</category><category>deepmind</category><category>apple</category><category>zep</category><category>perplexity-ai</category><category>github</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>notebooklm</category><category>simpleqa</category><category>recraft-v3</category><category>hamel-husain</category><category>swyx</category><category>critique-shadowing</category><category>llm-judging</category><category>domain-experts</category><category>dataset-creation</category><category>prompt-engineering</category><category>error-analysis</category><category>temporal-knowledge-graphs</category><category>memory-layer</category><category>ai-agent-memory</category><category>hallucination-reduction</category><category>integration</category></item><item><title>GitHub Copilot Strikes Back</title><link>https://news.smol.ai/issues/24-10-29-ainews-github-copilot-strikes-back/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-29-ainews-github-copilot-strikes-back/</guid><description>**GitHub&apos;s tenth annual Universe conference** introduced the **Multi-model Copilot** featuring **Anthropic&apos;s Claude 3.5 Sonnet**, **Google&apos;s Gemini 1.5 Pro**, and **OpenAI&apos;s o1-preview** models in a new picker UI, allowing developers to choose from multiple companies&apos; models. The event also showcased **GitHub Spark**, an AI-native tool for building natural language applications with deployment-free hosting and integrated model prompting. Additionally, GitHub updated its Copilot Workspace with new agents and security Autofix features. **Weights &amp; Biases** launched Weave with multimodal observability supporting audio, text, and images, integrating the OpenAI Realtime API. Twitter recaps highlighted **tinygrad&apos;s** codebase optimization and discussions on GenAI adoption and **Gemini Flash-8B&apos;s** cost efficiency at **$0.0375 per million tokens**.</description><pubDate>Wed, 30 Oct 2024 01:05:11 GMT</pubDate><category>github</category><category>anthropic</category><category>google-deepmind</category><category>openai</category><category>weights-biases</category><category>claude-3-5-sonnet</category><category>gemini-1.5-pro</category><category>o1-preview</category><category>gemini-flash-8b</category><category>cassidy-williams</category><category>fchollet</category><category>rohanpaul_ai</category><category>jxmnop</category><category>model-picker-ui</category><category>multi-model-integration</category><category>natural-language-applications</category><category>deployment-free-hosting</category><category>model-prompting</category><category>multimodal-observability</category><category>audio-tracing</category><category>codebase-optimization</category><category>price-performance-ratio</category></item><item><title>not much happened this weekend</title><link>https://news.smol.ai/issues/24-10-28-ainews-not-much-happened-this-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-28-ainews-not-much-happened-this-weekend/</guid><description>**Moondream**, a **1.6b vision language model**, secured seed funding, highlighting a trend in moon-themed tiny models alongside **Moonshine** (27-61m ASR model). **Claude 3.5 Sonnet** was used for AI Twitter recaps. Discussions included **pattern recognition** vs. **intelligence** in **LLMs**, **reinforcement learning** for prompt optimization, and **NotebookLlama**, an open-source **NotebookLM** variant using **LLaMA models** for tasks like **text-to-speech**. Advances in **model optimization** with **async-TP** in **PyTorch** for **tensor parallelism** and hyperparameter tuning were noted. **Mini-Omni 2** demonstrated multimodal capabilities across **image**, **audio**, and **text** for voice conversations with emphasis on **modal alignment** and **multimodal fine-tuning**. AI productivity tools like an **AI email writer** and **LlamaCloud**-based research assistants were introduced. Emphasis on practical skill development and privacy-conscious AI tool usage with **Llama3-8B** was highlighted. Generative AI tools such as **#AIPythonforBeginners** and **GenAI Agents** with **LangGraph** were shared. Business insights covered rapid execution in AI product development and emerging AI-related job roles. Challenges in enterprise-grade text-to-SQL and advanced retrieval methods were discussed with tutorials on **RAG** applications using **LangChain** and **MongoDB**.</description><pubDate>Mon, 28 Oct 2024 22:27:43 GMT</pubDate><category>moondream</category><category>openai</category><category>anthropic</category><category>hugging-face</category><category>mistral-ai</category><category>google-deepmind</category><category>langchain</category><category>deepmind</category><category>microsoft</category><category>claude-3.5-sonnet</category><category>llama-3</category><category>llama-3-8b</category><category>notebookllama</category><category>min-omni-2</category><category>amanda-askell</category><category>philschmid</category><category>stasbekman</category><category>francois-fleuret</category><category>mervenoyann</category><category>reach_vb</category><category>dzhng</category><category>aravsrinivas</category><category>sama</category><category>lateinteraction</category><category>andrew-y-ng</category><category>bindureddy</category><category>jerryjliu0</category><category>pattern-recognition</category><category>reinforcement-learning</category><category>prompt-optimization</category><category>text-to-speech</category><category>model-optimization</category><category>tensor-parallelism</category><category>hyperparameters</category><category>multimodal</category><category>modal-alignment</category><category>multimodal-fine-tuning</category><category>ai-productivity</category><category>privacy</category><category>generative-ai</category><category>rag</category><category>retrieval-augmentation</category><category>enterprise-text-to-sql</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-25-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-25-ainews-not-much-happened-today/</guid><description>**Liquid AI** held a launch event introducing new foundation models. **Anthropic** shared follow-up research on social bias and feature steering with their &quot;Golden Gate Claude&quot; feature. **Cohere** released multimodal Embed 3 embeddings models following Aya Expanse. There was misinformation about **GPT-5/Orion** debunked by **Sam Altman**. **Meta AI FAIR** announced **Open Materials 2024** with new models and datasets for inorganic materials discovery using the EquiformerV2 architecture. **Anthropic AI** demonstrated feature steering to balance social bias and model capabilities. **NVIDIA**&apos;s **Llama-3.1-Nemotron-70B** ranked highly on the Arena leaderboard with style control. **Perplexity AI** expanded to 100M weekly queries with new finance and reasoning modes. **LangChain** emphasized real application integration with interactive frame interpolation. **Kestra** highlighted scalable event-driven workflows with open-source YAML-based orchestration. **OpenFLUX** optimized inference speed by doubling it through guidance LoRA training. Discussions on AI safety included trust dynamics between humans and AI, economic impacts of AI automation, and the White House AI National Security memo addressing cyber and biological risks. **LlamaIndex** showcased knowledge-backed agents for enhanced AI applications.</description><pubDate>Sat, 26 Oct 2024 00:52:03 GMT</pubDate><category>liquid-ai</category><category>anthropic</category><category>cohere</category><category>openai</category><category>meta-ai-fair</category><category>nvidia</category><category>perplexity-ai</category><category>langchain</category><category>kestra</category><category>ostrisai</category><category>llamaindex</category><category>llama-3.1-nemotron-70b</category><category>golden-gate-claude</category><category>embed-3</category><category>sam-altman</category><category>lmarena_ai</category><category>aravsrinivas</category><category>svpino</category><category>richardmcngo</category><category>ajeya_cotra</category><category>tamaybes</category><category>danhendrycks</category><category>jerryjliu0</category><category>feature-steering</category><category>social-bias</category><category>multimodality</category><category>model-optimization</category><category>workflow-orchestration</category><category>inference-speed</category><category>event-driven-workflows</category><category>knowledge-backed-agents</category><category>economic-impact</category><category>ai-national-security</category><category>trust-dynamics</category></item><item><title>s{imple|table|calable} Consistency Models</title><link>https://news.smol.ai/issues/24-10-24-ainews-simpleortableorcalable-consistency-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-24-ainews-simpleortableorcalable-consistency-models/</guid><description>**Model distillation** significantly accelerates diffusion models, enabling near real-time image generation with only 1-4 sampling steps, as seen in **BlinkShot** and **Flux Schnell**. Research led by **Yang Song** introduced **simplified continuous-time consistency models (sCMs)**, achieving under 10% FID difference in just 2 steps and scaling up to **1.5B parameters** for higher quality. On AI hardware, **Tesla** is deploying a **50k H100 cluster** potentially capable of completing **GPT-4** training in under three weeks, while **Cerebras Systems** set a new inference speed record on **Llama 3.1 70B** with their wafer-scale AI chips. **Stability AI** released **Stable Diffusion 3.5** and its Turbo variant, and **Cohere** launched new multilingual models supporting **23 languages** with state-of-the-art performance. **LangChain** also announced ecosystem updates.</description><pubDate>Fri, 25 Oct 2024 02:36:02 GMT</pubDate><category>stability-ai</category><category>tesla</category><category>cerebras</category><category>cohere</category><category>langchain</category><category>llama-3-70b</category><category>llama-3-405b</category><category>llama-3-1</category><category>stable-diffusion-3.5</category><category>gpt-4</category><category>yang-song</category><category>model-distillation</category><category>diffusion-models</category><category>continuous-time-consistency-models</category><category>image-generation</category><category>ai-hardware</category><category>inference-speed</category><category>multilingual-models</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-23-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-23-ainews-not-much-happened-today/</guid><description>**Anthropic** released upgraded **Claude 3.5 Sonnet** and **Claude 3.5 Haiku** models featuring a new **computer use capability** that allows interaction with computer interfaces via screenshots and actions like mouse movement and typing. The **Claude 3.5 Sonnet** achieved state-of-the-art coding performance on SWE-bench Verified with a **49% score**, surpassing OpenAI&apos;s **o1-preview**. **Anthropic** focuses on teaching general computer skills rather than task-specific tools, with expected rapid improvements. Other releases include **Mochi 1**, an open-source video generation model, **Stable Diffusion 3.5** with Large and Medium variants, and **Embed 3** by **Cohere**, a multimodal embedding model for text and image search. **KerasHub** was launched by **François Chollet**, unifying KerasNLP and KerasCV with 37 pretrained models. Microsoft introduced the **Differential Transformer** to reduce attention noise via differential attention maps, and research on transformer attention layers was shared by **Rasbt**.</description><pubDate>Thu, 24 Oct 2024 00:39:59 GMT</pubDate><category>anthropic</category><category>openai</category><category>cohere</category><category>microsoft</category><category>claude-3.5-sonnet</category><category>claude-3.5-haiku</category><category>o1-preview</category><category>mochi-1</category><category>stable-diffusion-3.5</category><category>embed-3</category><category>kerashub</category><category>differential-transformer</category><category>alexalbert</category><category>fchollet</category><category>rasbt</category><category>computer-use</category><category>coding-performance</category><category>video-generation</category><category>fine-tuning</category><category>multimodality</category><category>transformers</category><category>attention-mechanisms</category><category>model-optimization</category></item><item><title>Claude 3.5 Sonnet (New) gets Computer Use</title><link>https://news.smol.ai/issues/24-10-22-ainews-claude-35-sonnet-new-gets-computer-use/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-22-ainews-claude-35-sonnet-new-gets-computer-use/</guid><description>**Anthropic** announced new Claude 3.5 models: **3.5 Sonnet** and **3.5 Haiku**, improving coding performance significantly, with Sonnet topping several coding benchmarks like **Aider** and **Vectara**. The new **Computer Use API** enables controlling computers via vision, scoring notably higher than other AI systems, showcasing progress in AI-driven computer interaction. **Zep** launched a cloud edition for AI agents memory management, highlighting challenges in **multimodal memory**. The update also mentions **Llama 3.1** and **Nemotron** models from **NVIDIA**.</description><pubDate>Wed, 23 Oct 2024 02:08:12 GMT</pubDate><category>anthropic</category><category>zep</category><category>nvidia</category><category>claude-3.5-sonnet</category><category>claude-3.5-haiku</category><category>llama-3.1</category><category>nemotron</category><category>philschmid</category><category>swyx</category><category>coding</category><category>benchmarks</category><category>computer-use</category><category>vision</category><category>multimodal-memory</category><category>model-updates</category><category>ai-integration</category></item><item><title>DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing</title><link>https://news.smol.ai/issues/24-10-21-ainews-docetl-agentic-query-rewriting-and-evaluation-for-complex-document-processing/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-21-ainews-docetl-agentic-query-rewriting-and-evaluation-for-complex-document-processing/</guid><description>**UC Berkeley&apos;s EPIC lab** introduces innovative LLM data operators with projects like **LOTUS** and **DocETL**, focusing on effective programming and computation over large data corpora. This approach contrasts GPU-rich big labs like **Deepmind** and **OpenAI** with GPU-poor compound AI systems. **Microsoft** open-sourced **BitNet b1.58**, a 1-bit ternary parameter LLM enabling **4-20x faster training** and on-device inference at human reading speeds. Nvidia released **Llama-3.1-Nemotron-70B-Instruct**, a fine-tuned open-source model outperforming **GPT-4o** and **Claude-3.5-sonnet**. These developments highlight advances in **model-optimization**, **on-device-ai**, and **fine-tuning**.</description><pubDate>Tue, 22 Oct 2024 00:04:21 GMT</pubDate><category>uc-berkeley</category><category>deepmind</category><category>openai</category><category>microsoft</category><category>nvidia</category><category>archetype-ai</category><category>boston-dynamics</category><category>toyota-research</category><category>google</category><category>adobe</category><category>openai</category><category>mistral</category><category>tesla</category><category>meta-ai-fair</category><category>bitnet-b1.58</category><category>llama-3.1-nemotron-70b-instruct</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>rohanpaul_ai</category><category>adcock_brett</category><category>david-patterson</category><category>model-optimization</category><category>on-device-ai</category><category>fine-tuning</category><category>large-corpus-processing</category><category>gpu-acceleration</category><category>frameworks</category><category>model-benchmarking</category></item><item><title>DeepSeek Janus and Meta SpiRit-LM: Decoupled Image and Expressive Voice Omnimodality</title><link>https://news.smol.ai/issues/24-10-18-ainews-deepseek-janus-and-meta-spirit-lm-decoupled-image-and-expressive-voice-omnimodality/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-18-ainews-deepseek-janus-and-meta-spirit-lm-decoupled-image-and-expressive-voice-omnimodality/</guid><description>**DeepSeek Janus** and **Meta SpiRit-LM** are two notable multimodality AI models recently released, showcasing advances in image generation and speech synthesis respectively. DeepSeek Janus separates vision encoders for image understanding and generation, achieving better results in both tasks. Meta&apos;s SpiRit-LM introduces an expressive speech and writing model generating pitch and style units, improving over standard TTS. Additionally, **W&amp;B Weave** offers comprehensive LLM observability and multimodality fine-tuning tools. Industry updates include Nvidia&apos;s Nemotron 70b model underperforming, Meta open-sourcing Movie Gen Bench for media generation benchmarking, Perplexity launching internal search with multi-step reasoning, and Anthropic updating Claude apps. Open source progress includes Hugging Face&apos;s gradient accumulation fix in transformers and advocacy for open source AI to prevent Big Tech dominance. *&quot;Model merging for combining skills of multiple models&quot;* is also highlighted.</description><pubDate>Fri, 18 Oct 2024 22:46:38 GMT</pubDate><category>deepseek</category><category>meta-ai-fair</category><category>wandb</category><category>nvidia</category><category>anthropic</category><category>hugging-face</category><category>perplexity-ai</category><category>nemotron-70b</category><category>claude</category><category>claude-3.5-sonnet</category><category>gpt-4o</category><category>bindureddy</category><category>aravsrinivas</category><category>danielhanchen</category><category>clementdelangue</category><category>cwolferesearch</category><category>multimodality</category><category>image-generation</category><category>speech-synthesis</category><category>fine-tuning</category><category>model-merging</category><category>benchmarking</category><category>open-source</category><category>model-optimization</category><category>reinforcement-learning</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-17-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-17-ainews-not-much-happened-today/</guid><description>**Answer.ai** launched **fastdata**, a synthetic data generation library using `claudette` and Tencent&apos;s Billion Persona paper. **NotebookLM** became customizable, and **Motherduck** introduced notable LLMs in SQL implementations. **Perplexity** and **Dropbox** announced competitors to **Glean**. **OpenAI** unveiled audio chat completions priced at 24 cents per minute. **Meta AI** released **Llama 3.1**, powering Lenovo AI Now&apos;s on-device agent. **Yi-Lightning** model ranked #6 globally, surpassing **GPT-4o**. **Zyphra AI** released the large **Zyda-2** dataset with 5 trillion tokens. **François Chollet** clarified transformer architecture as set-processing, not sequence-processing. Research suggests memorization aids LLM reasoning. **Anthropic** updated its Responsible Scaling Policy for AI safety. Tools like **Perplexity Finance**, **Open Canvas** by **LangChain**, and **AlphaCodium** code generation tool were highlighted. Approximately $500 million was raised for AI agent startups, with ongoing discussions on AI&apos;s job market impact. Combining prompt caching with the Batches API can yield a 95% discount on **Claude 3.5 Sonnet** tokens.</description><pubDate>Fri, 18 Oct 2024 01:13:21 GMT</pubDate><category>answer-ai</category><category>tencent</category><category>notebooklm</category><category>motherduck</category><category>perplexity</category><category>dropbox</category><category>openai</category><category>meta-ai-fair</category><category>yi-ai</category><category>zyphra-ai</category><category>anthropic</category><category>langchain</category><category>openai</category><category>claudette</category><category>llama-3-1</category><category>yi-lightning</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>fchollet</category><category>aravsrinivas</category><category>svpino</category><category>swyx</category><category>synthetic-data</category><category>fine-tuning</category><category>sql</category><category>audio-processing</category><category>on-device-ai</category><category>dataset-release</category><category>transformer</category><category>llm-reasoning</category><category>ai-safety</category><category>code-generation</category><category>ai-pricing</category><category>ai-job-market</category></item><item><title>Did Nvidia&apos;s Nemotron 70B train on test?</title><link>https://news.smol.ai/issues/24-10-16-ainews-did-nvidias-nemotron-70b-train-on-test/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-16-ainews-did-nvidias-nemotron-70b-train-on-test/</guid><description>**NVIDIA&apos;s Nemotron-70B** model has drawn scrutiny despite strong benchmark performances on **Arena Hard**, **AlpacaEval**, and **MT-Bench**, with some standard benchmarks like **GPQA** and **MMLU Pro** showing no improvement over the base **Llama-3.1-70B**. The new **HelpSteer2-Preference dataset** improves some benchmarks with minimal losses elsewhere. Meanwhile, **Mistral** released **Ministral 3B and 8B** models featuring **128k context length** and outperforming **Llama-3.1** and **GPT-4o** on various benchmarks under the **Mistral Commercial License**. **NVIDIA&apos;s Nemotron 70B** also surpasses **GPT-4o** and **Claude-3.5-Sonnet** on key benchmarks using **RLHF (REINFORCE)** training. Additionally, **Zep** introduced **Graphiti**, an open-source temporal knowledge graph memory layer for AI agents, built on **Neo4j**.</description><pubDate>Thu, 17 Oct 2024 00:44:43 GMT</pubDate><category>nvidia</category><category>mistral-ai</category><category>hugging-face</category><category>zep</category><category>nemotron-70b</category><category>llama-3.1-70b</category><category>llama-3.1</category><category>ministral-3b</category><category>ministral-8b</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>reach_vb</category><category>philschmid</category><category>swyx</category><category>benchmarking</category><category>reinforcement-learning</category><category>reward-models</category><category>temporal-knowledge-graphs</category><category>memory-layers</category><category>context-windows</category><category>model-releases</category><category>open-source</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-15-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-15-ainews-not-much-happened-today/</guid><description>**Vertical SaaS agents** are gaining rapid consensus as the future of AI applications, highlighted by **Decagon&apos;s $100m funding** and **Sierra&apos;s $4b round**. **OpenAI alumni** are actively raising venture capital and forming new startups, intensifying competition in the AI market. **Demis Hassabis** celebrated the **Nobel Prize** recognition for **AlphaFold2**, a breakthrough in protein structure prediction. Advances in AI models include techniques like **LoRA projectors** and **annealing on high-quality data**, while discussions emphasize the need for **high-bandwidth sensory inputs** beyond language for common sense learning. New methods like **LoLCATs** aim to optimize transformer models such as **Llama** and **Mistral** for efficiency. Ethical concerns about AI agents performing harmful tasks remain under investigation. The AI community continues to explore model evaluation challenges and optimization frameworks like **LPZero** for neural architecture search.</description><pubDate>Tue, 15 Oct 2024 21:33:05 GMT</pubDate><category>openai</category><category>decagon</category><category>sierra</category><category>togethercompute</category><category>llama</category><category>mistral</category><category>mira-murati</category><category>demis-hassabis</category><category>clement-delangue</category><category>john-o-whitaker</category><category>yann-lecun</category><category>francois-chollet</category><category>ajeya-cotra</category><category>rohan-paul</category><category>adcock-brett</category><category>vertical-saas</category><category>funding</category><category>protein-structure-prediction</category><category>lora</category><category>self-supervised-learning</category><category>model-optimization</category><category>neural-architecture-search</category><category>model-evaluation</category><category>ethics</category><category>transformers</category><category>multi-agent-systems</category><category>long-context</category></item><item><title>Not much (in AI) happened this weekend</title><link>https://news.smol.ai/issues/24-10-14-ainews-not-much-in-ai-happened-this-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-14-ainews-not-much-in-ai-happened-this-weekend/</guid><description>**OpenAI** introduced an &quot;edit this area&quot; feature for image generation, praised by **Sam Altman**. **Yann LeCun** highlighted a NYU paper improving pixel generation with feature prediction loss using pre-trained visual encoders like DINOv2. Long-context LLMs such as **llama-3.1-8b** and **llama-3.2** variants now support up to **131k tokens**, offering alternatives to RAG systems. **Bindu Reddy** announced AI agents capable of building and deploying code from English instructions, signaling AI&apos;s replacement of SQL and potential impact on Python. SpaceX&apos;s successful **Starship rocket catch** was celebrated by **Andrej Karpathy** and others, with **Soumith Chintala** praising SpaceX&apos;s efficient, low-bureaucracy research approach. Privacy concerns arose from **Harvard** students&apos; AI glasses, I-XRAY, which can reveal personal information. **Meta AI FAIR**&apos;s Movie Gen model advances media foundation models with high-quality text-to-image and video generation, including synced audio. Humanoid robots like **Ameca** and **Azi** now engage in expressive conversations using **ChatGPT**. **xAI** rapidly deployed **100K Nvidia H100 GPUs** in 19 days, with CEO Jensen Huang commending Elon Musk. Leading AI research labs compared include **Meta-FAIR**, **Google DeepMind**, and **Microsoft Research**. Skepticism about LLM intelligence was voiced by **Sam Pino**, emphasizing limitations in novel problem-solving despite strong memorization.</description><pubDate>Mon, 14 Oct 2024 22:52:37 GMT</pubDate><category>openai</category><category>meta-ai-fair</category><category>google-deepmind</category><category>microsoft</category><category>x-ai</category><category>spacex</category><category>harvard</category><category>nvidia</category><category>llama-3.1-8b</category><category>llama-3.2</category><category>chatgpt</category><category>movie-gen</category><category>sam-altman</category><category>yann-lecun</category><category>rasbt</category><category>bindureddy</category><category>andrej-karpathy</category><category>soumithchintala</category><category>svpino</category><category>adcock_brett</category><category>rohanpaul_ai</category><category>long-context</category><category>feature-prediction-loss</category><category>ai-agents</category><category>privacy</category><category>text-to-video</category><category>text-to-image</category><category>humanoid-robots</category><category>gpu-deployment</category><category>media-foundation-models</category><category>ai-research-labs</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-11-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-11-ainews-not-much-happened-today/</guid><description>**Rhymes AI** released **Aria**, a new **25.3B** parameter multimodal MoE model supporting text, code, image, and video with a **64k token context window** and Apache-2.0 license. **OpenAI**&apos;s **o1-preview** and **o1-mini** models show consistent improvement over **Anthropic** and **Google Gemini 1.5 Pro/Flash** on long context RAG benchmarks up to **128k tokens**, while **Google Gemini 1.5** models excel at extreme context lengths up to **2 million tokens**. **Meta AI** expanded rollout to 21 countries with new language support but remains unavailable in the EU. The one-year anniversary of **SWE-bench** benchmark for software engineering tasks was celebrated, alongside the introduction of SWE-bench Multimodal. New AI tools include **OxyCopilot** by Oxylabs for web scraping, **Taipy** for Python-based production apps, and **Latitude** for prompt engineering. Industry insights highlight changing AI funding dynamics and OpenAI&apos;s strategic focus on consumer products like ChatGPT. *&quot;all recaps done by Claude 3.5 Sonnet, best of 4 runs.&quot;*</description><pubDate>Fri, 11 Oct 2024 23:00:43 GMT</pubDate><category>rhymes-ai</category><category>openai</category><category>anthropic</category><category>google</category><category>meta-ai-fair</category><category>oxylabs</category><category>aria</category><category>o1-preview</category><category>o1-mini</category><category>gemini-1.5-pro</category><category>gemini-1.5-flash</category><category>gemini-1.5</category><category>claude-3.5-sonnet</category><category>mervenoyann</category><category>osanseviero</category><category>dbrxmosaicai</category><category>ylecun</category><category>ofirpress</category><category>clefourrier</category><category>omarsar0</category><category>rohanpaul_ai</category><category>svpino</category><category>finbarrtimbers</category><category>_philschmid</category><category>multimodality</category><category>mixture-of-experts</category><category>long-context</category><category>retrieval-augmented-generation</category><category>benchmarking</category><category>software-engineering</category><category>llm-evaluation</category><category>prompt-engineering</category><category>web-scraping</category><category>python</category><category>production-applications</category></item><item><title>State of AI 2024</title><link>https://news.smol.ai/issues/24-10-10-ainews-state-of-ai-2024/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-10-ainews-state-of-ai-2024/</guid><description>**Nathan Benaich&apos;s State of AI Report** in its 7th year provides a comprehensive overview of AI research and industry trends, including highlights like **BitNet** and the synthetic data debate. **Cerebras** is preparing for an IPO, reflecting growth in AI compute. A hackathon hosted by **Daily** and the **Pipecat** community focuses on conversational voice AI and multimodal experiences with $20,000 in prizes. Nobel Prizes in Physics and Chemistry were awarded for AI research: **Geoffrey Hinton** and **John Hopfield** for neural networks and statistical mechanics, and **Demis Hassabis**, **John Jumper**, and **David Baker** for AlphaFold and protein structure prediction. **Meta** released **Llama 3.2** with multimodal capabilities, accompanied by educational resources and performance updates. *&quot;This recognizes the impact of deep neural networks on society&quot;* and *&quot;tremendous impact of AlphaFold and ML-powered protein structure prediction&quot;* were noted by experts.</description><pubDate>Thu, 10 Oct 2024 22:35:38 GMT</pubDate><category>cerebras</category><category>daily</category><category>pipecat</category><category>meta-ai-fair</category><category>anthropic</category><category>llama-3-2</category><category>bitnet</category><category>geoffrey-hinton</category><category>john-hopfield</category><category>demis-hassabis</category><category>john-jumper</category><category>david-baker</category><category>multimodality</category><category>synthetic-data</category><category>protein-structure-prediction</category><category>neural-networks</category><category>statistical-mechanics</category><category>conversational-ai</category><category>voice-ai</category><category>hackathon</category><category>ipo</category><category>model-release</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-10-09-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-09-ainews-not-much-happened-today/</guid><description>**Geoffrey Hinton** and **John Hopfield** won the **Nobel Prize in Physics** for foundational work on neural networks linking AI and physics. **Meta AI** introduced a **13B parameter audio generation model** as part of Meta Movie Gen for video-synced audio. **Anthropic** launched the **Message Batches API** enabling asynchronous processing of up to 10,000 queries at half the cost. **Together Compute** released **Flux Schnell**, a free model for 3 months. New techniques like **PrefixQuant** quantization and **Prompt Caching** for low-latency inference were highlighted by **rohanpaul_ai**. **LangGraph** added long-term memory support for persistent document storage. **Hex-LLM** framework was introduced for TPU-based low-cost, high-throughput LLM serving from Hugging Face models. Discussions on AI safety emphasized gender equality in science, and concerns about premature AI regulation by media and Hollywood were raised.</description><pubDate>Thu, 10 Oct 2024 01:02:45 GMT</pubDate><category>meta-ai-fair</category><category>anthropic</category><category>togethercompute</category><category>hugging-face</category><category>flux-schnell</category><category>geoffrey-hinton</category><category>john-hopfield</category><category>demis-hassabis</category><category>rohanpaul_ai</category><category>svpino</category><category>hwchase17</category><category>shreyar</category><category>philschmid</category><category>mmitchell_ai</category><category>bindureddy</category><category>audio-generation</category><category>quantization</category><category>prompt-caching</category><category>long-term-memory</category><category>llm-serving-framework</category><category>hallucination-detection</category><category>ai-safety</category><category>ai-governance</category></item><item><title>The AI Nobel Prize</title><link>https://news.smol.ai/issues/24-10-08-ainews-the-ai-nobel-prize/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-08-ainews-the-ai-nobel-prize/</guid><description>**Geoff Hinton** and **John Hopfield** won the **Nobel Prize in Physics** for their work on **Artificial Neural Networks**. The award citation spans **14 pages** highlighting their contributions. **Zep** released a new community edition of their low-latency memory layer for AI agents, emphasizing knowledge graphs for memory. At OpenAI&apos;s DevDay, new features like real-time voice API, vision model fine-tuning, and prompt caching with a **50% discount** on reused tokens were introduced. **Anthropic&apos;s Claude 3.5 Sonnet** was recognized as the best model currently. **Reka AI Labs** updated their **Reka Flash** model with enhanced multimodal and function calling capabilities. The **GOT (Generic OCR Transformer)** achieved **98.79% accuracy** on OCR benchmarks. Discussions on open-source AI models highlighted their role in fostering competition and decentralization. Software development insights included the importance of Single Sign-On (SSO), thorough testing, and AI-assisted coding workflows. Ethical and societal topics covered critiques of tax policies and the appointment of France&apos;s first Minister of AI.</description><pubDate>Wed, 09 Oct 2024 01:33:48 GMT</pubDate><category>openai</category><category>anthropic</category><category>reka-ai</category><category>zep</category><category>claude-3.5-sonnet</category><category>reka-flash</category><category>got</category><category>geoff-hinton</category><category>john-hopfield</category><category>philschmid</category><category>alexalbert</category><category>mervenoyann</category><category>clementdelangue</category><category>svpino</category><category>bindureddy</category><category>ylecun</category><category>rohanpaul_ai</category><category>artificial-neural-networks</category><category>nobel-prize</category><category>knowledge-graphs</category><category>memory-layers</category><category>real-time-voice-api</category><category>vision</category><category>fine-tuning</category><category>prompt-caching</category><category>multimodality</category><category>function-calling</category><category>ocr</category><category>open-source</category><category>single-sign-on</category><category>software-testing</category><category>ai-assisted-coding</category><category>ai-ethics</category></item><item><title>not much happened this weekend</title><link>https://news.smol.ai/issues/24-10-07-ainews-not-much-happened-this-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-07-ainews-not-much-happened-this-weekend/</guid><description>**AI news from 10/4/2024 to 10/7/2024** highlights several developments: **OpenAI&apos;s o1-preview** shows strong performance on complex tasks but struggles with simpler ones, while **Claude 3.5 Sonnet** can match its reasoning through advanced prompting techniques. **Meta** introduced **Movie Gen**, a cutting-edge media foundation model for text-to-video generation and editing. **Reka** updated their 21B Flash Model with temporal video understanding, native audio, and tool use capabilities. Interest grows in &quot;open o1&quot; reproductions focusing on prompting and finetuning, with **Entropix** exploring entropy-based sampling. **LangChainAI** demonstrated a Retrieval Agent for complex Q&amp;A, and synthetic data generation research surveyed 417 models. A resurgence in RNNs shows efficient parallel training making them competitive with Transformers. Biologically-inspired AI safety approaches were also noted. *&quot;A quiet weekend and air conditioning is all you need.&quot;*</description><pubDate>Tue, 08 Oct 2024 02:36:09 GMT</pubDate><category>openai</category><category>meta-ai-fair</category><category>reka</category><category>langchainai</category><category>entropix</category><category>o1-preview</category><category>claude-3.5-sonnet</category><category>21b-flash-model</category><category>lex-fridman</category><category>imrat</category><category>jjitsev</category><category>giffmana</category><category>_philschmid</category><category>karpathy</category><category>rasbt</category><category>adcock_brett</category><category>glennko</category><category>rohanpaul_ai</category><category>labenz</category><category>prompting-techniques</category><category>finetuning</category><category>entropy-based-sampling</category><category>temporal-understanding</category><category>native-audio</category><category>tool-use</category><category>instruction-chaining</category><category>multimodality</category><category>retrieval-augmented-generation</category><category>synthetic-data-generation</category><category>rnn</category><category>parallel-training</category><category>biologically-inspired-ai-safety</category><category>text-to-video-generation</category><category>video-editing</category></item><item><title>Contextual Document Embeddings: `cde-small-v1`</title><link>https://news.smol.ai/issues/24-10-04-ainews-contextual-document-embeddings-cde-small-v1/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-04-ainews-contextual-document-embeddings-cde-small-v1/</guid><description>**Meta** announced a new text-to-video model, **Movie Gen**, claiming superior adaptation of **Llama 3** to video generation compared to OpenAI&apos;s Sora Diffusion Transformers, though no release is available yet. Researchers Jack Morris and Sasha Rush introduced the **cde-small-v1** model with a novel **contextual batching** training technique and **contextual embeddings**, achieving strong performance with only **143M parameters**. **OpenAI** launched Canvas, a collaborative interface for ChatGPT with synthetic data training. **Google DeepMind** welcomed Tim Brooks to work on video generation and world simulators. Google released **Gemini 1.5 Flash-8B**, improving cost and rate limits with algorithmic efficiency.</description><pubDate>Sat, 05 Oct 2024 01:38:06 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>google-deepmind</category><category>weights-biases</category><category>togethercompute</category><category>llama-3</category><category>cde-small-v1</category><category>gemini-1.5-flash-8b</category><category>chatgpt</category><category>jack-morris</category><category>sasha-rush</category><category>tim-brooks</category><category>demis-hassabis</category><category>karina-nguyen</category><category>contextual-embeddings</category><category>contextual-batching</category><category>video-generation</category><category>synthetic-data</category><category>model-efficiency</category><category>training-techniques</category><category>rag</category><category>algorithmic-efficiency</category></item><item><title>Canvas: OpenAI&apos;s answer to Claude Artifacts</title><link>https://news.smol.ai/issues/24-10-03-ainews-canvas-openais-answer-to-claude-artifacts/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-03-ainews-canvas-openais-answer-to-claude-artifacts/</guid><description>**OpenAI** released **Canvas**, an enhanced writing and coding tool based on **GPT-4o**, featuring inline suggestions, seamless editing, and a collaborative environment. Early feedback compares it to **Cursor** and **Claude Artifacts**, noting strengths and some execution issues. OpenAI also sponsors **Marijn Haverbeke**, creator of **ProseMirror** and **CodeMirror**, which are used in Canvas. The integration involved training a detector to trigger Canvas appropriately, achieving **83% accuracy** in correct triggers. Unlike Claude Artifacts, Canvas currently lacks Mermaid Diagrams and HTML preview support. Additionally, **Daily** is sponsoring a **$20,000** voice AI hackathon in San Francisco, highlighting voice AI as a key emerging skill.</description><pubDate>Thu, 03 Oct 2024 23:22:37 GMT</pubDate><category>openai</category><category>cursor_ai</category><category>daily</category><category>gpt-4o</category><category>claude-artifacts</category><category>marijn-haverbeke</category><category>karina-nguyen</category><category>vicente-silveira</category><category>swyx</category><category>inline-suggestions</category><category>collaborative-editing</category><category>code-editing</category><category>model-training</category><category>model-integration</category><category>feature-detection</category><category>accuracy-evaluation</category><category>voice-ai</category><category>hackathon</category><category>open-source-libraries</category></item><item><title>Not much technical happened today</title><link>https://news.smol.ai/issues/24-10-02-ainews-not-much-technical-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-02-ainews-not-much-technical-happened-today/</guid><description>**OpenAI** announced raising **$6.6B** in new funding at a **$157B valuation**, with ChatGPT reaching *250M weekly active users*. **Poolside** raised **$500M** to advance AGI development. **LiquidAI** introduced three new MoE models (1B, 3B, 40B) with a **32k context window** and efficient token handling. **OpenAI** released Whisper V3 Turbo, an open-source multilingual model with significant speed improvements. **Meta AI FAIR** is hiring research interns focusing on **LLM reasoning, alignment, synthetic data, and novel architectures**. **Cohere** partnered with Fujitsu to launch Takane, a custom Japanese model. Technical discussions included challenges in **LoRA fine-tuning**, **float8 quantization** in Keras, and new tools like **create-llama** for agent templates. Industry commentary raised concerns about AI development priorities and highlighted freelancing opportunities in AI.</description><pubDate>Wed, 02 Oct 2024 22:45:37 GMT</pubDate><category>openai</category><category>poolside</category><category>liquidai</category><category>perplexity-ai</category><category>meta-ai-fair</category><category>cohere</category><category>fujitsu</category><category>whisper-v3-turbo</category><category>llama-3</category><category>llamaindex</category><category>nick-turley</category><category>arav-srinivas</category><category>francois-fleuret</category><category>finbarr-timbers</category><category>lewtun</category><category>francois-chollet</category><category>jerry-j-liu</category><category>mmitchell-ai</category><category>jxnlco</category><category>mixture-of-experts</category><category>context-windows</category><category>model-optimization</category><category>fine-tuning</category><category>quantization</category><category>model-training</category><category>alignment</category><category>synthetic-data</category><category>model-architecture</category><category>agentic-ai</category></item><item><title>OpenAI Realtime API and other Dev Day Goodies</title><link>https://news.smol.ai/issues/24-10-01-ainews-openai-realtime-api-and-other-dev-day-goodies/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-10-01-ainews-openai-realtime-api-and-other-dev-day-goodies/</guid><description>**OpenAI** launched the **gpt-4o-realtime-preview** Realtime API featuring text and audio token processing with pricing details and future plans including vision and video support. The API supports voice activity detection modes, function calling, and ephemeral sessions with auto-truncation for context limits. Partnerships with **LiveKit**, **Agora**, and **Twilio** enhance audio components and AI virtual agent voice calls. Additionally, OpenAI introduced vision fine-tuning with only 100 examples improving mapping accuracy for **Grab** and RPA success for **Automat**. Model distillation and prompt caching features were also announced, including free eval inference for users opting to share data.</description><pubDate>Wed, 02 Oct 2024 06:06:20 GMT</pubDate><category>openai</category><category>livekit</category><category>agora</category><category>twilio</category><category>grab</category><category>automat</category><category>gpt-4o-realtime-preview</category><category>gpt-4o</category><category>voice-activity-detection</category><category>function-calling</category><category>ephemeral-sessions</category><category>auto-truncation</category><category>vision-fine-tuning</category><category>model-distillation</category><category>prompt-caching</category><category>audio-processing</category></item><item><title>Liquid Foundation Models: A New Transformers alternative + AINews Pod 2</title><link>https://news.smol.ai/issues/24-09-30-ainews-liquid-foundation-models-a-new-transformers-alternative-ainews-pod-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-30-ainews-liquid-foundation-models-a-new-transformers-alternative-ainews-pod-2/</guid><description>**Liquid.ai** emerged from stealth with three subquadratic foundation models demonstrating superior efficiency compared to state space models and Apple’s on-device and server models, backed by a $37M seed round. **Meta AI** announced **Llama 3.2** with multimodal vision-enabled models and lightweight text-only variants for mobile. **Google DeepMind** introduced production-ready **Gemini-1.5-Pro-002** and **Gemini-1.5-Flash-002** models with improved pricing and rate limits, alongside **AlphaChip**, an AI-driven chip design system using reinforcement learning for rapid superhuman layouts. **OpenAI** enhanced ChatGPT Plus and Teams with Advanced Voice Mode featuring Custom Instructions, Memory, and new nature-inspired voices. California Governor vetoed SB-1047 AI regulation bill, celebrated by AI community figures like **ylecun** and **svpino** as a win for open-source AI. Google upgraded **NotebookLM** with audio overviews supporting YouTube and audio files, turning documents into AI-generated podcasts. *&quot;Open source in AI is thriving,&quot;* noted **ylecun**, highlighting 1 million models on Github and HuggingFace.</description><pubDate>Tue, 01 Oct 2024 01:34:19 GMT</pubDate><category>liquid-ai</category><category>meta-ai-fair</category><category>google-deepmind</category><category>openai</category><category>llama-3-2</category><category>gemini-1.5-pro-002</category><category>gemini-1.5-flash-002</category><category>ylecun</category><category>svpino</category><category>reinforcement-learning</category><category>multimodality</category><category>model-efficiency</category><category>foundation-models</category><category>audio-processing</category><category>model-deployment</category><category>open-source</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-09-27-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-27-ainews-not-much-happened-today/</guid><description>**Meta** released **Llama 3.2**, including lightweight 1B and 3B models for on-device AI with capabilities like summarization and retrieval-augmented generation. **Molmo**, a new multimodal model, was introduced with a large dense captioning dataset. **Google DeepMind** announced **AlphaChip**, an AI-driven chip design method improving TPU and CPU designs. **Hugging Face** surpassed 1 million free public models, highlighting the value of smaller specialized models. Discussions covered challenges in scaling RAG applications, the future of on-device AI running ChatGPT-level models, reliability issues in larger LLMs, and new Elo benchmarking accepted at NeurIPS 2024. AI ethics and regulation topics included free speech responsibilities and California&apos;s SB-1047 bill potentially affecting open-source AI. *&quot;AlphaChip transformed computer chip design,&quot;* and *&quot;ChatGPT-level AI on mobile devices predicted within a year.&quot;*</description><pubDate>Fri, 27 Sep 2024 21:53:11 GMT</pubDate><category>meta-ai-fair</category><category>google-deepmind</category><category>hugging-face</category><category>llama-3-2</category><category>llama-3</category><category>molmo</category><category>demis-hassabis</category><category>clementdelangue</category><category>svpino</category><category>awnihannun</category><category>osanseviero</category><category>omarsar0</category><category>sarahookr</category><category>ylecun</category><category>on-device-ai</category><category>multimodality</category><category>chip-design</category><category>retrieval-augmented-generation</category><category>rag</category><category>benchmarking</category><category>reliability</category><category>ai-regulation</category><category>free-speech</category><category>pytorch-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-09-26-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-26-ainews-not-much-happened-today/</guid><description>**Meta AI** released **Llama 3.2** models including **1B, 3B text-only** and **11B, 90B vision** variants with **128K token context length** and adapter layers for image-text integration. These models outperform competitors like **Gemma 2** and **Phi 3.5-mini**, and are supported on major platforms including **AWS, Azure, and Google Cloud**. **OpenAI CTO Mira Murati** announced her departure. **Allen AI** released **Molmo**, an open-source multimodal model family outperforming proprietary systems. **Google** improved **Gemini 1.5** with Flash and Pro models. **Meta** showcased **Project Orion AR glasses** and hinted at a **Quest 3S** priced at $300. Discussions covered new benchmarks for multimodal models, model optimization, and AI safety and alignment.</description><pubDate>Thu, 26 Sep 2024 22:52:11 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>allenai</category><category>google-deepmind</category><category>llama-3-2</category><category>llama-3</category><category>gemma-2</category><category>phi-3-5-mini</category><category>claude-3-haiku</category><category>gpt-4o-mini</category><category>molmo</category><category>gemini-1.5</category><category>gemini</category><category>mira-murati</category><category>demis-hassabis</category><category>ylecun</category><category>sama</category><category>multimodality</category><category>model-optimization</category><category>benchmarks</category><category>ai-safety</category><category>model-distillation</category><category>pruning</category><category>adapter-layers</category><category>open-source-models</category><category>performance</category><category>context-windows</category></item><item><title>Llama 3.2: On-device 1B/3B, and Multimodal 11B/90B (with AI2 Molmo kicker)</title><link>https://news.smol.ai/issues/24-09-25-ainews-llama-32-on-device-1b3b-and-multimodal-11b90b-with-ai2-molmo-kicker/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-25-ainews-llama-32-on-device-1b3b-and-multimodal-11b90b-with-ai2-molmo-kicker/</guid><description>**Meta** released **Llama 3.2** with new multimodal versions including **3B** and **20B** vision adapters on a frozen Llama 3.1, showing competitive performance against **Claude Haiku** and **GPT-4o-mini**. **AI2** launched multimodal **Molmo 72B** and **7B** models outperforming Llama 3.2 in vision tasks. Meta also introduced new **128k-context 1B and 3B models** competing with **Gemma 2** and **Phi 3.5**, with collaborations hinted with **Qualcomm**, **Mediatek**, and **Arm** for on-device AI. The release includes a **9 trillion token count** for Llama 1B and 3B. Partner launches include **Ollama**, **Together AI** offering free 11B model access, and **Fireworks AI**. Additionally, a new **RAG++ course** from **Weights &amp; Biases**, **Cohere**, and **Weaviate** offers systematic evaluation and deployment guidance for retrieval-augmented generation systems based on extensive production experience.</description><pubDate>Wed, 25 Sep 2024 23:54:30 GMT</pubDate><category>meta-ai-fair</category><category>ai2</category><category>qualcomm</category><category>mediatek</category><category>arm</category><category>ollama</category><category>together-ai</category><category>fireworks-ai</category><category>weights-biases</category><category>cohere</category><category>weaviate</category><category>llama-3-2</category><category>llama-3-1</category><category>claude-3-haiku</category><category>gpt-4o-mini</category><category>molmo-72b</category><category>molmo-7b</category><category>gemma-2</category><category>phi-3-5</category><category>llama-3-2-vision</category><category>llama-3-2-3b</category><category>llama-3-2-20b</category><category>mira-murati</category><category>daniel-han</category><category>multimodality</category><category>vision</category><category>context-windows</category><category>quantization</category><category>model-release</category><category>tokenization</category><category>model-performance</category><category>model-optimization</category><category>rag</category><category>model-training</category><category>instruction-following</category></item><item><title>ChatGPT Advanced Voice Mode</title><link>https://news.smol.ai/issues/24-09-24-ainews-chatgpt-advanced-voice-mode/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-24-ainews-chatgpt-advanced-voice-mode/</guid><description>**OpenAI** rolled out **ChatGPT Advanced Voice Mode** with 5 new voices and improved accent and language support, available widely in the US. Ahead of rumored updates for **Llama 3** and **Claude 3.5**, **Gemini Pro** saw a significant price cut aligning with the new intelligence frontier pricing. **OpenAI&apos;s o1-preview model** showed promising planning task performance with 52.8% accuracy on Randomized Mystery Blocksworld. **Anthropic** is rumored to release a new model, generating community excitement. **Qwen 2.5** was released with models up to 32B parameters and support for 128K tokens, matching GPT-4 0613 benchmarks. Research highlights include PlanBench evaluation of o1-preview, OpenAI&apos;s release of a multilingual MMMLU dataset covering 14 languages, and RAGLAB framework standardizing Retrieval-Augmented Generation research. New AI tools include PDF2Audio for converting PDFs to audio, an open-source AI starter kit for local model deployment, and **Moshi**, a speech-based AI assistant from Kyutai. Industry updates feature **Scale AI** nearing $1B ARR with 4x YoY growth and **Together Compute&apos;s** enterprise platform offering faster inference and cost reductions. Insights from **Sam Altman**&apos;s blog post were also shared.</description><pubDate>Wed, 25 Sep 2024 01:31:24 GMT</pubDate><category>openai</category><category>anthropic</category><category>scale-ai</category><category>togethercompute</category><category>kyutai-labs</category><category>o1-preview</category><category>qwen-2.5</category><category>llama-3</category><category>claude-3.5</category><category>sam-altman</category><category>omarsar0</category><category>bindureddy</category><category>rohanpaul_ai</category><category>_philschmid</category><category>alexandr_wang</category><category>svpino</category><category>ylecun</category><category>_akhaliq</category><category>voice-synthesis</category><category>planning</category><category>multilingual-datasets</category><category>retrieval-augmented-generation</category><category>open-source</category><category>speech-assistants</category><category>enterprise-ai</category><category>price-cuts</category><category>benchmarking</category><category>model-performance</category></item><item><title>a calm before the storm</title><link>https://news.smol.ai/issues/24-09-23-ainews-a-calm-before-the-storm/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-23-ainews-a-calm-before-the-storm/</guid><description>**Anthropic** is raising funds at a valuation up to **$40 billion** ahead of anticipated major releases. **OpenAI** launched new reasoning models **o1** and **o1-mini**, with increased rate limits and a multilingual MMLU benchmark. **Alibaba** released the open-source **Qwen2.5** model supporting 29+ languages, showing competitive performance to **gpt-4** at lower cost. **Microsoft** and **Blackrock** plan to invest **$30 billion** in AI data centers, with **Groq** partnering with Aramco to build the world&apos;s largest AI inference center. Robotics advances include Disney Research and ETH Zurich&apos;s diffusion-based motion generation for robots and Pudu Robotics&apos; semi-humanoid robot. Slack and Microsoft introduced AI-powered agents integrated into their platforms. Research highlights include long-context scaling for **llama-2-70b** using Dual Chunk Attention and KV cache quantization enabling 1 million token context on **llama-7b** models.</description><pubDate>Mon, 23 Sep 2024 23:33:49 GMT</pubDate><category>anthropic</category><category>openai</category><category>alibaba</category><category>microsoft</category><category>blackrock</category><category>groq</category><category>aramco</category><category>disney</category><category>eth-zurich</category><category>pudu-robotics</category><category>slack</category><category>o1</category><category>o1-mini</category><category>qwen2.5</category><category>gpt-4</category><category>llama-2-70b</category><category>llama-7b</category><category>adcock_brett</category><category>philschmid</category><category>rohanpaul_ai</category><category>jvnixon</category><category>kateclarktweets</category><category>sama</category><category>long-context</category><category>kv-cache-quantization</category><category>diffusion-models</category><category>reinforcement-learning</category><category>robotics</category><category>ai-integration</category><category>multilinguality</category><category>model-benchmarking</category><category>model-performance</category><category>model-optimization</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-09-20-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-20-ainews-not-much-happened-today/</guid><description>**Anthropic** introduced a RAG technique called Contextual Retrieval that reduces retrieval failure rates by 67% using prompt caching. **Meta** is teasing multimodal **Llama 3** ahead of Meta Connect. **OpenAI** is hiring for a multi-agent research team focusing on improved AI reasoning with their **o1 models**, which have sparked mixed reactions. **DeepSeek 2.5** is noted as a cost-effective alternative to **GPT-4** and **Claude 3.5 sonnet**. New models like **3DTopia-XL** for 3D asset generation and **CogVideoX** for image-to-video conversion were highlighted. Techniques to boost reasoning by re-reading questions and combining retrieval with prompt caching were shared. Industry insights emphasize the necessity of AI adoption in enterprises and the disruption of traditional ML businesses. Tools like **LangChainAI&apos;s LangGraph Templates** and **LlamaIndex&apos;s LlamaParse Premium** enhance agentic applications and multimodal content extraction. Discussions on LLM evals and caching highlight production challenges and improvements. *&quot;Companies not allowing developers to use AI are unlikely to succeed&quot;* was a key sentiment.</description><pubDate>Sat, 21 Sep 2024 01:37:46 GMT</pubDate><category>anthropic</category><category>meta-ai-fair</category><category>openai</category><category>deepseek-ai</category><category>llamaindex</category><category>langchainai</category><category>llama-3</category><category>o1</category><category>deepseek-2.5</category><category>gpt-4</category><category>claude-3.5-sonnet</category><category>3dtopia-xl</category><category>cogvideox</category><category>retrieval-augmented-generation</category><category>prompt-caching</category><category>multimodality</category><category>multi-agent-systems</category><category>reasoning</category><category>diffusion-models</category><category>image-to-video</category><category>prompting</category><category>enterprise-ai</category><category>agentic-ai</category><category>long-context</category><category>model-evaluation</category><category>caching</category><category>model-cost-efficiency</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-09-19-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-19-ainews-not-much-happened-today/</guid><description>**OpenAI&apos;s o1-preview and o1-mini models** lead benchmarks in Math, Hard Prompts, and Coding. **Qwen 2.5 72B** model shows strong performance close to **GPT-4o**. **DeepSeek-V2.5** tops Chinese LLMs, rivaling **GPT-4-Turbo-2024-04-09**. **Microsoft&apos;s GRIN MoE** achieves good results with 6.6B active parameters. **Moshi voice model** from Kyutai Labs runs locally on Apple Silicon Macs. **Perplexity app** introduces voice mode with push-to-talk. **LlamaCoder** by Together.ai uses **Llama 3.1 405B** for app generation. **Google DeepMind&apos;s Veo** is a new generative video model for YouTube Shorts. The **2024 ARC-AGI competition** increases prize money and plans a university tour. A survey on model merging covers 50+ papers for LLM alignment. The **Kolmogorov–Arnold Transformer (KAT)** paper proposes replacing MLP layers with KAN layers for better expressiveness. **Hugging Face Hub** integrates with **Google Cloud Vertex AI Model Garden** for easier open-source model deployment. **Agent.ai** is introduced as a professional network for AI agents. *&quot;Touching grass is all you need.&quot;*</description><pubDate>Fri, 20 Sep 2024 01:00:56 GMT</pubDate><category>openai</category><category>qwen</category><category>deepseek-ai</category><category>microsoft</category><category>kyutai-labs</category><category>perplexity-ai</category><category>together-ai</category><category>meta-ai-fair</category><category>google-deepmind</category><category>hugging-face</category><category>google</category><category>anthropic</category><category>o1-preview</category><category>o1-mini</category><category>qwen-2.5</category><category>gpt-4o</category><category>deepseek-v2.5</category><category>gpt-4-turbo-2024-04-09</category><category>grin</category><category>llama-3-1-405b</category><category>veo</category><category>kat</category><category>hyung-won-chung</category><category>noam-brown</category><category>bindureddy</category><category>akhaliq</category><category>karpathy</category><category>aravsrinivas</category><category>fchollet</category><category>cwolferesearch</category><category>philschmid</category><category>labenz</category><category>ylecun</category><category>benchmarking</category><category>math</category><category>coding</category><category>instruction-following</category><category>model-merging</category><category>model-expressiveness</category><category>moe</category><category>voice</category><category>voice-models</category><category>generative-video</category><category>competition</category><category>open-source</category><category>model-deployment</category><category>ai-agents</category></item><item><title>o1 destroys Lmsys Arena, Qwen 2.5, Kyutai Moshi release</title><link>https://news.smol.ai/issues/24-09-18-ainews-o1-destroys-lmsys-arena-qwen-25-kyutai-moshi-release/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-18-ainews-o1-destroys-lmsys-arena-qwen-25-kyutai-moshi-release/</guid><description>**OpenAI&apos;s o1-preview** model has achieved a milestone by fully matching top daily AI news stories without human intervention, consistently outperforming other models like **Anthropic**, **Google**, and **Llama 3** in vibe check evaluations. **OpenAI** models dominate the top 4 slots on **LMsys** benchmarks, with rate limits increasing to **500-1000 requests per minute**. In open source, **Alibaba&apos;s Qwen 2.5** suite surpasses **Llama 3.1** at the 70B scale and updates its closed **Qwen-Plus** models to outperform **DeepSeek V2.5** but still lag behind leading American models. **Kyutai Moshi** released its open weights realtime voice model featuring a unique streaming neural architecture with an &quot;inner monologue.&quot; **Weights &amp; Biases** introduced **Weave**, an LLM observability toolkit that enhances experiment tracking and evaluation, turning prompting into a more scientific process. The news also highlights upcoming events like the **WandB LLM-as-judge hackathon** in San Francisco. *&quot;o1-preview consistently beats out our vibe check evals&quot;* and *&quot;OpenAI models are gradually raising rate limits by the day.&quot;*</description><pubDate>Wed, 18 Sep 2024 21:51:26 GMT</pubDate><category>openai</category><category>anthropic</category><category>google</category><category>alibaba</category><category>deepseek</category><category>kyutai</category><category>weights-biases</category><category>mistral-ai</category><category>o1-preview</category><category>o1-mini</category><category>qwen-2.5</category><category>qwen-plus</category><category>llama-3-1</category><category>deepseek-v2.5</category><category>sama</category><category>guillaumelample</category><category>chain-of-thought</category><category>multimodality</category><category>model-benchmarking</category><category>model-performance</category><category>streaming-neural-architecture</category><category>llm-observability</category><category>experiment-tracking</category><category>rate-limiting</category></item><item><title>nothing much happened today</title><link>https://news.smol.ai/issues/24-09-17-ainews-nothing-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-17-ainews-nothing-much-happened-today/</guid><description>**OpenAI&apos;s o1 model** faces skepticism about open-source replication due to its extreme restrictions and unique training advances like RL on CoT. **ChatGPT-4o** shows significant performance improvements across benchmarks. **Llama-3.1-405b** fp8 and bf16 versions perform similarly with cost benefits for fp8. A new open-source benchmark &quot;Humanity&apos;s Last Exam&quot; offers $500K in prizes to challenge LLMs. Model merging benefits from neural network sparsity and linear mode connectivity. Embedding-based toxic prompt detection achieves high accuracy with low compute. **InstantDrag** enables fast, optimization-free drag-based image editing. **LangChain v0.3** releases with improved dependency management. Automated code review tool **CodeRabbit** adapts to team coding styles. Visual search advances integrate multimodal data for better product search. Experts predict AI will be default software by 2030.</description><pubDate>Wed, 18 Sep 2024 00:27:31 GMT</pubDate><category>openai</category><category>lmsys</category><category>scale-ai</category><category>cognition</category><category>langchain</category><category>qdrant</category><category>rohanpaul_ai</category><category>o1</category><category>chatgpt-4o</category><category>llama-3-1-405b</category><category>denny_zhou</category><category>svpino</category><category>alexandr_wang</category><category>cwolferesearch</category><category>rohanpaul_ai</category><category>_akhaliq</category><category>kylebrussell</category><category>reinforcement-learning</category><category>model-merging</category><category>embedding-models</category><category>toxicity-detection</category><category>image-editing</category><category>dependency-management</category><category>automated-code-review</category><category>visual-search</category><category>benchmarking</category></item><item><title>a quiet weekend</title><link>https://news.smol.ai/issues/24-09-16-ainews-a-quiet-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-16-ainews-a-quiet-weekend/</guid><description>**OpenAI** released the new **o1** model, leveraging reinforcement learning and chain-of-thought prompting to excel in reasoning benchmarks, achieving an IQ-like score of **120**. **Google DeepMind** introduced **DataGemma** to reduce hallucinations by connecting LLMs with real-world data, and unveiled **ALOHA** and **DemoStart** for robot dexterity using diffusion methods. **Adobe** previewed its **Firefly AI Video Model** with text-to-video and generative extend features. **Mistral** launched the multimodal **Pixtral 12B** model, and **Tencent** presented the **GameGen-O** open-world video game generation model. Several research papers from **Stanford**, **OpenAI**, **Microsoft**, **Mila**, and **Notre Dame** focus on advanced reasoning, self-verification, and reflection tuning techniques. Experts like **Terence Tao** and **George Hotz** have shared mixed but optimistic views on o1&apos;s capabilities. Seed funding rounds include **Supermaven** ($12M) and **11x** ($24M).</description><pubDate>Tue, 17 Sep 2024 00:28:09 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>adobe</category><category>mistral-ai</category><category>tencent</category><category>supermaven</category><category>11x</category><category>cohere</category><category>anthropic</category><category>latent-space-university</category><category>stanford</category><category>microsoft</category><category>mila</category><category>notre-dame</category><category>o1</category><category>datagemma</category><category>aloha</category><category>demostart</category><category>firefly-ai-video-model</category><category>pixtral-12b</category><category>gamegen-o</category><category>george-hotz</category><category>terence-tao</category><category>adcock_brett</category><category>rohanpaul_ai</category><category>bindureddy</category><category>fchollet</category><category>philschmid</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>reasoning</category><category>robotics</category><category>diffusion-models</category><category>multimodality</category><category>video-generation</category><category>model-training</category><category>reflection-tuning</category><category>mathematical-reasoning</category><category>model-benchmarking</category><category>fine-tuning</category></item><item><title>Learnings from o1 AMA</title><link>https://news.smol.ai/issues/24-09-13-ainews-learnings-from-o1-ama/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-13-ainews-learnings-from-o1-ama/</guid><description>**OpenAI** released the **o1 model series**, touted as their &quot;most capable and aligned models yet,&quot; trained with reinforcement learning to enhance reasoning. The **o1-preview** model scored **21% on ARC-AGI**, **~80% on aider code editing** (surpassing Claude 3.5 Sonnet&apos;s 77%), and **~52% on Cognition-Golden**, showcasing a shift from memorizing answers to memorizing reasoning. The model employs a unique chain-of-thought approach enabling &quot;System II thinking&quot; for better problem-solving. Experts like **Andrew Mayne** advise framing o1 as a smart friend providing thoughtful explanations. Additionally, an advanced RAG course sponsored by **Weights &amp; Biases**, **Cohere**, and **Weaviate** offers strategies for hybrid search and prompting to optimize AI solutions.</description><pubDate>Sat, 14 Sep 2024 00:55:34 GMT</pubDate><category>openai</category><category>weights-biases</category><category>cohere</category><category>weaviate</category><category>o1-preview</category><category>o1-mini</category><category>claude-3.5-sonnet</category><category>gpt-4o</category><category>sama</category><category>rohanpaul_ai</category><category>gdb</category><category>andrew-mayne</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>reasoning</category><category>model-performance</category><category>prompting</category><category>code-editing</category><category>rag</category><category>hybrid-search</category></item><item><title>o1: OpenAI&apos;s new general reasoning models</title><link>https://news.smol.ai/issues/24-09-12-ainews-o1-openais-new-general-reasoning-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-12-ainews-o1-openais-new-general-reasoning-models/</guid><description>**OpenAI** has released the **o1** model family, including **o1-preview** and **o1-mini**, focusing on test-time reasoning with extended output token limits over 30k tokens. The models show strong performance, ranking in the 89th percentile on competitive programming, excelling in USA Math Olympiad qualifiers, and surpassing PhD-level accuracy on physics, biology, and chemistry benchmarks. Notably, **o1-mini** performs impressively despite its smaller size compared to **gpt-4o**. The release highlights new scaling laws for test-time compute that scale loglinearly. Additionally, **Nvidia** is reportedly losing AI chip market share to startups, with a shift in developer preference from CUDA to **llama** models for web development, though Nvidia remains dominant in training. This news reflects significant advances in reasoning-focused models and shifts in AI hardware competition.</description><pubDate>Fri, 13 Sep 2024 01:18:57 GMT</pubDate><category>openai</category><category>nvidia</category><category>o1</category><category>o1-preview</category><category>o1-mini</category><category>gpt-4o</category><category>llama</category><category>jason-wei</category><category>jim-fan</category><category>test-time-reasoning</category><category>reasoning-tokens</category><category>token-limit</category><category>competitive-programming</category><category>benchmarking</category><category>scaling-laws</category><category>ai-chip-competition</category><category>inference</category><category>training</category><category>model-performance</category></item><item><title>Pixtral 12B: Mistral beats Llama to Multimodality</title><link>https://news.smol.ai/issues/24-09-11-ainews-pixtral-12b-mistral-beats-llama-to-multimodality/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-11-ainews-pixtral-12b-mistral-beats-llama-to-multimodality/</guid><description>**Mistral AI** released **Pixtral 12B**, an open-weights **vision-language model** with a **Mistral Nemo 12B** text backbone and a 400M vision adapter, featuring a large vocabulary of **131,072 tokens** and support for **1024x1024 pixel images**. This release notably beat **Meta AI** in launching an open multimodal model. At the Mistral AI Summit, architecture details and benchmark performances were shared, showing strong OCR and screen understanding capabilities. Additionally, **Arcee AI** announced **SuperNova**, a distilled **Llama 3.1 70B &amp; 8B** model outperforming Meta&apos;s Llama 3.1 70B instruct on benchmarks. **DeepSeek** released **DeepSeek-V2.5**, scoring **89 on HumanEval**, surpassing **GPT-4-Turbo**, Opus, and Llama 3.1 in coding tasks. **OpenAI** plans to release **Strawberry** as part of ChatGPT soon, though its capabilities are debated. **Anthropic** introduced Workspaces for managing multiple Claude deployments with enhanced access controls.</description><pubDate>Thu, 12 Sep 2024 00:30:22 GMT</pubDate><category>mistral-ai</category><category>meta-ai-fair</category><category>hugging-face</category><category>arcee-ai</category><category>deepseek-ai</category><category>openai</category><category>anthropic</category><category>pixtral-12b</category><category>mistral-nemo-12b</category><category>llama-3-1-70b</category><category>llama-3-1-8b</category><category>deeps-eek-v2-5</category><category>gpt-4-turbo</category><category>llama-3-1</category><category>strawberry</category><category>claude</category><category>reach_vb</category><category>devendra_chapilot</category><category>_philschmid</category><category>rohanpaul_ai</category><category>vision</category><category>multimodality</category><category>ocr</category><category>benchmarking</category><category>model-release</category><category>model-architecture</category><category>model-performance</category><category>fine-tuning</category><category>model-deployment</category><category>reasoning</category><category>code-generation</category><category>api</category><category>access-control</category></item><item><title>not much happened today + AINews Podcast?</title><link>https://news.smol.ai/issues/24-09-10-ainews-not-much-happened-today-ainews-podcast/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-10-ainews-not-much-happened-today-ainews-podcast/</guid><description>**Glean** doubled its valuation again. **Dan Hendrycks&apos; Superforecaster AI** generates plausible election forecasts with interesting prompt engineering. A **Stanford** study found that **LLM-generated research ideas** are statistically more novel than those by expert humans. **SambaNova** announced faster inference for **llama-3** models, surpassing **Cerebras**. **Benjamin Clavie** gave a notable talk on retrieval-augmented generation techniques. **Strawberry** is reported to launch in two weeks. **Google Illuminate** offers AI-generated podcast discussions about papers and books. **Apple** unveiled new AI features in iOS 18, including visual intelligence and improved Siri, with on-device and cloud processing for camera-based event additions. The **Reflection 70B** model sparked controversy over performance claims. Experts highlighted the unreliability of traditional benchmarks like MMLU and HumanEval, recommending alternative evaluation methods such as LMSys Chatbot Arena and Hugging Face&apos;s open-sourced **Lighteval** suite. The AI research community continues to explore AI&apos;s role in generating novel research ideas and improving benchmarking.</description><pubDate>Wed, 11 Sep 2024 02:24:16 GMT</pubDate><category>glean</category><category>sambanova</category><category>cerebras</category><category>stanford</category><category>google</category><category>apple</category><category>hugging-face</category><category>lmsys</category><category>superforecaster-ai</category><category>llama-3</category><category>reflection-70b</category><category>danhendrycks</category><category>benjamin-clavie</category><category>bclavie</category><category>bindureddy</category><category>swyx</category><category>borismpower</category><category>corbtt</category><category>drjimfan</category><category>clementdelangue</category><category>rohanpaul_ai</category><category>prompt-engineering</category><category>research-ideas</category><category>inference-speed</category><category>retrieval-augmented-generation</category><category>evaluation-methods</category><category>visual-intelligence</category><category>on-device-ai</category><category>model-performance</category><category>benchmarking</category><category>novelty-detection</category></item><item><title>AIPhone 16: the Visual Intelligence Phone</title><link>https://news.smol.ai/issues/24-09-09-ainews-aiphone-16-the-visual-intelligence-phone/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-09-ainews-aiphone-16-the-visual-intelligence-phone/</guid><description>**Apple** announced the new **iPhone 16** lineup featuring **Visual Intelligence**, a new AI capability integrated with Camera Control, Apple Maps, and Siri, emphasizing privacy and default service use over third-party AI like OpenAI. **Apple Photos** now includes advanced video understanding with timestamp recognition. Meanwhile, **Reflection-70B** claims to be a top open-source model but benchmarks show it performs close to **Llama 3 70B** and slightly worse than **Qwen 2 72B**. **Yann LeCun** highlighted ongoing challenges with LLM planning abilities, noting models like **Llama-3.1-405b** and **Claude** show some skill, while **GPT-4** and **Gemini** lag behind. **Weights &amp; Biases** is sponsoring an event to advance LLM evaluation techniques with prizes and API access.</description><pubDate>Mon, 09 Sep 2024 23:00:14 GMT</pubDate><category>apple</category><category>openai</category><category>weights-biases</category><category>reflection-70b</category><category>llama-3-70b</category><category>qwen-2-72b</category><category>llama-3-1-405b</category><category>claude</category><category>gpt-4</category><category>gemini</category><category>yann-lecun</category><category>vision</category><category>video-understanding</category><category>benchmarking</category><category>planning</category><category>model-evaluation</category><category>privacy</category><category>ai-integration</category><category>instruction-following</category></item><item><title>Reflection 70B, by Matt from IT Department</title><link>https://news.smol.ai/issues/24-09-06-ainews-reflection-70b-by-matt-from-it-department/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-06-ainews-reflection-70b-by-matt-from-it-department/</guid><description>**Reflection Tuning** technique has been used by a two-person team from **Hyperwrite** and **Glaive** to finetune **llama-3.1-70b**, showing strong performance improvements with minimal synthetic data. The approach builds on the concept of adding `thinking` and `reflection` steps to outputs, related to the **Chain of Thought** method. Despite some criticisms like contamination concerns, worse coding performance, and reliance on system prompts, the model has received positive reception and comparisons to **claude-3.5-sonnet**. The work highlights efficient instruction tuning and synthetic data generation for large models.</description><pubDate>Sat, 07 Sep 2024 01:17:07 GMT</pubDate><category>hyperwrite</category><category>glaive</category><category>llama-3.1-70b</category><category>llama-3</category><category>claude-3.5-sonnet</category><category>matt-shumer</category><category>sahil-chaudhary</category><category>fine-tuning</category><category>chain-of-thought</category><category>instruction-following</category><category>synthetic-data</category><category>quantization</category><category>model-evaluation</category><category>prompt-engineering</category></item><item><title>Replit Agent - How did everybody beat Devin to market?</title><link>https://news.smol.ai/issues/24-09-05-ainews-replit-agent-how-did-everybody-beat-devin-to-market/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-05-ainews-replit-agent-how-did-everybody-beat-devin-to-market/</guid><description>**Replit Agent** launched as a fully integrated Web IDE enabling text-to-app generation with planning and self-healing, available immediately to paid users without a waitlist. Other notable developments include **Melodio**, a new text-to-music model, and **Together AI**&apos;s kernel and speculative decoding work. **Anthropic AI** announced a new enterprise plan featuring a **500K context window** and enhanced security. Discussions on **JPEG-LM** and **AVC-LM** models for improved image and video generation, and GPU market trends around the **H100 GPU** pricing were highlighted. Influential voices like **Andrej Karpathy** shared insights on AI agents and automation.</description><pubDate>Fri, 06 Sep 2024 01:54:59 GMT</pubDate><category>replit</category><category>anthropic</category><category>togethercompute</category><category>jpeg-lm</category><category>avc-lm</category><category>andrej-karpathy</category><category>mervenoyann</category><category>bindureddy</category><category>rohanpaul_ai</category><category>leptonai</category><category>teortaxestex</category><category>document-retrieval</category><category>retrieval-augmented-generation</category><category>ai-agents</category><category>image-generation</category><category>video-generation</category><category>context-windows</category><category>gpu-pricing</category><category>enterprise-ai</category><category>self-healing</category><category>text-to-music</category></item><item><title>$1150m for SSI, Sakana, You.com + Claude 500m context</title><link>https://news.smol.ai/issues/24-09-04-ainews-dollar1150m-for-ssi-sakana-youcom-claude-500m-context/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-04-ainews-dollar1150m-for-ssi-sakana-youcom-claude-500m-context/</guid><description>**Safe Superintelligence** raised **$1 billion** at a **$5 billion** valuation, focusing on safety and search approaches as hinted by Ilya Sutskever. **Sakana AI** secured a **$100 million Series A** funding round, emphasizing nature-inspired collective intelligence. **You.com** pivoted to a ChatGPT-like productivity agent after a **$50 million Series B** round, while **Perplexity AI** raised over **$250 million** this summer. **Anthropic** launched Claude for Enterprise with a **500 million token context window**. **AI2** released a **64-expert Mixture-of-Experts (MoE) model** called OLMo, outperforming Llama2-13B-Chat. Key AI research trends include efficient MoE architectures, challenges in AI alignment and GPU costs, and emerging AI agents for autonomous tasks. Innovations in AI development feature command and control for video generation, Retrieval-Augmented Generation (RAG) efficiency, and GitHub integration under Anthropic&apos;s Enterprise plan. *&quot;Our logo is meant to invoke the idea of a school of fish coming together and forming a coherent entity from simple rules as we want to make use of ideas from nature such as evolution and collective intelligence in our research.&quot;*</description><pubDate>Thu, 05 Sep 2024 03:25:36 GMT</pubDate><category>safe-superintelligence</category><category>sakana-ai</category><category>you-com</category><category>perplexity-ai</category><category>anthropic</category><category>ai2</category><category>olmo</category><category>llama2-13b-chat</category><category>claude</category><category>claude-3.5-sonnet</category><category>ilya-sutskever</category><category>mervenoyann</category><category>yuchenj_uw</category><category>rohanpaul_ai</category><category>ctojunior</category><category>omarsar0</category><category>mixture-of-experts</category><category>model-architecture</category><category>model-training</category><category>gpu-costs</category><category>retrieval-augmented-generation</category><category>video-generation</category><category>ai-alignment</category><category>enterprise-ai</category><category>agentic-ai</category><category>command-and-control</category></item><item><title>Everybody shipped small things this holiday weekend</title><link>https://news.smol.ai/issues/24-09-03-ainews-everybody-shipped-small-things-this-holiday-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-09-03-ainews-everybody-shipped-small-things-this-holiday-weekend/</guid><description>**xAI** announced the **Colossus 100k H100 cluster** capable of training an FP8 GPT-4 class model in 4 days. **Google** introduced **Structured Output** for **Gemini**. **Anthropic** discussed **Claude**&apos;s performance issues possibly due to API prompt modifications. **OpenAI** enhanced controls for File Search in their Assistants API. **Cognition** and **Anthropic** leaders appeared on podcasts. The viral **Kwai-Kolors** virtual try-on model and the open-source real-time audio conversational model **Mini-Omni** (similar to **gpt-4o-voice**) were released. Tutorials on parameter-efficient fine-tuning with LoRA and QLoRA, long-context embedding challenges, and Claude&apos;s LaTeX rendering feature were highlighted. **AI21 Labs** released **Jamba 1.5** models with a 256K context window and faster long-context performance. **NVIDIA** debuted **Mistral-Nemo-Minitron-8B** on the Open LLM Leaderboard. **LangChain** introduced resource tags for workspace organization, and a low-code AI app toolkit was shared by **svpino**. Legal AI agents and financial agent evaluations using LangSmith were also featured.</description><pubDate>Wed, 04 Sep 2024 01:35:37 GMT</pubDate><category>xai</category><category>google</category><category>anthropic</category><category>openai</category><category>cognition</category><category>ai21-labs</category><category>nvidia</category><category>langchain</category><category>gpt-4o-voice</category><category>gemini</category><category>claude</category><category>jamba-1.5</category><category>mistral-nemo-minitron-8b</category><category>dario-amodei</category><category>scott-wu</category><category>fchollet</category><category>svpino</category><category>fine-tuning</category><category>long-context</category><category>parameter-efficient-fine-tuning</category><category>latex-rendering</category><category>real-time-audio</category><category>virtual-try-on</category><category>resource-tags</category><category>low-code</category><category>ai-agents</category><category>workspace-organization</category><category>model-benchmarking</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-30-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-30-ainews-not-much-happened-today/</guid><description>**Meta** announced significant adoption of **LLaMA 3.1** with nearly **350 million downloads** on Hugging Face. **Magic AI Labs** introduced **LTM-2-Mini**, a long context model with a **100 million token context window**, and a new evaluation method called HashHop. **LMSys** added style control to their Chatbot Arena leaderboard, improving rankings for models like **Claude 3.5 Sonnet** and **LLaMA 3.1 405B**. **Alibaba** released **Qwen2-VL**, a multimodal LLM under Apache 2.0 license, competitive with **GPT-4o mini**. **OpenAI** CEO **Sam Altman** announced collaboration with the US AI Safety Institute for pre-release model testing. Discussions on AI safety and potential AI takeover risks were highlighted by **Ajeya Cotra**. Tools like **firecrawl** for web crawling and challenges in PDF processing were noted. AI hype cycles and market trends were discussed by **François Chollet**, and potential AI disruption in call centers was shared by **Rohan Paul**.</description><pubDate>Sat, 31 Aug 2024 00:41:42 GMT</pubDate><category>meta-ai-fair</category><category>hugging-face</category><category>magic-ai-labs</category><category>lmsys</category><category>alibaba</category><category>openai</category><category>llama-3-1</category><category>claude-3-5-sonnet</category><category>llama-3-1-405b</category><category>ltm-2-mini</category><category>qwen2-vl</category><category>gpt-4o-mini</category><category>sam-altman</category><category>ajeya-cotra</category><category>fchollet</category><category>rohanpaul_ai</category><category>philschmid</category><category>long-context</category><category>style-control</category><category>multimodality</category><category>ai-safety</category><category>model-evaluation</category><category>web-crawling</category><category>pdf-processing</category><category>ai-hype-cycles</category><category>call-center-automation</category></item><item><title>Summer of Code AI: $1.6b raised, 1 usable product</title><link>https://news.smol.ai/issues/24-08-29-ainews-summer-of-code-ai-dollar16b-raised-1-usable-product/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-29-ainews-summer-of-code-ai-dollar16b-raised-1-usable-product/</guid><description>**Code + AI** is emphasized as a key modality in AI engineering, highlighting productivity and verifiability benefits. Recent major funding rounds include **Cognition AI raising $175M**, **Poolside raising $400M**, **Codeium AI raising $150M**, and **Magic raising $320M**. Magic announced their **LTM-2** model with a **100 million token context window**, boasting efficiency improvements over **Llama 3.1 405B** by about **1000x cheaper** in sequence-dimension algorithm and drastically lower memory requirements. Magic&apos;s stack is built from scratch with custom CUDA and no open-source foundations, partnered with **Google Cloud** and powered by **NVIDIA H100** and **GB200 GPUs**, aiming to scale to tens of thousands of GPUs. Google DeepMind revealed updates to **Gemini Advanced** with customizable expert &quot;Gems.&quot; Neural Game Engines like **GameNGen** can run DOOM in a diffusion model trained on **0.9B frames**. The content also references **LLM quantization** research by Rohan Paul.</description><pubDate>Fri, 30 Aug 2024 00:01:06 GMT</pubDate><category>cognition</category><category>poolside</category><category>codeium</category><category>magic</category><category>google-deepmind</category><category>nvidia</category><category>google-cloud</category><category>ltm-2</category><category>llama-3-1-405b</category><category>gemini-advanced</category><category>nat-friedman</category><category>ben-chess</category><category>rohan-paul</category><category>long-context</category><category>model-efficiency</category><category>custom-hardware</category><category>cuda</category><category>training-stack</category><category>gpu-scaling</category><category>neural-world-models</category><category>diffusion-models</category><category>quantization</category></item><item><title>Cerebras Inference: Faster, Better, AND Cheaper</title><link>https://news.smol.ai/issues/24-08-28-ainews-cerebras-inference-faster-better-and-cheaper/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-28-ainews-cerebras-inference-faster-better-and-cheaper/</guid><description>**Groq** led early 2024 with superfast LLM inference speeds, achieving ~450 tokens/sec for Mixtral 8x7B and 240 tokens/sec for Llama 2 70B. **Cursor** introduced a specialized code edit model hitting 1000 tokens/sec. Now, **Cerebras** claims the fastest inference with their wafer-scale chips, running **Llama3.1-8b** at 1800 tokens/sec and **Llama3.1-70B** at 450 tokens/sec at full precision, with competitive pricing and a generous free tier. **Google&apos;s Gemini 1.5** models showed significant benchmark improvements, especially Gemini-1.5-Flash and Gemini-1.5-Pro. New open-source models like **CogVideoX-5B** and **Mamba-2 (Rene 1.3B)** were released, optimized for consumer hardware. **Anthropic&apos;s Claude** now supports prompt caching, improving speed and cost efficiency. *&quot;Cerebras Inference runs Llama3.1 20x faster than GPU solutions at 1/5 the price.&quot;*</description><pubDate>Thu, 29 Aug 2024 00:59:27 GMT</pubDate><category>groq</category><category>cerebras</category><category>cursor</category><category>google-deepmind</category><category>anthropic</category><category>llama-3.1-8b</category><category>llama-3.1-70b</category><category>gemini-1.5-flash</category><category>gemini-1.5-pro</category><category>cogvideox-5b</category><category>mamba-2</category><category>rene-1.3b</category><category>llama-3.1</category><category>gemini-1.5</category><category>claude</category><category>jeremyphoward</category><category>sam-altman</category><category>nat-friedman</category><category>daniel-gross</category><category>swyx</category><category>inference-speed</category><category>wafer-scale-chips</category><category>prompt-caching</category><category>model-merging</category><category>benchmarking</category><category>open-source-models</category><category>code-editing</category><category>model-optimization</category></item><item><title>CogVideoX: Zhipu&apos;s Open Source Sora</title><link>https://news.smol.ai/issues/24-08-27-ainews-cogvideox-zhipus-open-source-sora/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-27-ainews-cogvideox-zhipus-open-source-sora/</guid><description>**Zhipu AI**, Alibaba&apos;s AI arm and China&apos;s 3rd largest AI lab, released the open 5B video generation model **CogVIdeoX**, which can run without GPUs via their ChatGLM web and desktop apps. **Meta AI** announced trust &amp; safety research and CyberSecEval 3 alongside the release of **Llama 3.1**, with **Llama 3 405B** now available serverless on Google Cloud Vertex AI and Hugging Face x NVIDIA NIM API. Updates include **Moondream**, an open vision-language model improving DocVQA and TextVQA tasks, and the lightweight MoE chat model **Phi-3.5** with 16x3.8B parameters. **Together Compute** introduced the Rerank API featuring Salesforce&apos;s **LlamaRank** model for document and code ranking. Research highlights include superposition prompting for RAG without fine-tuning, the AgentWrite pipeline for long-form content generation over 20,000 words, and a comparison showing Long Context methods outperform RAG at higher costs. Tools include Not Diamond, an AI model router, AI command line interfaces, and an open-source WebGPU background removal tool. *&quot;You don&apos;t even need GPUs to run it,&quot;* referring to CogVIdeoX.</description><pubDate>Wed, 28 Aug 2024 01:26:46 GMT</pubDate><category>zhipu-ai</category><category>alibaba</category><category>meta-ai-fair</category><category>google</category><category>hugging-face</category><category>nvidia</category><category>togethercompute</category><category>salesforce</category><category>cogvideox</category><category>llama-3-1</category><category>llama-3-405b</category><category>moondream</category><category>phi-3.5</category><category>llama-rank</category><category>rohanpaul_ai</category><category>philschmid</category><category>vikhyatk</category><category>algo_diver</category><category>jayalammar</category><category>davidsholz</category><category>video-generation</category><category>serverless-computing</category><category>vision</category><category>document-vqa</category><category>text-vqa</category><category>mixture-of-experts</category><category>retrieval-augmented-generation</category><category>long-context</category><category>model-routing</category><category>webgpu</category><category>background-removal</category><category>long-form-generation</category><category>superposition-prompting</category></item><item><title>not much happened this weekend</title><link>https://news.smol.ai/issues/24-08-26-ainews-not-much-happened-this-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-26-ainews-not-much-happened-this-weekend/</guid><description>**Nous Research** announced **DisTrO**, a new optimizer that drastically reduces inter-GPU communication by 1000x to 10,000x enabling efficient training on slow networks, offering an alternative to **GDM&apos;s DiLoCo**. **Cursor AI** gained viral attention from an 8-year-old user and announced a new fundraise, with co-host Aman returning to their podcast. **George Hotz** launched **tinybox** for sale. In robotics, **AGIBOT** revealed 5 new humanoid robots with open-source plans, and **Unitree** showcased its G1 humanoid robot nearing mass production at $16,000. **ETH Zurich** and **Disney** developed an AI system for physics-based robot motion generation from text or images. **UC San Diego** released **ACE**, an open-source teleoperation system for controlling multiple robots. AI21 Labs unveiled **Jamba 1.5**, a multilingual model with 256k context length and permissive licensing. **Luma Labs** released **Dream Machine 1.5** for improved text-to-video generation. **Ideogram** launched **v2** of its text-to-image model with near-perfect text generation. **Nvidia** and **Mistral** released **Mistral-NeMo-Minitron 8B**, a small model outperforming **Mistral-7B** and **llama-3-8b** on the Open LLM leaderboard.</description><pubDate>Tue, 27 Aug 2024 00:09:52 GMT</pubDate><category>nous-research</category><category>cursor-ai</category><category>gdm</category><category>george-hotz</category><category>agibot</category><category>unitree</category><category>eth-zurich</category><category>disney</category><category>uc-san-diego</category><category>ai21-labs</category><category>luma-labs</category><category>ideogram</category><category>nvidia</category><category>mistral-ai</category><category>meta-ai-fair</category><category>jamba-1.5</category><category>dream-machine-1.5</category><category>ideogram-v2</category><category>mistral-nemo-minitron-8b</category><category>mistral-7b</category><category>llama-3-8b</category><category>george-hotz</category><category>adcock_brett</category><category>aman</category><category>distributed-ai</category><category>optimizer</category><category>inter-gpu-communication</category><category>low-latency-training</category><category>open-source</category><category>humanoid-robots</category><category>robotics</category><category>physics-based-motion</category><category>teleoperation</category><category>multilingual-models</category><category>long-context</category><category>text-to-video</category><category>text-to-image</category><category>model-performance</category></item><item><title>Nvidia Minitron: LLM Pruning and Distillation updated for Llama 3.1</title><link>https://news.smol.ai/issues/24-08-23-ainews-nvidia-minitron-llm-pruning-and-distillation-updated-for-llama-31/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-23-ainews-nvidia-minitron-llm-pruning-and-distillation-updated-for-llama-31/</guid><description>**Nvidia** and **Meta** researchers updated their **Llama 3** results with a paper demonstrating the effectiveness of combining **weight pruning** and **knowledge distillation** to reduce training costs by training only the largest model from scratch and deriving smaller models via pruning and distillation. The process involves teacher correction, activation-based pruning (favoring width pruning), and retraining with distillation using KL Divergence loss, resulting in better-performing models at comparable sizes. However, distillation incurs some accuracy tradeoffs. Additionally, **AI21 Labs** launched **Jamba 1.5**, a hybrid SSM-Transformer MoE model with large context windows and multilingual support. **Anthropic** updated **Claude 3** with LaTeX rendering and prompt caching. An open-source coding-focused LLM, **Dracarys**, was released in 70B and 72B sizes, showing improved coding performance. The **Mistral Nemo Minitron 8B** model outperforms **Llama 3.1 8B** and **Mistral 7B** on the Hugging Face leaderboard, highlighting pruning and distillation benefits. Research on prompt optimization reveals the complexity of prompt search spaces and the surprising effectiveness of simple algorithms like AutoPrompt/GCG.</description><pubDate>Fri, 23 Aug 2024 22:14:15 GMT</pubDate><category>nvidia</category><category>meta-ai-fair</category><category>ai21-labs</category><category>anthropic</category><category>hugging-face</category><category>llama-3-1-8b</category><category>llama-3-1</category><category>jamba-1.5</category><category>claude-3</category><category>dracarys-70b</category><category>dracarys-72b</category><category>mistral-nemo-minitron-8b</category><category>mistral-7b</category><category>pruning</category><category>knowledge-distillation</category><category>weight-pruning</category><category>activation-based-pruning</category><category>width-pruning</category><category>kl-divergence</category><category>teacher-correction</category><category>prompt-optimization</category><category>multilinguality</category><category>long-context</category><category>mixture-of-experts</category><category>model-fine-tuning</category></item><item><title>super quiet day</title><link>https://news.smol.ai/issues/24-08-22-ainews-super-quiet-day/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-22-ainews-super-quiet-day/</guid><description>**AI21 Labs** released **Jamba 1.5**, a scaled-up State Space Model optimized for long context windows with **94B parameters** and up to **2.5X faster inference**, outperforming models like **Llama 3.1 70B** on benchmarks. The **Phi-3.5** model was praised for its safety and performance, while **Dracarys**, a new **70B open-source coding model** announced by **Bindu Reddy**, claims superior benchmarks over Llama 3.1 70B. Discussions on **California&apos;s SB 1047** AI safety legislation involve **Stanford** and **Anthropic**, highlighting a balance between precaution and industry growth. Innovations include **uv virtual environments** for rapid setup, **LangChain&apos;s LangSmith** resource tags for project management, and multi-agent systems in **Qdrant** enhancing data workflows. Community events like the **RAG workshop** by **AWS**, **LangChain**, and **Elastic** continue to support AI learning and collaboration. Memes remain a popular way to engage with AI industry culture.</description><pubDate>Fri, 23 Aug 2024 00:55:37 GMT</pubDate><category>ai21-labs</category><category>anthropic</category><category>stanford</category><category>hugging-face</category><category>langchain</category><category>qdrant</category><category>aws</category><category>elastic</category><category>jamba-1.5</category><category>phi-3.5</category><category>dracarys</category><category>llama-3-1-70b</category><category>llama-3-1</category><category>bindu-reddy</category><category>rohanpaul_ai</category><category>jackclarksf</category><category>danhendrycks</category><category>reach_vb</category><category>iqdotgraph</category><category>state-space-models</category><category>long-context</category><category>benchmarking</category><category>ai-safety</category><category>virtual-environments</category><category>multi-agent-systems</category><category>resource-management</category><category>community-engagement</category><category>model-performance</category></item><item><title>Ideogram 2 + Berkeley Function Calling Leaderboard V2</title><link>https://news.smol.ai/issues/24-08-21-ainews-ideogram-2-berkeley-function-calling-leaderboard-v2/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-21-ainews-ideogram-2-berkeley-function-calling-leaderboard-v2/</guid><description>**Ideogram** returns with a new image generation model featuring **color palette control**, a fully controllable API, and an iOS app, reaching a milestone of **1 billion images created**. Meanwhile, **Midjourney** released a Web UI but still lacks an API. In function calling, the **Berkeley Function Calling Leaderboard (BFCL)** updated to **BFCL V2 • Live**, adding **2251 live, user-contributed function documentation and queries** to improve evaluation quality. **GPT-4** leads the leaderboard, but the open-source **Functionary Llama 3-70B finetune** from Kai surpasses **Claude**. On AI model releases, **Microsoft** launched three **Phi-3.5** models with impressive reasoning and context window capabilities, while **Meta AI FAIR** introduced **UniBench**, a unified benchmark suite for over **50 vision-language model tasks**. **Baseten** improved **Llama 3** inference speed by up to **122%** using Medusa. A new cybersecurity benchmark, **Cyberbench**, featuring **40 CTF tasks**, was released. Additionally, **Codegen** was introduced as a tool for programmatic codebase analysis and AI-assisted development. *&quot;Multiple functions &gt; parallel functions&quot;* was highlighted as a key insight in function calling.</description><pubDate>Thu, 22 Aug 2024 00:05:05 GMT</pubDate><category>ideogram</category><category>midjourney</category><category>berkeley</category><category>openai</category><category>hugging-face</category><category>microsoft</category><category>meta-ai-fair</category><category>baseten</category><category>kai</category><category>claude</category><category>functionary</category><category>llama-3-70b</category><category>gpt-4</category><category>phi-3.5</category><category>functionary-llama-3-70b</category><category>llama-3</category><category>function-calling</category><category>benchmarking</category><category>image-generation</category><category>model-optimization</category><category>vision</category><category>multimodality</category><category>model-performance</category><category>fine-tuning</category><category>context-windows</category><category>cybersecurity</category><category>code-analysis</category><category>ai-assisted-development</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-20-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-20-ainews-not-much-happened-today/</guid><description>**OpenAI** launched **GPT-4o finetuning** with a case study on Cosine. **Anthropic** released **Claude 3.5 Sonnet** with 8k token output. **Microsoft Phi** team introduced **Phi-3.5** in three variants: Mini (3.8B), MoE (16x3.8B), and Vision (4.2B), noted for sample efficiency. **Meta** released **Llama 3.1 405B**, deployable on Google Cloud Vertex AI, offering GPT-4 level capabilities. **Qwen2-Math-72B** achieved state-of-the-art math benchmark performance with a Gradio demo. Discussions included model comparisons like ViT vs CNN and Mamba architecture. Tools updates featured **DSPy** roadmap, **Flux Schnell** improving diffusion speed on M1 Max, and **LangChain** community events. Research highlights zero-shot DUP prompting for math reasoning and fine-tuning best practices. AI ethics covered California&apos;s AI Safety Bill SB 1047 and regulatory concerns from **Yann LeCun**. Commentary on AI engineer roles by **Swyx**. *&quot;Chat with PDF&quot;* feature now available for Box Enterprise Plus users.</description><pubDate>Wed, 21 Aug 2024 00:22:36 GMT</pubDate><category>openai</category><category>anthropic</category><category>microsoft</category><category>meta-ai-fair</category><category>hugging-face</category><category>langchain</category><category>box</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>phi-3.5-mini</category><category>phi-3.5-moe</category><category>phi-3.5-vision</category><category>llama-3-1-405b</category><category>qwen2-math-72b</category><category>swyx</category><category>ylecun</category><category>fine-tuning</category><category>benchmarking</category><category>model-comparison</category><category>model-performance</category><category>diffusion-models</category><category>reinforcement-learning</category><category>zero-shot-learning</category><category>math</category><category>model-efficiency</category><category>ai-regulation</category><category>ai-safety</category><category>ai-engineering</category><category>prompt-engineering</category></item><item><title>The DSPy Roadmap</title><link>https://news.smol.ai/issues/24-08-19-ainews-the-dspy-roadmap/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-19-ainews-the-dspy-roadmap/</guid><description>**Omar Khattab** announced joining **Databricks** before his MIT professorship and outlined the roadmap for **DSPy 2.5 and 3.0+**, focusing on improving core components like LMs, signatures, optimizers, and assertions with features such as adopting **LiteLLM** to reduce code and enhance caching and streaming. The roadmap also includes developing more accurate, cost-effective optimizers, building tutorials, and enabling interactive optimization tracking. On AI Twitter, **Google** launched **Gemini Live**, a mobile conversational AI with voice and 10 voices, alongside **Pixel Buds Pro 2** with a custom Tensor A1 chip. **OpenAI** updated **ChatGPT-4o**, reclaiming the top spot on LMSYS Arena. **xAI** released **Grok-2** in beta, achieving SOTA in image generation with FLUX 1. **Nous Research** released open-source **Hermes 3** models in 8B, 70B, and 405B sizes, with the 405B model achieving SOTA. Robotics updates include **Astribot**&apos;s humanoid robot and **Apple**&apos;s tabletop robot with Siri voice commands. **Sakana AI** introduced &quot;The AI Scientist,&quot; an autonomous AI research system.</description><pubDate>Tue, 20 Aug 2024 05:06:22 GMT</pubDate><category>databricks</category><category>mit</category><category>google</category><category>openai</category><category>x-ai</category><category>nous-research</category><category>astribot</category><category>apple</category><category>sakana-ai</category><category>dspy</category><category>litel-lm</category><category>gemini</category><category>chatgpt-4o</category><category>grok-2</category><category>hermes-3</category><category>omar-khattab</category><category>giffmana</category><category>model-optimization</category><category>fine-tuning</category><category>optimizers</category><category>interactive-optimization</category><category>robotics</category><category>autonomous-systems</category><category>voice</category><category>image-generation</category><category>open-source-models</category><category>scientific-research</category><category>streaming</category><category>caching</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-16-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-16-ainews-not-much-happened-today/</guid><description>**Anthropic** rolled out **prompt caching** in its API, reducing input costs by up to **90%** and latency by **80%**, enabling instant fine-tuning with longer prompts. **xAI** released **Grok-2**, a new model competing with frontier models from **Google DeepMind**, **OpenAI**, **Anthropic**, **Mistral AI**, and **Meta AI Fair**, supporting vision and text inputs and integrating external image generation models. **Claude 3.5 Sonnet** is reported to outperform **GPT-4** in coding and reasoning, while **ChatGPT-4o-latest** shows reasoning improvements. **François Chollet** proposed a theory defining intelligence as the efficiency of operationalizing past information for future tasks. The **Aya project** involves 3000 collaborators building multilingual AI datasets. **Demis Hassabis** discussed AI hype and safe AI development in a podcast. Tools like **Dora AI** for Figma and **Box&apos;s AI API** enhance design automation and document processing. **Salesforce** released **DEI**, an open AI software engineering agents framework with a 55% resolve rate on SWE-Bench Lite. Industry trends highlight rapid AI integration, networking importance in the AI job market, and potential OpenAI GPT-4 expansion in response to competitors. Memes include humor about Apple Vision Pro.</description><pubDate>Sat, 17 Aug 2024 03:43:03 GMT</pubDate><category>anthropic</category><category>x-ai</category><category>google-deepmind</category><category>openai</category><category>mistral-ai</category><category>meta-ai-fair</category><category>salesforce</category><category>box</category><category>grok-2</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>gpt-4</category><category>chatgpt-4o-latest</category><category>demis-hassabis</category><category>francois-chollet</category><category>prompt-caching</category><category>model-performance</category><category>vision</category><category>fine-tuning</category><category>multilinguality</category><category>ai-safety</category><category>design-automation</category><category>document-processing</category><category>ai-agents</category><category>ai-integration</category><category>ai-job-market</category><category>ai-acceleration</category><category>humor</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-15-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-15-ainews-not-much-happened-today/</guid><description>**GPT-5** delayed again amid a quiet news day. **Nous Research** released Hermes 3 finetune of **Llama 3** base models, rivaling FAIR&apos;s instruct tunes but sparking debate over emergent existential crisis behavior with 6% roleplay data. **Nvidia** introduced Minitron finetune of **Llama 3.1**. **Salesforce** launched a DEI agent scoring 55% on SWE-Bench Lite. **Goodfire AI** secured $7M seed funding for mechanistic interpretability work. **Anthropic** rolled out prompt caching in their API, cutting input costs by up to 90% and latency by 80%, aiding coding assistants and large document processing. **xAI** released **Grok-2**, matching **Claude 3.5 Sonnet** and **GPT-4 Turbo** on LMSYS leaderboard with vision+text inputs and image generation integration. **Claude 3.5 Sonnet** reportedly outperforms **GPT-4** in coding and reasoning. **François Chollet** defined intelligence as efficient operationalization of past info for future tasks. **Salesforce&apos;s** DEI framework surpasses individual agent performance. **Google DeepMind&apos;s** Demis Hassabis discussed AGI&apos;s role in scientific discovery and safe AI development. **Dora AI** plugin generates landing pages in under 60 seconds, boosting web team efficiency. **Box AI API** beta enables document chat, data extraction, and content summarization. **LangChain** updated Python &amp; JavaScript integration docs.</description><pubDate>Fri, 16 Aug 2024 04:05:53 GMT</pubDate><category>nous-research</category><category>nvidia</category><category>salesforce</category><category>goodfire-ai</category><category>anthropic</category><category>x-ai</category><category>google-deepmind</category><category>box</category><category>langchain</category><category>llama-3</category><category>llama-3-1</category><category>grok-2</category><category>claude-3.5-sonnet</category><category>gpt-4-turbo</category><category>fchollet</category><category>demis-hassabis</category><category>fine-tuning</category><category>prompt-caching</category><category>mechanistic-interpretability</category><category>model-performance</category><category>multimodality</category><category>agent-frameworks</category><category>software-engineering-agents</category><category>api</category><category>document-processing</category><category>text-generation</category><category>model-releases</category><category>vision</category><category>image-generation</category><category>efficiency</category><category>scientific-discovery</category></item><item><title>Grok 2! and ChatGPT-4o-latest confuses everybody</title><link>https://news.smol.ai/issues/24-08-14-ainews-grok-2-and-chatgpt-4o-latest-confuses-everybody/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-14-ainews-grok-2-and-chatgpt-4o-latest-confuses-everybody/</guid><description>**OpenAI** quietly released a new **GPT-4o** model in ChatGPT, distinct from the API version, reclaiming the #1 spot on Lmsys arena benchmarks across multiple categories including math, coding, and instruction-following. Meanwhile, **X.ai** launched **Grok 2**, outperforming **Claude 3.5 Sonnet** and previous GPT-4o versions, with plans for enterprise API release. Grok 2 integrates **Black Forest Labs&apos; Flux.1**, an open-source text-to-image model surpassing **Stable Diffusion 3**. **Google DeepMind** announced **Gemini Advanced** with enhanced conversational features and Pixel device integration. AI researcher **ylecun** highlighted LLM limitations in learning and creativity, while **rohanpaul_ai** discussed an AI Scientist system generating publishable ML research at low cost. **karpathy** warned of security risks in LLM tokenizers akin to SQL injection.</description><pubDate>Thu, 15 Aug 2024 00:51:40 GMT</pubDate><category>openai</category><category>x-ai</category><category>black-forest-labs</category><category>google-deepmind</category><category>gpt-4o</category><category>grok-2</category><category>claude-3.5-sonnet</category><category>flux-1</category><category>stable-diffusion-3</category><category>gemini-advanced</category><category>ylecun</category><category>rohanpaul_ai</category><category>karpathy</category><category>benchmarking</category><category>model-performance</category><category>tokenization</category><category>security-vulnerabilities</category><category>multi-agent-systems</category><category>research-automation</category><category>text-to-image</category><category>conversational-ai</category><category>model-integration</category></item><item><title>Gemini Live</title><link>https://news.smol.ai/issues/24-08-13-ainews-gemini-live/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-13-ainews-gemini-live/</guid><description>**Google** launched **Gemini Live** on Android for **Gemini Advanced** subscribers during the Pixel 9 event, featuring integrations with Google Workspace apps and other Google services. The rollout began on 8/12/2024, with iOS support planned. **Anthropic** released **Genie**, an AI software engineering system achieving a **57%** improvement on SWE-Bench. **TII** introduced **Falcon Mamba**, a 7B attention-free open-access model scalable to long sequences. Benchmarking showed that longer context lengths do not always improve Retrieval-Augmented Generation. **Supabase** launched an AI-powered Postgres service dubbed the &quot;ChatGPT of databases,&quot; fully open source. **Perplexity AI** partnered with Polymarket to integrate real-time probability predictions into search results. A tutorial demonstrated a multimodal recipe recommender using **Qdrant**, **LlamaIndex**, and **Gemini**. An OpenAI engineer shared success tips emphasizing debugging and hard work. The connection between matrices and graphs in linear algebra was highlighted for insights into nonnegative matrices and strongly connected components. **Keras 3.5.0** was released with Hugging Face Hub integration for model saving and loading.</description><pubDate>Wed, 14 Aug 2024 01:23:26 GMT</pubDate><category>google</category><category>anthropic</category><category>tii</category><category>supabase</category><category>perplexity-ai</category><category>llamaindex</category><category>openai</category><category>hugging-face</category><category>gemini-1.5-pro</category><category>genie</category><category>falcon-mamba</category><category>gemini-1.5</category><category>llamaindex</category><category>omarsar0</category><category>osanseviero</category><category>dbrxmosaicai</category><category>alphasignalai</category><category>perplexity_ai</category><category>_jasonwei</category><category>svpino</category><category>multimodality</category><category>benchmarking</category><category>long-context</category><category>retrieval-augmented-generation</category><category>open-source</category><category>model-releases</category><category>model-integration</category><category>model-performance</category><category>software-engineering</category><category>linear-algebra</category><category>hugging-face-hub</category><category>debugging</category></item><item><title>a quiet weekend</title><link>https://news.smol.ai/issues/24-08-12-ainews-a-quiet-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-12-ainews-a-quiet-weekend/</guid><description>**Figure** unveiled **Figure 02**, claimed as the most advanced humanoid robot, operating autonomously at BMW&apos;s Plant Spartanburg. **DeepMind** developed a table tennis robot achieving **100% wins against beginners** and **55% against intermediates**. **Boston Dynamics** showcased the dexterity of its fully-electric **Atlas** robot performing pushups and burpees. An autonomous dental robot performed the world&apos;s first dental procedure on a human, reducing a 2-hour process to 15 minutes using a **3D volumetric scanner**. **SAM 2** was introduced as an open model for real-time object segmentation without custom adaptation. **Alibaba** released **Qwen2-Math**, outperforming **GPT-4** and **Claude 3.5** in math capabilities. A new Listening-While-Speaking Language Model (LSLM) enables simultaneous listening and speaking in real-time. Researchers developed a disease prediction AI with **95% accuracy** for diseases like coronary artery disease, type 2 diabetes, and breast cancer. Tools like **LlamaParse CLI** and **MLX Whisper package** enhance PDF parsing and speech recognition, with the latter running **40X faster than realtime** on M1 Max. The news highlights significant advancements in robotics, AI models, and practical AI tools.</description><pubDate>Mon, 12 Aug 2024 22:36:30 GMT</pubDate><category>figure</category><category>deepmind</category><category>boston-dynamics</category><category>alibaba</category><category>llamaindex</category><category>sam-2</category><category>qwen2-math</category><category>gpt-4</category><category>claude-3.5</category><category>adcock_brett</category><category>rasbt</category><category>hamel-husain</category><category>rohanpaul_ai</category><category>robotics</category><category>object-segmentation</category><category>real-time-processing</category><category>disease-prediction</category><category>speech-recognition</category><category>cli-tools</category><category>model-performance</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-09-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-09-ainews-not-much-happened-today/</guid><description>**Qwen2-Math-72B** outperforms **GPT-4o**, **Claude-3.5-Sonnet**, **Gemini-1.5-Pro**, and **Llama-3.1-405B** on math benchmarks using synthetic data and advanced optimization techniques. **Google AI** cuts pricing for **Gemini 1.5 Flash** by up to 78%. **Anthropic** expands its bug bounty program targeting universal jailbreaks in next-gen safety systems. Tutorial on **QLoRA** fine-tuning of **IDEFICS3-Llama 8B** for visual question answering released. A Chinese open weights model surpasses previous MATH benchmark records. Surveys on **Mamba** models and LLM-based agents for software engineering highlight advancements and applications. Open-source tools like **R2R RAG engine** and **LlamaIndex Workflows** simplify building complex AI applications. **Mistral AI** introduces customizable AI agents. Concerns raised about California bill SB 1047&apos;s focus on existential risk and debates on banning open-source AI. Memes and humor continue in AI communities.</description><pubDate>Sat, 10 Aug 2024 05:51:12 GMT</pubDate><category>anthropic</category><category>google</category><category>mistral-ai</category><category>llamaindex</category><category>qwen2-math-72b</category><category>gpt-4o</category><category>claude-3.5-sonnet</category><category>gemini-1.5-pro</category><category>llama-3.1-405b</category><category>idefics3-llama-8b</category><category>rohanpaul_ai</category><category>anthropicai</category><category>mervenoyann</category><category>jeremyphoward</category><category>omarsar0</category><category>ylecun</category><category>bindureddy</category><category>math</category><category>fine-tuning</category><category>synthetic-data</category><category>reinforcement-learning</category><category>bug-bounty</category><category>visual-question-answering</category><category>open-source</category><category>retrieval-augmented-generation</category><category>agentic-ai</category><category>ai-safety</category><category>policy</category></item><item><title>Too Cheap To Meter: AI prices cut 50-70% in last 30 days</title><link>https://news.smol.ai/issues/24-08-08-ainews-too-cheap-to-meter-ai-prices-cut-50-70percent-in-last-30-days/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-08-ainews-too-cheap-to-meter-ai-prices-cut-50-70percent-in-last-30-days/</guid><description>**Gemini 1.5 Flash** has cut prices by approximately **70%**, offering a highly competitive free tier of **1 million tokens per minute** at **$0.075/mtok**, intensifying the AI model price war. Other significant price reductions include **GPT-4o** (~50% cut to **$2.50/mtok**), **GPT-4o mini** (70-98.5% cut to **$0.15/mtok**), **Llama 3.1 405b** (46% cut to **$2.7/mtok**), and **Mistral Large 2** (62% cut to **$3/mtok**). **Deepseek v2** introduced context caching, reducing input token costs by up to **90%** to **$0.014/mtok**. New model releases include **Llama 3.1 405b**, **Sonnet 3.5**, **EXAONE-3.0** (7.8B instruction-tuned by LG AI Research), and **MiniCPM V 2.6** (vision-language model combining SigLIP 400M and Qwen2-7B). Benchmarks show **Mistral Large** performing well on ZebraLogic and **Claude-3.5** leading LiveBench. **FlexAttention**, a new PyTorch API, simplifies and optimizes attention mechanisms. **Andrej Karpathy** analyzed RLHF, highlighting its limitations compared to traditional reinforcement learning. Google DeepMind research on compute-optimal scaling was also summarized.</description><pubDate>Fri, 09 Aug 2024 04:27:56 GMT</pubDate><category>llamaindex</category><category>together-ai</category><category>deepinfra</category><category>deepseek-ai</category><category>mistral-ai</category><category>google-deepmind</category><category>lg-ai-research</category><category>llamaindex</category><category>llamaindex</category><category>llamaindex</category><category>gpt-4o</category><category>gpt-4o-mini</category><category>llama-3-1-405b</category><category>mistral-large-2</category><category>gemini-1.5-flash</category><category>deepseek-v2</category><category>sonnet-3.5</category><category>exaone-3.0</category><category>minicpm-v-2.6</category><category>claude-3.5</category><category>gpt-4o-2024-08-06</category><category>rohanpaul_ai</category><category>akhaliq</category><category>mervenoyann</category><category>sophiamyang</category><category>chhillee</category><category>karpathy</category><category>price-cuts</category><category>context-caching</category><category>instruction-tuning</category><category>vision</category><category>benchmarks</category><category>pytorch</category><category>attention-mechanisms</category><category>reinforcement-learning-from-human-feedback</category><category>compute-optimal-scaling</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-08-07-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-07-ainews-not-much-happened-today/</guid><description>**OpenAI** introduced structured outputs in their API with a new &quot;strict&quot; mode and a &quot;response_format&quot; parameter, supporting models like **gpt-4-0613**, **gpt-3.5-turbo-0613**, and the new **gpt-4o-2024-08-06**. They also halved the price of **gpt-4o** to $2.50 per million tokens. **Mistral Large 2** outperforms **gpt4-turbo** and **claude-3-opus** on hard benchmarks and coding tasks. **Idefics3-Llama** offers multimodal capabilities with a 10k token context window. **BigLlama-3.1-1T-Instruct** is an upscaled version of **llama-3-120b-instruct**. New benchmark &quot;big_model_smell&quot; measures creativity and reliability. **Figure 02** robot features advanced AI hardware with onboard vision language model, enhanced battery, and speech-to-speech reasoning. **Yann LeCun** expressed concerns about California&apos;s SB1047 regulation.</description><pubDate>Thu, 08 Aug 2024 01:50:11 GMT</pubDate><category>openai</category><category>mistral-ai</category><category>meta-ai-fair</category><category>gpt-4-0613</category><category>gpt-3.5-turbo-0613</category><category>gpt-4o-2024-08-06</category><category>mistral-large-2</category><category>gpt4-turbo</category><category>claude-3-opus</category><category>idefics3-llama</category><category>bigllama-3.1-1t-instruct</category><category>llama-3-120b-instruct</category><category>sama</category><category>rohanpaul_ai</category><category>corbtt</category><category>guillaumelample</category><category>mervenoyann</category><category>maximelabonne</category><category>aidan_mclau</category><category>adcock_brett</category><category>ylecun</category><category>structured-outputs</category><category>function-calling</category><category>json-schema</category><category>benchmarking</category><category>multimodality</category><category>context-windows</category><category>model-scaling</category><category>ai-hardware</category><category>vision</category><category>speech-processing</category><category>robotics</category><category>ai-regulation</category></item><item><title>GPT4o August + 100% Structured Outputs for All (GPT4o mini edition)</title><link>https://news.smol.ai/issues/24-08-06-ainews-gpt4o-august-100percent-structured-outputs-for-all-gpt4o-mini-edition/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-06-ainews-gpt4o-august-100percent-structured-outputs-for-all-gpt4o-mini-edition/</guid><description>**Stability.ai** users are leveraging **LoRA** and **ControlNet** for enhanced line art and artistic style transformations, while facing challenges with **AMD GPUs** due to the discontinuation of **ZLUDA**. Community tensions persist around the **r/stablediffusion** subreddit moderation. **Unsloth AI** users report fine-tuning difficulties with **LLaMA3** models, especially with PPO trainer integration and prompt formatting, alongside anticipation for **multi-GPU** support and cost-effective cloud computing on **RunPod**. **Google** released the lightweight **Gemma 2 2B** model optimized for on-device use with **2.6B** parameters, featuring safety and sparse autoencoder tools, and announced **Diffusers** integration for efficient text-to-image generation on limited resources.</description><pubDate>Wed, 07 Aug 2024 02:55:03 GMT</pubDate><category>stability-ai</category><category>unsloth-ai</category><category>google</category><category>hugging-face</category><category>gpt-4o-mini</category><category>gpt-4o-2024-08-06</category><category>llama-3</category><category>bigllama-3.1-1t-instruct</category><category>meta-llama-3-120b-instruct</category><category>gemma-2-2b</category><category>lora</category><category>controlnet</category><category>line-art</category><category>gpu-performance</category><category>multi-gpu-support</category><category>fine-tuning</category><category>prompt-formatting</category><category>cloud-computing</category><category>text-to-image-generation</category><category>model-integration</category></item><item><title>GPT4o August + 100% Structured Outputs for All (GPT4o August edition)</title><link>https://news.smol.ai/issues/24-08-06-ainews-gpt4o-august-100percent-structured-outputs-for-all-gpt4o-august-edition/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-06-ainews-gpt4o-august-100percent-structured-outputs-for-all-gpt4o-august-edition/</guid><description>**OpenAI** released the new **gpt-4o-2024-08-06** model with **16k context window** and **33-50% lower pricing** than the previous 4o-May version, featuring a new Structured Output API that improves output quality and reduces retry costs. **Meta AI** launched **Llama 3.1**, a **405-billion parameter** model surpassing **GPT-4** and **Claude 3.5 Sonnet** on benchmarks, alongside expanding the **Llama Impact Grant** program. **Google DeepMind** quietly released **Gemini 1.5 Pro**, outperforming **GPT-4o**, **Claude-3.5**, and **Llama 3.1** on LMSYS benchmarks and leading the Vision Leaderboard. **Yi-Large Turbo** was introduced as a cost-effective upgrade priced at $0.19 per million tokens. In hardware, **NVIDIA H100 GPUs** were highlighted by **John Carmack** for their massive AI workload power, and **Groq** announced plans to deploy **108,000 LPUs** by Q1 2025. New AI tools and techniques include **RAG (Retrieval-Augmented Generation)**, the **JamAI Base** platform for Mixture of Agents systems, and **LangSmith**&apos;s enhanced filtering capabilities. Google DeepMind also introduced **PEER (Parameter Efficient Expert Retrieval)** architecture.</description><pubDate>Wed, 07 Aug 2024 02:40:09 GMT</pubDate><category>openai</category><category>meta-ai-fair</category><category>google-deepmind</category><category>yi-large</category><category>nvidia</category><category>groq</category><category>langchain</category><category>jamai</category><category>langsmith</category><category>gpt-4o-2024-08-06</category><category>llama-3-1-405b</category><category>llama-3</category><category>claude-3.5-sonnet</category><category>gemini-1.5-pro</category><category>gpt-4o</category><category>yi-large-turbo</category><category>john-carmack</category><category>jonathan-ross</category><category>rohanpaul_ai</category><category>structured-output</category><category>context-windows</category><category>model-pricing</category><category>benchmarking</category><category>parameter-efficient-expert-retrieval</category><category>retrieval-augmented-generation</category><category>mixture-of-experts</category><category>model-performance</category><category>ai-hardware</category><category>model-deployment</category><category>filtering</category><category>multi-lingual</category><category>vision</category></item><item><title>How Carlini Uses AI</title><link>https://news.smol.ai/issues/24-08-05-ainews-how-carlini-uses-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-05-ainews-how-carlini-uses-ai/</guid><description>**Groq&apos;s** shareholders&apos; net worth rises while others fall, with **Intel&apos;s CEO** expressing concern. **Nicholas Carlini** of **DeepMind** gains recognition and criticism for his extensive AI writings, including an 80,000-word treatise on AI use and a benchmark for large language models. **Chris Dixon** comments on AI Winter skepticism, emphasizing long-term impact. **Box** introduces an AI API for extracting structured data from documents, highlighting potential and risks of LLM-driven solutions. Recent AI developments include **Figure AI** launching the advanced humanoid robot Figure 02, **OpenAI** rolling out Advanced Voice Mode for ChatGPT with emotion detection, **Google** open-sourcing **Gemma 2 2B** model matching GPT-3.5-Turbo-0613 performance, **Meta AI Fair** releasing Segment Anything Model 2 (SAM 2) for real-time object tracking, **NVIDIA** showcasing Project GR00T for humanoid teleoperation with Apple Vision Pro, **Stability AI** launching Stable Fast 3D for rapid 3D asset generation, and **Runway** unveiling Gen-3 Alpha for AI text-to-video generation.</description><pubDate>Mon, 05 Aug 2024 23:43:14 GMT</pubDate><category>groq</category><category>intel</category><category>deepmind</category><category>box</category><category>figure-ai</category><category>openai</category><category>google</category><category>meta-ai-fair</category><category>nvidia</category><category>stability-ai</category><category>runway</category><category>gemma-2-2b</category><category>gpt-3.5-turbo-0613</category><category>mixtral-8x7b</category><category>gen-3-alpha</category><category>segment-anything-model-2</category><category>stable-fast-3d</category><category>nicholas-carlini</category><category>chris-dixon</category><category>rasbt</category><category>benchmarking</category><category>adversarial-attacks</category><category>large-language-models</category><category>text-generation</category><category>multimodality</category><category>robotics</category><category>emotion-detection</category><category>structured-data-extraction</category><category>real-time-processing</category><category>teleoperation</category><category>3d-generation</category><category>text-to-video</category></item><item><title>Execuhires: Tempting The Wrath of Khan</title><link>https://news.smol.ai/issues/24-08-02-ainews-execuhires-tempting-the-wrath-of-khan/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-02-ainews-execuhires-tempting-the-wrath-of-khan/</guid><description>**Character.ai&apos;s $2.5b execuhire to Google** marks a significant leadership move alongside **Adept&apos;s $429m execuhire to Amazon** and **Inflection&apos;s $650m execuhire to Microsoft**. Despite strong user growth and content momentum, Character.ai&apos;s CEO Noam Shazeer returns to Google, signaling shifting vibes in the AI industry. **Google DeepMind&apos;s Gemini 1.5 Pro** tops Chatbot Arena benchmarks, outperforming **GPT-4o** and **Claude-3.5**, excelling in multilingual, math, and coding tasks. The launch of **Black Forest Labs&apos; FLUX.1** text-to-image model and **LangGraph Studio** agent IDE highlight ongoing innovation. **Llama 3.1 405B** is released as the largest open-source model, fostering developer use and competition with closed models. The industry is focusing increasingly on post-training and data as key competitive factors, raising questions about acquisition practices and regulatory scrutiny.</description><pubDate>Sat, 03 Aug 2024 01:48:48 GMT</pubDate><category>character.ai</category><category>google</category><category>adept</category><category>amazon</category><category>inflection</category><category>microsoft</category><category>stability-ai</category><category>black-forest-labs</category><category>schelling</category><category>google-deepmind</category><category>openai</category><category>anthropic</category><category>meta-ai-fair</category><category>lmsys</category><category>langchainai</category><category>gemini-1.5-pro</category><category>gpt-4o</category><category>claude-3.5</category><category>flux-1</category><category>llama-3-1-405b</category><category>noam-shazeer</category><category>mostafa-mostaque</category><category>david-friedman</category><category>rob-rombach</category><category>alexandr-wang</category><category>svpino</category><category>rohanpaul_ai</category><category>execuhire</category><category>model-benchmarking</category><category>multilinguality</category><category>math</category><category>coding</category><category>text-to-image</category><category>agent-ide</category><category>open-source-models</category><category>post-training</category><category>data-driven-performance</category></item><item><title>Rombach et al: FLUX.1 [pro|dev|schnell], $31m seed for Black Forest Labs</title><link>https://news.smol.ai/issues/24-08-01-ainews-rombach-et-al-flux1-proordevorschnell-dollar31m-seed-for-black-forest-labs/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-08-01-ainews-rombach-et-al-flux1-proordevorschnell-dollar31m-seed-for-black-forest-labs/</guid><description>**Stability AI** co-founder Rombach launched **FLUX.1**, a new text-to-image model with three variants: pro (API only), dev (open-weight, non-commercial), and schnell (Apache 2.0). FLUX.1 outperforms **Midjourney** and **Ideogram** based on Black Forest Labs&apos; ELO score and plans to expand into text-to-video. **Google DeepMind** released **Gemma-2 2B**, a 2 billion parameter open-source model that outperforms larger models like **GPT-3.5-Turbo-0613** and **Mixtral-8x7b** on Chatbot Arena, optimized with NVIDIA TensorRT-LLM. The release includes safety classifiers (ShieldGemma) and sparse autoencoder analysis (Gemma Scope). Discussions highlight benchmarking discrepancies and US government support for open-weight AI models. Critiques of AI coding tools&apos; productivity gains were also noted.</description><pubDate>Fri, 02 Aug 2024 01:05:39 GMT</pubDate><category>stability-ai</category><category>google-deepmind</category><category>nvidia</category><category>gemma-2-2b</category><category>gpt-3.5-turbo-0613</category><category>mixtral-8x7b</category><category>flux-1</category><category>rohanpaul_ai</category><category>fchollet</category><category>bindureddy</category><category>clementdelangue</category><category>ylecun</category><category>svpino</category><category>text-to-image</category><category>text-to-video</category><category>model-benchmarking</category><category>open-weight-models</category><category>model-distillation</category><category>safety-classifiers</category><category>sparse-autoencoders</category><category>ai-coding-tools</category></item><item><title>Gemma 2 2B + Scope + Shield</title><link>https://news.smol.ai/issues/24-07-31-ainews-gemma-2-2b-scope-shield/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-31-ainews-gemma-2-2b-scope-shield/</guid><description>**Gemma 2B**, a 2 billion parameter model trained on **2 trillion tokens** and distilled from a larger unnamed LLM, has been released by **Google DeepMind** and shows strong leaderboard performance despite weaknesses in math. The Gemma series, including 9B and 27B models, has gained popularity since its June release. The team also released 400 SAEs for interpretability, inspired by **Anthropic**&apos;s research. A finetuned classifier called ShieldGemma outperforms Meta&apos;s LlamaGuard in harm detection. Meanwhile, **Meta AI** announced **Llama-3.1-405B** reaching #3 on the Overall Arena leaderboard, and released **SAM 2**, a video and image segmentation model with significant speed improvements. **OpenAI** is rolling out an advanced Voice Mode to Plus users. **Perplexity AI** launched a Publishers Program with major media partners and a status page. **NVIDIA** introduced Project GR00T for scaling robot data using Apple Vision Pro and generative simulation. Interest in quantization for compressing LLMs is growing, and LLM-as-a-Judge implementations from Vicuna, AlpacaEval, and G-Eval highlight the effectiveness of simple prompts and domain-specific evaluation.</description><pubDate>Thu, 01 Aug 2024 01:33:32 GMT</pubDate><category>google-deepmind</category><category>anthropic</category><category>meta-ai-fair</category><category>openai</category><category>perplexity-ai</category><category>nvidia</category><category>lmsys</category><category>gemma-2b</category><category>gemma-2-9b</category><category>gemma-2-27b</category><category>llama-3-1-405b</category><category>sam-2</category><category>gpt-3.5</category><category>vicuna</category><category>alpacaeval</category><category>g-eval</category><category>knowledge-distillation</category><category>leaderboards</category><category>model-interpretability</category><category>finetuning</category><category>harm-detection</category><category>video-segmentation</category><category>voice</category><category>publishers-program</category><category>robotics-data-scaling</category><category>quantization</category><category>llm-evaluation</category><category>prompt-engineering</category></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-07-31-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-31-ainews-not-much-happened-today/</guid><description>**Meta** released **SAM 2**, a unified model for real-time object segmentation with a new dataset 4.5x larger and 53x more annotated than previous ones. **FastHTML**, a new Python web framework by **Jeremy Howard**, enables easy creation and deployment of interactive web apps. **Scale AI** launched the SEAL Leaderboard on adversarial robustness, topped by **Gemini 1.5 Pro** from **Google DeepMind**. **Apple** published a technical report on their Intelligence Foundation Language Models for on-device and server use. **Yann LeCun** emphasized the importance of open source AI in an article co-authored with Martin Casado and Ion Stoica. **Maarten Grootendorst**&apos;s &quot;Visual Guide to Quantization&quot; on efficient LLM inference went viral. **ChatGPT** started rolling out advanced voice and vision-enabled modes to select users. **Leonardo AI** was acquired by **Canva**. **Jim Fan** shared insights on Project Groot augmenting human demonstration data for robotics. **Midjourney v6.1** was released.</description><pubDate>Wed, 31 Jul 2024 07:04:15 GMT</pubDate><category>meta-ai-fair</category><category>google-deepmind</category><category>scale-ai</category><category>apple</category><category>canva</category><category>hugging-face</category><category>sam-2</category><category>gemini-1.5-pro</category><category>chatgpt</category><category>midjourney-v6.1</category><category>jeremyphoward</category><category>demis-hassabis</category><category>ylecun</category><category>maartengrootendorst</category><category>jimfan</category><category>object-segmentation</category><category>quantization</category><category>web-development-framework</category><category>adversarial-robustness</category><category>on-device-ai</category><category>open-source</category><category>robotics</category><category>voice</category><category>vision</category></item><item><title>Apple Intelligence Beta + Segment Anything Model 2</title><link>https://news.smol.ai/issues/24-07-29-ainews-apple-intelligence-beta-segment-anything-model-2/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-29-ainews-apple-intelligence-beta-segment-anything-model-2/</guid><description>**Meta** advanced its open source AI with a sequel to the **Segment Anything Model**, enhancing image segmentation with memory attention for video applications using minimal data and compute. **Apple Intelligence** delayed its official release to iOS 18.1 in October but launched developer previews on **MacOS Sequoia**, **iOS 18**, and **iPadOS 18**, accompanied by a detailed 47-page paper revealing extensive pretraining on **6.3T tokens** and use of **Cloud TPUs** rather than Apple Silicon. The paper highlights improvements in instruction following, reasoning, and writing through post-training and synthetic data. Benchmarks show Apple’s model scores lower than **Llama 3**, but with trusted human evaluations. Additionally, **Meta** released **Llama 3.1** with a 405B parameter model, marking a significant open-source frontier model release.</description><pubDate>Tue, 30 Jul 2024 02:45:55 GMT</pubDate><category>meta-ai-fair</category><category>apple</category><category>llama-3-405b</category><category>llama-3</category><category>segment-anything-model</category><category>bindureddy</category><category>maximelabonne</category><category>reach_vb</category><category>image-segmentation</category><category>memory-attention</category><category>video-processing</category><category>pretraining</category><category>cloud-tpus</category><category>post-training</category><category>synthetic-data</category><category>instruction-following</category><category>reasoning</category><category>writing</category><category>benchmarking</category></item><item><title>AlphaProof + AlphaGeometry2 reach 1 point short of IMO Gold</title><link>https://news.smol.ai/issues/24-07-25-ainews-alphaproof-alphageometry2-reach-1-point-short-of-imo-gold/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-25-ainews-alphaproof-alphageometry2-reach-1-point-short-of-imo-gold/</guid><description>**Search+Verifier** highlights advances in neurosymbolic AI during the 2024 Math Olympics. **Google DeepMind**&apos;s combination of **AlphaProof** and **AlphaGeometry 2** solved four out of six IMO problems, with AlphaProof being a finetuned **Gemini** model using an AlphaZero approach, and AlphaGeometry 2 trained on significantly more synthetic data with a novel knowledge-sharing mechanism. Despite impressive results, human judges noted the AI required much longer time than human competitors. Meanwhile, **Meta AI** released **Llama 3.1** with a 405B parameter model and smaller variants, and **Mistral AI** launched **Mistral Large 2** with 123B parameters and 128k context windows, outperforming Llama 3.1 on coding tasks and multilingual benchmarks. This marks significant progress in AI mathematical reasoning, model scaling, and multilingual capabilities.</description><pubDate>Fri, 26 Jul 2024 01:15:56 GMT</pubDate><category>google-deepmind</category><category>meta-ai-fair</category><category>mistral-ai</category><category>gemini</category><category>alphageometry-2</category><category>alphaproof</category><category>llama-3-1-405b</category><category>llama-3-70b</category><category>llama-3-8b</category><category>mistral-large-2</category><category>tim-gowers</category><category>guillaume-lample</category><category>osanseviero</category><category>neurosymbolic-ai</category><category>mathematical-reasoning</category><category>synthetic-data</category><category>knowledge-sharing</category><category>model-fine-tuning</category><category>alpha-zero</category><category>multilinguality</category><category>context-windows</category><category>model-scaling</category><category>benchmarking</category><category>performance-comparison</category></item><item><title>Mistral Large 2 + RIP Mistral 7B, 8x7B, 8x22B</title><link>https://news.smol.ai/issues/24-07-24-ainews-mistral-large-2-rip-mistral-7b-8x7b-8x22b/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-24-ainews-mistral-large-2-rip-mistral-7b-8x7b-8x22b/</guid><description>**Mistral Large 2** introduces **123B parameters** with **Open Weights** under a Research License, focusing on **code generation**, **math performance**, and a massive **128k context window**, improving over Mistral Large 1&apos;s 32k context. It claims better **function calling** capabilities than **GPT-4o** and enhanced reasoning. Meanwhile, **Meta** officially released **Llama-3.1** models including **Llama-3.1-70B** and **Llama-3.1-8B** with detailed pre-training and post-training insights. The **Llama-3.1 8B** model&apos;s 128k context performance was found underwhelming compared to **Mistral Nemo** and **Yi 34B 200K**. Mistral is deprecating older Apache open-source models, focusing on Large 2 and **Mistral Nemo 12B**. The news also highlights community discussions and benchmarking comparisons.</description><pubDate>Wed, 24 Jul 2024 23:44:31 GMT</pubDate><category>mistral-ai</category><category>meta-ai-fair</category><category>groq</category><category>togethercompute</category><category>mistral-large-2</category><category>mistral-nemo-12b</category><category>llama-3.1-8b</category><category>llama-3.1-70b</category><category>llama-3.1</category><category>llama-3-405b</category><category>yi-34b-200k</category><category>gpt-4o</category><category>code-generation</category><category>math</category><category>function-calling</category><category>reasoning</category><category>context-windows</category><category>model-deprecation</category><category>pretraining</category><category>posttraining</category><category>benchmarking</category></item><item><title>Llama 3.1: The Synthetic Data Model</title><link>https://news.smol.ai/issues/24-07-23-ainews-llama-31-the-synthetic-data-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-23-ainews-llama-31-the-synthetic-data-model/</guid><description>**Meta AI** has released **Llama 3.1**, including a **405B parameter model** that triggers regulatory considerations like the **EU AI Act** and **SB 1047**. The model incorporates extensive **synthetic data** techniques for **code**, **math**, **multilinguality**, **long context**, and **tool use** fine-tuning, with **RLHF** using synthetic preference data from **Llama 2**. The launch was coordinated across major inference providers, with **Groq** demonstrating **750 tokens per second** inference speed and **Fireworks** leading in pricing. The updated license explicitly allows synthetic data generation, marking a significant step in open frontier-class LLMs and cost-efficiency improvements since March.</description><pubDate>Wed, 24 Jul 2024 00:13:31 GMT</pubDate><category>meta-ai-fair</category><category>groq</category><category>fireworks</category><category>llama-3-405b</category><category>llama-3-1</category><category>llama-3</category><category>bindureddy</category><category>thomas</category><category>synthetic-data</category><category>fine-tuning</category><category>reinforcement-learning</category><category>multilinguality</category><category>long-context</category><category>tool-use</category><category>code-generation</category><category>math</category><category>model-licensing</category><category>inference-speed</category><category>model-deployment</category></item><item><title>Llama 3.1 Leaks: big bumps to 8B, minor bumps to 70b, and SOTA OSS 405b model</title><link>https://news.smol.ai/issues/24-07-22-ainews-llama-31-leaks-big-bumps-to-8b-minor-bumps-to-70b-and-sota-oss-405b-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-22-ainews-llama-31-leaks-big-bumps-to-8b-minor-bumps-to-70b-and-sota-oss-405b-model/</guid><description>**Llama 3.1** leaks reveal a **405B dense model** with **128k context length**, trained on **39.3M GPU hours** using H100-80GB GPUs, and fine-tuned with **over 25M synthetic examples**. The model shows significant benchmark improvements, especially for the 8B and 70B variants, with some evals suggesting the 70B outperforms **GPT-4o**. **GPT-4o Mini** launched as a cost-efficient variant with strong performance but some reasoning weaknesses. Synthetic datasets like **NuminaMath** enable models such as **Alibaba Qwen 2** to surpass GPT-4o and Claude 3.5 in math competitions. Discussions include reasoning task benchmarks and dataset building for improved reasoning.</description><pubDate>Tue, 23 Jul 2024 01:12:50 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>alibaba</category><category>llama-3-1-405b</category><category>llama-3-8b</category><category>llama-3-70b</category><category>llama-3-1-8b</category><category>gpt-4o</category><category>gpt-4o-mini</category><category>claude-3-5</category><category>qwen-2</category><category>swyx</category><category>philschmid</category><category>jjitsev</category><category>lewtun</category><category>teknium1</category><category>adcock_brett</category><category>multilinguality</category><category>code-generation</category><category>context-windows</category><category>model-training</category><category>synthetic-data</category><category>benchmarking</category><category>reasoning</category><category>fine-tuning</category><category>model-performance</category><category>dataset-release</category></item><item><title>DataComp-LM: the best open-data 7B model/benchmark/dataset</title><link>https://news.smol.ai/issues/24-07-19-ainews-datacomp-lm-the-best-open-data-7b-modelbenchmarkdataset/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-19-ainews-datacomp-lm-the-best-open-data-7b-modelbenchmarkdataset/</guid><description>**DataComp team** released a competitive **7B open data language model** trained on only **2.5T tokens** from the massive **DCLM-POOL dataset** of **240 trillion tokens**, showing superior scaling trends compared to FineWeb. **OpenAI** launched **GPT-4o mini**, a cost-effective model with **82% MMLU** and performance near GPT-4-Turbo, aimed at developers for broad applications. **NVIDIA and Mistral** jointly released the **Mistral NeMo 12B** model featuring a **128k token context window**, FP8 checkpoint, multilingual support, and Apache 2.0 licensing. **DeepSeek** announced **DeepSeek-V2-0628** as the top open-source model on the LMSYS Chatbot Arena leaderboard with strong rankings in coding, math, and hard prompts. This news highlights advances in dataset design, model efficiency, and open-source contributions in the AI community.</description><pubDate>Sat, 20 Jul 2024 02:08:36 GMT</pubDate><category>datacomp</category><category>hugging-face</category><category>openai</category><category>nvidia</category><category>mistral-ai</category><category>deepseek</category><category>mistral-nemo-12b</category><category>gpt-4o-mini</category><category>deepseek-v2-0628</category><category>mistral-7b</category><category>llama-3</category><category>gemma-2</category><category>qwen-2</category><category>sam-altman</category><category>guillaume-lample</category><category>philschmid</category><category>miramurati</category><category>dataset-design</category><category>scaling-laws</category><category>model-benchmarking</category><category>model-performance</category><category>fine-tuning</category><category>multilinguality</category><category>function-calling</category><category>context-windows</category><category>open-source-models</category><category>model-optimization</category><category>cost-efficiency</category><category>benchmarking</category></item><item><title>Mini, Nemo, Turbo, Lite - Smol models go brrr (GPT4o-mini version)</title><link>https://news.smol.ai/issues/24-07-18-ainews-mini-nemo-turbo-lite-smol-models-go-brrr-gpt4o-mini-version/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-18-ainews-mini-nemo-turbo-lite-smol-models-go-brrr-gpt4o-mini-version/</guid><description>**OpenAI** launched the **GPT-4o Mini**, a cost-efficient small model priced at **$0.15 per million input tokens** and **$0.60 per million output tokens**, aiming to replace **GPT-3.5 Turbo** with enhanced intelligence but some performance limitations. **DeepSeek** open-sourced **DeepSeek-V2-0628**, topping the LMSYS Chatbot Arena Leaderboard and emphasizing their commitment to contributing to the AI ecosystem. **Mistral AI** and **NVIDIA** released the **Mistral NeMo**, a **12B parameter** multilingual model with a record **128k token context window** under an **Apache 2.0 license**, sparking debates on benchmarking accuracy against models like **Meta Llama 8B**. Research breakthroughs include the **TextGrad** framework for optimizing compound AI systems via textual feedback differentiation and the **STORM** system improving article writing by **25%** through simulating diverse perspectives and addressing source bias. Developer tooling trends highlight **LangChain**&apos;s evolving context-aware reasoning applications and the **Modular** ecosystem&apos;s new official GPU support, including discussions on **Mojo** and **Keras 3.0** integration.</description><pubDate>Fri, 19 Jul 2024 00:13:31 GMT</pubDate><category>openai</category><category>deepseek-ai</category><category>mistral-ai</category><category>nvidia</category><category>meta-ai-fair</category><category>hugging-face</category><category>langchain</category><category>keras</category><category>gpt-4o-mini</category><category>deepseek-v2-0628</category><category>mistral-nemo</category><category>llama-8b</category><category>liang-wenfeng</category><category>cost-efficiency</category><category>context-windows</category><category>open-source</category><category>benchmarking</category><category>neural-networks</category><category>model-optimization</category><category>text-generation</category><category>fine-tuning</category><category>developer-tools</category><category>gpu-support</category><category>parallelization</category><category>cuda-integration</category><category>multilinguality</category><category>long-context</category><category>article-generation</category></item><item><title>Mini, Nemo, Turbo, Lite - Smol models go brrr (GPT4o version)</title><link>https://news.smol.ai/issues/24-07-18-ainews-mini-nemo-turbo-lite-smol-models-go-brrr-gpt4o-version/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-18-ainews-mini-nemo-turbo-lite-smol-models-go-brrr-gpt4o-version/</guid><description>**GPT-4o-mini** launches with a **99% price reduction** compared to text-davinci-003, offering **3.5% the price of GPT-4o** and matching Opus-level benchmarks. It supports **16k output tokens**, is faster than previous models, and will soon support **text, image, video, and audio inputs and outputs**. **Mistral Nemo**, a **12B parameter model** developed with **Nvidia**, features a **128k token context window**, FP8 checkpoint, and strong benchmark performance. **Together Lite and Turbo** offer fp8/int4 quantizations of **Llama 3** with up to **4x throughput** and significantly reduced costs. **DeepSeek V2** is now open-sourced. Upcoming releases include at least **5 unreleased models** and **Llama 4** leaks ahead of ICML 2024.</description><pubDate>Fri, 19 Jul 2024 00:00:39 GMT</pubDate><category>openai</category><category>nvidia</category><category>mistral-ai</category><category>togethercompute</category><category>deepseek-ai</category><category>lmsys</category><category>gpt-4o-mini</category><category>mistral-nemo</category><category>llama-3</category><category>llama-3-400b</category><category>deepseek-v2</category><category>sam-altman</category><category>model-quantization</category><category>context-windows</category><category>instruction-following</category><category>model-performance</category><category>cost-efficiency</category><category>multimodality</category><category>benchmarking</category><category>open-source</category><category>model-release</category></item><item><title>Gemma 2 tops /r/LocalLlama vibe check</title><link>https://news.smol.ai/issues/24-07-17-ainews-gemma-2-tops-rlocalllama-vibe-check/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-17-ainews-gemma-2-tops-rlocalllama-vibe-check/</guid><description>**Gemma 2 (9B, 27B)** is highlighted as a top-performing local LLM, praised for its speed, multilingual capabilities, and efficiency on consumer GPUs like the 2080ti. It outperforms models like **Llama 3** and **Mistral 7B** in various tasks, including non-English text processing and reasoning. The community discussion on /r/LocalLlama reflects strong preference for Gemma 2, with **18 mentions**, compared to **10 mentions** for Llama 3 and **9 mentions** for Mistral. Other models like **Phi 3** and **Qwen** also received mentions but are considered surpassed by Gemma 2. Additionally, **Andrej Karpathy** announced the launch of **Eureka Labs**, an AI+Education startup aiming to create an AI-native school with AI Teaching Assistants, starting with the **LLM101n** course to teach AI training fundamentals. This initiative is seen as a significant development in AI education.</description><pubDate>Wed, 17 Jul 2024 22:57:14 GMT</pubDate><category>gemma</category><category>llamaindex</category><category>mistral-ai</category><category>cohere</category><category>deepseek-ai</category><category>nous-research</category><category>eureka-labs</category><category>gemma-2-9b</category><category>gemma-2-27b</category><category>llama-3</category><category>mistral-7b</category><category>phi-3</category><category>qwen</category><category>andrej-karpathy</category><category>model-comparison</category><category>local-llms</category><category>multilinguality</category><category>model-efficiency</category><category>fine-tuning</category><category>ai-education</category><category>ai-teaching-assistants</category></item><item><title>SciCode: HumanEval gets a STEM PhD upgrade</title><link>https://news.smol.ai/issues/24-07-16-ainews-scicode-humaneval-gets-a-stem-phd-upgrade/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-16-ainews-scicode-humaneval-gets-a-stem-phd-upgrade/</guid><description>**PhD-level benchmarks** highlight the difficulty of coding scientific problems for LLMs, with **GPT-4** and **Claude 3.5 Sonnet** scoring under 5% on the new **SciCode** benchmark. **Anthropic** doubled the max output token limit for Claude 3.5 Sonnet to 8192 tokens. The **Q-GaLore** method enables training **LLaMA-7B** on a single 16GB GPU. The **Mosaic compiler** now generates efficient code for NVIDIA H100 GPUs. The **Dolphin 2.9.3-Yi-1.5-34B-32k-GGUF** model on Hugging Face has over 111k downloads. **Llama 3** shows strong performance, achieving 90% zero-shot accuracy on the MATH dataset. Discussions continue on the limitations and forms of synthetic data for model training.</description><pubDate>Wed, 17 Jul 2024 02:04:35 GMT</pubDate><category>anthropic</category><category>hugging-face</category><category>nvidia</category><category>gpt-4</category><category>claude-3.5-sonnet</category><category>llama-3-7b</category><category>llama-3</category><category>dolphin-2.9.3-yi-1.5-34b-32k-gguf</category><category>yi-tay</category><category>rohanpaul_ai</category><category>alexalbert__</category><category>tri_dao</category><category>abacaj</category><category>benchmarks</category><category>coding</category><category>model-training</category><category>gpu-optimization</category><category>model-performance</category><category>synthetic-data</category><category>compiler-optimization</category><category>zero-shot-learning</category></item><item><title>Microsoft AgentInstruct + Orca 3</title><link>https://news.smol.ai/issues/24-07-15-ainews-microsoft-agentinstruct-orca-3/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-15-ainews-microsoft-agentinstruct-orca-3/</guid><description>**Microsoft Research** released **AgentInstruct**, the third paper in its **Orca** series, introducing a generative teaching pipeline that produces **25.8 million** synthetic instructions to fine-tune **mistral-7b**, achieving significant performance gains: +40% AGIEval, +19% MMLU, +54% GSM8K, +38% BBH, +45% AlpacaEval, and a 31.34% reduction in hallucinations. This synthetic data approach follows the success of **FineWeb** and **Apple&apos;s Rephrasing research** in improving dataset quality. Additionally, **Tencent** claims to have generated **1 billion** diverse personas for synthetic data. On AI Twitter, notable discussions included a shooting incident at a Trump rally and recent ML research highlights such as **FlashAttention-3**, **RankRAG**, and **Mixture of A Million Experts**.</description><pubDate>Tue, 16 Jul 2024 00:42:03 GMT</pubDate><category>microsoft-research</category><category>apple</category><category>tencent</category><category>hugging-face</category><category>mistral-7b</category><category>orca-2.5</category><category>philschmid</category><category>sama</category><category>bindureddy</category><category>rohanpaul_ai</category><category>zachtratar</category><category>dair_ai</category><category>synthetic-data</category><category>fine-tuning</category><category>instruction-following</category><category>transformers</category><category>model-performance</category><category>hallucination-detection</category><category>dataset-quality</category><category>flashattention</category><category>mixture-of-experts</category></item><item><title>We Solved Hallucinations</title><link>https://news.smol.ai/issues/24-07-12-ainews-we-solved-hallucinations/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-12-ainews-we-solved-hallucinations/</guid><description>**Reddit&apos;s URL structure causes link errors in AI-generated summaries, especially with NSFW content affecting models like Claude and GPT-4.** The team fixed this glitch while still leveraging LLMs for summarizing Reddit content. **GPT-2 training costs have dramatically dropped to ~$672 using H100 GPUs and software improvements like CUDA and FlashAttention.** **FlashAttention-3 was released, achieving up to 740 TFLOPS on H100 GPUs, with FP8 nearing 1.2 PFLOPS, developed collaboratively by Meta, NVIDIA, Princeton, and Colfax.** Hopper GPUs enable major speedups with new hardware features. **Synthetic data may not improve vision tasks, as shown in recent research.** The **Avocado360 benchmark evaluates vision-language models&apos; ability to detect avocados in images.** **Lynx, a hallucination detection model for LLMs, was introduced for real-world healthcare and fintech applications, trained by Patronus AI on Databricks Mosaic AI using Composer.**</description><pubDate>Sat, 13 Jul 2024 02:52:26 GMT</pubDate><category>meta-ai-fair</category><category>nvidia</category><category>princeton</category><category>colfax</category><category>patronus-ai</category><category>databricks</category><category>mosaic-ai</category><category>openai</category><category>gpt-2</category><category>flashattention-3</category><category>lynx</category><category>karpathy</category><category>tri_dao</category><category>giffmana</category><category>vikhyatk</category><category>dbrxmosaicai</category><category>compute-hardware</category><category>gpu-optimization</category><category>flashattention</category><category>llm-evaluation</category><category>hallucination-detection</category><category>vision</category><category>benchmarking</category><category>synthetic-data</category><category>model-training</category></item><item><title>FlashAttention 3, PaliGemma, OpenAI&apos;s 5 Levels to Superintelligence</title><link>https://news.smol.ai/issues/24-07-12-ainews-flashattention-3-paligemma-openais-5-levels-to-superintelligence/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-12-ainews-flashattention-3-paligemma-openais-5-levels-to-superintelligence/</guid><description>**FlashAttention-3** introduces fast and accurate attention optimized for **H100 GPUs**, advancing native **FP8 training**. **PaliGemma**, a versatile **3B Vision-Language Model (VLM)** combining a SigLIP-So400m ViT encoder with the **Gemma-2B** language model, emphasizes a prefix-LM architecture for improved image-query interaction. **OpenAI** reveals a framework on levels of superintelligence, signaling progress toward Level 2 and highlighting internal safety disagreements. On Reddit, **NuminaMath 7B**, fine-tuned from **DeepSeekMath-7B**, wins the AI Math Olympiad by solving 29 problems using iterative supervised fine-tuning and tool-integrated reasoning. Open-source LLMs like **CodeLlama-34b** and **WizardCoder-Python-34B-V1.0** are closing the coding performance gap with closed models such as **ChatGPT-3.5**.</description><pubDate>Fri, 12 Jul 2024 09:31:43 GMT</pubDate><category>openai</category><category>together-ai</category><category>google</category><category>hugging-face</category><category>deepseek</category><category>code-llama</category><category>flashattention-3</category><category>paligemma-3b</category><category>gemma-2b</category><category>numinamath-7b</category><category>deepseekmath-7b</category><category>codellama-34b</category><category>wizardcoder-python-34b-v1.0</category><category>chatgpt-3.5</category><category>ilya-sutskever</category><category>lucas-giffman</category><category>attention-mechanisms</category><category>fp8-training</category><category>vision</category><category>prefix-lm</category><category>superintelligence</category><category>fine-tuning</category><category>chain-of-thought</category><category>tool-integrated-reasoning</category><category>self-consistency-decoding</category><category>python</category><category>coding-capabilities</category><category>elo-ratings</category></item><item><title>Nothing much happened today</title><link>https://news.smol.ai/issues/24-07-10-ainews-nothing-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-10-ainews-nothing-much-happened-today/</guid><description>**HuggingFace** released a browser-based timestamped Whisper using transformers.js. A Twitter bot by **truth_terminal** became the first &quot;semiautonomous&quot; bot to secure VC funding. **Microsoft** and **Apple** abruptly left the **OpenAI** board amid regulatory scrutiny. **Meta** is finalizing a major upgrade to Reddit comments addressing hallucination issues. The **Yi model** gained popularity on GitHub with 7.4K stars and 454 forks, with potential integration with **Axolotl** for pregeneration and preprocessing. **AMD** technologies enable household/small business AI appliances. **Meta** released **Chameleon-7b** and **Chameleon-30b** models on HuggingFace supporting unified text and image tokenization. **Salesforce**&apos;s **xLAM-1b** model outperforms **GPT-3.5** in function calling despite its smaller size. **Anole** pioneered open-source multimodal text-image-video generation up to 720p 144fps. **Phi-3 Mini** expanded from 3.8B to 4.7B parameters with function calling, competing with **Mistral-7b v3**. *&quot;System 2 distillation&quot;* in humans relates to automaticity and procedural memory.</description><pubDate>Thu, 11 Jul 2024 01:15:43 GMT</pubDate><category>huggingface</category><category>truth_terminal</category><category>microsoft</category><category>apple</category><category>openai</category><category>meta-ai-fair</category><category>yi</category><category>axolotl</category><category>amd</category><category>salesforce</category><category>chameleon-7b</category><category>chameleon-30b</category><category>xlam-1b</category><category>gpt-3.5</category><category>phi-3-mini</category><category>mistral-7b-v3</category><category>function-calling</category><category>multimodality</category><category>model-releases</category><category>model-updates</category><category>model-integration</category><category>automaticity</category><category>procedural-memory</category><category>text-image-video-generation</category></item><item><title>Test-Time Training, MobileLLM, Lilian Weng on Hallucination (Plus: Turbopuffer)</title><link>https://news.smol.ai/issues/24-07-09-ainews-test-time-training-mobilellm-lilian-weng-on-hallucination-plus-turbopuffer/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-09-ainews-test-time-training-mobilellm-lilian-weng-on-hallucination-plus-turbopuffer/</guid><description>**Lilian Weng** released a comprehensive literature review on **hallucination detection** and **anti-hallucination methods** including techniques like FactualityPrompt, SelfCheckGPT, and WebGPT. **Facebook AI Research (FAIR)** published **MobileLLM**, a sub-billion parameter on-device language model architecture achieving performance comparable to **llama-2-7b** with innovations like thin and deep models and shared weights. A new **RNN-based LLM architecture** with expressive hidden states was introduced, replacing attention mechanisms and scaling better than Mamba and Transformer models for long-context modeling. Additionally, **Tsinghua University** open sourced **CodeGeeX4-ALL-9B**, a multilingual code generation model excelling in code assistance.</description><pubDate>Wed, 10 Jul 2024 05:57:13 GMT</pubDate><category>facebook-research</category><category>meta-ai-fair</category><category>tsinghua-university</category><category>llama-2-7b</category><category>codegeex4-all-9b</category><category>mamba</category><category>lilian-weng</category><category>yann-lecun</category><category>hallucination-detection</category><category>anti-hallucination-methods</category><category>on-device-ai</category><category>model-architecture</category><category>rnn</category><category>long-context-modeling</category><category>model-scaling</category><category>expressive-hidden-states</category><category>code-generation</category></item><item><title>Problems with MMLU-Pro</title><link>https://news.smol.ai/issues/24-07-08-ainews-problems-with-mmlu-pro/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-08-ainews-problems-with-mmlu-pro/</guid><description>**MMLU-Pro** is gaining attention as the successor to MMLU on the **Open LLM Leaderboard V2** by **HuggingFace**, despite community concerns about evaluation discrepancies and prompt sensitivity affecting model performance, notably a **10-point improvement** in **Llama-3-8b-q8** with simple prompt tweaks. **Meta&apos;s MobileLLM** research explores running sub-billion parameter LLMs on smartphones using shared weights and deeper architectures. **Salesforce&apos;s APIGen** introduces an automated dataset generation system for function-calling tasks outperforming larger models. **Runway Gen-3 Alpha** launches an AI video generator for paid users creating realistic 10-second clips. **Nomic AI&apos;s GPT4All 3.0** offers an open-source desktop app supporting thousands of local models. AI assistants with multimodal capabilities and affordable access to multiple LLMs like ChatGPT, Claude, Llama, and Gemini are emerging. **Meta 3D Gen** advances text-to-3D asset generation, while Argil AI enables deepfake video creation from text threads. Research on transformer grokking and reasoning highlights advances in robust reasoning capabilities.</description><pubDate>Tue, 09 Jul 2024 00:20:51 GMT</pubDate><category>huggingface</category><category>meta-ai-fair</category><category>salesforce</category><category>runway</category><category>nomic-ai</category><category>pineapple</category><category>argil-ai</category><category>mmlu-pro</category><category>llama-3-8b-q8</category><category>gpt4all-3.0</category><category>chatgpt</category><category>claude</category><category>llama</category><category>gemini</category><category>mobilellm</category><category>runway-gen-3-alpha</category><category>meta-3d-gen</category><category>wenhu-chen</category><category>danhendrycks</category><category>clementine</category><category>ylecun</category><category>adcock_brett</category><category>svpino</category><category>rohanpaul_ai</category><category>benchmarking</category><category>prompt-engineering</category><category>model-evaluation</category><category>model-performance</category><category>multimodality</category><category>automated-dataset-generation</category><category>video-generation</category><category>open-source-models</category><category>ai-assistants</category><category>text-to-3d</category><category>deepfake</category><category>transformers</category><category>reasoning</category></item><item><title>Qdrant&apos;s BM42: &quot;Please don&apos;t trust us&quot;</title><link>https://news.smol.ai/issues/24-07-05-ainews-qdrants-bm42-please-dont-trust-us/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-05-ainews-qdrants-bm42-please-dont-trust-us/</guid><description>**Qdrant** attempted to replace BM25 and SPLADE with a new method called &quot;BM42&quot; combining transformer attention and collection-wide statistics for semantic and keyword search, but their evaluation using the Quora dataset was flawed. **Nils Reimers** from **Cohere** reran BM42 on better datasets and found it underperformed. Qdrant acknowledged the errors but still ran a suboptimal BM25 implementation. This highlights the importance of dataset choice and evaluation sanity checks in search model claims. Additionally, **Stripe** faced criticism for AI/ML model failures causing account and payment issues, prompting calls for alternatives. **Anthropic** revealed that **Claude 3.5 Sonnet** suppresses some answer parts with backend tags, sparking debate. **Gemma 2** model optimizations allow 2x faster fine-tuning with 63% less memory and longer context windows, running up to 34B parameters on consumer GPUs. **nanoLLaVA-1.5** was announced as a compact 1B parameter vision model with significant improvements.</description><pubDate>Sat, 06 Jul 2024 02:25:00 GMT</pubDate><category>qdrant</category><category>cohere</category><category>stripe</category><category>anthropic</category><category>hugging-face</category><category>stablequan_ai</category><category>claude-3.5-sonnet</category><category>gemma-2</category><category>nano-llava-1.5</category><category>nils-reimers</category><category>jeremyphoward</category><category>hamelhusain</category><category>rohanpaul_ai</category><category>semantic-search</category><category>benchmarking</category><category>dataset-quality</category><category>model-evaluation</category><category>model-optimization</category><category>vision</category><category>fine-tuning</category><category>context-windows</category></item><item><title>Not much happened today.</title><link>https://news.smol.ai/issues/24-07-03-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-03-ainews-not-much-happened-today/</guid><description>**Meta** introduced **Meta 3D Gen**, a system for end-to-end generation of 3D assets from text in under 1 minute, producing high-quality 3D assets with detailed textures. **Perplexity AI** updated Pro Search to handle deeper research with multi-step reasoning and code execution. **Microsoft** improved **Phi-3 Mini** with better long-context understanding and instruction following. **GPT4All 3.0** launched with support for thousands of models and major OS compatibility, featuring local file chat. **Yi-Large** model launched on Fireworks AI Playground. Research highlights include the evolution of **reinforcement learning from human feedback (RLHF)**, persona-driven data synthesis using a billion diverse personas, meta-tuning for few-shot generalization, and steering vectors for model behavior control. Tools updates include **LangSmith** improving memory retrieval and **Qdrant Engine v1.10** adding universal query API and multivector search.</description><pubDate>Wed, 03 Jul 2024 22:39:42 GMT</pubDate><category>meta</category><category>perplexity-ai</category><category>microsoft</category><category>gpt4all</category><category>langchainai</category><category>qdrant-engine</category><category>phi-3-mini</category><category>gpt4all-3.0</category><category>yi-large</category><category>meta-3d-gen</category><category>rohanpaul_ai</category><category>andriy_mulyar</category><category>cwolferesearch</category><category>sarahookr</category><category>3d-generation</category><category>long-context</category><category>instruction-following</category><category>reinforcement-learning-from-human-feedback</category><category>persona-driven-data-synthesis</category><category>meta-tuning</category><category>model-steering</category><category>memory-retrieval</category><category>multivector-search</category><category>universal-query-api</category></item><item><title>GraphRAG: The Marriage of Knowledge Graphs and RAG</title><link>https://news.smol.ai/issues/24-07-02-ainews-graphrag-the-marriage-of-knowledge-graphs-and-rag/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-02-ainews-graphrag-the-marriage-of-knowledge-graphs-and-rag/</guid><description>**Microsoft Research** open sourced **GraphRAG**, a retrieval augmented generation (RAG) technique that extracts knowledge graphs from sources and clusters them for improved LLM answers, though it increases token usage and inference time. **Gemma 2** models were released focusing on efficient small LLMs with innovations like sliding window attention and RMS norm, nearly matching the larger **Llama 3 70B**. **Anthropic&apos;s Claude 3.5 Sonnet** leads in instruction following and coding benchmarks, while **Nvidia&apos;s Nemotron 340B** model was released in June. **Qwen2-72B** tops the HuggingFace Open LLM leaderboard excelling in math and long-range reasoning. Discussions on RAG highlighted its limitations and improvements in context usage via function calls. A persona-driven synthetic data generation approach introduced 1 billion personas, with a fine-tuned model matching GPT-4 performance on math benchmarks at 7B scale. The **200GB AutoMathText dataset** was also noted for math data synthesis.</description><pubDate>Wed, 03 Jul 2024 01:30:30 GMT</pubDate><category>microsoft-research</category><category>anthropic</category><category>nvidia</category><category>hugging-face</category><category>gemma-2</category><category>llama-3-70b</category><category>claude-3.5-sonnet</category><category>nemotron-340b</category><category>qwen2-72b</category><category>llama-3</category><category>travis-fischer</category><category>rasbt</category><category>alexandr-wang</category><category>osanseviero</category><category>rohanpaul_ai</category><category>hamelhusain</category><category>svpino</category><category>aaaazzam</category><category>omarsar0</category><category>retrieval-augmented-generation</category><category>knowledge-graphs</category><category>token-usage</category><category>inference-time</category><category>attention-mechanisms</category><category>instruction-following</category><category>coding</category><category>math</category><category>long-range-reasoning</category><category>synthetic-data</category><category>dataset-release</category><category>fine-tuning</category><category>context-windows</category><category>function-calling</category></item><item><title>RouteLLM: RIP Martian? (Plus: AINews Structured Summaries update)</title><link>https://news.smol.ai/issues/24-07-01-ainews-routellm-rip-martian-plus-ainews-structured-summaries-update/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-07-01-ainews-routellm-rip-martian-plus-ainews-structured-summaries-update/</guid><description>**LMSys** introduces RouteLLM, an open-source router framework trained on **preference data** from Chatbot Arena, achieving **cost reductions over 85% on MT Bench, 45% on MMLU, and 35% on GSM8K** while maintaining **95% of GPT-4&apos;s performance**. This approach surpasses previous task-specific routing by using syntax-based Mixture of Experts (MoE) routing and data augmentation, beating commercial solutions by 40%. The update highlights advances in **LLM routing**, **cost-efficiency**, and **model performance optimization** across multiple models rather than single-model or MoE-level improvements. Additionally, the AI Twitter recap notes the **Gemma 2 model family** as a top open model, the **Block Transformer architecture** for improved inference throughput, and a proposal for a fully Software 2.0 computer vision system by **karpathy**.</description><pubDate>Tue, 02 Jul 2024 00:23:08 GMT</pubDate><category>lmsys</category><category>openai</category><category>gpt-4</category><category>gemma-2-27b</category><category>gemma-2-9b</category><category>karpathy</category><category>bindureddy</category><category>armand-joulin</category><category>llm-routing</category><category>cost-efficiency</category><category>model-performance</category><category>model-optimization</category><category>data-augmentation</category><category>syntax-based-routing</category><category>mixture-of-experts</category><category>inference-throughput</category><category>software-2.0</category><category>computer-vision</category></item><item><title>That GPT-4o Demo</title><link>https://news.smol.ai/issues/24-06-28-ainews-that-gpt-4o-demo/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-28-ainews-that-gpt-4o-demo/</guid><description>**Romain Huet** demonstrated an unreleased version of **GPT-4o** on ChatGPT Desktop showcasing capabilities like low latency voice generation, whisper tone moderation, camera mode streaming video to GPT-4o, rapid OCR, screen sharing with ChatGPT for programming help, clipboard reading, and vision-based code conversation. OpenAI&apos;s four investment areas highlighted include textual intelligence, efficiency/cost, model customization, and multimodal agents. **Google DeepMind** released **Gemma 2** models in 9B and 27B sizes trained on 8T and 13T tokens respectively, using SFT, distillation, RLHF, and model merging, optimized for TPUv5e with strong performance and safety measures. **Meta AI** announced the Meta LLM Compiler built on Meta Code Llama with enhanced code optimization and compiler features.</description><pubDate>Sat, 29 Jun 2024 00:48:47 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>meta-ai-fair</category><category>gpt-4o</category><category>gemma-2</category><category>meta-code-llama</category><category>romain-huet</category><category>fchollet</category><category>voice-generation</category><category>ocr</category><category>screen-sharing</category><category>vision</category><category>code-understanding</category><category>model-customization</category><category>efficiency</category><category>textual-intelligence</category><category>multimodal-agents</category><category>sft</category><category>distillation</category><category>rlhf</category><category>model-merging</category><category>model-optimization</category><category>safety</category></item><item><title>Gemma 2: The Open Model for Everyone</title><link>https://news.smol.ai/issues/24-06-27-ainews-gemma-2-the-open-model-for-everyone/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-27-ainews-gemma-2-the-open-model-for-everyone/</guid><description>**Gemma 2**, a **27B** parameter model from **google-deepmind**, was released with innovations like 1:1 local-global attention alternation and logit soft-capping, leveraging **knowledge distillation** to train smaller models on over 50× the compute-optimal token quantity. The model supports multilingual and multimodal capabilities, with fine-tuning success on over 200 Indic language variants. The **Open LLM Leaderboard** highlights **alibaba&apos;s Qwen 72B** as the top model, with **mistral-ai&apos;s Mixtral-8x22B-Instruct** also ranking highly. **Anthropic** launched **Claude 3.5 Sonnet**, improving intelligence at mid-tier cost and speed. Research on eliminating matrix multiplication in LLMs promises significant memory savings without performance loss. *Kathleen Kenealy* and *Daniel Han* provided insights on Gemma 2&apos;s tokenizer and attention scaling respectively.</description><pubDate>Fri, 28 Jun 2024 06:21:39 GMT</pubDate><category>google-deepmind</category><category>alibaba</category><category>mistral-ai</category><category>anthropic</category><category>gemma-2</category><category>qwen-72b</category><category>mixtral-8x22b-instruct</category><category>claude-3.5-sonnet</category><category>kathleen-kenealy</category><category>daniel-han</category><category>knowledge-distillation</category><category>attention-mechanisms</category><category>multilingual-models</category><category>multimodality</category><category>model-training</category><category>model-optimization</category><category>memory-optimization</category><category>fine-tuning</category></item><item><title>Mozilla&apos;s AI Second Act</title><link>https://news.smol.ai/issues/24-06-26-ainews-mozillas-ai-second-act/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-26-ainews-mozillas-ai-second-act/</guid><description>**Mozilla** showcased detailed live demos of **llamafile** and announced **sqlite-vec** for vector search integration at the AIE World&apos;s Fair. **LlamaIndex** launched **llama-agents**. **Anthropic** introduced new UI features and **Projects** for **Claude** with a 200K context window. **Etched AI** revealed a specialized inference chip claiming **500k tokens/sec**, though benchmark claims are questioned. **Sohu** chip enables **15 agent trajectories/sec**. **Tim Dettmers** shared theoretical GPU inference limits of ~300k tokens/sec for 8xB200 NVLink on 70B Llama. **Deepseek Coder v2** outperforms **Gemini** and GPT-4 variants in coding and reasoning. The **PyTorch documentary** launched to little attention.</description><pubDate>Thu, 27 Jun 2024 01:37:35 GMT</pubDate><category>mozilla</category><category>llamaindex</category><category>anthropic</category><category>etched-ai</category><category>sohu</category><category>deepseek</category><category>openai</category><category>llama-3</category><category>claude-3-opus</category><category>gemini-1.5</category><category>deepseek-coder-v2</category><category>gpt-4</category><category>justine-tunney</category><category>stephen-hood</category><category>tim-dettmers</category><category>bindureddy</category><category>vector-search</category><category>inference-speed</category><category>hardware-benchmarks</category><category>context-windows</category><category>open-source-models</category><category>coding</category><category>reasoning</category><category>model-benchmarking</category><category>gpu-inference</category><category>agentic-ai</category></item><item><title>Shall I compare thee to a Sonnet&apos;s day?</title><link>https://news.smol.ai/issues/24-06-25-ainews-shall-i-compare-thee-to-a-sonnets-day/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-25-ainews-shall-i-compare-thee-to-a-sonnets-day/</guid><description>**Claude 3.5 Sonnet** from **Anthropic** achieves top rankings in coding and hard prompt arenas, surpassing **GPT-4o** and competing with **Gemini 1.5 Pro** at lower cost. **Glif** demonstrates a fully automated **Wojak meme generator** using Claude 3.5 for JSON generation and ComfyUI for images, showcasing new JSON extractor capabilities. **Artifacts** enables rapid creation of niche apps, exemplified by a dual monitor visualizer made in under 5 minutes. **François Chollet** highlights that fusion energy is not a near-term solution compared to existing nuclear fission plants. **Mustafa Suleyman** notes that 75% of desk workers now use AI, marking a shift toward AI-assisted productivity.</description><pubDate>Wed, 26 Jun 2024 00:39:44 GMT</pubDate><category>anthropic</category><category>lmsys</category><category>glif</category><category>comfyui</category><category>claude-3.5-sonnet</category><category>claude-3.5</category><category>gpt-4o</category><category>gemini-1.5-pro</category><category>fchollet</category><category>mustafasuleyman</category><category>hard-prompts</category><category>json</category><category>json-extraction</category><category>meme-generation</category><category>instruction-following</category><category>app-development</category><category>fusion-energy</category><category>nuclear-fission</category><category>productivity</category></item><item><title>Gemini Nano: 50-90% of Gemini Pro, &lt;100ms inference, on device, in Chrome Canary</title><link>https://news.smol.ai/issues/24-06-25-ainews-gemini-nano-50-90percent-of-gemini-pro-less100ms-inference-on-device-in-chrome-canary/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-25-ainews-gemini-nano-50-90percent-of-gemini-pro-less100ms-inference-on-device-in-chrome-canary/</guid><description>The latest **Chrome Canary** now includes a feature flag for **Gemini Nano**, offering a prompt API and on-device optimization guide, with models Nano 1 and 2 at **1.8B** and **3.25B** parameters respectively, showing decent performance relative to Gemini Pro. The base and instruct-tuned model weights have been extracted and posted to **HuggingFace**. In AI model releases, **Anthropic** launched **Claude 3.5 Sonnet**, which outperforms **GPT-4o** on some benchmarks, is twice as fast as Opus, and is free to try. **DeepSeek-Coder-V2** achieves **90.2%** on HumanEval and **75.7%** on MATH, surpassing GPT-4-Turbo-0409, with models up to **236B** parameters and **128K** context length. **GLM-0520** from **Zhipu AI/Tsinghua** ranks highly in coding and overall benchmarks. **NVIDIA** announced **Nemotron-4 340B**, an open model family for synthetic data generation. Research highlights include **TextGrad**, a framework for automatic differentiation on textual feedback; **PlanRAG**, an iterative plan-then-RAG decision-making technique; a paper on **goldfish loss** to mitigate memorization in LLMs; and a tree search algorithm for language model agents.</description><pubDate>Tue, 25 Jun 2024 07:02:13 GMT</pubDate><category>google</category><category>gemini</category><category>huggingface</category><category>anthropic</category><category>deepseek</category><category>zhipu-ai</category><category>tsinghua</category><category>nvidia</category><category>gemini-nano</category><category>gemini-pro</category><category>claude-3.5-sonnet</category><category>gpt-4o</category><category>deepseek-coder-v2</category><category>glm-0520</category><category>nemotron-4-340b</category><category>gpt-4-turbo-0409</category><category>adcock_brett</category><category>dair_ai</category><category>lmsysorg</category><category>model-quantization</category><category>prompt-api</category><category>optimization</category><category>model-weights</category><category>benchmarking</category><category>code-generation</category><category>math</category><category>synthetic-data</category><category>automatic-differentiation</category><category>retrieval-augmented-generation</category><category>mitigating-memorization</category><category>tree-search</category><category>inference-time-algorithms</category></item><item><title>Shazeer et al (2024): you are overpaying for inference &gt;13x</title><link>https://news.smol.ai/issues/24-06-21-ainews-shazeer-et-al-2024-you-are-overpaying-for-inference-greater13x/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-21-ainews-shazeer-et-al-2024-you-are-overpaying-for-inference-greater13x/</guid><description>**Noam Shazeer** explains how **Character.ai** serves **20% of Google Search Traffic** for LLM inference while reducing serving costs by a factor of **33** compared to late 2022, with leading commercial APIs costing at least **13.5X more**. Key memory-efficiency techniques include **MQA &gt; GQA** reducing KV cache size by 8X, hybrid attention horizons, cross-layer KV-sharing, stateful caching with a 95% cache rate, and native int8 precision with custom kernels. **Anthropic** released **Claude 3.5 Sonnet**, which outperforms **Claude 3 Opus** at twice the speed and one-fifth the cost, passing **64%** of internal pull request tests and introducing new features like Artifacts for real-time doc and code generation. Discussions on LLM architecture highlight the dominance of transformers, challenges in scaling and overfitting, and the importance of architecture work for progress.</description><pubDate>Sat, 22 Jun 2024 00:48:48 GMT</pubDate><category>character.ai</category><category>anthropic</category><category>claude-3.5-sonnet</category><category>claude-3-opus</category><category>noam-shazeer</category><category>kevin-a-fischer</category><category>sebastien-bubeck</category><category>_aidan_clark_</category><category>andrej-karpathy</category><category>memory-efficiency</category><category>kv-cache</category><category>attention-mechanisms</category><category>stateful-caching</category><category>int8-precision</category><category>transformer-architecture</category><category>scaling</category><category>overfitting</category><category>architecture</category></item><item><title>Claude Crushes Code - 92% HumanEval and Claude.ai Artifacts</title><link>https://news.smol.ai/issues/24-06-21-ainews-claude-crushes-code-92percent-humaneval-and-claudeai-artifacts/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-21-ainews-claude-crushes-code-92percent-humaneval-and-claudeai-artifacts/</guid><description>**Claude 3.5 Sonnet**, released by **Anthropic**, is positioned as a Pareto improvement over Claude 3 Opus, operating at **twice the speed** and costing **one-fifth** as much. It achieves state-of-the-art results on benchmarks like **GPQA, MMLU, and HumanEval**, surpassing even **GPT-4o** and Claude 3 Opus on vision tasks. The model demonstrates significant advances in coding capabilities, passing **64% of test cases** compared to 38% for Claude 3 Opus, and is capable of autonomously fixing pull requests. Anthropic also introduced the **Artifacts** feature, enabling users to interact with AI-generated content such as code snippets and documents in a dynamic workspace, similar to OpenAI&apos;s Code Interpreter. This release highlights improvements in performance, cost-efficiency, and coding proficiency, signaling a growing role for LLMs in software development.</description><pubDate>Fri, 21 Jun 2024 07:27:45 GMT</pubDate><category>anthropic</category><category>openai</category><category>cognition</category><category>claude-3.5-sonnet</category><category>claude-3-opus</category><category>gpt-4o</category><category>alex-albert</category><category>benchmarking</category><category>model-performance</category><category>coding</category><category>model-optimization</category><category>fine-tuning</category><category>instruction-following</category><category>model-efficiency</category><category>model-release</category><category>api</category><category>performance-optimization</category></item><item><title>There&apos;s Ilya!</title><link>https://news.smol.ai/issues/24-06-19-ainews-theres-ilya/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-19-ainews-theres-ilya/</guid><description>**Ilya Sutskever** has co-founded **Safe Superintelligence Inc** shortly after leaving **OpenAI**, while **Jan Leike** moved to **Anthropic**. **Meta** released new models including **Chameleon 7B** and **34B** with mixed-modal input and unified token space quantization. **DeepSeek-Coder-V2** shows code capabilities comparable to **GPT-4 Turbo**, supporting **338 programming languages** and **128K context length**. **Consistency Large Language Models (CLLMs)** enable parallel decoding generating multiple tokens per step. **Grokked Transformers** demonstrate reasoning through training dynamics affecting memory formation and generalization. **VoCo-LLaMA** compresses vision tokens with LLMs improving video temporal correlation understanding. The **BigCodeBench** benchmark evaluates LLMs on **1,140 coding tasks** across **139 Python libraries**, topped by DeepSeek-Coder-V2 and Claude 3 Opus. **PixelProse** is a large **16M image-caption dataset** with reduced toxicity.</description><pubDate>Thu, 20 Jun 2024 00:18:00 GMT</pubDate><category>safe-superintelligence-inc</category><category>openai</category><category>anthropic</category><category>meta</category><category>deepseek</category><category>google-deepmind</category><category>chameleon-7b</category><category>chameleon-34b</category><category>deepseek-coder-v2</category><category>gpt-4-turbo</category><category>claude-3-opus</category><category>voco-llama</category><category>ilya-sutskever</category><category>jan-leike</category><category>ylecun</category><category>akhaliq</category><category>philschmid</category><category>rohanpaul_ai</category><category>mervenoyann</category><category>fchollet</category><category>parallel-decoding</category><category>code-generation</category><category>quantization</category><category>training-dynamics</category><category>vision</category><category>benchmarks</category><category>datasets</category><category>image-captioning</category><category>reasoning</category><category>memory-optimization</category></item><item><title>Gemini launches context caching... or does it?</title><link>https://news.smol.ai/issues/24-06-18-ainews-gemini-launches-context-caching-or-does-it/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-18-ainews-gemini-launches-context-caching-or-does-it/</guid><description>**Nvidia&apos;s Nemotron** ranks #1 open model on LMsys and #11 overall, surpassing **Llama-3-70b**. **Meta AI** released **Chameleon 7B/34B** models after further post-training. **Google&apos;s Gemini** introduced context caching, offering a cost-efficient middle ground between RAG and finetuning, with a minimum input token count of 33k and no upper limit on cache duration. **DeepSeek** launched **DeepSeek-Coder-V2**, a 236B parameter model outperforming **GPT-4 Turbo**, **Claude-3-Opus**, and **Gemini-1.5-Pro** in coding tasks, supporting 338 programming languages and extending context length to 128K. It was trained on 6 trillion tokens using the **Group Relative Policy Optimization (GRPO)** algorithm and is available on Hugging Face with a commercial license. These developments highlight advances in model performance, context caching, and large-scale coding models.</description><pubDate>Tue, 18 Jun 2024 21:26:50 GMT</pubDate><category>nvidia</category><category>meta-ai-fair</category><category>google</category><category>deepseek</category><category>hugging-face</category><category>nemotron</category><category>llama-3-70b</category><category>chameleon-7b</category><category>chameleon-34b</category><category>gemini-1.5-pro</category><category>deepseek-coder-v2</category><category>gpt-4-turbo</category><category>claude-3-opus</category><category>gemini-1.5-pro</category><category>rohanpaul_ai</category><category>_philschmid</category><category>aman-sanger</category><category>context-caching</category><category>model-performance</category><category>fine-tuning</category><category>reinforcement-learning</category><category>group-relative-policy-optimization</category><category>large-context</category><category>model-training</category><category>coding</category><category>model-release</category></item><item><title>Is this... OpenQ*?</title><link>https://news.smol.ai/issues/24-06-17-ainews-is-this-openq/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-17-ainews-is-this-openq/</guid><description>**DeepSeekCoder V2** promises GPT4T-beating performance at a fraction of the cost. **Anthropic** released new research on reward tampering. **Runway** launched their Sora response and Gen-3 Alpha video generation model. A series of papers explore &quot;test-time&quot; search techniques improving mathematical reasoning with models like **LLaMa-3 8B**. **Apple** announced Apple Intelligence with smarter Siri and image/document understanding, partnered with **OpenAI** to integrate ChatGPT into iOS 18, and released 20 new CoreML models with LoRA fine-tuning for specialization. **NVIDIA** released **Nemotron-4 340B**, an open model matching GPT-4 performance. **DeepSeek-Coder-V2** excels in coding and math with 338 programming languages and 128K context length. **Stability AI** released Stable Diffusion 3 Medium weights. **Luma Labs** launched Dream Machine for 5-second video generation from text and images.</description><pubDate>Tue, 18 Jun 2024 00:38:33 GMT</pubDate><category>deepseek_ai</category><category>anthropic</category><category>runwayml</category><category>openai</category><category>apple</category><category>nvidia</category><category>stability-ai</category><category>luma-labs</category><category>deepseek-coder-v2</category><category>llama-3-8b</category><category>nemotron-4-340b</category><category>stable-diffusion-3-medium</category><category>adcock_brett</category><category>clementdelangue</category><category>svpino</category><category>reward-tampering</category><category>test-time-search</category><category>mathematical-reasoning</category><category>process-supervision</category><category>fine-tuning</category><category>on-device-ai</category><category>video-generation</category><category>cost-efficiency</category><category>context-length</category><category>coding</category><category>image-understanding</category><category>multimodality</category></item><item><title>Nemotron-4-340B: NVIDIA&apos;s new large open models, built on syndata, great for syndata</title><link>https://news.smol.ai/issues/24-06-14-ainews-nemotron-4-340b-nvidias-new-large-open-models-built-on-syndata-great-for-syndata/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-14-ainews-nemotron-4-340b-nvidias-new-large-open-models-built-on-syndata-great-for-syndata/</guid><description>**NVIDIA** has scaled up its **Nemotron-4** model from **15B** to a massive **340B** dense model, trained on **9T tokens**, achieving performance comparable to **GPT-4**. The model alignment process uses over **98% synthetic data**, with only about **20K human-annotated samples** for fine-tuning and reward model training. The synthetic data generation pipeline is open-sourced, including synthetic prompts and preference data generation. The base and instruct versions outperform **Mixtral** and **Llama 3**, while the reward model ranks better than **Gemini 1.5**, **Cohere**, and **GPT-4o**. Other notable models include **Mamba-2-Hybrid 8B**, which is up to **8x faster** than Transformers and excels on long-context tasks, **Samba-3.8B-instruct** for infinite context length with linear complexity, **Dolphin-2.9.3** tiny models optimized for low-resource devices, and **Faro Yi 9B DPO** with a **200K context window** running efficiently on **16GB VRAM**. The Mixture-of-Agents technique boosts open-source LLMs beyond GPT-4 Omni on AlpacaEval 2.0.</description><pubDate>Fri, 14 Jun 2024 21:06:38 GMT</pubDate><category>nvidia</category><category>hugging-face</category><category>mistral-ai</category><category>llamaindex</category><category>cohere</category><category>gemini</category><category>mistral</category><category>nemotron-4-340b</category><category>mixtral</category><category>llama-3</category><category>gemini-1.5</category><category>gpt-4o</category><category>mamba-2-hybrid-8b</category><category>samba-3.8b-instruct</category><category>dolphin-2.9.3</category><category>faro-yi-9b-dpo</category><category>philipp-schmid</category><category>bryan-catanzaro</category><category>oleksii-kuchaiev</category><category>rohanpaul_ai</category><category>cognitivecompai</category><category>_philschmid</category><category>01ai_yi</category><category>synthetic-data</category><category>model-alignment</category><category>reward-models</category><category>fine-tuning</category><category>long-context</category><category>model-scaling</category><category>inference-speed</category><category>mixture-of-agents</category><category>open-source-models</category><category>model-training</category><category>instruction-following</category><category>context-windows</category></item><item><title>Hybrid SSM/Transformers &gt; Pure SSMs/Pure Transformers</title><link>https://news.smol.ai/issues/24-06-13-ainews-hybrid-ssmtransformers-greater-pure-ssmspure-transformers/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-13-ainews-hybrid-ssmtransformers-greater-pure-ssmspure-transformers/</guid><description>**NVIDIA**&apos;s Bryan Catanzaro highlights a new paper on **Mamba models**, showing that mixing Mamba and Transformer blocks outperforms either alone, with optimal attention below **20%**. **Mixture-of-Agents (MoA)** architecture improves LLM generation quality, scoring **65.1% on AlpacaEval 2.0** versus **GPT-4 Omni&apos;s 57.5%**. The **LiveBench AI benchmark** evaluates reasoning, coding, writing, and data analysis. A hybrid **Mamba-2-Hybrid** model with **7% attention** surpasses a Transformer on MMLU accuracy, jumping from **50% to 53.6%**. **GPT-4** performs better at temperature=1. **Qwen 72B** leads open-source models on LiveBench AI. **LaminiAI Memory Tuning** achieves **95% accuracy** on a SQL agent task, improving over instruction fine-tuning. **Sakana AI Lab** uses evolutionary strategies for preference optimization. **Luma Labs Dream Machine** demonstrates advanced text-to-video generation. The **MMWorld benchmark** evaluates multimodal video understanding, and **Table-LLaVa 7B** competes with GPT-4V on multimodal table tasks.</description><pubDate>Thu, 13 Jun 2024 20:52:25 GMT</pubDate><category>nvidia</category><category>lamini-ai</category><category>sakana-ai</category><category>luma-labs</category><category>mamba-2-hybrid</category><category>gpt-4</category><category>qwen-72b</category><category>table-llava-7b</category><category>bryan-catanzaro</category><category>bindureddy</category><category>ylecun</category><category>ctnzr</category><category>corbtt</category><category>realsharonzhou</category><category>andrew-n-carr</category><category>karpathy</category><category>_akhaliq</category><category>omarsar0</category><category>mixture-of-experts</category><category>benchmarking</category><category>fine-tuning</category><category>multimodality</category><category>text-to-video</category><category>model-performance</category><category>memory-optimization</category><category>preference-optimization</category><category>video-understanding</category><category>multimodal-tables</category></item><item><title>The Last Hurrah of Stable Diffusion?</title><link>https://news.smol.ai/issues/24-06-12-ainews-the-last-hurrah-of-stable-diffusion/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-12-ainews-the-last-hurrah-of-stable-diffusion/</guid><description>**Stability AI** launched **Stable Diffusion 3 Medium** with models ranging from **450M to 8B parameters**, featuring the MMDiT architecture and T5 text encoder for image text rendering. The community has shown mixed reactions following the departure of key researchers like Emad Mostaque. On AI models, **Llama 3 8B Instruct** shows strong evaluation correlation with **GPT-4**, while **Qwen 2 Instruct** surpasses Llama 3 on MMLU benchmarks. The **Mixture of Agents (MoA)** framework outperforms GPT-4o on AlpacaEval 2.0. Techniques like **Spectrum** and **QLoRA** enable efficient fine-tuning with less VRAM. Research on **grokking** reveals transformers can transition from memorization to generalization through extended training. Benchmark initiatives include the **$1M ARC Prize Challenge** for AGI progress and **LiveBench**, a live LLM benchmark to prevent dataset contamination. The **Character Codex Dataset** offers open data on over **15,000 characters** for RAG and synthetic data. The **MLX 0.2** tool enhances LLM experience on Apple Silicon Macs with improved UI and faster retrieval-augmented generation.</description><pubDate>Wed, 12 Jun 2024 22:08:29 GMT</pubDate><category>stability-ai</category><category>togethercompute</category><category>llama-3-8b</category><category>llama-3</category><category>qwen-2</category><category>gpt-4</category><category>gpt-4o</category><category>emad-mostaque</category><category>rohanpaul_ai</category><category>fchollet</category><category>mikeknoop</category><category>micahgoldblum</category><category>teknium1</category><category>rasbt</category><category>percyliang</category><category>model-architecture</category><category>fine-tuning</category><category>benchmarks</category><category>dataset-release</category><category>model-evaluation</category><category>reasoning</category><category>model-training</category><category>retrieval-augmented-generation</category><category>multimodality</category></item><item><title>Francois Chollet launches $1m ARC Prize</title><link>https://news.smol.ai/issues/24-06-11-ainews-francois-chollet-launches-dollar1m-arc-prize/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-11-ainews-francois-chollet-launches-dollar1m-arc-prize/</guid><description>**François Chollet** critiques current paths to **AGI**, emphasizing the importance of benchmarks that resist saturation and focus on skill acquisition and open-ended problem solving. The **ARC-AGI** puzzles exemplify &quot;easy for humans, hard for AI&quot; challenges to measure progress toward AGI. Meanwhile, **Apple** announces integration of **ChatGPT** into iOS, iPadOS, and macOS through a partnership with **OpenAI**, enabling AI-powered features like document summarization and photo analysis with privacy-preserving measures. Discussions highlight Apple&apos;s focus on deep AI integration and on-device models optimized with techniques like mixed-precision quantization, though some skepticism remains about their AI capabilities compared to **GPT-4**. Additionally, **Together Compute** introduces a Mixture of Agents approach achieving strong performance on **AlpacaEval 2.0**.</description><pubDate>Tue, 11 Jun 2024 23:42:03 GMT</pubDate><category>openai</category><category>apple</category><category>togethercompute</category><category>gpt-4</category><category>chatgpt</category><category>francois-chollet</category><category>karpathy</category><category>svpino</category><category>philschmid</category><category>clementdelangue</category><category>sama</category><category>gdb</category><category>miramurati</category><category>kevin-weil</category><category>sarah-friar</category><category>benchmarking</category><category>agi</category><category>pattern-recognition</category><category>skill-acquisition</category><category>privacy</category><category>on-device-ai</category><category>mixed-precision-quantization</category><category>mixture-of-experts</category><category>multimodality</category><category>agentic-ai</category></item><item><title>Talaria: Apple&apos;s new MLOps Superweapon</title><link>https://news.smol.ai/issues/24-06-10-ainews-talaria-apples-new-mlops-superweapon/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-10-ainews-talaria-apples-new-mlops-superweapon/</guid><description>**Apple Intelligence** introduces a small (~3B parameters) on-device model and a larger server model running on Apple Silicon with Private Cloud Compute, aiming to surpass **Google Gemma**, **Mistral Mixtral**, **Microsoft Phi**, and **Mosaic DBRX**. The on-device model features a novel lossless quantization strategy using mixed 2-bit and 4-bit LoRA adapters averaging 3.5 bits-per-weight, enabling dynamic adapter hot-swapping and efficient memory management. Apple credits the **Talaria** tool for optimizing quantization and model latency, achieving about 0.6 ms time-to-first-token latency and 30 tokens per second generation rate on iPhone 15 Pro. Apple focuses on an &quot;adapter for everything&quot; strategy with initial deployment on SiriKit and App Intents. Performance benchmarks rely on human graders, emphasizing consumer-level adequacy over academic dominance. The Apple ML blog also mentions an Xcode code-focused model and a diffusion model for Genmoji.</description><pubDate>Tue, 11 Jun 2024 06:41:05 GMT</pubDate><category>apple</category><category>google</category><category>mistral-ai</category><category>microsoft</category><category>mosaic</category><category>gemma</category><category>mixtral</category><category>phi</category><category>dbrx</category><category>craig-federighi</category><category>andrej-karpathy</category><category>quantization</category><category>on-device-ai</category><category>adapter-models</category><category>model-optimization</category><category>model-latency</category><category>lossless-quantization</category><category>low-bit-palletization</category><category>token-generation</category><category>model-benchmarking</category><category>human-evaluation</category></item><item><title>HippoRAG: First, do know(ledge) Graph</title><link>https://news.smol.ai/issues/24-06-07-ainews-hipporag-first-do-knowledge-graph/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-07-ainews-hipporag-first-do-knowledge-graph/</guid><description>**Alibaba** released new open-source **Qwen2** models ranging from **0.5B to 72B parameters**, achieving SOTA results on benchmarks like MMLU and HumanEval. Researchers introduced **Sparse Autoencoders** to interpret **GPT-4** neural activity, improving feature representation. The **HippoRAG** paper proposes a hippocampus-inspired retrieval augmentation method using knowledge graphs and Personalized PageRank for efficient multi-hop reasoning. New techniques like **Stepwise Internalization** enable implicit chain-of-thought reasoning in LLMs, enhancing accuracy and speed. The **Buffer of Thoughts (BoT)** method improves reasoning efficiency with significant cost reduction. A novel scalable MatMul-free LLM architecture competitive with SOTA Transformers at billion-parameter scale was also presented. *&quot;Single-Step, Multi-Hop retrieval&quot;* is highlighted as a key advancement in retrieval speed and cost.</description><pubDate>Fri, 07 Jun 2024 23:55:52 GMT</pubDate><category>alibaba</category><category>openai</category><category>qwen-2</category><category>gpt-4</category><category>hipporag</category><category>rohanpaul_ai</category><category>omarsar0</category><category>nabla_theta</category><category>huybery</category><category>knowledge-graphs</category><category>personalized-pagerank</category><category>multi-hop-retrieval</category><category>chain-of-thought</category><category>implicit-reasoning</category><category>sparse-autoencoders</category><category>model-interpretability</category><category>model-efficiency</category><category>model-architecture</category><category>fine-tuning</category><category>reinforcement-learning</category></item><item><title>Qwen 2 beats Llama 3 (and we don&apos;t know how)</title><link>https://news.smol.ai/issues/24-06-06-ainews-qwen-2-beats-llama-3-and-we-dont-know-how/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-06-ainews-qwen-2-beats-llama-3-and-we-dont-know-how/</guid><description>**Alibaba** released **Qwen 2** models under Apache 2.0 license, claiming to outperform **Llama 3** in open models with multilingual support in **29 languages** and strong benchmark scores like **MMLU 82.3** and **HumanEval 86.0**. **Groq** demonstrated ultra-fast inference speed on **Llama-3 70B** at **40,792 tokens/s** and running 4 Wikipedia articles in 200ms. Research on **sparse autoencoders (SAEs)** for interpreting **GPT-4** neural activity showed new training methods, metrics, and scaling laws. **Meta AI** announced the **No Language Left Behind (NLLB)** model capable of high-quality translations between **200 languages**, including low-resource ones. *&quot;Our post-training phase is designed with the principle of scalable training with minimal human annotation,&quot;* highlighting techniques like rejection sampling for math and execution feedback for coding.</description><pubDate>Thu, 06 Jun 2024 22:33:41 GMT</pubDate><category>alibaba</category><category>groq</category><category>meta-ai-fair</category><category>qwen-2</category><category>llama-3</category><category>llama-3-70b</category><category>gpt-4</category><category>nllb</category><category>philschmid</category><category>huybery</category><category>jonathanross321</category><category>awnihannun</category><category>gdb</category><category>nabla_theta</category><category>ylecun</category><category>multilinguality</category><category>benchmarking</category><category>inference-speed</category><category>sparse-autoencoders</category><category>scaling-laws</category><category>post-training</category><category>instruction-following</category><category>rejection-sampling</category><category>execution-feedback</category><category>model-release</category><category>multilingual-models</category><category>model-training</category></item><item><title>5 small news items</title><link>https://news.smol.ai/issues/24-06-05-ainews-5-small-news-items/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-05-ainews-5-small-news-items/</guid><description>**OpenAI** announces that ChatGPT&apos;s voice mode is &quot;coming soon.&quot; **Leopold Aschenbrenner** launched a 5-part AGI timelines series predicting a **trillion dollar cluster** from current AI progress. **Will Brown** released a comprehensive GenAI Handbook. **Cohere** completed a **$450 million funding round** at a **$5 billion valuation**. DeepMind research on **uncertainty quantification in LLMs** and an **xLSTM model** outperforming transformers were highlighted. Studies on the **geometry of concepts in LLMs** and methods to **eliminate matrix multiplication** for efficiency gains were shared. Discussions on **parameter-efficient fine-tuning (PEFT)** and **automated alignment of LLMs** were noted. New tools include **LangGraph** for AI agents, **LlamaIndex** with longer context windows, and **Hugging Face&apos;s** integration with **NVIDIA NIM** for Llama3. **Mistral AI** released a fine-tuning API for their models.</description><pubDate>Thu, 06 Jun 2024 02:50:37 GMT</pubDate><category>openai</category><category>cohere</category><category>deepmind</category><category>hugging-face</category><category>nvidia</category><category>mistral-ai</category><category>llama-3</category><category>xLSTM</category><category>leopold-aschenbrenner</category><category>will-brown</category><category>rohanpaul_ai</category><category>richardmcngo</category><category>omarsar0</category><category>hwchase17</category><category>clementdelangue</category><category>sophiamyang</category><category>uncertainty-quantification</category><category>parameter-efficient-fine-tuning</category><category>automated-alignment</category><category>model-efficiency</category><category>long-context</category><category>agentic-ai</category><category>fine-tuning</category><category>inference-optimization</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-06-04-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-04-ainews-not-much-happened-today/</guid><description>**Twelve Labs** raised **$50m** in Series A funding co-led by NEA and **NVIDIA&apos;s NVentures** to advance multimodal AI. **Livekit** secured **$22m** in funding. **Groq** announced running at **800k tokens/second**. OpenAI saw a resignation from Daniel Kokotajlo. Twitter users highlighted **Gemini 1.5 FlashModel** for high performance at low cost and **Gemini Pro** ranking #2 in Japanese language tasks. **Mixtral** models can run up to 8x faster on NVIDIA RTX GPUs using TensorRT-LLM. **Mamba-2** model architecture introduces state space duality for larger states and faster training, outperforming previous models. **Phi-3 Medium (14B)** and **Small (7B)** models benchmark near GPT-3.5-Turbo-0613 and Llama 3 8B. Prompt engineering is emphasized for unlocking LLM capabilities. Data quality is critical for model performance, with upcoming masterclasses on data curation. Discussions on AI safety include a Frontier AI lab employee letter advocating whistleblower protections and debates on aligning AI to user intent versus broader humanity interests.</description><pubDate>Tue, 04 Jun 2024 23:53:47 GMT</pubDate><category>twelve-labs</category><category>livekit</category><category>groq</category><category>openai</category><category>nea</category><category>nvidia</category><category>lmsys</category><category>mistral-ai</category><category>gemini-1.5-flashmodel</category><category>gemini-pro</category><category>mixtral</category><category>mamba-2</category><category>phi-3-medium</category><category>phi-3-small</category><category>gpt-3.5-turbo-0613</category><category>llama-3-8b</category><category>llama-2-70b</category><category>mistral-finetune</category><category>daniel-kokotajlo</category><category>rohanpaul_ai</category><category>_arohan_</category><category>tri_dao</category><category>_albertgu</category><category>_philschmid</category><category>sarahcat21</category><category>hamelhusain</category><category>jachiam0</category><category>willdepue</category><category>teknium1</category><category>model-performance</category><category>prompt-engineering</category><category>data-curation</category><category>ai-safety</category><category>model-benchmarking</category><category>model-optimization</category><category>training</category><category>sequence-models</category><category>state-space-models</category></item><item><title>Mamba-2: State Space Duality</title><link>https://news.smol.ai/issues/24-06-03-ainews-mamba-2-state-space-duality/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-06-03-ainews-mamba-2-state-space-duality/</guid><description>**Mamba-2**, a new **state space model (SSM)**, outperforms previous models like Mamba and Transformer++ in **perplexity** and **wall-clock time**, featuring **8x larger states** and **50% faster training**. It introduces the concept of **state space duality (SSD)** connecting SSMs and linear attention. The **FineWeb-Edu dataset**, a high-quality subset of the **15 trillion token FineWeb dataset**, filtered using **llama-3-70b** for educational quality, enables better and faster LLM learning, potentially reducing tokens needed to surpass **GPT-3** performance. Additionally, perplexity-based data pruning using a **125M parameter model** improves downstream performance and reduces pretraining steps by up to **1.45x**. The **Video-MME benchmark** evaluates multi-modal LLMs on video analysis across multiple visual domains and video lengths.</description><pubDate>Mon, 03 Jun 2024 21:31:26 GMT</pubDate><category>hugging-face</category><category>mamba-2</category><category>mamba</category><category>transformer++</category><category>llama-3-70b</category><category>gpt-3</category><category>_albertgu</category><category>tri_dao</category><category>arankomatsuzaki</category><category>_akhaliq</category><category>clementdelangue</category><category>karpathy</category><category>state-space-models</category><category>perplexity</category><category>training-efficiency</category><category>data-pruning</category><category>benchmarking</category><category>multimodality</category><category>video-analysis</category></item><item><title>Ways to use Anthropic&apos;s Tool Use GA</title><link>https://news.smol.ai/issues/24-05-31-ainews-ways-to-use-anthropics-tool-use-ga/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-31-ainews-ways-to-use-anthropics-tool-use-ga/</guid><description>**Anthropic** launched general availability of tool use/function calling with support for streaming, forced use, and vision, alongside **Amazon** and **Google**. Alex Albert shared five architectures for agentic tool use: delegation, parallelization, debate, specialization, and tool suite experts. **Anthropic** also introduced a self-guided course on tool use. **Yann LeCun** emphasized ethical open science funding, gradual emergence of superintelligence with safety guardrails, and convolutional networks for image/video processing as competitive with vision transformers. He also noted growth in AI researchers across industry, academia, and government.</description><pubDate>Fri, 31 May 2024 20:31:29 GMT</pubDate><category>anthropic</category><category>amazon</category><category>google</category><category>claude-3-opus</category><category>haiku</category><category>opus</category><category>convnext</category><category>yann-lecun</category><category>alex-albert</category><category>sainingxie</category><category>tool-use</category><category>function-calling</category><category>agentic-ai</category><category>streaming</category><category>vision</category><category>parallelization</category><category>delegation</category><category>debate</category><category>specialization</category><category>open-science</category><category>superintelligence</category><category>convolutional-networks</category><category>self-attention</category><category>ai-research</category></item><item><title>Contextual Position Encoding (CoPE)</title><link>https://news.smol.ai/issues/24-05-30-ainews-contextual-position-encoding-cope/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-30-ainews-contextual-position-encoding-cope/</guid><description>**Meta AI** researcher **Jason Weston** introduced **CoPE**, a novel positional encoding method for transformers that incorporates *context* to create learnable gates, enabling improved handling of counting and copying tasks and better performance on language modeling and coding. The approach can potentially be extended with external memory for gate calculation. **Google DeepMind** released **Gemini 1.5 Flash** and **Pro** models optimized for fast inference. **Anthropic** announced general availability of tool use for **Claude**, enhancing its ability to orchestrate tools for complex tasks. **Alexandr Wang** launched **SEAL Leaderboards** for private, expert evaluations of frontier models. **Karpathy** reflected on the 4th anniversary of **GPT-3**, emphasizing scaling and practical improvements. **Perplexity AI** launched **Perplexity Pages** to convert research into visually appealing articles, described as an &quot;AI Wikipedia&quot; by **Arav Srinivas**.</description><pubDate>Fri, 31 May 2024 03:11:48 GMT</pubDate><category>meta-ai-fair</category><category>google-deepmind</category><category>anthropic</category><category>perplexity-ai</category><category>langchain</category><category>openai</category><category>cope</category><category>gemini-1.5-flash</category><category>gemini-1.5-pro</category><category>claude</category><category>gpt-3</category><category>jason-weston</category><category>alexandr-wang</category><category>karpathy</category><category>arav-srinivas</category><category>positional-encoding</category><category>transformers</category><category>counting</category><category>copying</category><category>language-modeling</category><category>coding</category><category>external-memory</category><category>tool-use</category><category>model-evaluation</category><category>inference-speed</category><category>model-benchmarking</category><category>scaling</category><category>research-synthesis</category></item><item><title>1 TRILLION token context, real time, on device?</title><link>https://news.smol.ai/issues/24-05-29-ainews-1-trillion-token-context-real-time-on-device/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-29-ainews-1-trillion-token-context-real-time-on-device/</guid><description>**Cartesia**, a startup specializing in **state space models (SSMs)**, launched a low latency voice model outperforming transformer-based models with **20% lower perplexity**, **2x lower word error**, and **1 point higher NISQA quality**. This breakthrough highlights the potential for models that can continuously process and reason over massive streams of multimodal data (text, audio, video) with a **trillion token context window** on-device. The news also covers recent AI developments including **Mistral&apos;s Codestral weights release**, **Schedule Free optimizers** paper release, and **Scale AI&apos;s** new elo-style eval leaderboards. Additionally, a debate between **yann-lecun** and **elon-musk** on the importance of publishing AI research versus engineering achievements was noted. The **Gemini 1.5 Pro/Advanced** models were mentioned for their strong performance.</description><pubDate>Wed, 29 May 2024 23:01:07 GMT</pubDate><category>cartesia</category><category>mistral-ai</category><category>scale-ai</category><category>gemini-1.5-pro</category><category>gemini-1.5</category><category>yann-lecun</category><category>elon-musk</category><category>state-space-models</category><category>voice-models</category><category>multimodality</category><category>model-performance</category><category>on-device-ai</category><category>long-context</category><category>evaluation-leaderboards</category><category>learning-rate-optimization</category><category>scientific-publishing</category><category>research-vs-engineering</category></item><item><title>Somebody give Andrej some H100s already</title><link>https://news.smol.ai/issues/24-05-28-ainews-somebody-give-andrej-some-h100s-already/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-28-ainews-somebody-give-andrej-some-h100s-already/</guid><description>**OpenAI**&apos;s GPT-2 sparked controversy five years ago for being &quot;too dangerous to release.&quot; Now, with **FineWeb** and **llm.c**, a tiny GPT-2 model can be trained in **90 minutes** for **$20** using **8xA100** GPUs, with the full 1.6B model estimated to take **1 week** and **$2.5k**. The project is notable for its heavy use of **CUDA** (75.8%) aiming to simplify the training stack. Meanwhile, a Twitter debate between **Yann LeCun** and **Elon Musk** highlighted the importance of **convolutional neural networks (CNNs)** in real-time image processing for autonomous driving, with LeCun emphasizing scientific research&apos;s role in technological progress. LeCun also criticized AI doomsday scenarios, arguing for cautious optimism about AI safety and regulation.</description><pubDate>Wed, 29 May 2024 01:24:27 GMT</pubDate><category>openai</category><category>fineweb</category><category>meta-ai-fair</category><category>nvidia</category><category>tesla</category><category>gpt-2</category><category>andrej-karpathy</category><category>yann-lecun</category><category>elon-musk</category><category>francois-chollet</category><category>svpino</category><category>mervenoyann</category><category>cuda</category><category>fine-tuning</category><category>training-time</category><category>gpu-acceleration</category><category>convolutional-neural-networks</category><category>real-time-processing</category><category>ai-safety</category><category>ai-regulation</category></item><item><title>Life after DPO (RewardBench)</title><link>https://news.smol.ai/issues/24-05-27-ainews-life-after-dpo-rewardbench/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-27-ainews-life-after-dpo-rewardbench/</guid><description>**xAI raised $6 billion at a $24 billion valuation**, positioning it among the most highly valued AI startups, with expectations to fund **GPT-5 and GPT-6 class models**. The **RewardBench** tool, developed by Nathan Lambert, evaluates reward models (RMs) for language models, showing Cohere&apos;s RMs outperforming open-source alternatives. The discussion highlights the evolution of language models from Claude Shannon&apos;s 1948 model to GPT-3 and beyond, emphasizing the role of **RLHF (Reinforcement Learning from Human Feedback)** and the newer **DPO (Direct Preference Optimization)** method. Notably, some **Llama 3 8B reward model-focused models** are currently outperforming GPT-4, Cohere, Gemini, and Claude on the RewardBench leaderboard, raising questions about reward hacking. Future alignment research directions include improving preference datasets, DPO techniques, and personalization in language models. The report also compares xAI&apos;s valuation with OpenAI, Mistral AI, and Anthropic, noting speculation about xAI&apos;s spending on Nvidia hardware.</description><pubDate>Tue, 28 May 2024 00:04:01 GMT</pubDate><category>x-ai</category><category>openai</category><category>mistral-ai</category><category>anthropic</category><category>cohere</category><category>meta-ai-fair</category><category>hugging-face</category><category>nvidia</category><category>gpt-3</category><category>gpt-4</category><category>gpt-5</category><category>gpt-6</category><category>llama-3-8b</category><category>llama-3</category><category>claude-3</category><category>gemini</category><category>nathan-lambert</category><category>chris-manning</category><category>elon-musk</category><category>bindureddy</category><category>rohanpaul_ai</category><category>nearcyan</category><category>reinforcement-learning-from-human-feedback</category><category>direct-preference-optimization</category><category>reward-models</category><category>rewardbench</category><category>language-model-history</category><category>model-evaluation</category><category>alignment-research</category><category>preference-datasets</category><category>personalization</category><category>transformer-architecture</category></item><item><title>Ten Commandments for Deploying Fine-Tuned Models</title><link>https://news.smol.ai/issues/24-05-24-ainews-ten-commandments-for-deploying-fine-tuned-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-24-ainews-ten-commandments-for-deploying-fine-tuned-models/</guid><description>**Gemini-in-Google-Slides** is highlighted as a useful tool for summarizing presentations. Kyle Corbitt&apos;s talk on deploying fine-tuned models in production emphasizes avoiding fine-tuning unless necessary, focusing on prompting, data quality, appropriate model choice, and thorough evaluation. **Anthropic** showcased feature alteration in **Claude AI**, demonstrating control over model behavior and increased understanding of large language models. Open-source models like **GPT-4o** are approaching closed-source performance on benchmarks like MMLU for simple tasks, though advanced models remain necessary for complex automation.</description><pubDate>Fri, 24 May 2024 22:12:57 GMT</pubDate><category>anthropic</category><category>google</category><category>openai</category><category>claude-3-opus</category><category>claude-3</category><category>gpt-4o</category><category>kyle-corbitt</category><category>bindureddy</category><category>alexalbert__</category><category>fine-tuning</category><category>prompt-engineering</category><category>model-evaluation</category><category>feature-alteration</category><category>benchmarking</category><category>model-performance</category><category>open-source-models</category></item><item><title>Clémentine Fourrier on LLM evals</title><link>https://news.smol.ai/issues/24-05-23-ainews-clementine-fourrier-on-llm-evals/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-23-ainews-clementine-fourrier-on-llm-evals/</guid><description>**Clémentine Fourrier** from **Huggingface** presented at **ICLR** about **GAIA** with **Meta** and shared insights on **LLM evaluation** methods. The blog outlines three main evaluation approaches: **Automated Benchmarking** using sample inputs/outputs and metrics, **Human Judges** involving grading and ranking with methods like **Vibe-checks**, **Arena**, and **systematic annotations**, and **Models as Judges** using generalist or specialist models with noted biases. Challenges include data contamination, subjectivity, and bias in scoring. These evaluations help prevent regressions, rank models, and track progress in the field.</description><pubDate>Thu, 23 May 2024 23:34:22 GMT</pubDate><category>huggingface</category><category>meta-ai-fair</category><category>claude-3-opus</category><category>clem_fourrier</category><category>llm-evaluation</category><category>automated-benchmarking</category><category>human-evaluation</category><category>model-bias</category><category>data-contamination</category><category>elo-ranking</category><category>systematic-annotations</category><category>preference-learning</category><category>evaluation-metrics</category><category>prompt-sensitivity</category></item><item><title>ALL of AI Engineering in One Place</title><link>https://news.smol.ai/issues/24-05-22-ainews-all-of-ai-engineering-in-one-place/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-22-ainews-all-of-ai-engineering-in-one-place/</guid><description>The upcoming **AI Engineer World&apos;s Fair** in San Francisco from **June 25-27** will feature a significantly expanded format with booths, talks, and workshops from **top model labs** like **OpenAI, DeepMind, Anthropic, Mistral, Cohere, HuggingFace**, and **Character.ai**. It includes participation from **Microsoft Azure, Amazon AWS, Google Vertex**, and major companies such as **Nvidia, Salesforce, Mastercard, Palo Alto Networks**, and more. The event covers **9 tracks** including **RAG, multimodality, evals/ops, open models, code generation, GPUs, agents, AI in Fortune 500**, and a new **AI leadership** track. Additionally, **Anthropic** shared interpretability research on **Claude 3 Sonnet**, revealing millions of interpretable features that can be steered to modify model behavior, including safety-relevant features related to bias and unsafe content, though more research is needed for practical applications. The event offers a discount code for AI News readers.</description><pubDate>Thu, 23 May 2024 01:22:53 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>mistral-ai</category><category>cohere</category><category>hugging-face</category><category>adept</category><category>midjourney</category><category>character-ai</category><category>microsoft</category><category>amazon</category><category>nvidia</category><category>salesforce</category><category>mastercard</category><category>palo-alto-networks</category><category>axa</category><category>novartis</category><category>discord</category><category>twilio</category><category>tinder</category><category>khan-academy</category><category>sourcegraph</category><category>mongodb</category><category>neo4j</category><category>hasura</category><category>modular</category><category>cognition</category><category>anysphere</category><category>perplexity-ai</category><category>groq</category><category>mozilla</category><category>nous-research</category><category>galileo</category><category>unsloth</category><category>langchain</category><category>llamaindex</category><category>instructor</category><category>weights-biases</category><category>lambda-labs</category><category>neptune</category><category>datastax</category><category>crusoe</category><category>covalent</category><category>qdrant</category><category>baseten</category><category>e2b</category><category>octo-ai</category><category>gradient-ai</category><category>lancedb</category><category>log10</category><category>deepgram</category><category>outlines</category><category>crew-ai</category><category>factory-ai</category><category>claude-3-sonnet</category><category>claude-3</category><category>interpretability</category><category>feature-steering</category><category>safety</category><category>multilinguality</category><category>multimodality</category><category>rag</category><category>evals-ops</category><category>open-models</category><category>code-generation</category><category>gpus</category><category>agents</category><category>ai-leadership</category></item><item><title>Anthropic&apos;s &quot;LLM Genome Project&quot;: learning &amp; clamping 34m features on Claude Sonnet</title><link>https://news.smol.ai/issues/24-05-21-ainews-anthropics-llm-genome-project-learning-and-clamping-34m-features-on-claude-sonnet/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-21-ainews-anthropics-llm-genome-project-learning-and-clamping-34m-features-on-claude-sonnet/</guid><description>**Anthropic** released their third paper in the MechInterp series, **Scaling Monosemanticity**, scaling interpretability analysis to **34 million features** on **Claude 3 Sonnet**. This work introduces the concept of **dictionary learning** to isolate recurring neuron activation patterns, enabling more interpretable internal states by combining features rather than neurons. The paper reveals abstract features related to code, errors, sycophancy, crime, self-representation, and deception, demonstrating intentional modifiability by clamping feature values. The research marks a significant advance in **model interpretability** and **neural network analysis** at frontier scale.</description><pubDate>Tue, 21 May 2024 22:47:46 GMT</pubDate><category>anthropic</category><category>scale-ai</category><category>suno-ai</category><category>microsoft</category><category>claude-3-sonnet</category><category>claude-3</category><category>emmanuel-ameisen</category><category>alex-albert</category><category>model-interpretability</category><category>dictionary-learning</category><category>neural-networks</category><category>feature-activation</category><category>intentional-modifiability</category><category>scaling</category><category>mechanistic-interpretability</category></item><item><title>Skyfall</title><link>https://news.smol.ai/issues/24-05-20-ainews-skyfall/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-20-ainews-skyfall/</guid><description>Between 5/17 and 5/20/2024, key AI updates include **Google DeepMind&apos;s Gemini 1.5 Pro and Flash models**, featuring sparse multimodal MoE architecture with up to **10M context** and a dense Transformer decoder that is **3x faster and 10x cheaper**. **Yi AI released Yi-1.5 models** with extended context windows of **32K and 16K tokens**. Other notable releases include **Kosmos 2.5 (Microsoft), PaliGemma (Google), Falcon 2, DeepSeek v2 lite, and HunyuanDiT diffusion model**. Research highlights feature an **Observational Scaling Laws paper** predicting model performance across families, a **Layer-Condensed KV Cache** technique boosting inference throughput by **up to 26×**, and the **SUPRA method** converting LLMs into RNNs for reduced compute costs. Hugging Face expanded local AI capabilities enabling on-device AI without cloud dependency. LangChain updated its v0.2 release with improved documentation. The community also welcomed a new LLM Finetuning Discord by Hamel Husain and Dan Becker for Maven course users. *&quot;Hugging Face is profitable, or close to profitable,&quot;* enabling $10 million in free shared GPUs for developers.</description><pubDate>Mon, 20 May 2024 23:02:42 GMT</pubDate><category>google-deepmind</category><category>yi-ai</category><category>microsoft</category><category>hugging-face</category><category>langchain</category><category>maven</category><category>gemini-1.5-pro</category><category>gemini-1.5-flash</category><category>yi-1.5</category><category>kosmos-2.5</category><category>paligemma</category><category>falcon-2</category><category>deepseek-v2</category><category>hunyuan-dit</category><category>gemini-1.5</category><category>gemini-1.5-flash</category><category>yi-1.5</category><category>hamel-husain</category><category>dan-becker</category><category>clement-delangue</category><category>philschmid</category><category>osanseviero</category><category>arankomatsuzaki</category><category>jason-wei</category><category>rohanpaul_ai</category><category>multimodality</category><category>mixture-of-experts</category><category>transformer</category><category>model-optimization</category><category>long-context</category><category>model-performance</category><category>model-inference</category><category>fine-tuning</category><category>local-ai</category><category>scaling-laws</category><category>causal-models</category><category>hallucination-detection</category><category>model-distillation</category><category>model-efficiency</category></item><item><title>Chameleon: Meta&apos;s (unreleased) GPT4o-like Omnimodal Model</title><link>https://news.smol.ai/issues/24-05-17-ainews-chameleon-metas-unreleased-gpt4o-like-omnimodal-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-17-ainews-chameleon-metas-unreleased-gpt4o-like-omnimodal-model/</guid><description>**Meta AI FAIR** introduced **Chameleon**, a new multimodal model family with **7B** and **34B** parameter versions trained on **10T tokens** of interleaved text and image data enabling &quot;early fusion&quot; multimodality that can natively output any modality. While reasoning benchmarks are modest, its &quot;omnimodality&quot; approach competes well with pre-GPT4o multimodal models. **OpenAI** launched **GPT-4o**, a model excelling in benchmarks like MMLU and coding tasks, with strong multimodal capabilities but some regression in ELO scores and hallucination issues. **Google DeepMind** announced **Gemini 1.5 Flash**, a small model with **1M context window** and flash performance, highlighting convergence trends between OpenAI and Google models. **Anthropic** updated **Claude 3** with streaming support, forced tool use, and vision tool integration for multimodal knowledge extraction. OpenAI also partnered with Reddit, raising industry attention.</description><pubDate>Fri, 17 May 2024 20:46:44 GMT</pubDate><category>meta-ai-fair</category><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>reddit</category><category>chameleon</category><category>gpt-4o</category><category>gemini-1.5-flash</category><category>claude-3</category><category>armen-aghajanyan</category><category>sama</category><category>alexandr-wang</category><category>abacaj</category><category>alexalbert__</category><category>multimodality</category><category>early-fusion</category><category>benchmarking</category><category>model-training</category><category>tokenization</category><category>streaming</category><category>tool-use</category><category>vision</category><category>coding</category><category>hallucination-detection</category><category>model-performance</category></item><item><title>Cursor reaches &gt;1000 tok/s finetuning Llama3-70b for fast file editing</title><link>https://news.smol.ai/issues/24-05-16-ainews-cursor-reaches-greater1000-toks-finetuning-llama3-70b-for-fast-file-editing/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-16-ainews-cursor-reaches-greater1000-toks-finetuning-llama3-70b-for-fast-file-editing/</guid><description>**Cursor**, an AI-native IDE, announced a **speculative edits** algorithm for code editing that surpasses **GPT-4** and **GPT-4o** in accuracy and latency, achieving speeds of over **1000 tokens/s** on a **70b** model. **OpenAI** released **GPT-4o** with multimodal capabilities including audio, vision, and text, noted to be **2x faster and 50% cheaper** than GPT-4 turbo, though with mixed coding performance. **Anthropic** introduced streaming, forced tool use, and vision features for developers. **Google DeepMind** unveiled **Imagen Video** and **Gemini 1.5 Flash**, a small model with a **1M-context** window. **HuggingFace** is distributing **$10M** in free GPUs for open-source AI models like **Llama**, **BLOOM**, and **Stable Diffusion**. Evaluation insights highlight challenges with LLMs on novel problems and benchmark saturation, with new benchmarks like **MMLU-Pro** showing significant drops in top model performance.</description><pubDate>Fri, 17 May 2024 00:50:41 GMT</pubDate><category>cursor</category><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>huggingface</category><category>gpt-4</category><category>gpt-4o</category><category>gpt-4-turbo</category><category>gpt-4o-mini</category><category>llama</category><category>bloom</category><category>stable-diffusion</category><category>sama</category><category>abacaj</category><category>imjaredz</category><category>erhartford</category><category>alexalbert</category><category>svpino</category><category>maximelabonne</category><category>_philschmid</category><category>speculative-decoding</category><category>code-edits</category><category>multimodality</category><category>image-generation</category><category>streaming</category><category>tool-use</category><category>fine-tuning</category><category>benchmarking</category><category>mmlu</category><category>model-performance</category><category>evaluation</category><category>synthetic-data</category><category>context-windows</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-05-15-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-15-ainews-not-much-happened-today/</guid><description>**Ilya Sutskever** steps down as Chief Scientist at **OpenAI** after nearly a decade, with **Jakub Pachocki** named as his successor. **Google DeepMind** announces **Gemini 1.5 Pro** and **Gemini 1.5 Flash** models featuring 2 million token context and improved multimodal capabilities, alongside demos of **Project Astra** AI assistant, **Imagen 3** text-to-image model, and **Veo** generative video model. **GPT-4o** tops the VHELM leaderboard and outperforms competitors on LMSYS Chatbot Arena. **Reka Core** multimodal model with 128K context and **Alibaba&apos;s Qwen1.5-110B** open-source model are released. **Salesforce** shares an online RLHF recipe.</description><pubDate>Wed, 15 May 2024 21:20:08 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>rekailabs</category><category>alibaba</category><category>salesforce</category><category>gpt-4o</category><category>gemini-1.5-pro</category><category>gemini-1.5-flash</category><category>imagen-3</category><category>veo</category><category>reka-core</category><category>qwen-1.5-110b</category><category>ilya-sutskever</category><category>jakub-pachocki</category><category>mike-krieger</category><category>sama</category><category>multimodality</category><category>long-context</category><category>model-releases</category><category>reinforcement-learning</category><category>model-benchmarking</category><category>text-to-image</category><category>video-generation</category><category>ai-assistants</category></item><item><title>Google I/O in 60 seconds</title><link>https://news.smol.ai/issues/24-05-14-ainews-google-io-in-60-seconds/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-14-ainews-google-io-in-60-seconds/</guid><description>**Google** announced updates to the **Gemini model family**, including **Gemini 1.5 Pro** with **2 million token support**, and the new **Gemini Flash** model optimized for speed with **1 million token capacity**. The Gemini suite now includes **Ultra**, **Pro**, **Flash**, and **Nano** models, with **Gemini Nano** integrated into **Chrome 126**. Additional Gemini features include **Gemini Gems** (custom GPTs), **Gemini Live** for voice conversations, and **Project Astra**, a live video understanding assistant. The **Gemma model family** was updated with **Gemma 2** at **27B parameters**, offering near-**llama-3-70b** performance at half the size, plus **PaliGemma**, a vision-language open model inspired by **PaLI-3**. Other launches include **DeepMind&apos;s Veo**, **Imagen 3** for photorealistic image generation, and a **Music AI Sandbox** collaboration with YouTube. **SynthID watermarking** now extends to text, images, audio, and video. The **Trillium TPUv6** codename was revealed. Google also integrated AI across its product suite including Workspace, Email, Docs, Sheets, Photos, Search, and Lens. *&quot;The world awaits Apple&apos;s answer.&quot;*</description><pubDate>Tue, 14 May 2024 22:01:01 GMT</pubDate><category>google</category><category>google-deepmind</category><category>youtube</category><category>gemini-1.5-pro</category><category>gemini-flash</category><category>gemini-ultra</category><category>gemini-pro</category><category>gemini-nano</category><category>gemma-2</category><category>llama-3-70b</category><category>paligemma</category><category>imagen-3</category><category>veo</category><category>tokenization</category><category>model-performance</category><category>fine-tuning</category><category>vision</category><category>multimodality</category><category>model-release</category><category>model-training</category><category>model-optimization</category><category>ai-integration</category><category>image-generation</category><category>watermarking</category><category>hardware-optimization</category><category>voice</category><category>video-understanding</category></item><item><title>GPT-4o: the new SOTA-EVERYTHING Frontier model (GPT4T version) </title><link>https://news.smol.ai/issues/24-05-13-ainews-gpt-4o-the-new-sota-everything-frontier-model-gpt4t-version/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-13-ainews-gpt-4o-the-new-sota-everything-frontier-model-gpt4t-version/</guid><description>**OpenAI** launched **GPT-4o**, a frontier model supporting real-time reasoning across **audio, vision, and text**, now free for all ChatGPT users with enhanced coding capabilities and upcoming advanced voice and video features. Discussions cover **open-source LLMs** like **Llama 3**, fine-tuning techniques including knowledge distillation for **GPT-3.5**, and hardware optimization strategies such as quantization. Emerging architectures include multimodal integrations with ChatGPT voice and Open Interpreter API, Mixture of Experts models combining autoregressive and diffusion approaches, and novel designs like the **YOCO architecture** and **ThunderKittens DSL** for efficient GPU use. Research advances in efficient attention methods like **Conv-Basis** using FFT and model scaling techniques such as depth upscaling were also highlighted.</description><pubDate>Mon, 13 May 2024 23:14:50 GMT</pubDate><category>openai</category><category>hugging-face</category><category>nous-research</category><category>eleutherai</category><category>hazyresearch</category><category>gpt-4o</category><category>gpt-3.5</category><category>llama-3</category><category>real-time-reasoning</category><category>coding-capabilities</category><category>fine-tuning</category><category>knowledge-distillation</category><category>hardware-optimization</category><category>quantization</category><category>multimodality</category><category>mixture-of-experts</category><category>efficient-attention</category><category>model-scaling</category><category>depth-upscaling</category><category>transformer-architecture</category><category>gpu-optimization</category><category>prompt-engineering</category></item><item><title>GPT-4o: the new SOTA-EVERYTHING Frontier model (GPT4O version)</title><link>https://news.smol.ai/issues/24-05-13-ainews-gpt-4o-the-new-sota-everything-frontier-model-gpt4o-version/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-13-ainews-gpt-4o-the-new-sota-everything-frontier-model-gpt4o-version/</guid><description>**OpenAI** has released **GPT-4o**, a new **multimodal** model capable of reasoning across text, audio, and video in real time with low latency (~300ms). It features voice and vision capabilities, improved non-English language performance with an expanded 200k vocabulary tokenizer, and is available to all ChatGPT users including free plans. GPT-4o is half the price and twice as fast as GPT-4-turbo with 5x rate limits. The model supports real-time voice and video input/output and shows strong coding capabilities. The release includes a new desktop app that can read screen and clipboard history, challenging existing desktop agent startups. The announcement was accompanied by demos including image generation and 3D object handling, with OpenAI achieving state-of-the-art performance in ASR and vision tasks. The update was widely discussed on social media, with comparisons to GPT-4T highlighting GPT-4o&apos;s speed and versatility. *&quot;GPT-4o is smart, fast, natively multimodal, and a step towards more natural human-computer interaction&quot;* and *&quot;extremely versatile and fun to play with&quot;*.</description><pubDate>Mon, 13 May 2024 22:58:05 GMT</pubDate><category>openai</category><category>lmsys</category><category>multion</category><category>adept</category><category>gpt-4o</category><category>gpt-4-turbo</category><category>sama</category><category>gdb</category><category>multimodality</category><category>vision</category><category>speech-recognition</category><category>tokenization</category><category>real-time-processing</category><category>coding</category><category>model-performance</category><category>model-optimization</category><category>desktop-agents</category></item><item><title>Quis promptum ipso promptiet?</title><link>https://news.smol.ai/issues/24-05-10-ainews-quis-promptum-ipso-promptiet/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-10-ainews-quis-promptum-ipso-promptiet/</guid><description>**Anthropic** released upgrades to their Workbench Console, introducing new prompt engineering features like chain-of-thought reasoning and prompt generators that significantly reduce development time, exemplified by their customer **Zoominfo**. **OpenAI** teased a &quot;magic&quot; new development coming soon, speculated to be a new LLM replacing GPT-3.5 in the free tier or a search competitor. The open-source community highlighted **Llama 3 70B** as &quot;game changing&quot; with new quantized weights for **Llama 3 120B** and CUDA graph support for **llama.cpp** improving GPU performance. **Neuralink** demonstrated a thought-controlled mouse, sparking interest in modeling consciousness from brain signals. The **ICLR 2024** conference is being held in Asia for the first time, generating excitement.</description><pubDate>Sat, 11 May 2024 06:34:12 GMT</pubDate><category>anthropic</category><category>openai</category><category>zoominfo</category><category>neuralink</category><category>llama-3-70b</category><category>llama-3-120b</category><category>llama-3</category><category>llama-cpp</category><category>sama</category><category>gdb</category><category>bindureddy</category><category>svpino</category><category>rohanpaul_ai</category><category>alexalbert__</category><category>abacaj</category><category>prompt-engineering</category><category>chain-of-thought</category><category>rag</category><category>quantization</category><category>cuda-graphs</category><category>gpu-optimization</category><category>thought-controlled-devices</category><category>modeling-consciousness</category><category>conference</category></item><item><title>LMSys advances Llama 3 eval analysis</title><link>https://news.smol.ai/issues/24-05-09-ainews-lmsys-advances-llama-3-eval-analysis/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-09-ainews-lmsys-advances-llama-3-eval-analysis/</guid><description>**LMSys** is enhancing LLM evaluation by categorizing performance across **8 query subcategories** and **7 prompt complexity levels**, revealing uneven strengths in models like **Llama-3-70b**. **DeepMind** released **AlphaFold 3**, advancing molecular structure prediction with holistic modeling of protein-DNA-RNA complexes, impacting biology and genetics research. **OpenAI** introduced the **Model Spec**, a public standard to clarify model behavior and tuning, inviting community feedback and aiming for models to learn directly from it. **Llama 3** has reached top leaderboard positions on LMSys, nearly matching **Claude-3-sonnet** in performance, with notable variations on complex prompts. The analysis highlights the evolving landscape of model benchmarking and behavior shaping.</description><pubDate>Fri, 10 May 2024 00:52:45 GMT</pubDate><category>lmsys</category><category>openai</category><category>google-deepmind</category><category>isomorphic-labs</category><category>llama-3-70b</category><category>llama-3</category><category>claude-3-sonnet</category><category>alphafold-3</category><category>demis-hassabis</category><category>sam-altman</category><category>miranda-murati</category><category>karina-nguyen</category><category>joanne-jang</category><category>john-schulman</category><category>benchmarking</category><category>model-behavior</category><category>prompt-complexity</category><category>model-specification</category><category>molecular-structure-prediction</category><category>performance-analysis</category><category>leaderboards</category></item><item><title>OpenAI&apos;s PR Campaign?</title><link>https://news.smol.ai/issues/24-05-08-ainews-openais-pr-campaign/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-08-ainews-openais-pr-campaign/</guid><description>**OpenAI** faces user data deletion backlash over its new partnership with StackOverflow amid GDPR complaints and US newspaper lawsuits, while addressing election year concerns with efforts like the Media Manager tool for content opt-in/out by 2025 and source link attribution. **Microsoft** develops a top-secret airgapped GPT-4 AI service for US intelligence agencies. OpenAI releases the Model Spec outlining responsible AI content generation policies, including NSFW content handling and profanity use, emphasizing clear distinctions between bugs and design decisions. **Google DeepMind** announces **AlphaFold 3**, a state-of-the-art model predicting molecular structures with high accuracy, showcasing cross-domain AI techniques. New research on **xLSTM** proposes scaling LSTMs to billions of parameters, competing with transformers in performance and scaling. Microsoft introduces **vAttention**, a dynamic memory management method for efficient large language model serving without PagedAttention.</description><pubDate>Thu, 09 May 2024 01:27:27 GMT</pubDate><category>openai</category><category>microsoft</category><category>google-deepmind</category><category>alphafold-3</category><category>xlstm</category><category>gpt-4</category><category>demis-hassabis</category><category>sama</category><category>joanne-jang</category><category>omarsar0</category><category>arankomatsuzaki</category><category>drjimfan</category><category>memory-management</category><category>model-spec</category><category>scaling</category><category>multimodality</category><category>performance</category><category>transformers</category><category>dynamic-memory</category><category>model-architecture</category></item><item><title>Kolmogorov-Arnold Networks: MLP killers or just spicy MLPs?</title><link>https://news.smol.ai/issues/24-05-07-ainews-kolmogorov-arnold-networks-mlp-killers-or-just-spicy-mlps/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-07-ainews-kolmogorov-arnold-networks-mlp-killers-or-just-spicy-mlps/</guid><description>**Ziming Liu**, a grad student of **Max Tegmark**, published a paper on **Kolmogorov-Arnold Networks (KANs)**, claiming they outperform **MLPs** in interpretability, inductive bias injection, function approximation accuracy, and scaling, despite being 10x slower to train but 100x more parameter efficient. KANs use learnable activation functions modeled by B-splines on edges rather than fixed activations on nodes. However, it was later shown that KANs can be mathematically rearranged back into MLPs with similar parameter counts, sparking debate on their interpretability and novelty. Meanwhile, on AI Twitter, there is speculation about a potential **GPT-5** release with mixed impressions, OpenAI&apos;s adoption of the **C2PA metadata standard** for detecting AI-generated images with high accuracy for **DALL-E 3**, and **Microsoft** training a large 500B parameter model called **MAI-1**, potentially previewed at Build conference, signaling increased competition with OpenAI. *&quot;OpenAI&apos;s safety testing for GPT-4.5 couldn&apos;t finish in time for Google I/O launch&quot;* was also noted.</description><pubDate>Tue, 07 May 2024 22:47:14 GMT</pubDate><category>openai</category><category>microsoft</category><category>gpt-5</category><category>gpt-4</category><category>dall-e-3</category><category>max-tegmark</category><category>ziming-liu</category><category>bindureddy</category><category>nptacek</category><category>zacharynado</category><category>rohanpaul_ai</category><category>svpino</category><category>learnable-activations</category><category>mlp</category><category>function-approximation</category><category>interpretability</category><category>inductive-bias-injection</category><category>b-splines</category><category>model-rearrangement</category><category>parameter-efficiency</category><category>ai-generated-image-detection</category><category>metadata-standards</category><category>large-model-training</category></item><item><title>DeepSeek-V2 beats Mixtral 8x22B with &gt;160 experts at HALF the cost</title><link>https://news.smol.ai/issues/24-05-06-ainews-deepseek-v2-beats-mixtral-8x22b-with-greater160-experts-at-half-the-cost/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-06-ainews-deepseek-v2-beats-mixtral-8x22b-with-greater160-experts-at-half-the-cost/</guid><description>**DeepSeek V2** introduces a new state-of-the-art MoE model with **236B parameters** and a novel Multi-Head Latent Attention mechanism, achieving faster inference and surpassing GPT-4 on AlignBench. **Llama 3 120B** shows strong creative writing skills, while Microsoft is reportedly developing a **500B parameter** LLM called **MAI-1**. Research from Scale AI highlights overfitting issues in models like **Mistral** and **Phi**, whereas **GPT-4**, **Claude**, **Gemini**, and **Llama** maintain benchmark robustness. In robotics, **Tesla Optimus** advances with superior data collection and teleoperation, **LeRobot** marks a move toward open-source robotics AI, and **Nvidia&apos;s DrEureka** automates robot skill training. Multimodal LLM hallucinations are surveyed with new mitigation strategies, and **Google&apos;s Med-Gemini** achieves SOTA on medical benchmarks with fine-tuned multimodal models.</description><pubDate>Mon, 06 May 2024 23:37:03 GMT</pubDate><category>deepseek-ai</category><category>mistral-ai</category><category>microsoft</category><category>openai</category><category>scale-ai</category><category>tesla</category><category>nvidia</category><category>google-deepmind</category><category>deepseek-v2</category><category>llama-3-120b</category><category>llama-3-400b</category><category>gpt-4</category><category>mistral</category><category>phi</category><category>claude</category><category>gemini</category><category>mai-1</category><category>med-gemini</category><category>erhartford</category><category>maximelabonne</category><category>bindureddy</category><category>adcock_brett</category><category>drjimfan</category><category>clementdelangue</category><category>omarsar0</category><category>rohanpaul_ai</category><category>mixture-of-experts</category><category>multi-head-attention</category><category>model-inference</category><category>benchmarking</category><category>overfitting</category><category>robotics</category><category>teleoperation</category><category>open-source</category><category>multimodality</category><category>hallucination-detection</category><category>fine-tuning</category><category>medical-ai</category><category>model-training</category></item><item><title>$100k to predict LMSYS human preferences in a Kaggle contest</title><link>https://news.smol.ai/issues/24-05-03-ainews-dollar100k-to-predict-lmsys-human-preferences-in-a-kaggle-contest/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-03-ainews-dollar100k-to-predict-lmsys-human-preferences-in-a-kaggle-contest/</guid><description>**Llama 3 models** are making breakthroughs with Groq&apos;s 70B model achieving record low costs per million tokens. A new **Kaggle competition** offers a $100,000 prize to develop models predicting human preferences from a dataset of over 55,000 user-LLM conversations. Open source evaluator LLMs like **Prometheus 2** outperform proprietary models such as **GPT-4** and **Claude 3 Opus** in judgment tasks. New datasets like **WildChat1M** provide over 1 million ChatGPT interaction logs with diverse and toxic examples. Techniques like **LoRA fine-tuning** show significant performance gains, and **NVIDIA&apos;s NeMo-Aligner** toolkit enables scalable LLM alignment across hundreds of GPUs. Factuality-aware alignment methods are proposed to reduce hallucinations in LLM outputs.</description><pubDate>Fri, 03 May 2024 22:09:28 GMT</pubDate><category>groq</category><category>openai</category><category>lmsys</category><category>scale-ai</category><category>ai2</category><category>nvidia</category><category>llama-3-70b</category><category>llama-3</category><category>gpt-4</category><category>claude-3-opus</category><category>prometheus-2</category><category>bindureddy</category><category>drjimfan</category><category>percyliang</category><category>seungonekim</category><category>mobicham</category><category>clefourrier</category><category>benchmarking</category><category>datasets</category><category>fine-tuning</category><category>reinforcement-learning</category><category>model-alignment</category><category>hallucination</category><category>parameter-efficient-fine-tuning</category><category>scalable-training</category><category>factuality</category><category>chatbot-performance</category></item><item><title>Evals: The Next Generation</title><link>https://news.smol.ai/issues/24-05-02-ainews-evals-the-next-generation/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-02-ainews-evals-the-next-generation/</guid><description>**Scale AI** highlighted issues with data contamination in benchmarks like **MMLU** and **GSM8K**, proposing a new benchmark where **Mistral** overfits and **Phi-3** performs well. **Reka** released the **VibeEval** benchmark for multimodal models addressing multiple choice benchmark limitations. **Sam Altman** of **OpenAI** discussed GPT-4 as &quot;dumb&quot; and hinted at **GPT-5** with AI agents as a major breakthrough. Researchers jailbroke **GPT-3.5** via fine-tuning. Global calls emerged to ban AI-powered weapons, with US officials urging human control over nuclear arms. Ukraine launched an AI consular avatar, while **Moderna** partnered with **OpenAI** for medical AI advancements. **Sanctuary AI** and **Microsoft** collaborate on AI for general-purpose robots. MIT introduced **Kolmogorov-Arnold networks** with improved neural network efficiency. **Meta AI** is training **Llama 3** models with over 400 billion parameters, featuring multimodality and longer context.</description><pubDate>Thu, 02 May 2024 23:54:22 GMT</pubDate><category>scale-ai</category><category>mistral-ai</category><category>reka-ai</category><category>openai</category><category>moderna</category><category>sanctuary-ai</category><category>microsoft</category><category>mit</category><category>meta-ai-fair</category><category>gpt-4</category><category>gpt-5</category><category>gpt-3.5</category><category>phi-3</category><category>mistral-7b</category><category>llama-3</category><category>sam-altman</category><category>jim-fan</category><category>benchmarking</category><category>data-contamination</category><category>multimodality</category><category>fine-tuning</category><category>ai-regulation</category><category>ai-safety</category><category>ai-weapons</category><category>neural-networks</category><category>model-architecture</category><category>model-training</category><category>model-performance</category><category>robotics</category><category>activation-functions</category><category>long-context</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-05-01-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-05-01-ainews-not-much-happened-today/</guid><description>**Anthropic** released a team plan and iOS app about 4 months after **OpenAI**. The **Command-R 35B** model excels at creative writing, outperforming larger models like **Goliath-120** and **Miqu-120**. The **Llama-3 8B** model now supports a 1 million token context window, improving long-context understanding with minimal training on a single 8xA800 GPU machine. **TensorRT-LLM** benchmarks show it is 30-70% faster than **llama.cpp** on consumer hardware. A benchmark suggests **GPT2-Chat** may have better reasoning than **GPT-4-Turbo**, though results are debated. Demos include a self-learning **Llama-3** voice agent running locally on Jetson Orin and a Self-Learning Large Action Model (LAM). **Amazon CodeWhisperer** was renamed to **Q Developer**, expanding its generative AI assistant capabilities. **Apple** plans an AI-enabled Safari browser with an on-device LLM in iOS 18 and macOS 15. Big Tech dominates AI lobbying in Washington, while major U.S. newspapers sued **OpenAI** and **Microsoft** for copyright infringement. **DeepMind&apos;s AlphaZero** became the greatest chess player in 9 hours, and their Naturalized Execution Tuning (NExT) method improves LLM code reasoning by 14-26%. **Stable Diffusion** is used for diverse image generation applications.</description><pubDate>Thu, 02 May 2024 00:47:12 GMT</pubDate><category>anthropic</category><category>openai</category><category>perplexity-ai</category><category>amazon</category><category>apple</category><category>microsoft</category><category>deepmind</category><category>command-r-35b</category><category>goliath-120</category><category>miqu-120</category><category>llama-3-8b</category><category>tensorrt-llm</category><category>llama-cpp</category><category>gpt2-chat</category><category>gpt-4-turbo</category><category>llama-3</category><category>deepmind-alphazero</category><category>creative-writing</category><category>context-windows</category><category>benchmarking</category><category>model-performance</category><category>self-learning</category><category>function-calling</category><category>retrieval-augmented-generation</category><category>ai-assistants</category><category>on-device-ai</category><category>ai-lobbying</category><category>copyright-infringement</category><category>code-reasoning</category><category>image-generation</category></item><item><title>LLMs-as-Juries</title><link>https://news.smol.ai/issues/24-04-30-ainews-llms-as-juries/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-30-ainews-llms-as-juries/</guid><description>**OpenAI** has rolled out the **memory feature** to all ChatGPT Plus users and partnered with the **Financial Times** to license content for AI training. Discussions on **OpenAI&apos;s profitability** arise due to paid training data licensing and potential **GPT-4 usage limit reductions**. Users report issues with ChatGPT&apos;s data cleansing after the memory update. Tutorials and projects include building AI voice assistants and interface agents powered by LLMs. In **Stable Diffusion**, users seek realistic **SDXL models** comparable to PonyXL, and new extensions like **Hi-diffusion** and **Virtuoso Nodes v1.1** enhance ComfyUI with advanced image generation and Photoshop-like features. Cohere finds that multiple agents outperform single agents in LLM judging tasks, highlighting advances in multi-agent systems.</description><pubDate>Wed, 01 May 2024 01:41:25 GMT</pubDate><category>openai</category><category>cohere</category><category>financial-times</category><category>gpt-4</category><category>gpt-3.5</category><category>sdxl</category><category>ponyxl</category><category>memory</category><category>training-data</category><category>model-usage-limits</category><category>data-cleansing</category><category>ai-voice-assistants</category><category>interface-agents</category><category>image-generation</category><category>model-extensions</category><category>multi-agent-systems</category></item><item><title>A quiet weekend</title><link>https://news.smol.ai/issues/24-04-29-ainews-a-quiet-weekend/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-29-ainews-a-quiet-weekend/</guid><description>**Yann LeCun** predicts a shift to **AR interfaces** with AI assistants in 10-15 years, moving away from smartphones. The **Dolphin-2.9 model** based on **Llama-3** was released, improving quality issues. **PixArt Sigma**, a **0.6B parameter** model, achieves **Stable Diffusion 3.0** level performance with complete prompt adherence and local usability. Research shows transformers can use meaningless filler tokens for algorithmic tasks with dense supervision. AI-generated restaurant reviews can pass the **Turing test**, fooling humans and AI detectors. **Uber** uses graph algorithms and learned embeddings for ETA prediction. **Coca-Cola** and **Microsoft** announced a 5-year AI partnership to accelerate cloud and generative AI initiatives. The **Llama-3 70B** model can run on a single 4GB GPU using **AirLLM** optimization without quantization but is slow. **Mistral.rs** is introduced as a fast LLM inference platform with quantization and OpenAI API compatibility. Only 5% of LLMs make it from prototype to production due to challenges, especially in enterprise. EXL2 and GGUF quantization methods for Llama models show similar perplexity vs model size, with Llama-3 and Llama-2 degrading more under quantization compared to full precision.</description><pubDate>Mon, 29 Apr 2024 22:10:15 GMT</pubDate><category>microsoft</category><category>coca-cola</category><category>uber</category><category>lmsys</category><category>nous-research</category><category>mistral-ai</category><category>llama-3</category><category>dolphin-2.9</category><category>pixart-sigma</category><category>llama-3-70b</category><category>yann-lecun</category><category>ar-interfaces</category><category>transformers</category><category>algorithmic-tasks</category><category>turing-test</category><category>graph-algorithms</category><category>embeddings</category><category>generative-ai</category><category>model-optimization</category><category>llm-inference</category><category>quantization</category><category>model-deployment</category></item><item><title>Apple&apos;s OpenELM beats OLMo with 50% of its dataset, using DeLighT</title><link>https://news.smol.ai/issues/24-04-26-ainews-apples-openelm-beats-olmo-with-50percent-of-its-dataset-using-delight/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-26-ainews-apples-openelm-beats-olmo-with-50percent-of-its-dataset-using-delight/</guid><description>**Apple** advances its AI presence with the release of **OpenELM**, its first relatively open large language model available in sizes from **270M to 3B** parameters, featuring a novel layer-wise scaling architecture inspired by the **DeLight** paper. Meanwhile, **Meta&apos;s LLaMA 3** family pushes context length boundaries with models supporting over **160K tokens** and an **8B-Instruct model with 262K context length** released on Hugging Face, alongside performance improvements in quantized versions. A new paper on AI alignment highlights **KTO** as the best-performing method, with sensitivity to training data volume noted. In AI ethics and regulation, former **Google** CEO **Eric Schmidt** warns about the risks of open-source AI empowering bad actors and geopolitical rivals, while a U.S. proposal aims to enforce &quot;Know Your Customer&quot; rules to end anonymous cloud usage.</description><pubDate>Fri, 26 Apr 2024 21:32:41 GMT</pubDate><category>apple</category><category>meta-ai-fair</category><category>google</category><category>openelm</category><category>llama-3</category><category>llama-3-8b-instruct</category><category>llama-3-70b</category><category>eric-schmidt</category><category>sebastian-raschka</category><category>layer-wise-scaling</category><category>context-length</category><category>quantization</category><category>ai-alignment</category><category>open-source</category><category>ai-regulation</category></item><item><title>Snowflake Arctic: Fully Open 10B+128x4B Dense-MoE Hybrid LLM</title><link>https://news.smol.ai/issues/24-04-25-ainews-snowflake-arctic-fully-open-10b128x4b-dense-moe-hybrid-llm/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-25-ainews-snowflake-arctic-fully-open-10b128x4b-dense-moe-hybrid-llm/</guid><description>**Snowflake Arctic** is a notable new foundation language model released under Apache 2.0, claiming superiority over **Databricks** in data warehouse AI applications and adopting a mixture-of-experts architecture inspired by **DeepSeekMOE** and **DeepSpeedMOE**. The model employs a 3-stage curriculum training strategy similar to the recent **Phi-3** paper. In AI image and video generation, **Nvidia** introduced the **Align Your Steps** technique improving image quality at low step counts, while **Stable Diffusion 3** and **SD3 Turbo** models were compared for prompt understanding and image quality. **Adobe** launched an AI video upscaling project enhancing blurry videos to HD, though with some high-resolution artifacts. **Apple** released open-source on-device language models with code and training logs, diverging from typical weight-only releases. The **Llama-3-70b** model ties for first place on the LMSYS leaderboard for English queries, and **Phi-3** (4B params) outperforms **GPT-3.5 Turbo** in the banana logic benchmark. Fast inference and quantization of **Llama 3** models were demonstrated on MacBook devices.</description><pubDate>Fri, 26 Apr 2024 01:33:53 GMT</pubDate><category>snowflake</category><category>databricks</category><category>deepseek</category><category>deepspeed</category><category>nvidia</category><category>stable-diffusion</category><category>adobe</category><category>apple</category><category>llamaindex</category><category>lmsys</category><category>openai</category><category>snowflake-arctic</category><category>phi-3</category><category>llama-3-70b</category><category>llama-3</category><category>stable-diffusion-3</category><category>sd3-turbo</category><category>gpt-3.5-turbo</category><category>mixture-of-experts</category><category>curriculum-learning</category><category>model-release</category><category>image-generation</category><category>video-upscaling</category><category>quantization</category><category>inference-speed</category><category>benchmarking</category><category>model-comparison</category><category>open-source</category><category>on-device-ai</category></item><item><title>OpenAI&apos;s Instruction Hierarchy for the LLM OS</title><link>https://news.smol.ai/issues/24-04-24-ainews-openais-instruction-hierarchy-for-the-llm-os/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-24-ainews-openais-instruction-hierarchy-for-the-llm-os/</guid><description>**OpenAI** published a paper introducing the concept of privilege levels for LLMs to address prompt injection vulnerabilities, improving defenses by 20-30%. **Microsoft** released the lightweight **Phi-3-mini** model with 4K and 128K context lengths. **Apple** open-sourced the **OpenELM** language model family with an open training and inference framework. An instruction accuracy benchmark compared 12 models, with **Claude 3 Opus**, **GPT-4 Turbo**, and **Llama 3 70B** performing best. The **Rho-1** method enables training state-of-the-art models using only 3% of tokens, boosting models like **Mistral**. **Wendy&apos;s** deployed AI-powered drive-thru ordering, and a study found **Gen Z** workers prefer generative AI for career advice. Tutorials on deploying **Llama 3** models on AWS EC2 highlight hardware requirements and inference server use.</description><pubDate>Thu, 25 Apr 2024 00:15:11 GMT</pubDate><category>openai</category><category>microsoft</category><category>apple</category><category>deepseek</category><category>mistral-ai</category><category>llamaindex</category><category>wendys</category><category>phi-3-mini</category><category>openelm</category><category>claude-3-opus</category><category>gpt-4-turbo</category><category>gpt-3.5-turbo</category><category>llama-3-70b</category><category>rho-1</category><category>mistral-7b</category><category>llama-3-8b</category><category>llama-3</category><category>prompt-injection</category><category>alignment</category><category>benchmarking</category><category>instruction-following</category><category>context-windows</category><category>model-training</category><category>model-deployment</category><category>inference</category><category>performance-optimization</category><category>ai-application</category><category>career-advice</category><category>drive-thru-ai</category></item><item><title>Perplexity, the newest AI unicorn</title><link>https://news.smol.ai/issues/24-04-23-ainews-perplexity-the-newest-ai-unicorn/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-23-ainews-perplexity-the-newest-ai-unicorn/</guid><description>**Perplexity** doubles its valuation shortly after its Series B with a Series B-1 funding round. Significant developments around **Llama 3** include context length extension to **16K tokens**, new multimodal **LLaVA models** outperforming Llama 2, and fine-tuning improvements like QDoRA surpassing QLoRA. The **Llama-3-70B** model is praised for instruction following and performance across quantization formats. **Phi-3 models** by **Meta AI** released in multiple sizes show competitive benchmark results, with the 14B model achieving **78% on MMLU** and the 3.8B model nearing **GPT-3.5** performance.</description><pubDate>Tue, 23 Apr 2024 22:48:23 GMT</pubDate><category>perplexity-ai</category><category>meta-ai-fair</category><category>hugging-face</category><category>groq</category><category>llama-3-8b</category><category>llama-3-70b</category><category>llama-3</category><category>llava-llama-3-8b-v1_1</category><category>phi-3</category><category>gpt-3.5</category><category>daniel-gross</category><category>aravind-srinivas</category><category>context-length</category><category>fine-tuning</category><category>quantization</category><category>instruction-following</category><category>model-comparison</category><category>multimodality</category><category>benchmarking</category><category>memory-optimization</category><category>model-performance</category></item><item><title>FineWeb: 15T Tokens, 12 years of CommonCrawl (deduped and filtered, you&apos;re welcome)</title><link>https://news.smol.ai/issues/24-04-22-ainews-fineweb-15t-tokens-12-years-of-commoncrawl-deduped-and-filtered-youre-welcome/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-22-ainews-fineweb-15t-tokens-12-years-of-commoncrawl-deduped-and-filtered-youre-welcome/</guid><description>**2024** has seen a significant increase in dataset sizes for training large language models, with **Redpajama 2** offering up to **30T tokens**, **DBRX** at **12T tokens**, **Reka Core/Flash/Edge** with **5T tokens**, and **Llama 3** trained on **15T tokens**. **Huggingface** released an open dataset containing **15T tokens** from **12 years** of filtered CommonCrawl data, enabling training of models like **Llama 3** if compute resources are available. On Reddit, **WizardLM-2-8x22b** outperformed other open LLMs including **Llama-3-70b-instruct** in reasoning and math benchmarks. **Claude Opus** demonstrated strong zero-shot code error spotting, surpassing **Llama 3**. Benchmarks revealed limitations in the **LMSYS chatbot leaderboard** due to instruction-tuned models gaming the system, and a new RAG benchmark showed **Llama 3 70B** underperforming compared to **GPT-4**, while **Mistral 8x7B** remained strong. Efficient quantized versions of **Llama 3** models are available on **Huggingface**, with users reporting token generation limits around **9600 tokens** on a 3090 GPU. Safety concerns include a UK sex offender banned from AI tool usage and **GPT-4** demonstrating an **87% success rate** exploiting real vulnerabilities, raising security concerns.</description><pubDate>Tue, 23 Apr 2024 00:03:58 GMT</pubDate><category>huggingface</category><category>meta-ai-fair</category><category>dbrx</category><category>reka-ai</category><category>mistral-ai</category><category>lmsys</category><category>openai</category><category>llama-3-70b</category><category>llama-3</category><category>wizardlm-2-8x22b</category><category>claude-opus</category><category>mistral-8x7b</category><category>gpt-4</category><category>datasets</category><category>benchmarking</category><category>quantization</category><category>zero-shot-learning</category><category>reasoning</category><category>code-error-detection</category><category>token-generation</category><category>security</category></item><item><title>Llama-3-70b is GPT-4-level Open Model</title><link>https://news.smol.ai/issues/24-04-19-ainews-llama-3-70b-is-gpt-4-level-open-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-19-ainews-llama-3-70b-is-gpt-4-level-open-model/</guid><description>**Meta** has released **Llama 3**, their most capable open large language model with **8B and 70B parameter versions** supporting **8K context length** and outperforming previous models including **Llama 2** and **Mistral 7B**. **Groq** serves the **Llama 3 70B** model at **500-800 tokens/second**, making it the fastest GPT-4-level token source. Discussions highlight AI scaling challenges with **Elon Musk** stating that training **Grok 3** will require **100,000 Nvidia H100 GPUs**, and **AWS** planning to acquire **20,000 B200 GPUs** for a **27 trillion parameter model**. Microsoft unveiled **VASA-1** for lifelike talking face generation, while **Stable Diffusion 3** and its extensions received mixed impressions. Concerns about AI energy usage and political bias in AI were also discussed.</description><pubDate>Sat, 20 Apr 2024 02:21:27 GMT</pubDate><category>meta-ai-fair</category><category>groq</category><category>nvidia</category><category>amazon</category><category>microsoft</category><category>llama-3-70b</category><category>llama-3-8b</category><category>llama-3</category><category>llama-2-70b</category><category>mistral-7b</category><category>grok-3</category><category>stable-diffusion-3</category><category>vasa-1</category><category>elon-musk</category><category>benchmarking</category><category>model-performance</category><category>fine-tuning</category><category>function-calling</category><category>arithmetic</category><category>image-generation</category><category>video-generation</category><category>energy-usage</category><category>gpu-demand</category><category>political-bias</category><category>ai-safety</category><category>scaling</category><category>context-windows</category><category>tokenization</category></item><item><title>Meta Llama 3 (8B, 70B)</title><link>https://news.smol.ai/issues/24-04-18-ainews-meta-llama-3-8b-70b/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-18-ainews-meta-llama-3-8b-70b/</guid><description>**Meta** partially released **Llama 3** models including **8B** and **70B** variants, with a **400B** variant still in training, touted as the first GPT-4 level open-source model. **Stability AI** launched **Stable Diffusion 3 API** with model weights coming soon, showing competitive realism against **Midjourney V6**. **Boston Dynamics** unveiled an electric humanoid robot **Atlas**, and **Microsoft** introduced the **VASA-1** model generating lifelike talking faces at 40fps on RTX 4090. **Mistral AI**, a European OpenAI rival, is seeking $5B funding with its **Mixtral-8x22B-Instruct-v0.1** model achieving 100% accuracy on 64K context benchmarks. AI safety discussions include calls from former OpenAI board member **Helen Toner** for audits of top AI companies, and the **Mormon Church** released AI usage principles. New AI development tools include **Ctrl-Adapter** for diffusion models, **Distilabel 1.0.0** for synthetic dataset pipelines, **Data Bonsai** for data cleaning with LLMs, and **Dendron** for building LLM agents with behavior trees. Memes highlight AI development humor and cultural references. The release of **Llama 3** models features improved reasoning, a 128K token vocabulary, 8K token sequences, and grouped query attention.</description><pubDate>Fri, 19 Apr 2024 04:28:01 GMT</pubDate><category>meta-ai-fair</category><category>stability-ai</category><category>boston-dynamics</category><category>microsoft</category><category>mistral-ai</category><category>hugging-face</category><category>llama-3-8b</category><category>llama-3-70b</category><category>llama-3-400b</category><category>stable-diffusion-3</category><category>mixtral-8x22b-instruct-v0.1</category><category>vasa-1</category><category>helen-toner</category><category>transformer</category><category>tokenization</category><category>model-training</category><category>benchmarking</category><category>robotics</category><category>natural-language-processing</category><category>real-time-processing</category><category>synthetic-data</category><category>dataset-cleaning</category><category>behavior-trees</category><category>ai-safety</category><category>model-accuracy</category><category>api</category><category>model-release</category><category>humor</category></item><item><title>Mixtral 8x22B Instruct sparks efficiency memes</title><link>https://news.smol.ai/issues/24-04-17-ainews-mixtral-8x22b-instruct-sparks-efficiency-memes/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-17-ainews-mixtral-8x22b-instruct-sparks-efficiency-memes/</guid><description>**Mistral** released an instruct-tuned version of their **Mixtral 8x22B** model, notable for using only **39B active parameters** during inference, outperforming larger models and supporting **5 languages** with **64k context window** and math/code capabilities. The model is available on **Hugging Face** under an **Apache 2.0 license** for local use. **Google** plans to invest over **$100 billion** in AI, with other giants like **Microsoft**, **Intel**, and **SoftBank** also making large investments. The UK criminalized non-consensual deepfake porn, raising enforcement debates. A former **Nvidia** employee claims Nvidia&apos;s AI chip lead is unmatchable this decade. AI companions could become a **$1 billion** market. AI has surpassed humans on several basic tasks but lags on complex ones. **Zyphra** introduced **Zamba**, a novel 7B parameter hybrid model outperforming **LLaMA-2 7B** and **OLMo-7B** with less training data, trained on 128 H100 GPUs over 30 days. **GroundX** API advances retrieval-augmented generation accuracy.</description><pubDate>Wed, 17 Apr 2024 21:02:34 GMT</pubDate><category>mistral-ai</category><category>hugging-face</category><category>google</category><category>microsoft</category><category>intel</category><category>softbank</category><category>nvidia</category><category>mixtral-8x22b</category><category>llama-2-7b</category><category>olmo-7b</category><category>guillaume-lample</category><category>osanseviero</category><category>_philschmid</category><category>svpino</category><category>multilinguality</category><category>math</category><category>code-generation</category><category>context-window</category><category>model-performance</category><category>model-release</category><category>retrieval-augmented-generation</category><category>deepfake</category><category>ai-investment</category><category>ai-chip</category><category>hybrid-architecture</category><category>training-data</category></item><item><title>Lilian Weng on Video Diffusion</title><link>https://news.smol.ai/issues/24-04-16-ainews-lilian-weng-on-video-diffusion/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-16-ainews-lilian-weng-on-video-diffusion/</guid><description>**OpenAI** expands with a launch in **Japan**, introduces a **Batch API**, and partners with **Adobe** to bring the **Sora video model** to Premiere Pro. **Reka AI** releases the **Reka Core multimodal language model**. **WizardLM-2** is released showing impressive performance, and **Llama 3** news is anticipated soon. Geoffrey Hinton highlights AI models exhibiting **intuition, creativity, and analogy recognition** beyond humans. The **Devin AI model** notably contributes to its own codebase. **Opus** demonstrates the ability to recognize its own generated outputs. **Sam Altman** warns startups about being steamrolled by OpenAI if they don&apos;t adapt quickly. **Yann LeCun** discusses AGI timelines, emphasizing it is inevitable but not imminent or solely from LLMs. Lilian Weng&apos;s blog on **diffusion models for video generation** highlights **training-free adaptation** as a breakthrough technique.</description><pubDate>Wed, 17 Apr 2024 02:15:37 GMT</pubDate><category>openai</category><category>adobe</category><category>reka-ai</category><category>wizardlm-2</category><category>llama-3</category><category>reka-core</category><category>devin</category><category>opus</category><category>sora</category><category>lilian-weng</category><category>sam-altman</category><category>geoffrey-hinton</category><category>yann-lecun</category><category>diffusion-models</category><category>video-generation</category><category>training-free-adaptation</category><category>multimodality</category><category>intuition</category><category>creativity</category><category>analogy-recognition</category><category>self-improving-ai</category><category>model-recognition</category><category>agi-timelines</category><category>model-performance</category><category>startup-competition</category></item><item><title>Multi-modal, Multi-Aspect, Multi-Form-Factor AI</title><link>https://news.smol.ai/issues/24-04-15-ainews-multi-modal-multi-aspect-multi-form-factor-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-15-ainews-multi-modal-multi-aspect-multi-form-factor-ai/</guid><description>Between April 12-15, **Reka Core** launched a new GPT4-class multimodal foundation model with a detailed technical report described as &quot;full Shazeer.&quot; **Cohere Compass** introduced a foundation embedding model for indexing and searching multi-aspect enterprise data like emails and invoices. The open-source **IDEFICS 2-8B** model continues Google&apos;s Flamingo multimodal model reproduction. **Rewind** pivoted to a multi-platform app called Limitless, moving away from spyware. Reddit discussions highlighted **Apple MLX** outperforming **Ollama** and **Mistral Instruct** on M2 Ultra GPUs, GPU choices for LLMs and Stable Diffusion, and AI-human comparisons by Microsoft Research&apos;s Chris Bishop. Former PayPal CEO Dan Schulman predicted **GPT-5** will drastically reduce job scopes by 80%. **Mistral** CEO Arthur Mensch criticized the obsession with AGI as &quot;creating God.&quot;</description><pubDate>Mon, 15 Apr 2024 22:42:55 GMT</pubDate><category>reka-ai</category><category>cohere</category><category>google</category><category>rewind</category><category>apple</category><category>mistral-ai</category><category>microsoft</category><category>paypal</category><category>gpt-4</category><category>idefics-2-8b</category><category>mistral-instruct</category><category>apple-mlx</category><category>gpt-5</category><category>arthur-mensch</category><category>dan-schulman</category><category>chris-bishop</category><category>multimodality</category><category>foundation-models</category><category>embedding-models</category><category>gpu-performance</category><category>model-comparison</category><category>enterprise-data</category><category>open-source</category><category>performance-optimization</category><category>job-impact</category><category>agi-criticism</category><category>technical-report</category></item><item><title>Zero to GPT in 1 Year</title><link>https://news.smol.ai/issues/24-04-12-ainews-zero-to-gpt-in-1-year/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-12-ainews-zero-to-gpt-in-1-year/</guid><description>**GPT-4 Turbo** reclaimed the top leaderboard spot with significant improvements in coding, multilingual, and English-only tasks, now rolled out in paid **ChatGPT**. Despite this, **Claude Opus** remains superior in creativity and intelligence. **Mistral AI** released powerful open-source models like **Mixtral-8x22B** and **Zephyr 141B** suited for fine-tuning. **LangChain** enhanced tool integration across models, and **Hugging Face** introduced Transformer.js for running transformers in browsers. Medical domain-focused **Medical mT5** was shared as an open-source multilingual text-to-text model. The community also highlighted research on LLMs as regressors and shared practical advice on OCR/PDF data modeling from **Vik Paruchuri**&apos;s journey.</description><pubDate>Fri, 12 Apr 2024 23:27:50 GMT</pubDate><category>openai</category><category>anthropic</category><category>mistral-ai</category><category>langchain</category><category>hugging-face</category><category>gpt-4-turbo</category><category>claude-3-opus</category><category>mixtral-8x22b</category><category>zephyr-141b</category><category>medical-mt5</category><category>vik-paruchuri</category><category>sam-altman</category><category>greg-brockman</category><category>miranda-murati</category><category>abacaj</category><category>mbusigin</category><category>akhaliq</category><category>clementdelangue</category><category>fine-tuning</category><category>multilinguality</category><category>tool-integration</category><category>transformers</category><category>model-evaluation</category><category>open-source-models</category><category>multimodal-llms</category><category>natural-language-processing</category><category>ocr</category><category>model-training</category></item><item><title>Mergestral, Meta MTIAv2, Cohere Rerank 3, Google Infini-Attention</title><link>https://news.smol.ai/issues/24-04-11-ainews-mergestral-meta-mtiav2-cohere-rerank-3-google-infini-attention/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-11-ainews-mergestral-meta-mtiav2-cohere-rerank-3-google-infini-attention/</guid><description>**Meta** announced their new **MTIAv2 chips** designed for training and inference acceleration with improved architecture and integration with PyTorch 2.0. **Mistral** released the **8x22B Mixtral** model, which was merged back into a dense model to effectively create a 22B Mistral model. **Cohere** launched **Rerank 3**, a foundation model enhancing enterprise search and retrieval-augmented generation (RAG) systems supporting 100+ languages. **Google** published a paper on **Infini-attention**, an ultra-scalable linear attention mechanism demonstrated on 1B and 8B models with 1 million sequence length. Additionally, **Meta&apos;s Llama 3** is expected to start rolling out soon. Other notable updates include **Command R+**, an open model surpassing GPT-4 in chatbot performance with 128k context length, and advancements in Stable Diffusion models and RAG pipelines.</description><pubDate>Thu, 11 Apr 2024 22:56:47 GMT</pubDate><category>meta-ai-fair</category><category>mistral-ai</category><category>cohere</category><category>google</category><category>stability-ai</category><category>hugging-face</category><category>ollama</category><category>mistral-8x22b</category><category>command-r-plus</category><category>rerank-3</category><category>infini-attention</category><category>llama-3</category><category>sd-1.5</category><category>cosxl</category><category>aidan_gomez</category><category>ylecun</category><category>swyx</category><category>model-merging</category><category>training-accelerators</category><category>retrieval-augmented-generation</category><category>linear-attention</category><category>long-context</category><category>foundation-models</category><category>image-generation</category><category>rag-pipelines</category><category>model-benchmarking</category><category>context-length</category><category>model-performance</category></item><item><title>Music&apos;s Dall-E moment</title><link>https://news.smol.ai/issues/24-04-10-ainews-musics-dall-e-moment/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-10-ainews-musics-dall-e-moment/</guid><description>**Google&apos;s Griffin architecture** outperforms transformers with faster inference and lower memory usage on long contexts. **Command R+** climbs to 6th place on the LMSYS Chatbot Arena leaderboard, surpassing **GPT-4-0613** and **GPT-4-0314**. **Mistral AI** releases an open-source **8x22B model** with a 64K context window and around 130B total parameters. **Google** open-sources **CodeGemma** models with pre-quantized 4-bit versions for faster downloads. **Ella weights** enhance Stable Diffusion 1.5 with LLM for semantic alignment. **Unsloth** enables 4x larger context windows and 80% memory reduction for finetuning. **Andrej Karpathy** releases LLMs implemented in pure C for potential performance gains. **Command R+** runs in realtime on M2 Max MacBook using iMat q1 quantization. **Cohere&apos;s Command R** model offers low API costs and strong leaderboard performance. **Gemini 1.5** impresses with audio capabilities recognizing speech tone and speaker identification from audio clips.</description><pubDate>Wed, 10 Apr 2024 22:07:48 GMT</pubDate><category>google</category><category>mistral-ai</category><category>lmsys</category><category>cohere</category><category>griffin</category><category>command-r-plus</category><category>gpt-4-0613</category><category>gpt-4-0314</category><category>mistral-8x22b</category><category>codegemma</category><category>stable-diffusion-1.5</category><category>command-r</category><category>gemini-1.5</category><category>andrej-karpathy</category><category>model-architecture</category><category>benchmarking</category><category>open-source</category><category>model-quantization</category><category>memory-optimization</category><category>inference-speed</category><category>multimodality</category><category>finetuning</category><category>performance-optimization</category><category>audio-processing</category></item><item><title>Gemini Pro and GPT4T Vision go GA on the same day by complete coincidence</title><link>https://news.smol.ai/issues/24-04-09-ainews-gemini-pro-and-gpt4t-vision-go-ga-on-the-same-day-by-complete-coincidence/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-09-ainews-gemini-pro-and-gpt4t-vision-go-ga-on-the-same-day-by-complete-coincidence/</guid><description>At **Google Cloud Next**, **Gemini 1.5 Pro** was released with a **million-token context window**, available in **180+ countries**, featuring **9.5 hours of audio understanding**, a new **File API** for nearly unlimited free uploads, and the **Gecko-1b-256/768 embedding model**. **GPT-4 Turbo with Vision** became generally available in the API with a major update improving reasoning capabilities. **Meta Platforms** plans to launch smaller versions of **Llama 3** next week. The **Orca 2.5 7B** model using Direct Nash Optimization outperforms older GPT-4 versions in AlpacaEval. New releases include **Functionary-V2.4** with enhanced function calling and code interpretation, and **CosXL** models for image editing. Research highlights include continuous U-Nets for diffusion models achieving up to **80% faster inference** and a massive multilingual dataset with **~5.6 trillion word tokens**. Creative applications include a no-code touch screen game made with Gemini 1.5 and AI-generated novel trailers.</description><pubDate>Wed, 10 Apr 2024 01:05:31 GMT</pubDate><category>google</category><category>openai</category><category>meta-ai-fair</category><category>hugging-face</category><category>cohere</category><category>gemini-1.5-pro</category><category>gpt-4-turbo</category><category>llama-3</category><category>orca-2.5-7b</category><category>functionary-v2.4</category><category>cosxl</category><category>million-token-context-window</category><category>audio-processing</category><category>file-api</category><category>text-embedding</category><category>function-calling</category><category>reasoning</category><category>direct-nash-optimization</category><category>contrastive-learning</category><category>code-interpreter</category><category>diffusion-models</category><category>neural-odes</category><category>inference-speed</category><category>multilingual-dataset</category><category>image-editing</category><category>no-code-development</category></item><item><title>Anime pfp anon eclipses $10k A::B prompting challenge</title><link>https://news.smol.ai/issues/24-04-08-ainews-anime-pfp-anon-eclipses-dollar10k-ab-prompting-challenge/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-08-ainews-anime-pfp-anon-eclipses-dollar10k-ab-prompting-challenge/</guid><description>**Victor Taelin** issued a $10k challenge to GPT models, initially achieving only **10% success** with state-of-the-art models, but community efforts surpassed **90% success** within 48 hours, highlighting GPT capabilities and common skill gaps. In Reddit AI communities, **Command R Plus (104B)** is running quantized on **M2 Max hardware** via **Ollama** and **llama.cpp** forks, with **GGUF quantizations** released on Huggingface. Streaming text-to-video generation is now available through the **st2v** GitHub repo. **WD Tagger v3** was released for mass auto-captioning datasets with a WebUI. Lesser-known prompting techniques like self-tagging and generational frameworks produced thought-provoking outputs in OpenAI discussions, including experiments with self-evolving system prompts. Stable Diffusion users discussed image composition importance for training character LoRAs and best checkpoints for video game character generation. Discussions also covered scarcity of **5B parameter models** and open(ish) licenses for open source AI. Memes included jokes about ChatGPT and Gemini training data differences.</description><pubDate>Tue, 09 Apr 2024 01:18:42 GMT</pubDate><category>openai</category><category>ollama</category><category>huggingface</category><category>command-r-plus-104b</category><category>stable-diffusion-1.5</category><category>victor-taelin</category><category>futuristfrog</category><category>quantization</category><category>model-optimization</category><category>streaming</category><category>prompt-engineering</category><category>self-prompting</category><category>image-composition</category><category>character-lora-training</category><category>model-size</category><category>open-source-licenses</category><category>memes</category><category>humor</category></item><item><title>Mixture of Depths: Dynamically allocating compute in transformer-based language models</title><link>https://news.smol.ai/issues/24-04-05-ainews-mixture-of-depths-dynamically-allocating-compute-in-transformer-based-language-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-05-ainews-mixture-of-depths-dynamically-allocating-compute-in-transformer-based-language-models/</guid><description>**DeepMind** introduces the Mixture-of-Depths (MoD) technique, dynamically allocating FLOPs across transformer layers to optimize compute usage, achieving over **50% faster** forward passes without training impact. MoD selectively processes tokens using top-k routing, improving efficiency and potentially enabling faster ultra-long context handling. The method can combine with Mixture-of-Experts (MoE) for decoupled routing of queries, keys, and values. Reddit discussions highlight concerns about **LLM hype** overshadowing other AI tech, improvements in transformer efficiency, a new Think-and-Execute framework boosting algorithmic reasoning by **10-20%**, and Visual Autoregressive modeling (VAR) surpassing diffusion models in image quality and speed. On-device model Octopus v2 outperforms GPT-4 in function calling accuracy and latency.</description><pubDate>Fri, 05 Apr 2024 22:44:29 GMT</pubDate><category>deepmind</category><category>octopus-v2</category><category>piotrpadlewski</category><category>transformer-efficiency</category><category>dynamic-compute-allocation</category><category>mixture-of-experts</category><category>mixture-of-depths</category><category>top-k-routing</category><category>algorithmic-reasoning</category><category>visual-autoregressive-modeling</category><category>on-device-models</category><category>function-calling</category><category>scaling-laws</category></item><item><title>Cohere Command R+, Anthropic Claude Tool Use, OpenAI Finetuning</title><link>https://news.smol.ai/issues/24-04-04-ainews-cohere-command-r-anthropic-claude-tool-use-openai-finetuning/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-04-ainews-cohere-command-r-anthropic-claude-tool-use-openai-finetuning/</guid><description>**Cohere** launched **Command R+**, a **104B dense model** with **128k context length** focusing on **RAG**, **tool-use**, and **multilingual** capabilities across **10 key languages**. It supports **Multi-Step Tool use** and offers open weights for research. **Anthropic** introduced **tool use in beta** for **Claude**, supporting over **250 tools** with new cookbooks for practical applications. **OpenAI** enhanced its fine-tuning API with new upgrades and case studies from Indeed, SK Telecom, and Harvey, promoting DIY fine-tuning and custom model training. **Microsoft** achieved a quantum computing breakthrough with an **800x error rate improvement** and the most usable qubits to date. **Stability AI** released **Stable Audio 2.0**, improving audio generation quality and control. The **Opera browser** added local inference support for large language models like **Meta&apos;s Llama**, **Google&apos;s Gemma**, and **Vicuna**. Discussions on Reddit highlighted **Gemini&apos;s large context window**, analysis of **GPT-3.5-Turbo** model size, and a battle simulation between **Claude 3** and **ChatGPT** using local 7B models like **Mistral** and **Gemma**.</description><pubDate>Thu, 04 Apr 2024 22:21:15 GMT</pubDate><category>cohere</category><category>anthropic</category><category>openai</category><category>microsoft</category><category>stability-ai</category><category>opera-software</category><category>meta-ai-fair</category><category>google-deepmind</category><category>mistral-ai</category><category>c4ai-command-r-plus</category><category>claude-3</category><category>gpt-3.5-turbo</category><category>gemini</category><category>mistral-7b</category><category>gemma-2</category><category>claude-3-5</category><category>llama-3</category><category>vicuna</category><category>tool-use</category><category>multilingual-models</category><category>rag</category><category>fine-tuning</category><category>quantum-computing</category><category>audio-generation</category><category>local-inference</category><category>context-windows</category><category>model-size-analysis</category><category>model-comparison</category></item><item><title>ReALM: Reference Resolution As Language Modeling</title><link>https://news.smol.ai/issues/24-04-03-ainews-realm-reference-resolution-as-language-modeling/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-03-ainews-realm-reference-resolution-as-language-modeling/</guid><description>**Apple** is advancing in AI with a new approach called **ReALM: Reference Resolution As Language Modeling**, which improves understanding of ambiguous references using three contexts and finetunes a smaller **FLAN-T5** model that outperforms **GPT-4** on this task. In Reddit AI news, an open-source coding agent **SWE-agent** achieves **12.29%** on the SWE-bench benchmark, and **RAGFlow** introduces a customizable retrieval-augmented generation engine. A new quantization method, **QuaRot**, enables efficient 4-bit inference. AI applications include a t-shirt design generator, **podgenai** for GPT-4 based podcast generation, and an open-source model from **HuggingFace** that runs without a GPU. Industry discussions focus on the impact of large language models on the AI field and efforts to decentralize AI development. **Takuto Takizawa** joins **Stability AI Japan** as Head of Sales &amp; Partnerships.</description><pubDate>Thu, 04 Apr 2024 00:00:20 GMT</pubDate><category>apple</category><category>openai</category><category>hugging-face</category><category>stability-ai</category><category>flan-t5</category><category>gpt-4</category><category>takuto-takizawa</category><category>reference-resolution</category><category>finetuning</category><category>quantization</category><category>retrieval-augmented-generation</category><category>open-source</category><category>coding-agents</category><category>podcast-generation</category><category>image-generation</category><category>ai-industry-trends</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-04-02-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-02-ainews-not-much-happened-today/</guid><description>**RAGFlow** open sourced, a deep document understanding RAG engine with **16.3k context length** and natural language instruction support. **Jamba v0.1**, a **52B parameter** MoE model by Lightblue, released but with mixed user feedback. **Command-R** from **Cohere** available on Ollama library. Analysis of **GPT-3.5-Turbo** architecture reveals about **7 billion parameters** and embedding size of **4096**, comparable to OpenChat-3.5-0106 and Mixtral-8x7B. AI chatbots, including **GPT-4**, outperform humans in debates on persuasion. **Mistral-7B** made amusing mistakes on a math riddle. Hardware highlights include a discounted **HGX H100 640GB** machine with 8 H100 GPUs bought for $58k, and CPU comparisons between **Epyc 9374F** and **Threadripper 1950X** for LLM inference. GPU recommendations for local LLMs focus on VRAM and inference speed, with users testing **4090 GPU** and **Midnight-miqu-70b-v1.0.q5_k_s** model. Stable Diffusion influences gaming habits and AI art evaluation shows bias favoring human-labeled art.</description><pubDate>Tue, 02 Apr 2024 21:04:12 GMT</pubDate><category>cohere</category><category>lightblue</category><category>openai</category><category>mistral-ai</category><category>nvidia</category><category>amd</category><category>hugging-face</category><category>ollama</category><category>jamba-v0.1</category><category>command-r</category><category>gpt-3.5-turbo</category><category>openchat-3.5-0106</category><category>mixtral-8x7b</category><category>mistral-7b</category><category>midnight-miqu-70b-v1.0.q5_k_s</category><category>rag</category><category>mixture-of-experts</category><category>model-architecture</category><category>model-analysis</category><category>debate-persuasion</category><category>hardware-performance</category><category>gpu-inference</category><category>cpu-comparison</category><category>local-llm</category><category>stable-diffusion</category><category>ai-art-bias</category></item><item><title>AdamW -&gt; AaronD?</title><link>https://news.smol.ai/issues/24-04-01-ainews-adamw-greater-aarond/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-04-01-ainews-adamw-greater-aarond/</guid><description>**Aaron Defazio** is gaining attention for proposing a potential tuning-free replacement of the long-standing **Adam optimizer**, showing promising experimental results across classic machine learning benchmarks like ImageNet ResNet-50 and CIFAR-10/100. On Reddit, **Claude 3 Opus** has surpassed all **OpenAI** models on the LMSys leaderboard, while a user pretrained a **LLaMA-based 300M** model outperforming **bert-large** on language modeling tasks with a modest budget. The new **MambaMixer** architecture demonstrates promising results in vision and time series forecasting. In image generation, **Stable Diffusion 1.5** with LoRAs achieves realistic outputs, and the **WDXL** release showcases impressive capabilities. AI applications include an AI-generated Nike spec ad and a chatbot built with OpenAI models that may resist prompt injections. OpenAI is reportedly planning a ban wave targeting policy violators and jailbreak users. *&quot;The high alpha seems to come from Aaron Defazio,&quot;* highlighting his impactful work in optimizer research.</description><pubDate>Mon, 01 Apr 2024 19:58:53 GMT</pubDate><category>openai</category><category>hugging-face</category><category>claude-3-opus</category><category>llama-3</category><category>llama-3-300m</category><category>bert-large</category><category>stable-diffusion-1.5</category><category>wdxl</category><category>aaron-defazio</category><category>optimizer</category><category>machine-learning-benchmarks</category><category>vision</category><category>time-series-forecasting</category><category>image-generation</category><category>prompt-injection</category><category>policy-enforcement</category></item><item><title>Evals-based AI Engineering</title><link>https://news.smol.ai/issues/24-03-29-ainews-evals-based-ai-engineering/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-29-ainews-evals-based-ai-engineering/</guid><description>**Hamel Husain** emphasizes the importance of comprehensive evals in AI product development, highlighting evaluation, debugging, and behavior change as key iterative steps. **OpenAI** released a voice engine demo showcasing advanced voice cloning from small samples, raising safety concerns. Reddit discussions introduced new models like **Jamba** (hybrid Transformer-SSM with MoE), **Bamboo** (7B LLM with high sparsity based on Mistral), **Qwen1.5-MoE** (efficient parameter activation), and **Grok 1.5** (128k context length, surpassing GPT-4 in code generation). Advances in quantization include **1-bit Llama2-7B** models outperforming full precision and the **QLLM** quantization toolbox supporting GPTQ/AWQ/HQQ methods.</description><pubDate>Fri, 29 Mar 2024 22:20:49 GMT</pubDate><category>openai</category><category>mistral-ai</category><category>x-ai</category><category>llamaindex</category><category>jamba</category><category>bamboo</category><category>qwen-1.5-moe</category><category>grok-1.5</category><category>llama2-7b</category><category>hamel-husain</category><category>alec-radford</category><category>evaluation</category><category>fine-tuning</category><category>prompt-engineering</category><category>voice-cloning</category><category>quantization</category><category>model-optimization</category><category>code-generation</category><category>context-windows</category></item><item><title>Jamba: Mixture of Architectures dethrones Mixtral</title><link>https://news.smol.ai/issues/24-03-28-ainews-jamba-mixture-of-architectures-dethrones-mixtral/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-28-ainews-jamba-mixture-of-architectures-dethrones-mixtral/</guid><description>**AI21 labs** released **Jamba**, a **52B parameter MoE model** with **256K context length** and open weights under Apache 2.0 license, optimized for single A100 GPU performance. It features a unique blocks-and-layers architecture combining transformer and MoE layers, competing with models like **Mixtral**. Meanwhile, **Databricks** introduced **DBRX**, a **36B active parameter MoE model** trained on **12T tokens**, noted as a new standard for open LLMs. In image generation, advancements include **Animatediff** for video-quality image generation and **FastSD CPU v1.0.0 beta 28** enabling ultra-fast image generation on CPUs. Other innovations involve style-content separation using **B-LoRA** and improvements in high-resolution image upscaling with **SUPIR**.</description><pubDate>Thu, 28 Mar 2024 23:43:23 GMT</pubDate><category>ai21-labs</category><category>databricks</category><category>together-ai</category><category>hugging-face</category><category>midjourney</category><category>jamba</category><category>dbrx</category><category>mixtral</category><category>animatediff</category><category>fastsd</category><category>sdxs512-0.9</category><category>b-lora</category><category>supir</category><category>mixture-of-experts</category><category>model-architecture</category><category>context-windows</category><category>model-optimization</category><category>fine-tuning</category><category>image-generation</category><category>video-generation</category><category>cpu-optimization</category><category>style-content-separation</category><category>high-resolution-upscaling</category></item><item><title>DBRX: Best open model (just not most efficient)</title><link>https://news.smol.ai/issues/24-03-27-ainews-dbrx-best-open-model-just-not-most-efficient/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-27-ainews-dbrx-best-open-model-just-not-most-efficient/</guid><description>**Databricks Mosaic** has released a new open-source model called **DBRX** that outperforms **Grok**, **Mixtral**, and **Llama2** on evaluations while being about **2x more efficient** than Llama2 and Grok. The model was trained on **12 trillion tokens** using **3,000 H100 GPUs** over 2 months, with an estimated compute cost of **$10 million**. It uses OpenAI&apos;s **100k tiktoken tokenizer** and shows strong zero-shot code generation performance, even beating **GPT-4** on the Humaneval benchmark. DBRX also upstreamed work to **MegaBlocks** open source. Despite its scale and efficiency, DBRX&apos;s performance on MMLU is only slightly better than Mixtral, raising questions about its scaling efficiency. The focus of DBRX is on enabling users to train models efficiently, with MoE training being about **2x more FLOP-efficient** than dense models, achieving similar quality with nearly **4x less compute** than previous MPT models. This release is part of the ongoing competition for open-source AI leadership, including models like **Dolly**, **MPT**, and **Mistral**. *&quot;If it activates 36B params, the model&apos;s perf should be equivalent to a 72B dense model or even 80B,&quot;* says Qwen&apos;s tech lead.</description><pubDate>Wed, 27 Mar 2024 22:33:19 GMT</pubDate><category>databricks</category><category>hugging-face</category><category>mistral-ai</category><category>mosaicml</category><category>openai</category><category>dbrx</category><category>grok</category><category>mixtral</category><category>llama-2</category><category>mpt-7b</category><category>gpt-4</category><category>mixture-of-experts</category><category>model-efficiency</category><category>tokenization</category><category>model-training</category><category>code-generation</category><category>model-architecture</category><category>open-source-models</category><category>benchmarking</category><category>fine-tuning</category></item><item><title>Claude 3 is officially America&apos;s Next Top Model</title><link>https://news.smol.ai/issues/24-03-26-ainews-claude-3-is-officially-americas-next-top-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-26-ainews-claude-3-is-officially-americas-next-top-model/</guid><description>**Claude 3 Opus** outperforms **GPT4T** and **Mistral Large** in blind Elo rankings, with **Claude 3 Haiku** marking a new cost-performance frontier. Fine-tuning techniques like **QLoRA** on **Mistral 7B** and evolutionary model merging on HuggingFace models are highlighted. Public opinion shows strong opposition to ASI development. Research supervision opportunities in AI alignment are announced. The **Stable Diffusion 3 (SD3)** release raises workflow concerns for tools like **ComfyUI** and **automatic1111**. **Opus** shows a 5% performance dip on **OpenRouter** compared to the **Anthropic API**. A new benchmark stresses LLM recall at long contexts, with **Mistral 7B** struggling and **Qwen 72b** performing well.</description><pubDate>Wed, 27 Mar 2024 00:11:55 GMT</pubDate><category>anthropic</category><category>mistral-ai</category><category>huggingface</category><category>openrouter</category><category>stable-diffusion</category><category>automatic1111</category><category>comfyui</category><category>claude-3-opus</category><category>claude-3-sonnet</category><category>claude-3-haiku</category><category>gpt-4o-mini</category><category>mistral-7b</category><category>qwen-72b</category><category>mark_riedl</category><category>ethanjperez</category><category>stuhlmueller</category><category>ylecun</category><category>aravsrinivas</category><category>fine-tuning</category><category>model-merging</category><category>alignment</category><category>ai-ethics</category><category>benchmarking</category><category>model-performance</category><category>long-context</category><category>cost-efficiency</category><category>model-evaluation</category></item><item><title>Andrew likes Agents</title><link>https://news.smol.ai/issues/24-03-25-ainews-andrew-likes-agents/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-25-ainews-andrew-likes-agents/</guid><description>**Andrew Ng&apos;s The Batch writeup on Agents** highlighted the significant improvement in coding benchmark performance when using an iterative agent workflow, with **GPT-3.5** wrapped in an agent loop achieving up to **95.1%** correctness on HumanEval, surpassing **GPT-4** zero-shot at **67.0%**. The report also covers new developments in **Stable Diffusion** models like **Cyberrealistic_v40**, **Platypus XL**, and **SDXL Lightning** for Naruto-style image generation, alongside innovations in LoRA and upscaling techniques. Discussions on **local LLM deployment** and optimization focus on hardware setups and finetuning strategies for efficient inference and multi-user serving. Emad&apos;s departure from **Stability AI** and new **Sora** videos from **OpenAI** were also noted.</description><pubDate>Tue, 26 Mar 2024 01:11:50 GMT</pubDate><category>openai</category><category>stability-ai</category><category>gpt-3.5</category><category>gpt-4</category><category>cyberrealistic_v40</category><category>platypus-xl</category><category>sdxl-lightning</category><category>andrew-ng</category><category>lilian-weng</category><category>emad</category><category>agents</category><category>human-eval-benchmark</category><category>fine-tuning</category><category>local-llm-deployment</category><category>inference-speed</category><category>image-generation</category><category>lora</category><category>upscaling</category><category>workflow-optimization</category></item><item><title>Astro Nano</title><link>https://news.smol.ai/projects/project-2/</link><guid isPermaLink="true">https://news.smol.ai/projects/project-2/</guid><description>Minimal portfolio and blog build with astro and no frameworks.</description><pubDate>Tue, 26 Mar 2024 00:00:00 GMT</pubDate></item><item><title>not much happened today</title><link>https://news.smol.ai/issues/24-03-22-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-22-ainews-not-much-happened-today/</guid><description>The Reddit community /r/LocalLlama discusses **fine-tuning and training LLMs**, including tutorials and questions on training models with specific data like dictionaries and synthetic datasets with **25B+ tokens**. Users explore **retrieval-augmented generation (RAG)** challenges with models like **mistral-7b** and embedding generation for EEG brain activity. Discussions include **hardware optimization** for running **llama-2-70b** locally under budget constraints, and performance benchmarks for **qwen-1.5** models. There is interest in extending LLM capabilities, such as converting **llama-2-7b** into a vision-capable model like **llava** and improving model memory for longer context retention.</description><pubDate>Fri, 22 Mar 2024 23:55:31 GMT</pubDate><category>microsoft</category><category>mistral-ai</category><category>ollama</category><category>llama-2-70b</category><category>llama-2-7b</category><category>mistral-7b</category><category>qwen-1.5</category><category>llava</category><category>fine-tuning</category><category>synthetic-data</category><category>retrieval-augmented-generation</category><category>embeddings</category><category>hardware-optimization</category><category>performance-benchmarks</category><category>model-memory</category><category>multimodality</category></item><item><title>Welcome /r/LocalLlama!</title><link>https://news.smol.ai/issues/24-03-21-ainews-welcome-rlocalllama/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-21-ainews-welcome-rlocalllama/</guid><description>**Sakana** released a paper on evolutionary model merging. **OpenInterpreter** launched their **O1 devkit**. Discussions highlight **Claude Haiku**&apos;s underrated performance with 10-shot examples. On **Reddit&apos;s IPO**, AINews introduces Reddit summaries starting with /r/LocalLlama, covering upcoming subreddits like r/machinelearning and r/openai. **Aether Research** released **Cerebrum 8x7b** based on **Mixtral**, matching **GPT-3.5 Turbo** and **Gemini Pro** on reasoning tasks, setting a new open-source reasoning SOTA. **Moistral 11B v1** finetuned model from Cream-Phi-2 creators was released. A creative writing benchmark uses **Claude Opus** as judge. Hobbyists explore **1.58 BitNet** ternary quantization and **1-bit LLMs** training. Nvidia&apos;s **Blackwell (h200)** chip supports **FP4 precision** quantization. **LMDeploy v0.2.6+** enables efficient vision-language model deployment with models like **Qwen-VL-Chat**. Users seek GUIs for LLM APIs with plugin and RAG support. Pipelines for synthetic training data generation and fine-tuning language models for chat are discussed.</description><pubDate>Thu, 21 Mar 2024 23:33:53 GMT</pubDate><category>sakana</category><category>openinterpreter</category><category>reddit</category><category>aether-research</category><category>mistral-ai</category><category>nvidia</category><category>lmdeploy</category><category>cerebrum-8x7b</category><category>mixtral-7b</category><category>gpt-3.5-turbo</category><category>gemini-pro</category><category>moistral-11b-v1</category><category>claude-opus</category><category>qwen-vl-chat</category><category>model-merging</category><category>benchmarking</category><category>quantization</category><category>performance-optimization</category><category>deployment</category><category>vision</category><category>fine-tuning</category><category>training-data</category><category>synthetic-data</category><category>rag</category><category>gui</category></item><item><title>Shipping and Dipping: Inflection + Stability edition</title><link>https://news.smol.ai/issues/24-03-20-ainews-shipping-and-dipping-inflection-stability-edition/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-20-ainews-shipping-and-dipping-inflection-stability-edition/</guid><description>**Inflection AI** and **Stability AI** recently shipped major updates (**Inflection AI 2.5** and **Stable Diffusion 3**) but are now experiencing significant executive departures, signaling potential consolidation in the GPU-rich startup space. **Mustafa Suleyman** has joined **Microsoft AI** as CEO, overseeing consumer AI products like Copilot, Bing, and Edge. **Microsoft Azure** is collaborating with **NVIDIA** on the Grace Blackwell 200 Superchip. **Google DeepMind** announced **TacticAI**, an AI assistant for football tactics developed with Liverpool FC, using geometric deep learning and achieving 90% expert approval in blind tests. **Anthropic** released **Claude 3 Haiku** and **Claude 3 Sonnet** on Google Cloud&apos;s Vertex AI, with **Claude 3 Opus** coming soon. Concerns about AI job displacement arise as **NVIDIA** introduces AI nurses that outperform humans at bedside manner at 90% lower cost.</description><pubDate>Thu, 21 Mar 2024 00:59:01 GMT</pubDate><category>inflection-ai</category><category>stability-ai</category><category>microsoft</category><category>nvidia</category><category>google-deepmind</category><category>anthropic</category><category>inflection-ai-2.5</category><category>stable-diffusion-3</category><category>claude-3-haiku</category><category>claude-3-sonnet</category><category>claude-3-opus</category><category>tacticai</category><category>mustafa-suleyman</category><category>executive-departures</category><category>gpu-acceleration</category><category>ai-assistants</category><category>geometric-deep-learning</category><category>ai-integration</category><category>ai-cost-reduction</category><category>ai-job-displacement</category><category>ai-healthcare</category><category>model-release</category></item><item><title>World_sim.exe</title><link>https://news.smol.ai/issues/24-03-19-ainews-worldsimexe/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-19-ainews-worldsimexe/</guid><description>**NVIDIA** announced **Project GR00T**, a foundation model for humanoid robot learning using multimodal instructions, built on their tech stack including Isaac Lab, OSMO, and Jetson Thor. They revealed the **DGX Grace-Blackwell GB200** with over **1 exaflop** compute, capable of training **GPT-4 1.8T parameters** in 90 days on 2000 Blackwells. Jensen Huang confirmed GPT-4 has **1.8 trillion parameters**. The new **GB200 GPU** supports float4/6 precision with ~3 bits per parameter and achieves **40,000 TFLOPs** on fp4 with 2x sparsity. 

Open source highlights include the release of **Grok-1**, a **340B parameter** model, and **Stability AI&apos;s SV3D**, an open-source text-to-video generation solution. **Nous Research** collaborated on implementing Steering Vectors in Llama.CPP. 

In Retrieval Augmented Generation (RAG), a new **5.5-hour tutorial** builds a pipeline using open-source HF models, and **LangChain** released a video on query routing and announced integration with **NVIDIA NIM** for GPU-optimized LLM inference. 

Prominent opinions include **Yann LeCun** distinguishing language from other cognitive abilities, **Sam Altman** predicting AGI arrival in 6 years with a leap from GPT-4 to GPT-5 comparable to GPT-3 to GPT-4, and discussions on the philosophical status of LLMs like Claude. There is also advice against training models from scratch for most companies.</description><pubDate>Wed, 20 Mar 2024 00:46:48 GMT</pubDate><category>nvidia</category><category>nous-research</category><category>stability-ai</category><category>hugging-face</category><category>langchain</category><category>anthropic</category><category>openai</category><category>gpt-4</category><category>gpt-4o</category><category>grok-1</category><category>llama-cpp</category><category>claude-3-opus</category><category>claude-3</category><category>gpt-5</category><category>jensen-huang</category><category>yann-lecun</category><category>sam-altman</category><category>multimodality</category><category>foundation-models</category><category>hardware-optimization</category><category>model-quantization</category><category>float4</category><category>float6</category><category>retrieval-augmented-generation</category><category>text-to-video</category><category>prompt-engineering</category><category>long-form-rag</category><category>gpu-optimization</category><category>philosophy-of-ai</category><category>agi-predictions</category></item><item><title>Grok-1 in Bio</title><link>https://news.smol.ai/issues/24-03-18-ainews-grok-1-in-bio/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-18-ainews-grok-1-in-bio/</guid><description>**Grok-1**, a **314B parameter Mixture-of-Experts (MoE) model** from **xAI**, has been released under an Apache 2.0 license, sparking discussions on its architecture, finetuning challenges, and performance compared to models like **Mixtral** and **Miqu 70B**. Despite its size, its **MMLU benchmark performance** is currently unimpressive, with expectations that **Grok-2** will be more competitive. The model&apos;s weights and code are publicly available, encouraging community experimentation. **Sam Altman** highlighted the growing importance of compute resources, while **Grok&apos;s** potential deployment on **Groq hardware** was noted as a possible game-changer. Meanwhile, **Anthropic&apos;s Claude** continues to attract attention for its &quot;spiritual&quot; interaction experience and consistent ethical framework. The release also inspired memes and humor within the AI community.</description><pubDate>Tue, 19 Mar 2024 00:07:45 GMT</pubDate><category>xai</category><category>mistral-ai</category><category>perplexity-ai</category><category>groq</category><category>anthropic</category><category>openai</category><category>grok-1</category><category>mixtral</category><category>miqu-70b</category><category>claude-3-opus</category><category>claude-3</category><category>claude-3-haiku</category><category>sam-altman</category><category>arthur-mensch</category><category>daniel-han</category><category>arav-srinivas</category><category>francis-yao</category><category>mixture-of-experts</category><category>model-release</category><category>model-performance</category><category>benchmarking</category><category>finetuning</category><category>compute</category><category>hardware-optimization</category><category>mmlu</category><category>model-architecture</category><category>open-source</category><category>memes</category></item><item><title>Astro Sphere</title><link>https://news.smol.ai/projects/project-1/</link><guid isPermaLink="true">https://news.smol.ai/projects/project-1/</guid><description>Portfolio and blog build with astro.</description><pubDate>Mon, 18 Mar 2024 00:00:00 GMT</pubDate></item><item><title>MM1: Apple&apos;s first Large Multimodal Model</title><link>https://news.smol.ai/issues/24-03-15-ainews-mm1-apples-first-large-multimodal-model/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-15-ainews-mm1-apples-first-large-multimodal-model/</guid><description>**Apple** announced the **MM1** multimodal LLM family with up to **30B parameters**, claiming performance comparable to **Gemini-1** and beating larger older models on VQA benchmarks. The paper targets researchers and hints at applications in embodied agents and business/education. **Yann LeCun** emphasized that human-level AI requires understanding the physical world, memory, reasoning, and hierarchical planning, while **Franois Chollet** cautioned that NLP is far from solved despite LLM advances. **Cohere** released **Command-R**, a model for Retrieval Augmented Generation, and **Anthropic** highlighted the **Claude 3** family (Opus, Sonnet, Haiku) for various application needs. Open-source hardware **DexCap** enables dexterous robot manipulation data collection affordably. Tools like **CopilotKit** simplify AI integration into React apps, and migration to **Keras 3** with JAX backend offers faster training. New projects improve reranking for retrieval and add financial agents to **LangChain**. The content includes insights on AI progress, new models, open-source tools, and frameworks.</description><pubDate>Fri, 15 Mar 2024 23:34:51 GMT</pubDate><category>apple</category><category>cohere</category><category>anthropic</category><category>hugging-face</category><category>langchain</category><category>mm1</category><category>gemini-1</category><category>command-r</category><category>claude-3-opus</category><category>claude-3-sonnet</category><category>claude-3-haiku</category><category>claude-3</category><category>yann-lecun</category><category>francois-chollet</category><category>multimodality</category><category>vqa</category><category>fine-tuning</category><category>retrieval-augmented-generation</category><category>open-source</category><category>robotics</category><category>model-training</category><category>react</category><category>reranking</category><category>financial-agents</category></item><item><title>Not much happened piday</title><link>https://news.smol.ai/issues/24-03-14-ainews-not-much-happened-piday/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-14-ainews-not-much-happened-piday/</guid><description>**DeepMind** announces **SIMA**, a generalist AI agent capable of following natural language instructions across diverse 3D environments and video games, advancing embodied AI agents. **Anthropic** releases **Claude 3 Haiku**, their fastest and most affordable model, now available via API and Perplexity. New research explores language model scaling laws, over-training, and introduces **Branch-Train-MiX (BTX)** for efficient training of large language models using mixture-of-experts. Predictions suggest software engineering jobs will grow to **30-35 million** in five years, aided by AI coding assistants like **Cohere&apos;s Command-R** focusing on retrieval-augmented generation and tool use. The **EU AI Act** is approved, mandating transparency in training data for GPAI systems. Privacy-preserving in-context learning with differential privacy is highlighted as promising work. Memes humorously discuss AI software engineers and notable figures like **Andrej Karpathy**.</description><pubDate>Thu, 14 Mar 2024 23:53:52 GMT</pubDate><category>deepmind</category><category>anthropic</category><category>cohere</category><category>claude-3-haiku</category><category>demis-hassabis</category><category>fchollet</category><category>abacaj</category><category>andrej-karpathy</category><category>embodied-ai-agents</category><category>natural-language-instructions</category><category>language-model-scaling</category><category>mixture-of-experts</category><category>retrieval-augmented-generation</category><category>software-engineering</category><category>ai-regulation</category><category>differential-privacy</category><category>privacy-preserving-learning</category><category>humor</category></item><item><title>DeepMind SIMA: one AI, 9 games, 600 tasks, vision+language ONLY</title><link>https://news.smol.ai/issues/24-03-13-ainews-deepmind-sima-one-ai-9-games-600-tasks-visionlanguage-only/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-13-ainews-deepmind-sima-one-ai-9-games-600-tasks-visionlanguage-only/</guid><description>**DeepMind SIMA** is a generalist AI agent for 3D virtual environments evaluated on **600 tasks** across **9 games** using only screengrabs and natural language instructions, achieving **34%** success compared to humans&apos; **60%**. The model uses a multimodal Transformer architecture. **Andrej Karpathy** outlines AI autonomy progression in software engineering, while **Arav Srinivas** praises Cognition Labs&apos; AI agent demo. **François Chollet** expresses skepticism about automating software engineering fully. **Yann LeCun** suggests moving away from generative models and reinforcement learning towards human-level AI. Meta&apos;s **Llama-3** training infrastructure with **24k H100 Cluster Pods** is shared by **Soumith Chintala** and **Yann LeCun**. **Deepgram&apos;s Aura** offers low-latency speech APIs, and **Modal Labs&apos; Devin AI** demonstrates document navigation and interaction with ComfyUI. Memes and humor circulate in the AI community.</description><pubDate>Thu, 14 Mar 2024 01:07:46 GMT</pubDate><category>deepmind</category><category>cognition-labs</category><category>deepgram</category><category>modal-labs</category><category>meta-ai-fair</category><category>anthropic</category><category>llama-3</category><category>claude-3-opus</category><category>claude-3</category><category>gpt-3.5-turbo</category><category>andrej-karpathy</category><category>arav-srinivas</category><category>francois-chollet</category><category>yann-lecun</category><category>soumith-chintala</category><category>john-carmack</category><category>multimodality</category><category>transformer</category><category>software-engineering</category><category>ai-agents</category><category>ai-infrastructure</category><category>training</category><category>text-to-speech</category><category>speech-to-text</category><category>real-time-processing</category><category>model-architecture</category><category>benchmarking</category></item><item><title>The world&apos;s first fully autonomous AI Engineer</title><link>https://news.smol.ai/issues/24-03-12-ainews-the-worlds-first-fully-autonomous-ai-engineer/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-12-ainews-the-worlds-first-fully-autonomous-ai-engineer/</guid><description>**Cognition Labs&apos;s Devin** is highlighted as a potentially groundbreaking AI software engineer agent capable of learning unfamiliar technologies, addressing bugs, deploying frontend apps, and fine-tuning its own AI models. It integrates **OpenAI&apos;s GPT-4** with reinforcement learning and features tools like asynchronous chat, browser, shell access, and an IDE. The system claims advanced long-term reasoning and planning abilities, attracting praise from investors like **Patrick Collison** and **Fred Ehrsam**. The technology is noted for its potential as one of the most advanced AI agents, sparking excitement about agents and AGI.</description><pubDate>Tue, 12 Mar 2024 23:05:08 GMT</pubDate><category>cognition-labs</category><category>openai</category><category>gpt-4</category><category>devin</category><category>patrick-collison</category><category>fred-ehrsam</category><category>tim-dettmers</category><category>reinforcement-learning</category><category>fine-tuning</category><category>long-term-reasoning</category><category>planning</category><category>ai-agents</category><category>software-engineering</category><category>model-integration</category><category>asynchronous-chat</category><category>ide</category><category>agentic-ai</category></item><item><title>Fixing Gemma</title><link>https://news.smol.ai/issues/24-03-11-ainews-fixing-gemma/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-11-ainews-fixing-gemma/</guid><description>**Google&apos;s Gemma model** was found unstable for finetuning until **Daniel Han from Unsloth AI** fixed 8 bugs, improving its implementation. **Yann LeCun** explained technical details of a pseudo-random bit sequence for adaptive equalizers, while **François Chollet** discussed the low information bandwidth of the human visual system. **Arav Srinivas** reported that **Claude 3 Opus** showed no hallucinations in extensive testing, outperforming **GPT-4** and **Mistral-Large** in benchmarks. Reflections from **Yann LeCun** highlight ongoing AI progress toward human-level intelligence. The community is shifting pipelines to work better with Claude models, and emotional experiences in ML development were shared by **Aidan Clark**.</description><pubDate>Tue, 12 Mar 2024 00:03:26 GMT</pubDate><category>google</category><category>unsloth</category><category>anthropic</category><category>mistral-ai</category><category>gemma</category><category>claude-3-opus</category><category>claude-3</category><category>mistral-large</category><category>gpt-4</category><category>daniel-han</category><category>yann-lecun</category><category>francois-chollet</category><category>arav-srinivas</category><category>_aidan_clark_</category><category>finetuning</category><category>numerical-precision</category><category>benchmarking</category><category>structured-data-extraction</category><category>adaptive-equalizer</category><category>information-theory</category><category>hallucination-detection</category><category>model-stability</category></item><item><title>FSDP+QLoRA: the Answer to 70b-scale AI for desktop class GPUs</title><link>https://news.smol.ai/issues/24-03-08-ainews-fsdpqlora-the-answer-to-70b-scale-ai-for-desktop-class-gpus/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-08-ainews-fsdpqlora-the-answer-to-70b-scale-ai-for-desktop-class-gpus/</guid><description>**Jeremy Howard** and collaborators released a new tool combining **FSDP**, **QLoRA**, and **HQQ** to enable training **70b-parameter** models on affordable consumer GPUs like **RTX 4090s** with only **24GB RAM**, overcoming traditional memory constraints that required expensive data center GPUs costing over $150k. The approach shards quantized models across multiple GPUs and uses techniques like gradient checkpointing and CPU offloading to achieve efficient training on desktop-class hardware. The blogpost details challenges and solutions integrating these methods, highlighting a significant cost reduction from $150k to under $2.5k for training large language models. Additionally, Twitter recaps mention **Inflection AI**&apos;s **Inflection-2.5** model rivaling **GPT-4** in benchmarks with less compute, and **Grok** improving speed by 3x. **Yann LeCun** discusses multi-step reasoning training for LLMs.</description><pubDate>Fri, 08 Mar 2024 23:21:13 GMT</pubDate><category>answer.ai</category><category>hugging-face</category><category>meta-ai-fair</category><category>nvidia</category><category>inflectionai</category><category>qlora</category><category>fsdp</category><category>inflection-2.5</category><category>gpt-4</category><category>jeremy_howard</category><category>tim_dettmers</category><category>yann_lecun</category><category>model-training</category><category>quantization</category><category>memory-optimization</category><category>gradient-checkpointing</category><category>cpu-offloading</category><category>fine-tuning</category><category>model-sharding</category><category>reinforcement-learning</category><category>chain-of-thought</category><category>benchmarking</category></item><item><title>Inflection-2.5 at 94% of GPT4, and Pi at 6m MAU</title><link>https://news.smol.ai/issues/24-03-07-ainews-inflection-25-at-94percent-of-gpt4-and-pi-at-6m-mau/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-07-ainews-inflection-25-at-94percent-of-gpt4-and-pi-at-6m-mau/</guid><description>**Mustafa Suleyman** announced **Inflection 2.5**, which achieves *more than 94% the average performance of GPT-4 despite using only 40% the training FLOPs*. **Pi**&apos;s user base is growing about 10% weekly, with new features like realtime web search. The community noted similarities between Inflection 2.5 and **Claude 3 Sonnet**. **Claude 3 Opus** outperformed **GPT-4** in a 1.5:1 vote and is now the default for **Perplexity Pro** users. **Anthropic** added experimental tool calling support for Claude 3 via **LangChain**. **LlamaIndex** released LlamaParse JSON Mode for structured PDF parsing and added video retrieval via VideoDB, enabling retrieval-augmented generation (RAG) pipelines. A paper proposed knowledge-augmented planning for LLM agents. New benchmarks like TinyBenchmarks and the **Yi-9B** model release show strong code and math performance, surpassing **Mistral**.</description><pubDate>Fri, 08 Mar 2024 02:11:17 GMT</pubDate><category>inflection</category><category>anthropic</category><category>perplexity-ai</category><category>llamaindex</category><category>mistral-ai</category><category>langchain</category><category>inflection-2.5</category><category>claude-3-sonnet</category><category>claude-3-opus</category><category>gpt-4</category><category>yi-9b</category><category>mistral</category><category>mustafa-suleyman</category><category>amanda-askell</category><category>jeremyphoward</category><category>abacaj</category><category>omarsar0</category><category>retrieval-augmented-generation</category><category>benchmarking</category><category>ocr</category><category>structured-output</category><category>video-retrieval</category><category>knowledge-augmentation</category><category>planning</category><category>tool-use</category><category>evaluation</category><category>code-benchmarks</category><category>math-benchmarks</category></item><item><title>Not much happened today</title><link>https://news.smol.ai/issues/24-03-06-ainews-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-06-ainews-not-much-happened-today/</guid><description>**Anthropic** released **Claude 3**, replacing Claude 2.1 as the default on Perplexity AI, with **Claude 3 Opus** surpassing **GPT-4** in capability. Debate continues on whether Claude 3&apos;s performance stems from emergent properties or pattern matching. **LangChain** and **LlamaIndex** added support for Claude 3 enabling multimodal and tool-augmented applications. Despite progress, current models still face challenges in out-of-distribution reasoning and robustness. **Cohere** partnered with **Accenture** for enterprise AI search, while **Mistral AI** and **Snowflake** collaborate to provide LLMs on Snowflake&apos;s platform. **Together AI Research** integrates **Deepspeed** innovations to accelerate generative AI infrastructure. **Hugging Face** and the **European Space Agency** released a large earth observation dataset, and **Google** open sourced **Gemma 2B**, optimized for smartphones via the MLC-LLM project. **GPT4All** improved model discoverability for open models. The AI community balances excitement over new models with concerns about limitations and robustness, alongside growing enterprise adoption and open-source contributions. Memes and humor continue to provide social commentary.</description><pubDate>Thu, 07 Mar 2024 01:15:26 GMT</pubDate><category>anthropic</category><category>perplexity</category><category>langchain</category><category>llamaindex</category><category>cohere</category><category>accenture</category><category>mistral-ai</category><category>snowflake</category><category>together-ai</category><category>hugging-face</category><category>european-space-agency</category><category>google</category><category>gpt4all</category><category>claude-3</category><category>claude-3-opus</category><category>claude-3-sonnet</category><category>gpt-4</category><category>gemma-2b</category><category>multimodality</category><category>instruction-following</category><category>out-of-distribution-reasoning</category><category>robustness</category><category>enterprise-ai</category><category>cloud-infrastructure</category><category>open-datasets</category><category>model-deployment</category><category>model-discoverability</category><category>generative-ai</category><category>image-generation</category></item><item><title>Stable Diffusion 3 — Rombach &amp; Esser did it again!</title><link>https://news.smol.ai/issues/24-03-05-ainews-stable-diffusion-3-rombach-and-esser-did-it-again/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-05-ainews-stable-diffusion-3-rombach-and-esser-did-it-again/</guid><description>**Over 2500 new community members joined following Soumith Chintala&apos;s shoutout, highlighting growing interest in SOTA LLM-based summarization. The major highlight is the detailed paper release of **Stable Diffusion 3 (SD3)**, showcasing advanced text-in-image control and complex prompt handling, with the model outperforming other SOTA image generation models in human-evaluated benchmarks. The SD3 model is based on an enhanced Diffusion Transformer architecture called **MMDiT**. Meanwhile, **Anthropic** released **Claude 3** models, noted for human-like responses and emotional depth, scoring 79.88% on HumanEval but costing over twice as much as GPT-4. Microsoft launched new Orca-based models and datasets, and Latitude released **DolphinCoder-StarCoder2-15b** with strong coding capabilities. Integration of image models by **Perplexity AI** and 3D CAD generation by **PolySpectra** powered by **LlamaIndex** were also highlighted. *&quot;SD3&apos;s win rate beats all other SOTA image gen models (except perhaps Ideogram)&quot;* and *&quot;Claude 3 models are very good at generating d3 visualizations from text descriptions.&quot;*</description><pubDate>Tue, 05 Mar 2024 22:30:03 GMT</pubDate><category>stability-ai</category><category>anthropic</category><category>microsoft</category><category>latitude</category><category>perplexity-ai</category><category>llamaindex</category><category>tripo-ai</category><category>stable-diffusion-3</category><category>claude-3</category><category>orca</category><category>dolphincoder-starcoder2-15b</category><category>soumith-chintala</category><category>bill-peebles</category><category>swyx</category><category>kevinafischer</category><category>jeremyphoward</category><category>akhaliq</category><category>karinanguyen_</category><category>aravsrinivas</category><category>diffusion-models</category><category>multimodality</category><category>benchmarking</category><category>human-evaluation</category><category>text-generation</category><category>image-generation</category><category>3d-modeling</category><category>fine-tuning</category><category>roleplay</category><category>coding</category><category>dataset-release</category></item><item><title>Claude 3 just destroyed GPT 4 (see for yourself)</title><link>https://news.smol.ai/issues/24-03-04-ainews-claude-3-just-destroyed-gpt-4-see-for-yourself/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-04-ainews-claude-3-just-destroyed-gpt-4-see-for-yourself/</guid><description>**Claude 3** from **Anthropic** launches in three sizes: Haiku (small, unreleased), Sonnet (medium, default on claude.ai, AWS, and GCP), and Opus (large, on Claude Pro). Opus outperforms **GPT-4** on key benchmarks like GPQA, impressing benchmark authors. All models support **multimodality** with advanced vision capabilities, including converting a 2-hour video into a blog post. Claude 3 offers improved alignment, fewer refusals, and extended context length up to **1 million tokens** with near-perfect recall. Haiku is noted for speed and cost-efficiency, processing dense research papers in under three seconds. The models excel at following complex instructions and producing structured outputs like JSON. Safety improvements reduce refusal rates, though some criticism remains from experts. Claude 3 is trained on synthetic data and shows strong domain-specific evaluation results in finance, medicine, and philosophy.</description><pubDate>Mon, 04 Mar 2024 23:59:02 GMT</pubDate><category>anthropic</category><category>amazon</category><category>google</category><category>claude-ai</category><category>claude-3</category><category>claude-3-opus</category><category>claude-3-sonnet</category><category>claude-3-haiku</category><category>gpt-4</category><category>mmitchell</category><category>connor-leahy</category><category>multimodality</category><category>vision</category><category>long-context</category><category>model-alignment</category><category>model-evaluation</category><category>synthetic-data</category><category>structured-output</category><category>instruction-following</category><category>model-speed</category><category>cost-efficiency</category><category>benchmarking</category><category>safety</category></item><item><title>The Era of 1-bit LLMs</title><link>https://news.smol.ai/issues/24-03-01-ainews-the-era-of-1-bit-llms/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-03-01-ainews-the-era-of-1-bit-llms/</guid><description>**The Era of 1-bit LLMs** research, including the **BitNet b1.58** model, introduces a ternary parameter approach that matches full-precision Transformer LLMs in performance while drastically reducing energy costs by **38x**. This innovation promises new scaling laws and hardware designs optimized for 1-bit LLMs. Discussions on AI Twitter highlight advances in **AGI societal impact**, **robotics with multimodal models**, **fine-tuning techniques like ResLoRA**, and **AI security efforts at Hugging Face**. Ethical considerations in generative AI and humor within the AI community are also prominent topics.</description><pubDate>Fri, 01 Mar 2024 22:33:03 GMT</pubDate><category>hugging-face</category><category>bitnet-b1.58</category><category>swyx</category><category>levelsio</category><category>gdb</category><category>npew</category><category>_akhaliq</category><category>osanseviero</category><category>mmitchell_ai</category><category>deliprao</category><category>nearcyan</category><category>clementdelangue</category><category>quantization</category><category>model-optimization</category><category>energy-efficiency</category><category>fine-tuning</category><category>robotics</category><category>multimodality</category><category>ai-security</category><category>ethics</category><category>humor</category></item><item><title>Dia de las Secuelas (StarCoder, The Stack, Dune, SemiAnalysis)</title><link>https://news.smol.ai/issues/24-02-29-ainews-dia-de-las-secuelas-starcoder-the-stack-dune-semianalysis/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-29-ainews-dia-de-las-secuelas-starcoder-the-stack-dune-semianalysis/</guid><description>**HuggingFace/BigCode** has released **StarCoder v2**, including the **StarCoder2-15B** model trained on over **600 programming languages** using the **The Stack v2** dataset. This release marks a state-of-the-art achievement for models of this size, with opt-out requests excluded from training data. A detailed technical report is available, highlighting the model&apos;s capabilities and training methodology. Additionally, a live event featuring **Dylan Patel** discussing GPU economics is announced for San Francisco.</description><pubDate>Fri, 01 Mar 2024 00:14:08 GMT</pubDate><category>hugging-face</category><category>bigcode</category><category>starcoder-2</category><category>starcoder2-15b</category><category>dylan-patel</category><category>code-generation</category><category>model-training</category><category>dataset-release</category><category>model-performance</category></item><item><title>... and welcome AI Twitter!</title><link>https://news.smol.ai/issues/24-02-28-ainews-and-welcome-ai-twitter/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-28-ainews-and-welcome-ai-twitter/</guid><description>The AI Twitter discourse from **2/27-28/2024** covers a broad spectrum including **ethical considerations** highlighted by **Margaret Mitchell** around **Google Gemini&apos;s** launch, and **John Carmack&apos;s** insights on evolving coding skills in the AI era. **Guillaume Lample** announced the release of the **Mistral Large** multilingual model. Discussions also touched on potential leadership changes at **Google** involving **Sundar Pichai**, and **OpenAI&apos;s** possible entry into the synthetic data market as noted by **Delip Rao**. Technological advancements include **Yann LeCun&apos;s** commentary on running LLMs on mobile devices and **Alex Wang&apos;s** praise for the **Apple Vision Pro**. Financial platform issues were raised by **Pieter Levels** regarding **Stripe&apos;s** payment policies. The cultural dynamics within big tech were discussed by **François Chollet** and **Dhéliat**. The lighter side of AI was represented by memes and humor from **Pieter Levels** and **AISafetyMemes**. This summary reflects the fast-evolving AI landscape blending technical innovation, corporate strategy, ethics, and community culture.</description><pubDate>Thu, 29 Feb 2024 00:50:17 GMT</pubDate><category>google</category><category>openai</category><category>apple</category><category>stripe</category><category>mistral-large</category><category>google-gemini</category><category>margaret-mitchell</category><category>john-carmack</category><category>guillaume-lample</category><category>sundar-pichai</category><category>delip-rao</category><category>santiago-l-valdarrama</category><category>alex-wang</category><category>yann-lecun</category><category>pieter-levels</category><category>francois-chollet</category><category>dheliat</category><category>ai-ethics</category><category>multilinguality</category><category>on-device-ai</category><category>convolutional-neural-networks</category><category>synthetic-data</category><category>financial-transaction-systems</category><category>corporate-culture</category><category>humor</category></item><item><title>Welcome Interconnects and OpenRouter</title><link>https://news.smol.ai/issues/24-02-27-ainews-welcome-interconnects-and-openrouter/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-27-ainews-welcome-interconnects-and-openrouter/</guid><description>**Discord communities** analyzed **22 guilds**, **349 channels**, and **12885 messages** revealing active discussions on **model comparisons and optimizations** involving **Mistral AI**, **Miqu**, and **GGUF quantized models**. Highlights include comparing **Mistral Large** with **GPT-4**, focusing on cost-effectiveness and performance, and exploring quantization techniques like **GPTQ** and **QLORA** to reduce VRAM usage. Advanced applications such as **role-playing**, **story-writing**, **code clarity**, and **AI-assisted decompilation** were emphasized, alongside development of tools like an **asynchronous summarization script** for **Mistral 7b**. The intersection of **quantum computing** and AI was discussed, including DARPA-funded projects and **encoder-based diffusion techniques** for image processing. Community efforts featured new Spanish LLM announcements, hardware experimentation, and open-source initiatives, with platforms like **Perplexity AI** and **LlamaIndex** noted for innovation and integration. Speculation about **Mistral AI**&apos;s open-source commitment and tools like **R2R** for rapid RAG deployment highlighted collaborative spirit.</description><pubDate>Tue, 27 Feb 2024 20:03:47 GMT</pubDate><category>mistral-ai</category><category>openai</category><category>perplexity-ai</category><category>llamaindex</category><category>qwen</category><category>langchain</category><category>mistral-large</category><category>miqu</category><category>mixtral</category><category>gpt-4</category><category>mistral-7b</category><category>nathan-lambert</category><category>alex-atallah</category><category>model-comparison</category><category>model-optimization</category><category>quantization</category><category>role-playing</category><category>story-writing</category><category>code-clarity</category><category>ai-assisted-decompilation</category><category>asynchronous-processing</category><category>quantum-computing</category><category>encoder-based-diffusion</category><category>open-source</category><category>hardware-experimentation</category><category>rag-systems</category></item><item><title>Mistral Large disappoints</title><link>https://news.smol.ai/issues/24-02-26-ainews-mistral-large-disappoints/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-26-ainews-mistral-large-disappoints/</guid><description>**Mistral** announced **Mistral Large**, a new language model achieving **81.2% accuracy on MMLU**, trailing **GPT-4 Turbo** by about 5 percentage points on benchmarks. The community reception has been mixed, with skepticism about open sourcing and claims that **Mistral Small** outperforms the open **Mixtral 8x7B**. Discussions in the **TheBloke** Discord highlighted performance and cost-efficiency comparisons between **Mistral Large** and **GPT-4 Turbo**, technical challenges with **DeepSpeed** and **DPOTrainer** for training, advances in AI deception for roleplay characters using **DreamGen Opus V1**, and complexities in model merging using linear interpolation and PEFT methods. Enthusiasm for AI-assisted decompilation was also expressed, emphasizing the use of open-source projects for training data.</description><pubDate>Mon, 26 Feb 2024 21:59:34 GMT</pubDate><category>mistral-ai</category><category>openai</category><category>hugging-face</category><category>mistral-large</category><category>mistral-small</category><category>mixtral-8x7b</category><category>gpt-4-turbo</category><category>dreamgen-opus-v1</category><category>timotheeee1</category><category>cogbuji</category><category>plasmator</category><category>jsarnecki</category><category>maldevide</category><category>spottyluck</category><category>mrjackspade</category><category>benchmarking</category><category>model-merging</category><category>fine-tuning</category><category>reinforcement-learning</category><category>model-training</category><category>tokenization</category><category>model-optimization</category><category>ai-assisted-decompilation</category><category>performance</category><category>cost-efficiency</category><category>deception</category><category>roleplay</category><category>deep-speed</category><category>dpo</category></item><item><title>One Year of Latent Space</title><link>https://news.smol.ai/issues/24-02-23-ainews-one-year-of-latent-space/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-23-ainews-one-year-of-latent-space/</guid><description>**Latent Space** podcast celebrated its first anniversary, reaching #1 in AI Engineering podcasts and 1 million unique readers on Substack. The **Gemini 1.5** image generator by **Google DeepMind** sparked controversy over bias and inaccurate representation, leading to community debates on AI ethics. Discussions in **TheBloke** and **LM Studio** Discords highlighted AI&apos;s growing role in creative industries, especially game development and text-to-3D tools. Fine-tuning and performance optimization of models like **Gemma 7B** and **Mistral-next** were explored in **Nous Research AI** and **Mistral** Discords, with shared solutions including learning rates and open-source tools. Emerging trends in AI hardware and application development were discussed in **CUDA MODE** and **LangChain AI** Discords, including critiques of **Nvidia&apos;s CUDA** by **Jim Keller** and advancements in reducing AI hallucinations hinted by **Richard Socher**.</description><pubDate>Sat, 24 Feb 2024 01:05:00 GMT</pubDate><category>google-deepmind</category><category>nous-research</category><category>mistral-ai</category><category>hugging-face</category><category>nvidia</category><category>langchain</category><category>jetbrains</category><category>gemini-1.5</category><category>gemma-7b</category><category>mistral-next</category><category>opus-v1</category><category>orca-2-13b</category><category>nous-hermes-2-dpo-7b</category><category>jim-keller</category><category>richard-socher</category><category>ai-ethics</category><category>bias-mitigation</category><category>fine-tuning</category><category>performance-optimization</category><category>model-merging</category><category>knowledge-transfer</category><category>text-to-3d</category><category>ai-hallucination</category><category>hardware-optimization</category><category>application-development</category><category>vulnerability-research</category></item><item><title>Ring Attention for &gt;1M Context</title><link>https://news.smol.ai/issues/24-02-22-ainews-ring-attention-for-greater1m-context/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-22-ainews-ring-attention-for-greater1m-context/</guid><description>**Google Gemini Pro** has sparked renewed interest in long context capabilities. The CUDA MODE Discord is actively working on implementing the **RingAttention** paper by Liu, Zaharia, and Abbeel, including extensions from the World Model RingAttention paper, with available PyTorch and CUDA implementations. TheBloke Discord discussed various topics including **LLM guessing game evaluation**, chatbot UX comparisons between **Nvidia&apos;s Chat with RTX** and **Polymind**, challenges in **retrieval-augmented generation (RAG)** integration, VRAM optimization, fine-tuning for character roleplay using **Dynamic Prompt Optimization (DPO)**, and model choices like **deepseek-coder-6.7B-instruct**. There was also discussion on ML workflows on Mac Studio, with preferences for **llama.cpp** over **ollama**, and scaling inference cost-effectively using GPUs like the **4090** on Runpod. LM Studio users face manual update requirements for version **0.2.16**, which includes support for **Gemma models** and bug fixes, especially for MacOS. The Gemma 7B model has had performance issues, while Gemma 2B received positive feedback.</description><pubDate>Fri, 23 Feb 2024 00:51:56 GMT</pubDate><category>google</category><category>cuda-mode</category><category>nvidia</category><category>polymind</category><category>deepseek</category><category>ollama</category><category>runpod</category><category>lmstudio</category><category>gemini-pro</category><category>gemma-7b</category><category>gemma-2b</category><category>deepseek-coder-6.7b-instruct</category><category>llama-cpp</category><category>liu</category><category>zaharia</category><category>abbeel</category><category>long-context</category><category>ringattention</category><category>pytorch</category><category>cuda</category><category>llm-guessing-game</category><category>chatbots</category><category>retrieval-augmented-generation</category><category>vram-optimization</category><category>fine-tuning</category><category>dynamic-prompt-optimization</category><category>ml-workflows</category><category>gpu-scaling</category><category>model-updates</category></item><item><title>Google AI: Win some (Gemma, 1.5 Pro), Lose some (Image gen)</title><link>https://news.smol.ai/issues/24-02-21-ainews-google-ai-win-some-gemma-15-pro-lose-some-image-gen/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-21-ainews-google-ai-win-some-gemma-15-pro-lose-some-image-gen/</guid><description>**Google&apos;s Gemma open models** (2-7B parameters) outperform **Llama 2** and **Mistral** in benchmarks but face criticism for an unusual license and poor image generation quality, which Google partially acknowledges. The upcoming **Gemini Pro 1.5** model features a 1 million token context window, excelling in video understanding and needle-in-haystack tasks. Discord communities like **TheBloke** and **LM Studio** discuss mixed reception of Gemma models, anticipation for **Llama 3** release, challenges in dataset editing, and hardware considerations such as **NVIDIA GeForce RTX 3090** and **RTX 4090** GPUs. LM Studio users report issues with version 0.2.15 Beta and ongoing integration of Gemma models, with resources shared on **Hugging Face**.</description><pubDate>Thu, 22 Feb 2024 02:21:19 GMT</pubDate><category>google</category><category>hugging-face</category><category>nvidia</category><category>gemma-2b</category><category>gemma-7b</category><category>gemma</category><category>gemini-pro-1.5</category><category>llama-2</category><category>llama-3</category><category>mistral</category><category>benchmarking</category><category>license-policies</category><category>image-generation</category><category>video-understanding</category><category>long-context</category><category>dataset-editing</category><category>model-integration</category><category>gpu-hardware</category><category>bug-fixes</category><category>quantization</category></item><item><title>Karpathy emerges from stealth?</title><link>https://news.smol.ai/issues/24-02-20-ainews-karpathy-emerges-from-stealth/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-20-ainews-karpathy-emerges-from-stealth/</guid><description>**Andrej Karpathy** released a comprehensive 2-hour tutorial on **tokenization**, detailing techniques up to **GPT-4**&apos;s tokenizer and noting the complexity of **Llama 2** tokenization with SentencePiece. Discussions in AI Discord communities covered **model optimization and efficiency**, focusing on **quantization** of models like **Mistral 7B** and **Zephyr-7B** to reduce memory usage for consumer GPUs, including Intel&apos;s new weight-only quantization algorithm. Efforts to improve computational efficiency included selective augmentation reducing costs by 57.76% and memory token usage versus kNN for Transformers. Challenges in hardware compatibility and software issues were shared, alongside fine-tuning techniques such as LoRA and model merging. Innovative applications of LLMs in retrieval-augmented generation (RAG), multi-model learning, and meta-reasoning were explored. The community emphasized dataset sharing, open-source releases like SDXL VAE encoded datasets and Audiogen AI codecs, and ethical AI use with censorship and guardrails. Collaboration and resource sharing remain strong in these AI communities.</description><pubDate>Wed, 21 Feb 2024 01:54:38 GMT</pubDate><category>intel</category><category>mistral-ai</category><category>audiogen</category><category>thebloke</category><category>mistral-7b</category><category>mixtral-8x7b</category><category>zephyr-7b</category><category>gpt-4</category><category>llama-2</category><category>andrej-karpathy</category><category>tokenization</category><category>quantization</category><category>model-optimization</category><category>fine-tuning</category><category>model-merging</category><category>computational-efficiency</category><category>memory-optimization</category><category>retrieval-augmented-generation</category><category>multi-model-learning</category><category>meta-reasoning</category><category>dataset-sharing</category><category>open-source</category><category>ethical-ai</category><category>community-collaboration</category></item><item><title>Companies liable for AI hallucination is Good Actually for AI Engineers</title><link>https://news.smol.ai/issues/24-02-19-ainews-companies-liable-for-ai-hallucination-is-good-actually-for-ai-engineers/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-19-ainews-companies-liable-for-ai-hallucination-is-good-actually-for-ai-engineers/</guid><description>**Air Canada** faced a legal ruling requiring it to honor refund policies communicated by its AI chatbot, setting a precedent for corporate liability in AI engineering accuracy. The tribunal ordered a refund of **$650.88 CAD** plus damages after the chatbot misled a customer about bereavement travel refunds. Meanwhile, AI community discussions highlighted innovations in **quantization techniques** for GPU inference, **Retrieval-Augmented Generation (RAG)** and fine-tuning of LLMs, and **CUDA** optimizations for PyTorch models. New prototype models like **Mistral-Next** and the **Large World Model (LWM)** were introduced, showcasing advances in handling large text contexts and video generation with models like **Sora**. Ethical and legal implications of AI autonomy were debated alongside challenges in dataset management. Community-driven projects such as the open-source TypeScript agent framework **bazed-af** emphasize collaborative AI development. Additionally, benchmarks like **BABILong** for up to **10M context evaluation** and tools from **karpathy** were noted.</description><pubDate>Tue, 20 Feb 2024 00:05:26 GMT</pubDate><category>air-canada</category><category>huggingface</category><category>mistral-ai</category><category>mistral-next</category><category>large-world-model</category><category>sora</category><category>babilong</category><category>andrej-karpathy</category><category>quantization</category><category>retrieval-augmented-generation</category><category>fine-tuning</category><category>cuda-optimization</category><category>video-generation</category><category>ai-ethics</category><category>dataset-management</category><category>open-source</category><category>community-driven-development</category></item><item><title>Sora pushes SOTA</title><link>https://news.smol.ai/issues/24-02-16-ainews-sora-pushes-sota/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-16-ainews-sora-pushes-sota/</guid><description>**Discord communities** analyzed over **20 guilds**, **312 channels**, and **10550 messages** reveal intense discussions on AI developments. Key highlights include the **Dungeon Master AI assistant** for Dungeons and Dragons using models like **H20 GPT**, GPU power supply debates involving **3090** and **3060 GPUs**, and excitement around **Google&apos;s Gemini 1.5** with its **1 million token context window** and **OpenAI&apos;s Sora** model. Challenges with **large world models (LWM)** multimodality, **GPT-assisted coding**, and **role-play model optimization** with **Yi models** and **Mixtral Instruct** were discussed. Technical issues like **model merging errors** with **MistralCasualML**, fine-tuning scripts like **AutoFineTune**, and cross-language engineering via **JSPyBridge** were also prominent. NVIDIA&apos;s **Chat with RTX** feature leveraging **retrieval-augmented generation (RAG)** on 30+ series GPUs was compared to LMStudio&apos;s support for **Mistral 7b** and **Llama 13b** models. The community is cautiously optimistic about these frontier models&apos; applications in media and coding.</description><pubDate>Fri, 16 Feb 2024 11:15:03 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>nvidia</category><category>mistral-ai</category><category>h2oai</category><category>gemini-1.5</category><category>sora</category><category>h20-gpt</category><category>mistral-7b</category><category>llama-13b</category><category>mistralcasualml</category><category>mixtral-instruct</category><category>yi-models</category><category>multimodality</category><category>gpu-power-management</category><category>long-context</category><category>model-merging</category><category>fine-tuning</category><category>retrieval-augmented-generation</category><category>role-play-model-optimization</category><category>cross-language-integration</category><category>training-loss</category><category>synthetic-data-generation</category><category>coding-support</category></item><item><title>AI gets Memory</title><link>https://news.smol.ai/issues/24-02-14-ainews-ai-gets-memory/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-14-ainews-ai-gets-memory/</guid><description>**AI Discords** analysis covered **20 guilds**, **312 channels**, and **6901 messages**. The report highlights the divergence of RAG style operations for context and memory, with implementations like **MemGPT** rolling out in **ChatGPT** and **LangChain**. The **TheBloke Discord** discussed **open-source large language models** such as the **Large World Model** with contexts up to **1 million tokens**, and the **Cohere aya model** supporting **101 languages**. Roleplay-focused models like **MiquMaid-v2-70B** were noted for performance improvements with enhanced hardware. Finetuning techniques like **Sequential Fine-Tuning (SFT)** and **Direct Preference Optimization (DPO)** were explained, with tools like **Unsloth AI&apos;s apply_chat_template** preferred over Alpaca. Integration of JavaScript and Python via **JSPyBridge** in the **SillyTavern** project was also discussed. Training challenges with **Mixtral 8x7b qlora** versus **Mistral 7b** were noted. The **LM Studio Discord** focused on hardware limitations affecting large model loading, medical LLMs like **medAlpaca**, and hardware discussions around GPU upgrades and overclocking. Anticipation for **IQ3_XSS** 1.5 bit quantization support in LM Studio was expressed.</description><pubDate>Thu, 15 Feb 2024 00:47:59 GMT</pubDate><category>openai</category><category>langchain</category><category>thebloke</category><category>cohere</category><category>unsloth-ai</category><category>mistral-ai</category><category>microsoft</category><category>miqumaid-v2-70b</category><category>mixtral-8x7b-qlora</category><category>mistral-7b</category><category>phi-2</category><category>medalpaca</category><category>aya</category><category>joanne-jang</category><category>rag</category><category>memory-modeling</category><category>context-windows</category><category>open-source</category><category>finetuning</category><category>sequential-fine-tuning</category><category>direct-preference-optimization</category><category>rlhf</category><category>ppo</category><category>javascript-python-integration</category><category>hardware-optimization</category><category>gpu-overclocking</category><category>quantization</category><category>model-training</category><category>large-context</category><category>multilinguality</category></item><item><title>The Dissection of Smaug (72B)</title><link>https://news.smol.ai/issues/24-02-12-ainews-the-dissection-of-smaug-72b/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-12-ainews-the-dissection-of-smaug-72b/</guid><description>**Abacus AI** launched **Smaug 72B**, a large finetune of **Qwen 1.0**, which remains unchallenged on the **Hugging Face Open LLM Leaderboard** despite skepticism from **Nous Research**. **LAION** introduced a local voice assistant model named **Bud-E** with a notable demo. The **TheBloke Discord** community discussed model performance trade-offs between large models like **GPT-4** and smaller quantized models, fine-tuning techniques using datasets like **WizardLM_evol_instruct_V2_196k** and **OpenHermes-2.5**, and challenges in web UI development and model merging involving **Mistral-7b** and **MiquMaid**. The **LM Studio Discord** highlighted issues with model conversion from PyTorch to gguf, hardware setups involving **Intel Xeon CPUs** and **Nvidia P40 GPUs**, privacy concerns, and limitations in image generation and web UI availability.</description><pubDate>Tue, 13 Feb 2024 01:40:29 GMT</pubDate><category>abacus-ai</category><category>hugging-face</category><category>nous-research</category><category>laion</category><category>thebloke</category><category>lm-studio</category><category>intel</category><category>nvidia</category><category>elevenlabs</category><category>smaug-72b</category><category>qwen-1.0</category><category>qwen-1.5</category><category>gpt-4</category><category>mistral-7b</category><category>miqumaid</category><category>wizardlm_evol_instruct_v2_196k</category><category>openhermes-2.5</category><category>bindureddy</category><category>fine-tuning</category><category>model-merging</category><category>quantization</category><category>web-ui</category><category>model-conversion</category><category>hardware-setup</category><category>privacy</category><category>image-generation</category><category>optical-character-recognition</category><category>prompt-engineering</category></item><item><title>Gemini Ultra is out, to mixed reviews</title><link>https://news.smol.ai/issues/24-02-08-ainews-gemini-ultra-is-out-to-mixed-reviews/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-08-ainews-gemini-ultra-is-out-to-mixed-reviews/</guid><description>**Google** released **Gemini Ultra** as a paid tier for &quot;Gemini Advanced with Ultra 1.0&quot; following the discontinuation of Bard. Reviews noted it is &quot;slightly faster/better than ChatGPT&quot; but with reasoning gaps. The **Steam Deck** was highlighted as a surprising AI workstation capable of running models like Solar 10.7B. Discussions in AI communities covered topics such as multi-GPU support for OSS Unsloth, training data contamination from OpenAI outputs, ethical concerns over model merging, and new alignment techniques like Listwise Preference Optimization (LiPO). The **Mojo** programming language was praised for high-performance computing. In research, the **Subformer** model uses sandwich-style parameter sharing and SAFE for efficiency, and **BiLLM** introduced 1-bit post-training quantization to reduce resource use. The **OpenHermes** dataset viewer tool was launched, and GPU scheduling with Slurm was discussed. Fine-tuning challenges for models like **OpenHermes-2.5-Mistral-7B** and VRAM requirements were also topics of interest.</description><pubDate>Fri, 09 Feb 2024 05:58:08 GMT</pubDate><category>google</category><category>openai</category><category>mistral-ai</category><category>hugging-face</category><category>gemini-ultra</category><category>gemini-advanced</category><category>solar-10.7b</category><category>openhermes-2.5-mistral-7b</category><category>subformer</category><category>billm</category><category>multi-gpu-support</category><category>training-data-contamination</category><category>model-merging</category><category>model-alignment</category><category>listwise-preference-optimization</category><category>high-performance-computing</category><category>parameter-sharing</category><category>post-training-quantization</category><category>dataset-viewer</category><category>gpu-scheduling</category><category>fine-tuning</category><category>vram-optimization</category></item><item><title>MetaVoice &amp; RIP Bard</title><link>https://news.smol.ai/issues/24-02-07-ainews-metavoice-and-rip-bard/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-07-ainews-metavoice-and-rip-bard/</guid><description>**Coqui**, a TTS startup that recently shut down, inspired a new **TTS model** supporting voice cloning and longform synthesis from a small startup called **MetaVoice**. **Google** discontinued the **Bard** brand in favor of **Gemini**. On **TheBloke Discord**, discussions focused on AI training with models like **Mixtral**, **Nous Mixtral DPO**, and **Miqu 70B**, comparing them to **OpenAI&apos;s GPT** models, and debated prompt engineering, lorebooks, and removing safety features via **LoRA fine-tuning** on models such as **Llama2 70B instruct**. Technical topics included transformer layer offloading limitations and adapting **LLaMa 2** for Apple Silicon. On **OpenAI Discord**, **DALL-E** images now include **C2PA metadata** for content authenticity, sparking debates on AI censorship, metadata manipulation, and open-source AI models versus commercial giants like **GPT-4**. Users discussed GPT-4 usability, limitations, and practical applications.</description><pubDate>Wed, 07 Feb 2024 22:41:50 GMT</pubDate><category>coqui</category><category>metavoice</category><category>google</category><category>openai</category><category>thebloke</category><category>mixtral</category><category>nous-mixtral-dpo</category><category>miqu-70b</category><category>gpt-4</category><category>llama-2-70b-instruct</category><category>llama-2</category><category>llama-2-70b</category><category>llama-2-70b-instruct</category><category>text-to-speech</category><category>voice-cloning</category><category>longform-synthesis</category><category>prompt-engineering</category><category>direct-preference-optimization</category><category>lora-fine-tuning</category><category>transformers</category><category>gpu-acceleration</category><category>apple-silicon</category><category>content-authenticity</category><category>metadata</category><category>ai-censorship</category><category>open-source-ai</category><category>model-comparison</category><category>usability</category><category>model-limitations</category></item><item><title>Qwen 1.5 Released</title><link>https://news.smol.ai/issues/24-02-06-ainews-qwen-15-released/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-06-ainews-qwen-15-released/</guid><description>**Chinese AI models Yi, Deepseek, and Qwen** are gaining attention for strong performance, with **Qwen 1.5** offering up to **32k token context** and compatibility with Hugging Face transformers and quantized models. The **TheBloke Discord** discussed topics like quantization of a **70B LLM**, the introduction of the **Sparse MoE model Sparsetral** based on **Mistral**, debates on merging vs fine-tuning, and Direct Preference Optimization (DPO) for character generation. The **Nous Research AI Discord** covered challenges in Japanese Kanji generation, AI scams on social media, and Meta&apos;s VR headset prototypes showcased at **SIGGRAPH 2023**. Discussions also included fine-tuning frozen networks and new models like **bagel-7b-v0.4**, **DeepSeek-Math-7b-instruct**, and **Sparsetral-16x7B-v2**.</description><pubDate>Tue, 06 Feb 2024 23:40:32 GMT</pubDate><category>deepseek</category><category>qwen</category><category>mistral-ai</category><category>hugging-face</category><category>meta-ai-fair</category><category>qwen-1.5</category><category>mistral-7b</category><category>sparsetral-16x7b-v2</category><category>bagel-7b-v0.4</category><category>deepseek-math-7b-instruct</category><category>quantization</category><category>token-context</category><category>multilinguality</category><category>retrieval-augmented-generation</category><category>agent-planning</category><category>code-generation</category><category>sparse-moe</category><category>model-merging</category><category>fine-tuning</category><category>direct-preference-optimization</category><category>character-generation</category><category>ascii-art</category><category>kanji-generation</category><category>vr</category><category>retinal-resolution</category><category>light-field-passthrough</category><category>frozen-networks</category><category>normalization-layers</category></item><item><title>Less Lazy AI</title><link>https://news.smol.ai/issues/24-02-05-ainews-less-lazy-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-05-ainews-less-lazy-ai/</guid><description>The AI Discord summaries for early 2024 cover various community discussions and developments. Highlights include **20** guilds, **308** channels, and **10449** messages analyzed, saving an estimated **780 minutes** of reading time. Key topics include **Polymind Plugin Puzzle** integrating PubMed API, roleplay with **HamSter v0.2**, VRAM challenges in **Axolotl** training, fine-tuning tips for **FLAN-T5**, and innovative **model merging** strategies. The **Nous Research AI** community discussed GPT-4&apos;s lyricism issues, quantization techniques using `llama.cpp`, **frankenmerging** with models like **miqu-1-120b-GGUF**, anticipation for **Qwen2**, and tools like `text-generation-webui` and **ExLlamaV2**. The **LM Studio** community reported a bug where the app continues running after UI closure, with a workaround to forcibly terminate the process. These discussions reflect ongoing challenges and innovations in AI model training, deployment, and interaction.</description><pubDate>Tue, 06 Feb 2024 00:50:28 GMT</pubDate><category>openai</category><category>hugging-face</category><category>nous-research</category><category>h2oai</category><category>apple</category><category>hamster-v0.2</category><category>flan-t5</category><category>miqu-1-120b-gguf</category><category>qwen2</category><category>axolotl</category><category>philschmid</category><category>model-merging</category><category>fine-tuning</category><category>quantization</category><category>vram-optimization</category><category>plugin-development</category><category>chatbot-memory</category><category>model-training</category><category>bug-reporting</category><category>api-compatibility</category></item><item><title>The Core Skills of AI Engineering</title><link>https://news.smol.ai/issues/24-02-03-ainews-the-core-skills-of-ai-engineering/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-03-ainews-the-core-skills-of-ai-engineering/</guid><description>**AI Discords for 2/2/2024** analyzed **21 guilds**, **312 channels**, and **4782 messages** saving an estimated **382 minutes** of reading time. Discussions included **Eugene Yan** initiating a deep dive into **AI engineering** challenges, highlighting overlaps between software engineering and data science skills. The **TheBloke Discord** featured talks on **MiquMaid**, **OLMo** (an open-source 65B LLM by **AI2** under Apache 2.0), **Aphrodite** model batching, **AWQ** quantization, and **LoRA** fine-tuning techniques like **QLoRA** and **LoftQ**. The **LAION Discord** discussed **SSD-1B** distillation issues, data quality optimization with captioning datasets like **BLIP**, **COCO**, and **LLaVA**, and tokenization strategies for prompt adherence in image generation. Other topics included AI security with watermarking, superconductors and carbon nanotubes for hardware, and deployment of LLMs via **Hugging Face** tools.</description><pubDate>Sun, 04 Feb 2024 00:54:29 GMT</pubDate><category>ai2</category><category>hugging-face</category><category>miqumaid</category><category>olmo</category><category>aphrodite</category><category>awq</category><category>exl2</category><category>mistral-medium</category><category>internlm</category><category>ssd-1b</category><category>lora</category><category>qlora</category><category>loftq</category><category>eugene-yan</category><category>ai-engineering</category><category>quantization</category><category>fine-tuning</category><category>open-source</category><category>model-deployment</category><category>data-quality</category><category>tokenization</category><category>prompt-adherence</category><category>distillation</category><category>ai-security</category><category>batching</category><category>hardware</category><category>role-playing</category></item><item><title>AI2 releases OLMo - the 4th open-everything LLM</title><link>https://news.smol.ai/issues/24-02-02-ainews-ai2-releases-olmo-the-4th-open-everything-llm/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-02-ainews-ai2-releases-olmo-the-4th-open-everything-llm/</guid><description>**AI2** is gaining attention in 2024 with its new **OLMo** models, including 1B and 7B sizes and a 65B model forthcoming, emphasizing open and reproducible research akin to **Pythia**. The **Miqu-70B** model, especially the Mistral Medium variant, is praised for self-correction and speed optimizations. Discussions in **TheBloke** Discord covered programming language preferences, VRAM constraints for large models, and fine-tuning experiments with **Distilbert-base-uncased**. The **Mistral** Discord highlighted challenges in the **GPU shortage** affecting semiconductor production involving **TSMC**, **ASML**, and **Zeiss**, debates on open-source versus proprietary models, and fine-tuning techniques including **LoRA** for low-resource languages. Community insights also touched on embedding chunking strategies and JSON output improvements.</description><pubDate>Sat, 03 Feb 2024 03:35:10 GMT</pubDate><category>ai2</category><category>allenai</category><category>mistral-ai</category><category>tsmc</category><category>asml</category><category>zeiss</category><category>olmo-1b</category><category>olmo-7b</category><category>olmo-65b</category><category>miqu-70b</category><category>mistral-medium</category><category>distilbert-base-uncased</category><category>nathan-lambert</category><category>lhc1921</category><category>mrdragonfox</category><category>yashkhare_</category><category>gbourdin</category><category>fine-tuning</category><category>gpu-shortage</category><category>embedding-chunking</category><category>json-generation</category><category>model-optimization</category><category>reproducible-research</category><category>self-correction</category><category>vram-constraints</category><category>programming-languages</category></item><item><title>Trust in GPTs at all time low</title><link>https://news.smol.ai/issues/24-02-01-ainews-trust-in-gpts-at-all-time-low/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-02-01-ainews-trust-in-gpts-at-all-time-low/</guid><description>**Discord communities** were analyzed with **21 guilds**, **312 channels**, and **8530 messages** reviewed, saving an estimated **628 minutes** of reading time. Discussions highlighted challenges with **GPTs** and the **GPT store**, including critiques of the **knowledge files capability** and context management issues. The **CUDA MODE Discord** was introduced for CUDA coding support. Key conversations in the **TheBloke Discord** covered **Xeon** GPU server cost-effectiveness, **Llama3** and **Mistral Medium** model comparisons, **LLaVA-1.6**&apos;s visual reasoning and OCR capabilities, and the leaked **Miqu** 70B model. Technical topics included fine-tuning **TinyLlama** and **MiquMaid+Euryale** models, and model merging with examples like **Harmony-4x7B-bf16** and **Smaug-34B-v0.1**. The **Nous Research AI Discord** discussed style influence in LLMs, quantization issues, **Bittensor** incentives for AI model improvements, and the identification of **MIQU** as **Mistral Medium**. The release of the **Open Hermes 2.5 dataset** on **Hugging Face** was also announced. *&quot;Discussions pointed towards the need for better context management in GPTs, contrasting with OpenAI&apos;s no-code approach.&quot;*</description><pubDate>Fri, 02 Feb 2024 03:25:24 GMT</pubDate><category>openai</category><category>hugging-face</category><category>mistral-ai</category><category>nous-research</category><category>bittensor</category><category>llama-3</category><category>mistral-medium</category><category>llava-1.6</category><category>miquella-120b-gguf</category><category>tinymodels</category><category>miqumaid</category><category>harmony-4x7b-bf16</category><category>smaug-34b-v0.1</category><category>nick-dobos</category><category>manojbh</category><category>teknium</category><category>arthurmensch</category><category>context-management</category><category>fine-tuning</category><category>model-merging</category><category>quantization</category><category>gpu-servers</category><category>visual-reasoning</category><category>ocr</category><category>dataset-release</category><category>incentive-structures</category></item><item><title>Miqu confirmed to be an early Mistral-medium checkpoint</title><link>https://news.smol.ai/issues/24-01-31-ainews-miqu-confirmed-to-be-an-early-mistral-medium-checkpoint/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-31-ainews-miqu-confirmed-to-be-an-early-mistral-medium-checkpoint/</guid><description>**Miqu**, an open access model, scores **74 on MMLU** and **84.5 on EQ-Bench**, sparking debates about its performance compared to **Mistral Medium**. The **CEO of Mistral** confirmed these results. Discussions in the **TheBloke Discord** highlight **Miqu&apos;s** superiority in instruction-following and sampling methods like dynatemp and min-p. Developers also explore browser preferences and Discord UI themes. Role-playing with models like **BagelMistery Tour v2** and **Psyfighter v2** is popular, alongside technical talks on **fp16 quantization** of **Miqu-1-70b**. Training and fine-tuning tips for models like **Unsloth** and **Mistral 7B** are shared. In the **Nous Research AI Discord**, the **Activation Beacon** method is discussed for extending LLM context length from 4K to 400K tokens. **SQLCoder-70B**, fine-tuned on **CodeLlama-70B**, leads in text-to-SQL generation and is available on Hugging Face. The **Miqu model** also impresses with an **83.5 EQ-Bench score**, fueling speculation about its capabilities.</description><pubDate>Wed, 31 Jan 2024 23:15:13 GMT</pubDate><category>mistral-ai</category><category>hugging-face</category><category>nous-research</category><category>aiatmeta</category><category>miqu-1-70b</category><category>mistral-medium</category><category>llama-2-70b-chat</category><category>mixtral</category><category>sqlcoder-70b</category><category>codellama-70b</category><category>bagelmistery-tour-v2</category><category>psyfighter-v2</category><category>intrstllrninja</category><category>instruction-following</category><category>sampling-methods</category><category>fp16-quantization</category><category>fine-tuning</category><category>model-training</category><category>context-length</category><category>text-to-sql</category><category>model-performance</category><category>model-optimization</category></item><item><title>CodeLLama 70B beats GPT4 on HumanEval</title><link>https://news.smol.ai/issues/24-01-30-ainews-codellama-70b-beats-gpt4-on-humaneval/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-30-ainews-codellama-70b-beats-gpt4-on-humaneval/</guid><description>**Meta AI** surprised the community with the release of **CodeLlama**, an open-source model now available on platforms like **Ollama** and **MLX** for local use. The **Miqu model** sparked debate over its origins, possibly linked to **Mistral Medium** or a fine-tuned **Llama-2-70b**, alongside discussions on **AI ethics** and alignment risks. The **Aphrodite engine** showed strong performance on **A6000 GPUs** with specific configurations. Role-playing AI models such as **Mixtral** and **Flatdolphinmaid** faced challenges with repetitiveness, while **Noromaid** and **Rpcal** performed better, with **ChatML** and **DPO** recommended for improved responses. Learning resources like fast.ai&apos;s course were highlighted for ML/DL beginners, and fine-tuning techniques with optimizers like *Paged 8bit lion* and *adafactor* were discussed. 

At **Nous Research AI**, the **Activation Beacon** project introduced a method for unlimited context length in LLMs using &quot;global state&quot; tokens, potentially transforming retrieval-augmented models. The **Eagle-7B** model, based on **RWKV-v5**, outperformed **Mistral** in benchmarks with efficiency and multilingual capabilities. **OpenHermes2.5** was recommended for consumer hardware due to its quantization methods. Multimodal and domain-specific models like **IMP v1-3b**, **Bakllava**, **Moondream**, and **Qwen-vl** were explored for classification and vision-language tasks. The community emphasized centralizing AI resources for collaborative research.</description><pubDate>Tue, 30 Jan 2024 21:10:01 GMT</pubDate><category>meta-ai-fair</category><category>ollama</category><category>nous-research</category><category>mistral-ai</category><category>hugging-face</category><category>codellama</category><category>miqu</category><category>mistral-medium</category><category>llama-2-70b</category><category>aphrodite-engine</category><category>mixtral</category><category>flatdolphinmaid</category><category>noromaid</category><category>rpcal</category><category>chatml</category><category>mistral-7b</category><category>activation-beacon</category><category>eagle-7b</category><category>rwkv-v5</category><category>openhermes2.5</category><category>nous-hermes-2-mixtral-8x7b-dpo</category><category>imp-v1-3b</category><category>bakllava</category><category>moondream</category><category>qwen-vl</category><category>ai-ethics</category><category>alignment</category><category>gpu-optimization</category><category>direct-prompt-optimization</category><category>fine-tuning</category><category>cuda-programming</category><category>optimizer-technology</category><category>quantization</category><category>multimodality</category><category>context-length</category><category>dense-retrieval</category><category>retrieval-augmented-generation</category><category>multilinguality</category><category>model-performance</category><category>open-source</category><category>code-generation</category><category>classification</category><category>vision</category></item><item><title>RWKV &quot;Eagle&quot; v5: Your move, Mamba</title><link>https://news.smol.ai/issues/24-01-29-ainews-rwkv-eagle-v5-your-move-mamba/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-29-ainews-rwkv-eagle-v5-your-move-mamba/</guid><description>**RWKV v5 Eagle** was released with better-than-**mistral-7b** evaluation results, trading some English performance for multilingual capabilities. The mysterious **miqu-1-70b** model sparked debate about its origins, possibly a leak or distillation of **Mistral Medium** or a fine-tuned **Llama 2**. Discussions highlighted fine-tuning techniques, including the effectiveness of **1,000 high-quality prompts** over larger mixed-quality datasets, and tools like **Deepspeed**, **Axolotl**, and **QLoRA**. The **Nous Research AI** community emphasized the impact of **Rotary Position Embedding (RoPE) theta settings** on LLM extrapolation, improving models like **Mistral Instruct v0.2**. Speed improvements in **Mistral Tuna** kernels reduced token processing costs, enhancing efficiency. The launch of **Eagle 7B** with 7.52B parameters showcased strong multilingual performance, surpassing other 7B class models.</description><pubDate>Tue, 30 Jan 2024 01:20:56 GMT</pubDate><category>eleutherai</category><category>mistral-ai</category><category>hugging-face</category><category>llamaindex</category><category>nous-research</category><category>rwkv</category><category>lmsys</category><category>rwkv-v5</category><category>mistral-7b</category><category>miqu-1-70b</category><category>mistral-medium</category><category>llama-2</category><category>mistral-instruct-v0.2</category><category>mistral-tuna</category><category>llama-2-13b</category><category>kunoichi-dpo-v2-7b</category><category>gpt-4</category><category>andrej-karpathy</category><category>fine-tuning</category><category>multilinguality</category><category>rotary-position-embedding</category><category>model-optimization</category><category>model-performance</category><category>quantization</category><category>speed-optimization</category><category>prompt-engineering</category><category>model-benchmarking</category><category>reinforcement-learning</category></item><item><title>GPT4Turbo A/B Test: gpt-4-0125-preview</title><link>https://news.smol.ai/issues/24-01-26-ainews-gpt4turbo-ab-test-gpt-4-0125-preview/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-26-ainews-gpt4turbo-ab-test-gpt-4-0125-preview/</guid><description>**OpenAI** released a new **GPT-4 Turbo** version in January 2024, prompting natural experiments in summarization and discussions on API performance and cost trade-offs. The **TheBloke** Discord highlighted **UnSloth&apos;s** upcoming limited multi-GPU support for Google Colab beginners, AI models like **Tiny Llama** and **Mistral** running on Nintendo Switch, and advanced model merging techniques such as DARE and SLERP. The **OpenAI** Discord noted issues with **GPT-4-1106-preview** processing delays, troubleshooting GPT model errors, and transcription challenges with **GPT-3.5** and **GPT-4 Turbo**. **Nous Research AI** focused on extending context windows, notably **LLaMA-2-7B-Chat** reaching **16,384** tokens, and fine-tuning alternatives like **SelfExtend**. Discussions also touched on chatbot persona creation, model configuration optimizations, and societal impacts of AI technology.</description><pubDate>Fri, 26 Jan 2024 22:48:31 GMT</pubDate><category>openai</category><category>thebloke</category><category>nous-research</category><category>hugging-face</category><category>gpt-4-turbo</category><category>gpt-4-1106-preview</category><category>gpt-3.5</category><category>llama-2-7b-chat</category><category>tiny-llama</category><category>mistral</category><category>multi-gpu-support</category><category>model-optimization</category><category>model-merging</category><category>fine-tuning</category><category>context-windows</category><category>chatbot-personas</category><category>api-performance</category><category>text-transcription</category><category>cost-considerations</category><category>model-troubleshooting</category></item><item><title>GPT4Turbo A/B Test: gpt-4-1106-preview</title><link>https://news.smol.ai/issues/24-01-26-ainews-gpt4turbo-ab-test-gpt-4-1106-preview/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-26-ainews-gpt4turbo-ab-test-gpt-4-1106-preview/</guid><description>**OpenAI** released a new **GPT-4 Turbo** version, prompting a natural experiment in summarization comparing the November 2023 and January 2024 versions. The **TheBloke** Discord discussed troubleshooting model loading errors with **OpenHermes-2.5-Mistral-7B-4.0bpw** and **exllamav2**, debates on **RHEL** in ML, dataset generation for understanding GPT flaws, and running LLMs like **Llama** and **Mistral** on consoles. **LangChain** fine-tuning challenges for **Llama2** were also noted. The **OpenAI** Discord highlighted **GPT-4** speed inconsistencies, API vs web performance, prompt engineering with **GPT-3.5** and **GPT-4 Turbo**, and **DALL-E** typo issues in image text. Discussions included NLP tools like *semantic-text-splitter* and collaboration concerns with **GPT-4 Vision** on **Azure**. The **Nous Research AI** Discord focused on extending context windows with **Mistral instruct v0.2**, **MistralLite**, and **LLaMA-2-7B-Chat** achieving 16,384 token context, plus alternatives like **SelfExtend** for context extension without fine-tuning. The societal impact of AI technology was also considered.</description><pubDate>Fri, 26 Jan 2024 22:07:42 GMT</pubDate><category>openai</category><category>huggingface</category><category>thebloke</category><category>nous-research</category><category>mistral-ai</category><category>langchain</category><category>microsoft</category><category>azure</category><category>gpt-4-turbo</category><category>gpt-4</category><category>gpt-3.5</category><category>openhermes-2.5-mistral-7b-4.0bpw</category><category>exllamav2</category><category>llama-2-7b-chat</category><category>mistral-instruct-v0.2</category><category>mistrallite</category><category>llama2</category><category>model-loading</category><category>rhel</category><category>dataset-generation</category><category>llm-on-consoles</category><category>fine-tuning</category><category>speed-optimization</category><category>api-performance</category><category>prompt-engineering</category><category>token-limits</category><category>memory-constraints</category><category>text-generation</category><category>nlp-tools</category><category>context-window-extension</category><category>sliding-windows</category><category>rope-theta</category><category>non-finetuning-context-extension</category><category>societal-impact</category></item><item><title>Adept Fuyu-Heavy: Multimodal model for Agents</title><link>https://news.smol.ai/issues/24-01-25-ainews-adept-fuyu-heavy-multimodal-model-for-agents/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-25-ainews-adept-fuyu-heavy-multimodal-model-for-agents/</guid><description>**Adept** launched **Fuyu-Heavy**, a multimodal model focused on UI understanding and visual QA, outperforming **Gemini Pro** on the MMMU benchmark. The model uses **DPO** (Direct Preference Optimization), gaining attention as a leading tuning method. The size of Fuyu-Heavy is undisclosed but estimated between **20B-170B** parameters, smaller than rumored frontier models like **Claude 2**, **GPT4V**, and **Gemini Ultra**. Meanwhile, **Mamba** was rejected at ICLR for quality concerns. In Discord discussions, **DeepSeek Coder 33B** was claimed to outperform **GPT-4** in coding tasks, and deployment strategies for large models like **Yi-34B-200K** and **Goliath-120B** were explored. Quantization debates highlighted mixed views on **Q8** and **EXL2 quants**. Fine-tuning and instruct-tuning of **Mistral 7B Instruct v0.2** were discussed, alongside insights on RMS optimization and heterogeneous AI architectures combining **Transformers** and **Selective SSM (Mamba)**. The potential of recurrent LLMs like **RWKV** and techniques like **Contrastive Preference Optimization (CPO)** were also noted.</description><pubDate>Thu, 25 Jan 2024 21:30:23 GMT</pubDate><category>adept</category><category>hugging-face</category><category>deepseek</category><category>mistral-ai</category><category>nous-research</category><category>fuyu-heavy</category><category>fuyu-8b</category><category>gemini-pro</category><category>claude-2</category><category>gpt4v</category><category>gemini-ultra</category><category>deepseek-coder-33b</category><category>yi-34b-200k</category><category>goliath-120b</category><category>mistral-7b-instruct-v0.2</category><category>mamba</category><category>rwkv</category><category>multimodality</category><category>visual-question-answering</category><category>direct-preference-optimization</category><category>benchmarking</category><category>model-size-estimation</category><category>quantization</category><category>model-merging</category><category>fine-tuning</category><category>instruct-tuning</category><category>rms-optimization</category><category>heterogeneous-ai-architectures</category><category>recurrent-llms</category><category>contrastive-preference-optimization</category></item><item><title>Google Solves Text to Video</title><link>https://news.smol.ai/issues/24-01-24-ainews-google-solves-text-to-video/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-24-ainews-google-solves-text-to-video/</guid><description>**Google Research** introduced **Lumiere**, a text-to-video model featuring advanced inpainting capabilities using a Space-Time diffusion process, surpassing previous models like Pika and Runway. Manveer from UseScholar.org compiled a comprehensive list of code evaluation benchmarks beyond HumanEval, including datasets from **Amazon Science**, **Hugging Face**, and others. Discord communities such as **TheBloke** discussed topics including running **Mistral-7B** via API, GPU rentals, and multimodal model integration with **LLava**. **Nous Research AI** highlighted learning rate strategies for LLM fine-tuning, issues with inference, and benchmarks like HumanEval and MBPP. **RestGPT** gained attention for controlling applications via RESTful APIs, showcasing LLM application capabilities.</description><pubDate>Thu, 25 Jan 2024 05:36:26 GMT</pubDate><category>google-research</category><category>amazon-science</category><category>huggingface</category><category>mistral-ai</category><category>together-ai</category><category>mistral-7b</category><category>llava</category><category>text-to-video</category><category>inpainting</category><category>space-time-diffusion</category><category>code-evaluation</category><category>fine-tuning</category><category>inference</category><category>gpu-rentals</category><category>multimodality</category><category>api</category><category>model-integration</category><category>learning-rates</category></item><item><title>RIP Latent Diffusion, Hello Hourglass Diffusion</title><link>https://news.smol.ai/issues/24-01-23-ainews-rip-latent-diffusion-hello-hourglass-diffusion/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-23-ainews-rip-latent-diffusion-hello-hourglass-diffusion/</guid><description>**Katherine Crowson** from **Stable Diffusion** introduces a hierarchical pure transformer backbone for diffusion-based image generation that efficiently scales to megapixel resolutions with under 600 million parameters, improving upon the original ~900M parameter model. This architecture processes local and global image phenomena separately, enhancing efficiency and resolution without latent steps. Additionally, Meta&apos;s Self Rewarding LM paper has inspired **lucidrains** to begin an implementation. Discord summaries highlight GPT-4&apos;s robustness against quantification tricks, discussions on open-source GPT-0 alternatives, challenges in DPO training on limited VRAM with suggestions like QLoRA and rmsprop, and efforts to improve roleplay model consistency through fine-tuning and merging. Philosophical debates on AI sentience and GPT-4 customization for markdown and translation tasks were also noted.</description><pubDate>Wed, 24 Jan 2024 01:38:15 GMT</pubDate><category>stable-diffusion</category><category>meta-ai-fair</category><category>openai</category><category>hugging-face</category><category>gpt-4</category><category>latent-diffusion</category><category>katherine-crowson</category><category>lucidrains</category><category>diffusion-models</category><category>transformers</category><category>image-generation</category><category>model-efficiency</category><category>fine-tuning</category><category>quantization</category><category>prompt-engineering</category><category>roleplay</category><category>training-optimization</category></item><item><title>Nightshade poisons AI art... kinda?</title><link>https://news.smol.ai/issues/24-01-22-ainews-nightshade-poisons-ai-art-kinda/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-22-ainews-nightshade-poisons-ai-art-kinda/</guid><description>Over the weekend of **1/19-20/2024**, discussions in **TheBloke Discord** covered key topics including **Mixture of Experts (MoE)** model efficiency, GPU parallelism, and quantization strategies. Users debated the effectiveness of AI detection tools like **GPTZero** and explored fine-tuning challenges with models such as **Mistral 7B** and **Falcon 7B**. Community interest was strong in developing simpler, community-powered quantization services and understanding model merging techniques. Ethical considerations around AI applications like AI girlfriend sites were also discussed.</description><pubDate>Mon, 22 Jan 2024 21:09:56 GMT</pubDate><category>mistral-ai</category><category>hugging-face</category><category>mistral-7b</category><category>falcon-7b</category><category>mixture-of-experts</category><category>gpu-parallelism</category><category>quantization</category><category>fine-tuning</category><category>model-merging</category><category>ai-detection</category><category>role-playing</category><category>benchmarking</category></item><item><title>Sama says: GPT-5 soon</title><link>https://news.smol.ai/issues/24-01-22-ainews-sama-says-gpt-5-soon/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-22-ainews-sama-says-gpt-5-soon/</guid><description>**Sam Altman** at Davos highlighted that his top priority is launching the new model, likely called **GPT-5**, while expressing uncertainty about **Ilya Sutskever**&apos;s employment status. **Itamar from Codium** introduced the concept of **Flow Engineering** with **AlphaCodium**, gaining attention from **Andrej Karpathy**. On the **TheBloke Discord**, engineers discussed a **multi-specialty mixture-of-experts (MOE) model** combining seven distinct 7 billion parameter models specialized in law, finance, and medicine. Debates on **8-bit fine-tuning** and the use of **bitsandbytes** with GPU support were prominent. Discussions also covered **model merging** using tools like **Mergekit** and compatibility with **Alpaca format**. Interest in optimizing AI models on **AMD** hardware using **AOCL blas and lapack libraries** with **llama.cpp** was noted. Users experimented with AI for command line tasks, and the **Mixtral MoE model** was refined to surpass larger models in coding ability. Comparisons among LLMs such as **GPT-3.5**, **Mixtral**, **Gemini Pro**, and **GPT-4** focused on knowledge depth, problem-solving, and speed, especially for coding tasks.</description><pubDate>Mon, 22 Jan 2024 20:51:23 GMT</pubDate><category>openai</category><category>codium</category><category>thebloke</category><category>amd</category><category>hugging-face</category><category>gpt-5</category><category>mixtral-7b</category><category>gpt-3.5</category><category>gemini-pro</category><category>gpt-4</category><category>llama-cpp</category><category>sam-altman</category><category>ilya-sutskever</category><category>itamar</category><category>andrej-karpathy</category><category>mixture-of-experts</category><category>fine-tuning</category><category>model-merging</category><category>8-bit-optimization</category><category>gpu-acceleration</category><category>performance-comparison</category><category>command-line-ai</category><category>vector-stores</category><category>embeddings</category><category>coding-capabilities</category></item><item><title>1/17/2024: Help crowdsource function calling datasets</title><link>https://news.smol.ai/issues/24-01-18-ainews-1172024-help-crowdsource-function-calling-datasets/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-18-ainews-1172024-help-crowdsource-function-calling-datasets/</guid><description>**LM Studio** updated its FAQ clarifying its **closed-source** status and perpetual freeness for personal use with no data collection. The new beta release includes fixes and hints at upcoming **2-bit quantization** support. For gaming, models like **Dolphin 2.7 Mixtral 8x7B**, **MegaDolphin**, and **Dolphin 2.6 Mistral 7B DPO** with **Q4_K_M** quantization were recommended. Discussions highlighted that single powerful GPUs outperform multi-GPU setups due to bottlenecks, with older GPUs like Tesla P40 being cost-effective. **Microsoft&apos;s AutoGen Studio** was introduced but has issues and requires **API fees** for open-source models. Linux users are advised to use **llama.cpp** over LM Studio due to lack of headless mode. Additional tools like **LLMFarm** for iOS and various Hugging Face repositories were also mentioned. *&quot;LM Studio must be running to use the local inference server as there is no headless mode available&quot;* and *&quot;matching model size to GPU memory is key for performance&quot;* were notable points.</description><pubDate>Thu, 18 Jan 2024 21:20:01 GMT</pubDate><category>lm-studio</category><category>mistral-ai</category><category>microsoft</category><category>hugging-face</category><category>apple</category><category>mistral-7b</category><category>dolphin-2.7-mixtral-8x7b</category><category>mega-dolphin</category><category>dolphin-2.6-mistral-7b-dpo</category><category>llama-cpp</category><category>yagilb</category><category>heyitsyorkie</category><category>function-calling</category><category>quantization</category><category>model-performance</category><category>gpu-optimization</category><category>model-selection</category><category>closed-source</category><category>memory-optimization</category><category>linux-server</category><category>api-fees</category><category>headless-mode</category></item><item><title>1/16/2024: ArtificialAnalysis - a new model/host benchmark site</title><link>https://news.smol.ai/issues/24-01-17-ainews-1162024-artificialanalysis-a-new-modelhost-benchmark-site/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-17-ainews-1162024-artificialanalysis-a-new-modelhost-benchmark-site/</guid><description>**Artificial Analysis** launched a new models and hosts comparison site, highlighted by **swyx**. **Nous Research AI** Discord discussed innovative summarization techniques using **NVIDIA 3090 and 2080ti GPUs** for processing around **100k tokens**, and adapting prompts for smaller models like **OpenChat 7B**. The availability of **Hermes 2 Mixtral** on **Huggingface&apos;s HuggingChat** was noted, alongside fine-tuning challenges with **Mixtral** using Axolotl. Discussions included byte-level tokenization experiments with **Byte Mistral**, multimodal training on **COCO image bytes**, and inference speed improvements using **vllm** and **llama.cpp**. Calls for transparency in data sharing and open-sourcing the **Hermes 2 Mixtral** dataset were emphasized, with comparisons of **dpo** and **sft** methods and quantized LLM use on **M1 MacBook Pro**.</description><pubDate>Wed, 17 Jan 2024 22:14:53 GMT</pubDate><category>nous-research</category><category>nvidia</category><category>hugging-face</category><category>mixtral</category><category>hermes-2-mixtral</category><category>openchat-7b</category><category>byte-mistral</category><category>swyx</category><category>gabriel_syme</category><category>manojbh</category><category>carsonpoole</category><category>fullstack6209</category><category>summarization</category><category>fine-tuning</category><category>byte-level-tokenization</category><category>multimodality</category><category>inference-speed-optimization</category><category>dataset-sharing</category><category>quantization</category></item><item><title>1/16/2024: TIES-Merging</title><link>https://news.smol.ai/issues/24-01-16-ainews-1162024-ties-merging/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-16-ainews-1162024-ties-merging/</guid><description>**TheBloke&apos;s Discord** community actively discusses **Mixture of Experts (MoE) models**, focusing on **random gate routing layers** for training and the challenges of immediate model use. There is a robust debate on **quantization methods**, comparing **GPTQ** and **EXL2 quants**, with EXL2 noted for faster execution on specialized hardware. A new model, **Nous Hermes 2**, based on **Mixtral 8x7B** and trained with **RLHF**, claims benchmark superiority but shows some inconsistencies. The **Frontier supercomputer** at Oak Ridge National Laboratory is highlighted for training a **trillion-parameter LLM** with **14TB RAM**, sparking discussions on open-sourcing government-funded AI research. Additionally, the application of **ghost attention** in the **academicat** model is explored, with mixed reactions from the community. *&quot;Random gate layer is good for training but not for immediate use,&quot;* and *&quot;EXL2 might offer faster execution on specialized hardware,&quot;* are key insights shared.</description><pubDate>Tue, 16 Jan 2024 20:51:01 GMT</pubDate><category>thebloke</category><category>hugging-face</category><category>nous-research</category><category>togethercompute</category><category>oak-ridge-national-laboratory</category><category>vast-ai</category><category>runpod</category><category>mixtral-8x7b</category><category>nous-hermes-2</category><category>frankendpo-4x7b-bf16</category><category>sanjiwatsuki</category><category>superking__</category><category>mrdragonfox</category><category>_dampf</category><category>kaltcit</category><category>rombodawg</category><category>technotech</category><category>mixture-of-experts</category><category>random-gate-routing</category><category>quantization</category><category>gptq</category><category>exl2-quants</category><category>reinforcement-learning-from-human-feedback</category><category>supercomputing</category><category>trillion-parameter-models</category><category>ghost-attention</category><category>model-fine-tuning</category><category>reward-models</category></item><item><title>1/13-14/2024: Don&apos;t sleep on #prompt-engineering </title><link>https://news.smol.ai/issues/24-01-15-ainews-113-142024-dont-sleep-on-prompt-engineering/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-15-ainews-113-142024-dont-sleep-on-prompt-engineering/</guid><description>The **OpenAI** Discord community engaged in diverse discussions including **prompt engineering** techniques like contrastive Chain of Thought and step back prompting, and explored **model merging** and **mixture-of-experts (MoE)** concepts. Philosophical debates on **AI consciousness** and the ethics of **AI-generated voices** highlighted concerns about AI sentience and copyright issues. Technical clarifications were made on **hyperdimensional vector space models** used in modern AI embeddings. Users also discussed **customizing GPT** with personality profiles and prompt personalization to overcome token limits, and proposed a **universal translator** feature for multilingual Discord interactions. Key contributors included longtime regular MadameArchitect and community members such as @darthgustav and @metaldrgn.</description><pubDate>Tue, 16 Jan 2024 00:58:42 GMT</pubDate><category>openai</category><category>madamearchitect</category><category>darthgustav</category><category>metaldrgn</category><category>prompt-engineering</category><category>model-merging</category><category>mixture-of-experts</category><category>ai-consciousness</category><category>ethics</category><category>hyperdimensional-vector-space</category><category>tokenization</category><category>multilinguality</category><category>prompt-personalization</category></item><item><title>1/12/2024: Anthropic coins Sleeper Agents</title><link>https://news.smol.ai/issues/24-01-13-ainews-1122024-anthropic-coins-sleeper-agents/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-13-ainews-1122024-anthropic-coins-sleeper-agents/</guid><description>**Anthropic** released a new paper exploring the persistence of deceptive alignment and backdoors in models through stages of training including supervised fine-tuning and reinforcement learning safety training. The study found that safety training and adversarial training did not eliminate backdoors, which can cause models to write insecure code or exhibit hidden behaviors triggered by specific prompts. Notable AI figures like **leo gao** and **andrej-karpathy** praised the work, highlighting its implications for future model security and the risks of sleeper agent LLMs. Additionally, the **Nous Research AI** Discord community discussed topics such as the trade-off between security and convenience, the **Hulk Dataset 0.1** for LLM fine-tuning, curiosity about a **120B model** and **Nous Mixtral**, debates on LLM leaderboard legitimacy, and the rise of Frankenmerge techniques for model merging and capacity enhancement.</description><pubDate>Sat, 13 Jan 2024 22:06:35 GMT</pubDate><category>anthropic</category><category>openai</category><category>nous-research</category><category>hugging-face</category><category>nous-mixtral</category><category>120b</category><category>leo-gao</category><category>andrej-karpathy</category><category>reinforcement-learning</category><category>fine-tuning</category><category>backdoors</category><category>model-security</category><category>adversarial-training</category><category>chain-of-thought</category><category>model-merging</category><category>dataset-release</category><category>security-vs-convenience</category></item><item><title>1/11/2024: Mixing Experts vs Merging Models</title><link>https://news.smol.ai/issues/24-01-12-ainews-1112024-mixing-experts-vs-merging-models/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-12-ainews-1112024-mixing-experts-vs-merging-models/</guid><description>**18 guilds**, **277 channels**, and **1342 messages** were analyzed with an estimated reading time saved of **187 minutes**. The community switched to **GPT-4 turbo** and discussed the rise of **Mixture of Experts (MoE) models** like **Mixtral**, **DeepSeekMOE**, and **Phixtral**. Model merging techniques, including naive linear interpolation and &quot;frankenmerges&quot; by **SOLAR** and **Goliath**, are driving new performance gains on open leaderboards. Discussions in the **Nous Research AI Discord** covered topics such as AI playgrounds supporting prompt and RAG parameters, security concerns about third-party cloud usage, debates on Discord bots and TOS, skepticism about **Teenage Engineering&apos;s** cloud LLM, and performance differences between **GPT-4 0613** and **GPT-4 turbo**. The community also explored fine-tuning strategies involving **DPO**, **LoRA**, and safetensors, integration of RAG with API calls, semantic differences between MoE and dense LLMs, and data frameworks like **llama index** and **SciPhi-AI&apos;s synthesizer**. Issues with anomalous characters in fine-tuning were also raised.</description><pubDate>Fri, 12 Jan 2024 18:49:15 GMT</pubDate><category>deepseek-ai</category><category>hugging-face</category><category>nous-research</category><category>teenage-engineering</category><category>discord</category><category>gpt-4-turbo</category><category>gpt-4-0613</category><category>mixtral</category><category>deepseekmoe</category><category>phixtral</category><category>ash_prabaker</category><category>shacrw</category><category>teknium</category><category>0xevil</category><category>everyoneisgross</category><category>ldj</category><category>pramod8481</category><category>mgreg_42266</category><category>georgejrjrjr</category><category>kenakafrosty</category><category>mixture-of-experts</category><category>model-merging</category><category>fine-tuning</category><category>rag</category><category>security</category><category>discord-tos</category><category>model-performance</category><category>prompt-engineering</category><category>function-calling</category><category>semantic-analysis</category><category>data-frameworks</category></item><item><title>1/10/2024: All the best papers for AI Engineers</title><link>https://news.smol.ai/issues/24-01-11-ainews-1102024-all-the-best-papers-for-ai-engineers/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-11-ainews-1102024-all-the-best-papers-for-ai-engineers/</guid><description>**OpenAI** launched the **GPT Store** featuring over **3 million** custom versions of **ChatGPT** accessible to Plus, Team, and Enterprise users, with weekly highlights of impactful GPTs like **AllTrails**. The new **ChatGPT Team** plan offers advanced models including **GPT-4** and **DALL·E 3**, alongside collaborative tools and enhanced data privacy. Discussions around AI-generated imagery favored **DALL·E** and **Stable Diffusion**, while users faced rate limit challenges and debated the GPT Store&apos;s SEO and categorization. Ethical considerations in prompt engineering were raised with a three-layer framework called &apos;The Sieve&apos;. Additionally, **DeepSeek-MoE** was noted for its range of Mixture of Experts (MoE) model sizes. *&quot;The Sieve,&quot; a three-layer ethical framework for AI,* was highlighted in prompt engineering discussions.</description><pubDate>Thu, 11 Jan 2024 08:35:15 GMT</pubDate><category>openai</category><category>deepseek-ai</category><category>chatgpt</category><category>gpt-4</category><category>dall-e-3</category><category>stable-diffusion</category><category>deepseek-moe</category><category>abdubs</category><category>darthgustav</category><category>prompt-engineering</category><category>model-release</category><category>rate-limiting</category><category>ethics</category><category>image-generation</category><category>moe</category><category>collaborative-workspaces</category><category>data-privacy</category></item><item><title>1/9/2024: Nous Research lands $5m for Open Source AI</title><link>https://news.smol.ai/issues/24-01-10-ainews-192024-nous-research-lands-dollar5m-for-open-source-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-10-ainews-192024-nous-research-lands-dollar5m-for-open-source-ai/</guid><description>**Nous Research** announced a **$5.2 million seed financing** focused on **Nous-Forge**, aiming to embed transformer architecture into chips for powerful servers supporting real-time voice agents and **trillion parameter models**. **Rabbit R1** launched a demo at CES with mixed reactions. **OpenAI** shipped the **GPT store** and briefly leaked an upcoming personalization feature. A new paper on **Activation Beacon** proposes a solution to extend LLMs&apos; context window significantly, with code to be released on GitHub. Discussions also covered **QLORA**, **fine-tuning**, **synthetic data**, and **custom architectures** for LLMs.</description><pubDate>Thu, 11 Jan 2024 00:53:13 GMT</pubDate><category>nous-research</category><category>openai</category><category>rabbit-tech</category><category>qlora</category><category>phi-3</category><category>mixtral</category><category>ollama</category><category>kenakafrosty</category><category>_stilic_</category><category>teknium</category><category>context-window</category><category>fine-tuning</category><category>synthetic-data</category><category>activation-beacon</category><category>transformer-architecture</category><category>seed-financing</category><category>real-time-voice-agents</category><category>trillion-parameter-models</category></item><item><title>1/8/2024: The Four Wars of the AI Stack</title><link>https://news.smol.ai/issues/24-01-08-ainews-182024-the-four-wars-of-the-ai-stack/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-08-ainews-182024-the-four-wars-of-the-ai-stack/</guid><description>The **Nous Research AI Discord** discussions highlighted several key topics including the use of **DINO**, **CLIP**, and **CNNs** in the **Obsidian Project**. A research paper on distributed models like **DistAttention** and **DistKV-LLM** was shared to address cloud-based **LLM** service challenges. Another paper titled &apos;Self-Extend LLM Context Window Without Tuning&apos; argued that existing **LLMs** can handle long contexts inherently. The community also discussed AI models like **Mixtral**, favored for its **32k context window**, and compared it with **Mistral** and **Marcoroni**. Other topics included hierarchical embeddings, agentic retrieval-augmented generation (**RAG**), synthetic data for fine-tuning, and the application of **LLMs** in the oil &amp; gas industry. The launch of the **AgentSearch-V1** dataset with one billion embedding vectors was also announced. The discussions covered **mixture-of-experts (MoE)** implementations and the performance of smaller models.</description><pubDate>Tue, 09 Jan 2024 07:39:51 GMT</pubDate><category>nous-research</category><category>openai</category><category>mistral-ai</category><category>hugging-face</category><category>mixtral</category><category>mistral</category><category>context-window</category><category>distributed-models</category><category>long-context</category><category>hierarchical-embeddings</category><category>agentic-rag</category><category>fine-tuning</category><category>synthetic-data</category><category>oil-and-gas</category><category>embedding-datasets</category><category>mixture-of-experts</category><category>model-comparison</category></item><item><title>1/6-7/2024: LlaMA Pro - an alternative to PEFT/RAG??</title><link>https://news.smol.ai/issues/24-01-07-ainews-16-72024-llama-pro-an-alternative-to-peftrag/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-07-ainews-16-72024-llama-pro-an-alternative-to-peftrag/</guid><description>New research papers introduce promising **Llama Extensions** including **TinyLlama**, a compact **1.1B** parameter model pretrained on about **1 trillion tokens** for 3 epochs, and **LLaMA Pro**, an **8.3B** parameter model expanding **LLaMA2-7B** with additional training on **80 billion tokens** of code and math data. LLaMA Pro adds layers to avoid catastrophic forgetting and balances language and code tasks but faces scrutiny for not using newer models like **Mistral** or **Qwen**. Meanwhile, **OpenAI** Discord discussions reveal insights on **GPT-4** token limits, privacy reassurances, fine-tuning for GPT-3.5, challenges with multi-language image recognition, custom GPT creation requiring **ChatGPT Plus**, and security concerns in GPT deployment. Users also share tips on dynamic image generation with **DALL-E** and logo creation.</description><pubDate>Mon, 08 Jan 2024 00:51:41 GMT</pubDate><category>openai</category><category>mistral-ai</category><category>llamaindex</category><category>langchain</category><category>llama-3</category><category>llama-3-1-1b</category><category>llama-3-8-3b</category><category>gpt-4</category><category>gpt-3.5</category><category>dall-e</category><category>yannic-kilcher</category><category>fine-tuning</category><category>model-expansion</category><category>token-limits</category><category>privacy</category><category>multilinguality</category><category>image-generation</category><category>security</category><category>custom-models</category><category>model-training</category></item><item><title>1/4/2024: Jeff Bezos backs Perplexity&apos;s $520m Series B.</title><link>https://news.smol.ai/issues/24-01-05-ainews-142024-jeff-bezos-backs-perplexitys-dollar520m-series-b/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-05-ainews-142024-jeff-bezos-backs-perplexitys-dollar520m-series-b/</guid><description>**Perplexity** announced their **Series B** funding round with notable investor **Jeff Bezos**, who previously invested in **Google** 25 years ago. **Anthropic** is raising **$750 million**, projecting at least **$850 million in annualized revenue** next year and implementing &quot;brutal&quot; changes to their Terms of Service. Discussions in **Nous Research AI Discord** cover topics such as **document recall limits from gigabytes of data**, **RNN memory and compute trade-offs**, **synthetic datasets**, and benchmarking of models like **WizardCoder-33B-V1.1**, **MobileLLaMA-1.4B-Base**, **ShearedLLaMA**, and **TinyLLaMA**. Other highlights include **UnsLOTH** optimizations for multi-GPU systems, **AI rap voice models**, **context-extending code**, and architectural innovations like applying **Detectron/ViT backbones to LLMs**, **sliding window attention** in **Mistral**, and parallelizing **Mixtral 8x7b** with **FSDP** and **HF Accelerate**.</description><pubDate>Fri, 05 Jan 2024 08:29:59 GMT</pubDate><category>perplexity</category><category>anthropic</category><category>google</category><category>nous-research</category><category>mistral-ai</category><category>hugging-face</category><category>wizardcoder-33b-v1.1</category><category>mobilellama-1.4b-base</category><category>shearedllama</category><category>tinyllama</category><category>mixtral-8x7b</category><category>jeff-bezos</category><category>document-recall</category><category>rnn-memory</category><category>synthetic-data</category><category>benchmarking</category><category>multi-gpu-support</category><category>context-length</category><category>model-architecture</category><category>sliding-window-attention</category><category>model-parallelism</category><category>gpu-optimization</category></item><item><title>1/3/2024: RIP Coqui</title><link>https://news.smol.ai/issues/24-01-03-ainews-132024-rip-coqui/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-03-ainews-132024-rip-coqui/</guid><description>**Coqui**, a prominent open source text-to-speech project from the Mozilla ML group, officially shut down. Discussions in the **HuggingFace** Discord highlighted skepticism about the claimed `3X faster` speed of **sdxl**, attributing improvements more to techniques like `torch.compile` and removal of `fp16` and `attention` rather than **diffusers 0.25** features. Users confirmed that a *HuggingFace user token* can be used across multiple machines, though distinct tokens are recommended for safety. The **Learning Loss Minimization (LLM) Leaderboard** briefly experienced issues but was later confirmed operational. A Kaggle notebook was shared demonstrating how to build Transformer architectures from scratch using PyTorch. Additionally, a new image dataset with 15k shoe, sandal, and boot images was introduced for multiclass classification tasks. Explanations about the workings of the Common Crawl web-crawling process were also shared.</description><pubDate>Thu, 04 Jan 2024 06:56:46 GMT</pubDate><category>coqui</category><category>mozilla</category><category>hugging-face</category><category>google</category><category>sdxl</category><category>diffusers-0.25</category><category>text-to-speech</category><category>performance-optimization</category><category>token-management</category><category>transformer-architecture</category><category>image-datasets</category><category>web-crawling</category><category>pytorch</category><category>leaderboards</category></item><item><title>1/2/2024: Smol tweaks to Smol Talk</title><link>https://news.smol.ai/issues/24-01-02-ainews-122024-smol-tweaks-to-smol-talk/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-02-ainews-122024-smol-tweaks-to-smol-talk/</guid><description>**OpenAI** Discord discussions highlight a detailed comparison of AI search engines including **Perplexity**, **Copilot**, **Bard**, and **Claude 2**, with Bard and Claude 2 trailing behind. **Meta AI** chatbot by Meta is introduced, available on Instagram and Whatsapp, featuring image generation likened to a free GPT version. Users report multiple browser issues with **ChatGPT**, including persistent captchas when using VPNs and plugin malfunctions. Debates cover prompt engineering, API usage, and data formats like **JSON**, **YAML**, and **Markdown**. Discussions also touch on ChatGPT&apos;s personality tuning and model capability variations. *&quot;Meta AI includes an image generation feature, which he likened to a free version of GPT.&quot;*</description><pubDate>Wed, 03 Jan 2024 07:38:24 GMT</pubDate><category>openai</category><category>meta-ai-fair</category><category>perplexity-ai</category><category>claude-2</category><category>bard</category><category>copilot</category><category>meta-ai</category><category>gemini-ultra</category><category>chatgpt</category><category>prompt-engineering</category><category>api</category><category>json</category><category>yaml</category><category>markdown</category><category>chatbot</category><category>image-generation</category><category>vpn</category><category>browser-compatibility</category><category>personality-tuning</category><category>plugin-issues</category></item><item><title>1/1/2024: How to start with Open Source AI</title><link>https://news.smol.ai/issues/24-01-02-ainews-112024-how-to-start-with-open-source-ai/</link><guid isPermaLink="true">https://news.smol.ai/issues/24-01-02-ainews-112024-how-to-start-with-open-source-ai/</guid><description>**OpenAI Discord** discussions revealed mixed sentiments about **Bing&apos;s AI** versus **ChatGPT** and **Perplexity AI**, and debated **Microsoft Copilot&apos;s** integration with **Office 365**. Users discussed **DALL-E 3** access within **ChatGPT Plus**, **ChatGPT&apos;s performance issues**, and ways to train a **GPT model** using book content via **OpenAI API** or custom GPTs. Anticipation for **GPT-4 turbo** in **Microsoft Copilot** was noted alongside conversations on **AI reasoning**, **prompt engineering**, and overcoming **Custom GPT** glitches. Advice for AI beginners included starting with **Python** and using YAML or Markdown for knowledge integration. The future of AI with multiple specialized GPTs and **Microsoft Copilot&apos;s** role was also explored.</description><pubDate>Wed, 03 Jan 2024 07:23:06 GMT</pubDate><category>openai</category><category>microsoft</category><category>perplexity-ai</category><category>gpt-4-turbo</category><category>dall-e-3</category><category>chatgpt</category><category>swyx</category><category>prompt-engineering</category><category>ai-reasoning</category><category>custom-gpt</category><category>performance</category><category>python</category><category>knowledge-integration</category></item><item><title>12/31/2023: Happy New Year</title><link>https://news.smol.ai/issues/23-12-31-ainews-12312023-happy-new-year/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-31-ainews-12312023-happy-new-year/</guid><description>**LM Studio** community discussions highlight variations and optimizations in **Dolphin** and **Mistral 7b** models, focusing on hardware-software configurations and GPU vRAM impact on processing speed. Challenges with **Mixtral** model deployment on local machines and workarounds for downloading models from **HuggingFace** in restricted regions were addressed. Users explored enhancing AI&apos;s emotional intelligence and personalities through extended prompts, referencing research on emotional stimuli in large language models. The community also discussed hardware setups for budget AI compute servers, integration issues with **ChromaDB** and **Autogen**, and shared positive feedback on LM Studio&apos;s usability and UI. Celebrations for the New Year added a social touch to the guild interactions.</description><pubDate>Mon, 01 Jan 2024 05:33:14 GMT</pubDate><category>lm-studio</category><category>mistral-ai</category><category>hugging-face</category><category>amd</category><category>mistral-7b</category><category>mixtral</category><category>fine-tuning</category><category>hardware-optimization</category><category>vram</category><category>emotional-intelligence</category><category>model-deployment</category><category>integration</category><category>gpu-optimization</category><category>software-updates</category></item><item><title>12/30/2023: Mega List of all LLMs</title><link>https://news.smol.ai/issues/23-12-31-ainews-12302023-mega-list-of-all-llms/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-31-ainews-12302023-mega-list-of-all-llms/</guid><description>**Stella Biderman**&apos;s tracking list of **LLMs** is highlighted, with resources shared for browsing. The **Nous Research AI** Discord discussed the **Local Attention Flax** module focusing on computational complexity, debating linear vs quadratic complexity and proposing chunking as a solution. Benchmark logs for various LLMs including **Deita v1.0** with its **SFT+DPO** training method were shared. Discussions covered model merging, graded modal types, function calling in AI models, and data contamination issues in **Mixtral**. Community insights were sought on **Amazon Titan Text Express** and **Amazon Titan Text Lite** LLMs, including a unique training strategy involving bad datasets. Several GitHub repositories and projects like **DRUGS**, **MathPile**, **CL-FoMo**, and **SplaTAM** were referenced for performance and data quality evaluations.</description><pubDate>Sun, 31 Dec 2023 10:23:31 GMT</pubDate><category>nous-research</category><category>hugging-face</category><category>amazon</category><category>mistral-ai</category><category>deita-v1.0</category><category>mixtral</category><category>amazon-titan-text-express</category><category>amazon-titan-text-lite</category><category>stella-biderman</category><category>euclaise</category><category>joey00072</category><category>local-attention</category><category>computational-complexity</category><category>benchmarking</category><category>model-merging</category><category>graded-modal-types</category><category>function-calling</category><category>data-contamination</category><category>training-methods</category></item><item><title>12/29/2023: TinyLlama on the way</title><link>https://news.smol.ai/issues/23-12-30-ainews-12292023-tinyllama-on-the-way/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-30-ainews-12292023-tinyllama-on-the-way/</guid><description>The **Nous/Axolotl community** is pretraining a **1.1B model on 3 trillion tokens**, showing promising results on **HellaSwag** for a small 1B model. The **LM Studio Discord** discussions cover extensive **GPU-related issues**, **Discord bot integration** with the **OpenAI API**, and **hardware limitations** affecting model usage. Community members also discuss **server hosting** for embeddings and LLMs, propose updates for **Discord channels** to improve model development collaboration, and address a **gibberish problem** in beta releases. The **Autogen** tool&apos;s installation and operational challenges are also clarified by users.</description><pubDate>Sat, 30 Dec 2023 11:06:56 GMT</pubDate><category>openai</category><category>hugging-face</category><category>tinyllama-1.1b</category><category>gpu-optimization</category><category>model-deployment</category><category>discord-bots</category><category>embedding-models</category><category>inference-server</category><category>hardware-compatibility</category><category>model-performance</category><category>beta-testing</category><category>autogen</category><category>context-window</category></item><item><title>12/28/2023: Smol Talk updates</title><link>https://news.smol.ai/issues/23-12-29-ainews-12282023-smol-talk-updates/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-29-ainews-12282023-smol-talk-updates/</guid><description>**Nous Research AI** Discord discussions covered topics such as AI placement charts, **ChatGPT**&apos;s issues with Latex math format compatibility with Obsidian, and performance metrics of the **TinyLlama 1.1B** model on various benchmarks. Users shared resources including the math-centric corpus **MathPile**, knowledge graph building methods, and open-source large language model repositories. Technical discussions included decentralized computation feasibility for models like **Mixtral**, philosophical debates on AI sentience, and strategies for model finetuning and token counting. The community also discussed the **Obsidian** model, vision model training, and the release of the multimodal **TinyGPT-V** model by Tyrannosaurus. *&quot;ChatGPT not generating Latex math format compatible with Obsidian&quot;* and *&quot;optimistic about human-level AI within our lifetime&quot;* were notable quotes.</description><pubDate>Fri, 29 Dec 2023 10:32:18 GMT</pubDate><category>nous-research</category><category>tyrannosaurus</category><category>tinyllama-1.1b</category><category>mixtral</category><category>tinygpt-v</category><category>gary-marcus</category><category>latex</category><category>benchmarking</category><category>knowledge-graphs</category><category>model-finetuning</category><category>tokenization</category><category>decentralized-computation</category><category>philosophy-of-ai</category><category>multimodality</category><category>vision</category><category>open-source-models</category></item><item><title>12/27/2023: NYT vs OpenAI</title><link>https://news.smol.ai/issues/23-12-29-ainews-12272023-nyt-vs-openai/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-29-ainews-12272023-nyt-vs-openai/</guid><description>The LM Studio Discord community extensively discussed **model performance** comparisons, notably between **Phi2** by **Microsoft Research** and **OpenHermes 2.5 Mistral 7b**, with focus on **U.S. history knowledge** and fine-tuning for improved accuracy. Technical challenges around **LLM API** usage, conversation history maintenance, and **GPU optimization** for inference speed were addressed. Hardware discussions covered **DDR4 vs DDR5**, multi-GPU setups, and potential of **Apple M1/M3** and **AMD AI CPUs** for AI workloads. The community also announced the **ChromaDB Plugin v3.0.2** release enabling image search in vector databases. Users shared practical tips on running multiple LM Studio instances and optimizing resource usage.</description><pubDate>Fri, 29 Dec 2023 10:14:01 GMT</pubDate><category>microsoft-research</category><category>mistral-ai</category><category>apple</category><category>amd</category><category>phi2</category><category>openhermes-2.5-mistral-7b</category><category>llama-2-7b</category><category>llama-2-13b</category><category>model-performance</category><category>fine-tuning</category><category>llm-api</category><category>gpu-optimization</category><category>hardware-configuration</category><category>multi-gpu</category><category>inference-speed</category><category>plugin-release</category><category>conversation-history</category></item><item><title>12/26/2023: not much happened today</title><link>https://news.smol.ai/issues/23-12-29-ainews-12262023-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-29-ainews-12262023-not-much-happened-today/</guid><description>**LM Studio** users extensively discussed its performance, installation issues on macOS, and upcoming features like **Exllama2 support** and multimodality with the **Llava model**. Conversations covered **GPU offloading**, **vRAM utilization**, **MoE model expert selection**, and **model conversion compatibility**. The community also addressed **inefficient help requests** referencing the blog &apos;Don&apos;t Ask to Ask, Just Ask&apos;. Technical challenges with **ChromaDB Plugin**, **server vs desktop hardware performance**, and **saving model states with Autogen** were highlighted. Discussions included comparisons with other chatbots and mentions of **AudioCraft** from **meta-ai-fair** and **MusicLM** from **google-deepmind** for music generation.</description><pubDate>Fri, 29 Dec 2023 10:07:18 GMT</pubDate><category>meta-ai-fair</category><category>google-deepmind</category><category>llava</category><category>exllama2</category><category>gpu-offloading</category><category>vram-utilization</category><category>model-conversion</category><category>moe-models</category><category>multimodality</category><category>model-performance</category><category>hardware-configuration</category><category>model-saving</category><category>chatml</category><category>installation-issues</category><category>music-generation</category></item><item><title>12/25/2023: Nous Hermes 2 Yi 34B for Christmas</title><link>https://news.smol.ai/issues/23-12-25-ainews-12252023-nous-hermes-2-yi-34b-for-christmas/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-25-ainews-12252023-nous-hermes-2-yi-34b-for-christmas/</guid><description>**Teknium** released **Nous Hermes 2** on **Yi 34B**, positioning it as a top open model compared to **Mixtral**, **DeepSeek**, and **Qwen**. **Apple** introduced **Ferret**, a new open-source multimodal LLM. Discussions in the **Nous Research AI Discord** focused on **AI model optimization** and **quantization** techniques like **AWQ**, **GPTQ**, and **AutoAWQ**, with insights on proprietary optimization and throughput metrics. Additional highlights include the addition of **NucleusX Model** to **transformers**, a **30B model with 80 MMLU**, and the **YAYI 2** language model by **Wenge Technology** trained on **2.65 trillion tokens**. *&quot;AutoAWQ outperforms vLLM up to batch size 8&quot;* was noted, and proprietary parallel decoding and tensor parallelization across GPUs were discussed for speed improvements.</description><pubDate>Tue, 26 Dec 2023 07:45:27 GMT</pubDate><category>teknim</category><category>nous-research</category><category>apple</category><category>mixtral</category><category>deepseek</category><category>qwen</category><category>huggingface</category><category>wenge-technology</category><category>nous-hermes-2</category><category>yi-34b</category><category>nucleusx</category><category>yayi-2</category><category>ferret</category><category>teknium</category><category>carsonpoole</category><category>casper_ai</category><category>pradeep1148</category><category>osanseviero</category><category>metaldragon01</category><category>quantization</category><category>model-optimization</category><category>throughput-metrics</category><category>batch-processing</category><category>parallel-decoding</category><category>tensor-parallelization</category><category>multimodality</category><category>language-model-pretraining</category><category>model-benchmarking</category></item><item><title>12/24/2023: Dolphin Mixtral 8x7b is wild</title><link>https://news.smol.ai/issues/23-12-25-ainews-12242023-dolphin-mixtral-8x7b-is-wild/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-25-ainews-12242023-dolphin-mixtral-8x7b-is-wild/</guid><description>**Mistral** models are recognized for being uncensored, and Eric Hartford&apos;s **Dolphin** series applies uncensoring fine-tunes to these models, gaining popularity on Discord and Reddit. The **LM Studio** Discord community discusses various topics including hardware compatibility, especially GPU performance with Nvidia preferred, fine-tuning and training models, and troubleshooting issues with LM Studio&apos;s local model hosting capabilities. Integration efforts with **GPT Pilot** and a beta release for ROCm integration are underway. Users also explore the use of **Autogen** for group chat features and share resources like the **Ollama** NexusRaven library. Discussions highlight challenges with running LM Studio on different operating systems, model performance issues, and external tools like **Google Gemini** and **ChatGLM3** compilation.</description><pubDate>Tue, 26 Dec 2023 07:23:04 GMT</pubDate><category>mistral-ai</category><category>ollama</category><category>google</category><category>openai</category><category>dolphin</category><category>glm3</category><category>chatglm3-ggml</category><category>eric-hartford</category><category>fine-tuning</category><category>hardware-compatibility</category><category>gpu-inference</category><category>local-model-hosting</category><category>model-integration</category><category>rocm-integration</category><category>performance-issues</category><category>autogen</category><category>linux</category><category>model-training</category></item><item><title>12/23/2023: NeurIPS Best Papers of 2023</title><link>https://news.smol.ai/issues/23-12-23-ainews-12232023-neurips-best-papers-of-2023/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-23-ainews-12232023-neurips-best-papers-of-2023/</guid><description>The **Latent Space Pod** released a **3-hour recap** of the **best NeurIPS 2023 papers**. The **Nous Research AI Discord** community discussed **optimizing AI performance** with shorter context lengths, **malware security concerns** linked to **HuggingFace**, and shared insights on **video and music content**. Technical discussions included the **DYAD research paper** proposing a faster alternative to linear layers, **Apple&apos;s ML Ferret** machine learning tool, and accessing **PALM2** via API. The community also explored **Large Language Models** focusing on specialized models, data scaling, embedding/vector databases, model merging, and interpretability, with mentions of **Hermes 2.5**, **GPT-4**, and **Mistral**. Additionally, there were conversations on the **Striped Hyena Architecture**, **quantization challenges**, and fixes related to **RMSNorm** and the **&quot;Attention is All You Need&quot;** paper.</description><pubDate>Sun, 24 Dec 2023 07:45:58 GMT</pubDate><category>nous-research</category><category>hugging-face</category><category>apple</category><category>gpt-4</category><category>palm2</category><category>hermes-2.5</category><category>mistral-7b</category><category>context-length</category><category>malware-security</category><category>video-content</category><category>music-content</category><category>linear-layers</category><category>api-access</category><category>large-language-models</category><category>embedding</category><category>vector-databases</category><category>model-merging</category><category>model-interpretability</category><category>striped-hyena-architecture</category><category>quantization</category><category>rmsnorm</category><category>attention-mechanisms</category></item><item><title>12/22/2023: Anyscale&apos;s Benchmark Criticisms</title><link>https://news.smol.ai/issues/23-12-22-ainews-12222023-anyscales-benchmark-criticisms/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-22-ainews-12222023-anyscales-benchmark-criticisms/</guid><description>**Anyscale** launched their **LLMPerf leaderboard** to benchmark large language model inference performance, but it faced criticism for lacking detailed metrics like cost per token and throughput, and for comparing public LLM endpoints without accounting for batching and load. In **OpenAI Discord** discussions, users reported issues with **Bard** and preferred **Microsoft Copilot** for storytelling, noting fewer hallucinations. There was debate on the value of upgrading from **GPT-3.5** to **GPT-4**, with many finding paid AI models worthwhile for coding productivity. Bugs and performance issues with OpenAI APIs were also highlighted, including slow responses and message limits. Future AI developments like **GPT-6** and concerns about OpenAI&apos;s transparency and profitability were discussed. Prompt engineering for image generation was another active topic, emphasizing clear positive prompts and the desire for negative prompts.</description><pubDate>Sat, 23 Dec 2023 01:16:52 GMT</pubDate><category>anyscale</category><category>openai</category><category>microsoft</category><category>gpt-4</category><category>gpt-3.5</category><category>bard</category><category>benchmarking</category><category>performance</category><category>api</category><category>prompt-engineering</category><category>bug-tracking</category><category>model-comparison</category><category>productivity</category><category>programming-languages</category><category>storytelling</category></item><item><title>12/21/2023: The State of AI (according to LangChain)</title><link>https://news.smol.ai/issues/23-12-21-ainews-12212023-the-state-of-ai-according-to-langchain/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-21-ainews-12212023-the-state-of-ai-according-to-langchain/</guid><description>**LangChain** launched their first report based on **LangSmith** stats revealing top charts for mindshare. On **OpenAI**&apos;s Discord, users raised issues about the **Mixtral model**, noting inconsistencies and comparing it to **Poe&apos;s Mixtral**. There were reports of declining output quality and unpredictable behavior in **GPT-4** and **ChatGPT**, with discussions on differences between **Playground GPT-4** and **ChatGPT GPT-4**. Users also reported anomalous behavior in **Bing** and **Bard AI** models, including hallucinations and strange assertions. Various user concerns included message limits on GPT-4, response completion errors, chat lags, voice setting inaccessibility, password reset failures, 2FA issues, and subscription restrictions. Techniques for guiding GPT-4 outputs and creative uses with **DALL-E** were also discussed. *Users highlighted financial constraints affecting subscriptions and queries about earning with ChatGPT and token costs.*</description><pubDate>Fri, 22 Dec 2023 00:20:28 GMT</pubDate><category>langchain</category><category>openai</category><category>perplexity-ai</category><category>microsoft</category><category>poe</category><category>mixtral</category><category>gpt-4</category><category>chatgpt</category><category>bard</category><category>dall-e</category><category>model-consistency</category><category>model-behavior</category><category>response-quality</category><category>chatgpt-usage-limitations</category><category>error-handling</category><category>user-experience</category><category>model-comparison</category><category>hallucination-detection</category><category>prompt-engineering</category><category>creative-ai</category></item><item><title>12/20/2023: Project Obsidian - Multimodal Mistral 7B from Nous</title><link>https://news.smol.ai/issues/23-12-20-ainews-12202023-project-obsidian-multimodal-mistral-7b-from-nous/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-20-ainews-12202023-project-obsidian-multimodal-mistral-7b-from-nous/</guid><description>**Project Obsidian** is a multimodal model being trained publicly, tracked by **Teknium** on the Nous Discord. Discussions include **4M: Massively Multimodal Masked Modeling** and **Reason.dev**, a TypeScript framework for LLM applications. The **OpenAI Discord** community discussed hardware specs for running **TensorFlow JS** for image detection, security API ideas for filtering inappropriate images, and concerns about racial and cultural bias in AI, especially in facial recognition and healthcare. Challenges with **GPT-3.5** and **GPT-4** in word puzzle games were noted, along with GPU recommendations prioritizing VRAM for AI inference. Users also debated **GPT-4**&apos;s vision capabilities, limitations of **DALL·E 3**, platform access issues, and prompting strategies for better outputs.</description><pubDate>Thu, 21 Dec 2023 03:20:57 GMT</pubDate><category>nous-research</category><category>teknim</category><category>openai</category><category>gpt-4</category><category>gpt-3.5</category><category>dall-e-3</category><category>multimodality</category><category>image-detection</category><category>security-api</category><category>bias</category><category>facial-recognition</category><category>healthcare-ai</category><category>gpu-optimization</category><category>prompt-engineering</category><category>vision</category></item><item><title>12/19/2023: Everybody Loves OpenRouter</title><link>https://news.smol.ai/issues/23-12-20-ainews-12192023-everybody-loves-openrouter/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-20-ainews-12192023-everybody-loves-openrouter/</guid><description>**OpenRouter** offers an easy OpenAI-compatible proxy for **Mixtral-8x7b-instruct**. Discord discussions highlight **GPT-4** performance and usability issues compared to **GPT-3.5**, including memory management and accessibility problems. Users debate local language models versus OpenAI API usage, with mentions of **Dolphin 2.0 Mistral 7B** and **Google&apos;s video generation project**. Prompt engineering and custom instructions for GPT models are also key topics. Concerns about censorship on models like **Gemini** and translation tool preferences such as **DeepL** were discussed.</description><pubDate>Wed, 20 Dec 2023 08:10:20 GMT</pubDate><category>openai</category><category>mistral-ai</category><category>google</category><category>hugging-face</category><category>gpt-4</category><category>gpt-3.5</category><category>mixtral-8x7b-instruct</category><category>dolphin-2.0-mistral-7b</category><category>gemini</category><category>performance</category><category>memory-management</category><category>api</category><category>prompt-engineering</category><category>local-language-models</category><category>translation</category><category>censorship</category><category>video-generation</category></item><item><title>12/18/2023: Gaslighting Mistral for fun and profit</title><link>https://news.smol.ai/issues/23-12-18-ainews-12182023-gaslighting-mistral-for-fun-and-profit/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-18-ainews-12182023-gaslighting-mistral-for-fun-and-profit/</guid><description>**OpenAI** Discord discussions reveal comparisons among language models including **GPT-4 Turbo**, **GPT-3.5 Turbo**, **Claude 2.1**, **Claude Instant 1**, and **Gemini Pro**, with **GPT-4 Turbo** noted for user-centric explanations. Rumors about **GPT-4.5** remain unconfirmed, with skepticism prevailing until official announcements. Users discuss technical challenges like slow responses and API issues, and explore role-play prompt techniques to enhance model performance. Ethical concerns about AI&apos;s impact on academia and employment are debated. Future features for **Dalle 3** and a proposed new GPT model are speculated upon, while a school project seeks help using the **OpenAI API**. The community also touches on AI glasses and job market implications of AI adoption.</description><pubDate>Tue, 19 Dec 2023 03:35:50 GMT</pubDate><category>openai</category><category>anthropic</category><category>google-deepmind</category><category>gpt-4-turbo</category><category>gpt-3.5-turbo</category><category>claude-2.1</category><category>claude-instant-1</category><category>gemini-pro</category><category>gpt-4.5</category><category>dalle-3</category><category>sam-altman</category><category>prompt-engineering</category><category>api</category><category>model-performance</category><category>ethics</category><category>role-play</category><category>user-experience</category><category>ai-impact-on-jobs</category><category>ai-translation</category><category>technical-issues</category></item><item><title>12/16/2023: ByteDance suspended by OpenAI</title><link>https://news.smol.ai/issues/23-12-16-ainews-12162023-bytedance-suspended-by-openai/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-16-ainews-12162023-bytedance-suspended-by-openai/</guid><description>The OpenAI Discord community discussed hardware options like **Mac racks** and the **A6000 GPU**, highlighting their value for AI workloads. They compared **Claude 2.1** and **GPT 4 Turbo** on coding tasks, with **GPT 4 Turbo** outperforming Claude 2.1. The benefits of the **Bard API** for **gemini pro** were noted, including a free quota of **60 queries per minute**. Users shared experiences with **ChatGPT Plus** membership issues, payment problems, and speculated about the upcoming **GPT-5** and the rumored **GPT-4.5**. Discussions also covered the confidentiality of the **Alpha feature**, AI art generation policies, and improvements in organizational work features. The community expressed mixed feelings about GPT-4&apos;s performance and awaited future model updates.</description><pubDate>Sat, 16 Dec 2023 19:41:52 GMT</pubDate><category>openai</category><category>google-deepmind</category><category>anthropic</category><category>claude-2.1</category><category>gpt-4-turbo</category><category>gemini-1.5-pro</category><category>gpt-5</category><category>gpt-4.5</category><category>gpt-4</category><category>hardware</category><category>gpu</category><category>api-costs</category><category>coding</category><category>model-comparison</category><category>subscription-issues</category><category>payment-processing</category><category>feature-confidentiality</category><category>ai-art-generation</category><category>organizational-productivity</category><category>model-speculation</category></item><item><title>12/15/2023: Mixtral-Instruct beats Gemini Pro (and matches GPT3.5)</title><link>https://news.smol.ai/issues/23-12-15-ainews-12152023-mixtral-instruct-beats-gemini-pro-and-matches-gpt35/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-15-ainews-12152023-mixtral-instruct-beats-gemini-pro-and-matches-gpt35/</guid><description>Thanks to a **karpathy** shoutout, **lmsys** now has enough data to rank **mixtral** and **gemini pro**. The discussion highlights the impressive performance of these state-of-the-art open-source models that can run on laptops. In the **openai** Discord, users compared AI tools like **perplexity** and **chatgpt&apos;s browsing tool**, favoring Perplexity for its superior data gathering, pricing, and usage limits. Interest was shown in AI&apos;s ability to convert large code files with **deepseek coder** recommended. Debates on privacy implications for AI advancement and challenges of running LLMs on local and cloud GPUs were prominent. Users reported issues with **chatgpt** including performance problems, loss of access to custom GPTs, and unauthorized access. Discussions also covered prompt engineering for large context windows and speculations about **gpt-4.5** and **gpt-4** future developments.</description><pubDate>Fri, 15 Dec 2023 22:33:20 GMT</pubDate><category>lmsys</category><category>openai</category><category>deepseek</category><category>cloudflare</category><category>huggingface</category><category>mixtral</category><category>gemini-pro</category><category>gpt-3.5</category><category>gpt-4.5</category><category>gpt-4</category><category>chatgpt</category><category>karpathy</category><category>performance</category><category>context-window</category><category>prompt-engineering</category><category>privacy</category><category>local-gpu</category><category>cloud-gpu</category><category>code-generation</category><category>model-comparison</category><category>model-usage</category><category>api-errors</category></item><item><title>12/14/2023: $1e7 for Superalignment</title><link>https://news.smol.ai/issues/23-12-14-ainews-12142023-dollar1e7-for-superalignment/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-14-ainews-12142023-dollar1e7-for-superalignment/</guid><description>**Jan Leike** is launching a new grant initiative inspired by **Patrick Collison&apos;s Fast Grants** to support AI research. **OpenAI** introduced a new developers Twitter handle @OpenAIDevs for community updates. Discussions on **OpenAI&apos;s Gemini** and **Bard** chatbots highlight their ability to read each other&apos;s instructions and offer unique coding solutions. Users reported various issues with **GPT-4**, including performance problems, customization difficulties, and a resolved bug in image recognition. There are ongoing conversations about **prompt engineering** challenges and new **JSON mode support** in Convo-lang for API use. Concerns about misuse of chatbots for illegal activities and alternatives like **Llama2** models and the **Perplexity chatbot** were also discussed.</description><pubDate>Thu, 14 Dec 2023 22:51:28 GMT</pubDate><category>openai</category><category>llamaindex</category><category>perplexity-ai</category><category>gemini</category><category>bard</category><category>gpt-4</category><category>gpt-4.5</category><category>llama-2</category><category>jan-leike</category><category>patrick-collison</category><category>prompt-engineering</category><category>api</category><category>custom-gpt</category><category>json</category><category>bug-fixes</category><category>chatbots</category><category>performance</category><category>tts</category><category>code-generation</category><category>image-recognition</category></item><item><title>12/13/2023 SOLAR10.7B upstages Mistral7B?</title><link>https://news.smol.ai/issues/23-12-13-ainews-12132023-solar107b-upstages-mistral7b/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-13-ainews-12132023-solar107b-upstages-mistral7b/</guid><description>**Upstage** released the **SOLAR-10.7B** model, which uses a novel Depth Up-Scaling technique built on the **llama-2** architecture and integrates **mistral-7b** weights, followed by continued pre-training. The **Nous** community finds it promising but not exceptional. Additionally, weights for the **phi-2** base model were released, trained on **1.4 trillion tokens** including synthetic texts created by GPT-3 and filtered by GPT-4, using **96 A100 GPUs** over 14 days. On **OpenAI&apos;s** Discord, users discussed challenges with various **GPT** models, including incoherent outputs, API usage limitations, and issues with **GPT-4 Vision API**. Conversations also covered understanding **AGI** and **ASI**, concerns about OpenAI&apos;s partnership with Axel Springer, and pricing changes for GPT Plus. Discussions included the **Gemini** chat model integrated into Bard and comparisons with GPT-4 performance.</description><pubDate>Wed, 13 Dec 2023 23:29:29 GMT</pubDate><category>upstage</category><category>nous-research</category><category>openai</category><category>mistral-ai</category><category>microsoft</category><category>solar-10.7b</category><category>llama-2</category><category>mistral-7b</category><category>phi-2</category><category>gpt-4</category><category>gemini</category><category>depth-up-scaling</category><category>pretraining</category><category>synthetic-data</category><category>gpu-training</category><category>api-usage</category><category>model-integration</category><category>agi</category><category>asi</category><category>chat-models</category><category>vision</category><category>model-performance</category><category>fine-tuning</category></item><item><title>12/12/2023: Towards LangChain 0.1</title><link>https://news.smol.ai/issues/23-12-12-ainews-12122023-towards-langchain-01/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-12-ainews-12122023-towards-langchain-01/</guid><description>The **Langchain rearchitecture** has been completed, splitting the repo for better maintainability and scalability, while remaining backwards compatible. **Mistral** launched a new Discord community, and **Anthropic** is rumored to be raising another **$3 billion**. On the **OpenAI Discord**, discussions covered **information leakage** in AI training, **mixture of experts (MoE) models** like **mixtral 8x7b**, advanced **prompt engineering techniques**, and issues with **ChatGPT** performance and API access. Users also explored AI applications in **logo generation**, **education**, and **gaming**, and shared solutions for **Oauth2 authentication** problems. A new small language model named **Phi-2** was mentioned from **Microsoft**.</description><pubDate>Wed, 13 Dec 2023 03:45:12 GMT</pubDate><category>langchain</category><category>mistral-ai</category><category>anthropic</category><category>openai</category><category>microsoft</category><category>mixtral-8x7b</category><category>phi-2</category><category>gpt-3</category><category>chatgpt</category><category>gpt-4</category><category>mixture-of-experts</category><category>information-leakage</category><category>prompt-engineering</category><category>oauth2</category><category>logo-generation</category><category>education-ai</category><category>gaming-ai</category><category>api-access</category><category>model-maintainability</category><category>scalability</category></item><item><title>12/11/2023: Mixtral beats GPT3.5 and Llama2-70B</title><link>https://news.smol.ai/issues/23-12-11-ainews-12112023-mixtral-beats-gpt35-and-llama2-70b/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-11-ainews-12112023-mixtral-beats-gpt35-and-llama2-70b/</guid><description>**Mistral AI** announced the **Mixtral 8x7B** model featuring a Sparse Mixture of Experts (SMoE) architecture, sparking discussions on its potential to rival **GPT-4**. The community debated GPU hardware options for training and fine-tuning transformer models, including **RTX 4070s**, **A4500**, **RTX 3090s with nvlink**, and **A100 GPUs**. Interest was expressed in fine-tuning Mixtral and generating quantized versions, alongside curating high-quality coding datasets. Resources shared include a YouTube video on open-source model deployment, an Arxiv paper, GitHub repositories, and a blog post on Mixture-of-Experts. Discussions also touched on potential open-source releases of **GPT-3.5 Turbo** and **llama-3**, and running **OpenHermes 2.5** on Mac M3 Pro with VRAM considerations.</description><pubDate>Mon, 11 Dec 2023 20:11:07 GMT</pubDate><category>mistral-ai</category><category>openai</category><category>huggingface</category><category>mixtral-8x7b</category><category>gpt-4</category><category>gpt-3.5-turbo</category><category>llama-3</category><category>openhermes-2.5</category><category>llava-v1.5-13b-gptq</category><category>sparse-mixture-of-experts</category><category>fine-tuning</category><category>quantization</category><category>gpu-hardware</category><category>transformers</category><category>model-deployment</category><category>open-source</category><category>coding-datasets</category></item><item><title>12/10/2023: not much happened today</title><link>https://news.smol.ai/issues/23-12-10-ainews-12102023-not-much-happened-today/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-10-ainews-12102023-not-much-happened-today/</guid><description>**Nous Research AI** Discord community discussed attending **NeurIPS** and organizing future AI events in Australia. Highlights include interest in open-source and decentralized AI projects, with **Richard Blythman** seeking co-founders. Users shared projects like **Photo GPT AI** and introduced **StableLM Zephyr 3B**. The **Mixtral** model, based on **Mistral**, sparked debate on performance and GPU requirements, with comparisons to **GPT-3.5** and potential competitiveness with **GPT-4** after fine-tuning. Tools like **Tensorboard**, **Wandb**, and **Llamahub** were noted for fine-tuning and evaluation. Discussions covered **Mixture of Experts (MoE)** architectures, fine-tuning with limited data, and inference optimization strategies for ChatGPT. Memes and community interactions referenced AI figures like **Andrej Karpathy** and **Yann LeCun**. The community also shared resources such as GitHub links and YouTube videos related to these models and tools.</description><pubDate>Sun, 10 Dec 2023 23:49:57 GMT</pubDate><category>nous-research</category><category>openai</category><category>mistral-ai</category><category>hugging-face</category><category>ollama</category><category>lm-studio</category><category>mixtral-8x7b-32kseqlen</category><category>mistral-7b</category><category>stablelm-zephyr-3b</category><category>openhermes-2.5-neural-chat-v3-3-slerp</category><category>gpt-3.5</category><category>gpt-4</category><category>andrej-karpathy</category><category>yann-lecun</category><category>richard-blythman</category><category>gabriel-syme</category><category>pradeep1148</category><category>cyborg_1552</category><category>fine-tuning</category><category>mixture-of-experts</category><category>model-benchmarking</category><category>inference-optimization</category><category>model-evaluation</category><category>open-source</category><category>decentralized-ai</category><category>gpu-optimization</category><category>community-engagement</category></item><item><title>12/9/2023: The Mixtral Rush</title><link>https://news.smol.ai/issues/23-12-09-ainews-1292023-the-mixtral-rush/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-09-ainews-1292023-the-mixtral-rush/</guid><description>**Mixtral&apos;s weights** were released without code, prompting the **Disco Research community** and **Fireworks AI** to implement it rapidly. Despite efforts, no significant benchmark improvements were reported, limiting its usefulness for local LLM usage but marking progress for the **small models community**. Discussions in the DiscoResearch Discord covered **Mixtral&apos;s performance** compared to models like **Hermes 2.5** and **Hermes 2**, with evaluations on benchmarks such as **winogrande**, **truthfulqa_mc2**, and **arc_challenge**. Technical topics included GPU requirements, multi-GPU setups, and quantization via **GPTQ**. Benchmarking strategies like grammar-based evaluation, chain of thought (CoT), and min_p sampling were explored, alongside model sampling techniques like Min P and Top P to enhance response stability and creativity. Users also discussed GPTs&apos; learning limitations and the adaptability of models under varying conditions, emphasizing min_p sampling&apos;s role in enabling higher temperature settings for creativity.</description><pubDate>Sat, 09 Dec 2023 23:30:00 GMT</pubDate><category>discoresearch</category><category>fireworks-ai</category><category>hugging-face</category><category>mistral-ai</category><category>mixtral</category><category>hermes-2.5</category><category>hermes-2</category><category>mistral-yarn</category><category>ultrachat</category><category>bjoernp</category><category>the_bloke</category><category>rtyax</category><category>kalomaze</category><category>solbus</category><category>calytrix</category><category>benchmarking</category><category>gpu-requirements</category><category>multi-gpu</category><category>quantization</category><category>gptq</category><category>chain-of-thought</category><category>min-p-sampling</category><category>top-p-sampling</category><category>model-sampling</category><category>model-merging</category><category>model-performance</category><category>small-models</category><category>reasoning-consistency</category><category>temperature-sampling</category></item><item><title>12/8/2023 - Mamba v Mistral v Hyena</title><link>https://news.smol.ai/issues/23-12-08-ainews-1282023-mamba-v-mistral-v-hyena/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-08-ainews-1282023-mamba-v-mistral-v-hyena/</guid><description>Three new AI models are highlighted: **Mistral&apos;s 8x7B MoE model (Mixtral)**, **Mamba models** up to 3B by Together, and **StripedHyena 7B**, a competitive subquadratic attention model from Stanford&apos;s Hazy Research. Discussions on **Anthropic&apos;s Claude 2.1** focus on its prompting technique and alignment challenges. The **Gemini AI** from Google is noted as potentially superior to **GPT-4**. The community also explores **Dreambooth** for image training and shares resources like the **DialogRPT-human-vs-machine** model on Hugging Face. Deployment challenges for large language models, including CPU performance and GPU requirements, are discussed with references to **Falcon 180B** and transformer batching techniques. User engagement includes meme sharing and humor.</description><pubDate>Fri, 08 Dec 2023 22:40:04 GMT</pubDate><category>mistral-ai</category><category>togethercompute</category><category>stanford</category><category>anthropic</category><category>google</category><category>hugging-face</category><category>mistral-8x7b-moe</category><category>mamba-3b</category><category>stripedhyena-7b</category><category>claude-2.1</category><category>gemini</category><category>gpt-4</category><category>dialogrpt-human-vs-machine</category><category>cybertron-7b-v2-gguf</category><category>falcon-180b</category><category>andrej-karpathy</category><category>tri-dao</category><category>maxwellandrews</category><category>raddka</category><category>mixture-of-experts</category><category>attention-mechanisms</category><category>prompt-engineering</category><category>alignment</category><category>image-training</category><category>model-deployment</category><category>gpu-requirements</category><category>cpu-performance</category><category>model-inference</category><category>long-context</category><category>model-evaluation</category><category>open-source</category><category>chatbots</category></item><item><title>12/7/2023: Anthropic says &quot;skill issue&quot;</title><link>https://news.smol.ai/issues/23-12-07-ainews-1272023-anthropic-says-skill-issue/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-07-ainews-1272023-anthropic-says-skill-issue/</guid><description>**Anthropic** fixed a glitch in their **Claude 2.1** model&apos;s needle in a haystack test by adding a prompt. Discussions on **OpenAI&apos;s** Discord compared **Google&apos;s Gemini Pro and Gemini Ultra** models with **OpenAI&apos;s GPT-4** and **GPT-3.5**, with some users finding GPT-4 superior in benchmarks. Rumors about a **GPT-4.5** release circulated without official confirmation. Concerns were raised about &quot;selective censorship&quot; affecting language model performance. The EU&apos;s potential regulation of AI, including **ChatGPT**, was highlighted. Users reported issues with **ChatGPT Plus** message limits and subscription upgrades, and shared experiences with **BingChat** and **DALL-E**. The community discussed prompt engineering techniques and future applications like image generation and MIDI sequence analysis, expressing hopes for **GPT-5**.</description><pubDate>Thu, 07 Dec 2023 20:49:01 GMT</pubDate><category>anthropic</category><category>openai</category><category>google</category><category>claude-2.1</category><category>gpt-4</category><category>gpt-3.5</category><category>gemini-pro</category><category>gemini-ultra</category><category>gpt-4.5</category><category>chatgpt</category><category>bingchat</category><category>dall-e</category><category>gpt-5</category><category>prompt-engineering</category><category>model-performance</category><category>regulation</category><category>language-model-performance</category><category>image-generation</category><category>audio-processing</category><category>midi-sequence-analysis</category><category>subscription-issues</category><category>network-errors</category></item><item><title>Is Google&apos;s Gemini... legit?</title><link>https://news.smol.ai/issues/23-12-06-ainews-is-googles-gemini-legit/</link><guid isPermaLink="true">https://news.smol.ai/issues/23-12-06-ainews-is-googles-gemini-legit/</guid><description>**Google&apos;s Gemini** AI model is generating significant discussion and skepticism, especially regarding its **32-shot chain of thought** MMLU claim and **32k context window**. The community is comparing Gemini&apos;s performance and capabilities with **OpenAI&apos;s GPT-4** and **GPT-3.5**, highlighting the upcoming **Gemini Pro** and **Gemini Ultra** models on the Bard platform. Users report various **OpenAI service issues** including chatbot errors and subscription problems. Discussions also cover **prompt engineering techniques**, AI model evaluation comparing **GPT-4**, **Claude 2.1**, and **PaLM2**, and improvements in speech and multimodal capabilities. The bot now supports reading and summarizing links from platforms like arXiv, Twitter, and YouTube, enhancing user interaction.</description><pubDate>Wed, 06 Dec 2023 22:22:18 GMT</pubDate><category>google</category><category>openai</category><category>gemini</category><category>gemini-pro</category><category>gemini-ultra</category><category>gpt-4</category><category>gpt-3.5</category><category>claude-2.1</category><category>palm2</category><category>swyx</category><category>chain-of-thought</category><category>context-windows</category><category>prompt-engineering</category><category>model-evaluation</category><category>multimodality</category><category>speech-processing</category><category>chatbot-errors</category><category>subscription-management</category></item></channel></rss>