All tags
Person: "koraykv"
not much happened today
gemini-3.7-flash google-deepmind google deepseek arcee agentic-workflows coding knowledge-work benchmarking runtime-systems open-source long-running-processes asynchronous-computation software-architecture developer-tools price-performance _philschmid koraykv officiallogank tianyi eliebakouch bookwormengr 0xlogicrw teortaxestex latkins stochasticchasm code_star fujikanaeda
Google rapidly released Gemini 3.7 Flash just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory price cut and improved benchmark scores like DeepSWE 65.3% and Code Arena Elo 1588. The update quickly integrated across multiple platforms including Gemini API and Android Studio, with independent benchmarks confirming performance gains. Meanwhile, DeepSeek open-sourced DeepSeek Harness under MIT license as a developer preview, focusing on architecture innovations like KV-cache-aware append-only history semantics and treating the harness as an OS/runtime substrate for recursive improvement. Arcee also open-sourced NAC under Apache 2.0, designed for long-running asynchronous tasks and powering significant code pipelines, enabling orchestration from phones or delegation via Codex/Claude.
Gemini 3.1 Pro: 2x 3.0 on ARC-AGI 2
gemini-3.1-pro gemini-3-deep-think google google-deepmind geminiapp reasoning benchmarking agentic-ai cost-efficiency hallucination code-generation model-release developer-tools sundarpichai demishassabis jeffdean koraykv noamshazeer joshwoodward artificialanlys arena oriolvinyalsml scaling01
Google released Gemini 3.1 Pro, a developer preview integrated across the Gemini app, NotebookLM, Gemini API / AI Studio, and Vertex AI, highlighting a significant reasoning improvement with ARC-AGI-2 = 77.1% and strong coding and agentic-tool benchmarks like SWE-Bench Verified = 80.6%. Independent evaluators such as Artificial Analysis and Arena confirmed top-tier performance and cost efficiency, though community reactions included excitement about practical gains, skepticism about benchmark targeting, and concerns over rollout inconsistencies. The release emphasizes the same core intelligence powering Gemini 3 Deep Think scaled for practical use, with notable mentions from leaders like @sundarpichai, @demishassabis, and @JeffDean.