All tags
Topic: "model-compression"
not much happened today
deepseek-v4.1-flash glm-5.3-flash deepseek baseten ollama causal-encoder-decoder inference-efficiency model-architecture multimodality model-optimization vision model-quantization model-compression context-windows sebastian_raschka
DeepSeek launched V4.1-Flash, a new open-weight flagship model focused on extreme inference efficiency and low cost, featuring a 763B total-parameter causal encoder-decoder architecture with 8B active input and 16B active output parameters and 1M-token context. It scored 40 on the Artificial Analysis Intelligence Index, outperforming its predecessor and ranking just below GLM-5.3-Flash. The model supports text and image input, is available under an MIT license, and is accessible via US/API. The architecture introduces a novel causal encoder-decoder design aimed at reducing active compute and KV/cache costs, with a hybrid sparse/local approach and a unique vision encoder differing from recent Chinese models. Early layers use a SWA-only pattern, and the model has an effective depth of about 40 layers with 20 decoder layers. Baseten and Ollama have begun supporting and rolling out the model to users.
not much happened today
ornith-1.5 qwen3.8-27b claude-opus-5 kimi-k3 glm-5.2 grok-4.5 gpt-5.6-luna grok-4.6 glm-5.3 trueforge ornith vllm ollama unsloth qwen arena valsai deepseek truefoundry claude model-compression quantization reinforcement-learning agent-evaluation plugin-architecture open-agent-runtime cost-efficiency session-management tooling benchmarking ornith_ unslothai danielhanchen arena valsai zhihufrontier theturingpost truefoundry omarsar0 kimmonismus bradenjhancock dbreunig rseroter claudedevs
Ornith-1.5 launches as a new open-weight model family with 9B dense, 35B MoE, and 397B MoE variants under MIT license, featuring quantized formats like FP8, GGUF, MLX, and NVFP4 and showcasing end-to-end self-improvement capabilities. Compression techniques improve accuracy and efficiency, with Qwen3.8-27B GGUFs using Dynamic V3 achieving 10% higher accuracy and 1-bit quantization retaining 77% BF16 accuracy on 8GB RAM. Agent evaluation boards highlight models like Claude Opus 5 (High), Kimi K3, GLM 5.2, Grok 4.5, and GPT-5.6 Luna leading in quality and value. DeepSeek Harness (DSH) introduces a plugin-based open agent runtime architecture optimized for extensibility and tooling. TrueFoundry open-sources TrueForge, a self-hostable, vendor-neutral agent harness that reduces token usage by 30% and cuts costs by 75% while maintaining accuracy, emphasizing the growing importance of session, environment, memory, and tools layers in agent platforms.
not much happened today
qwen3.8-27b deepseek-v4-pro gpt-5.6-luna openai nvidia stripe openrouter vercel cursor langchain vanta deepseek ai-infrastructure power-management model-routing api-pricing developer-platforms agentic-coding multi-agent-systems evaluation-tools harness-level-evaluation sandboxing permissioning model-compression local-models markchen90 kimmonismus hamelhusain tonbistudio teknium omarsar0 cline
OpenAI is advancing its power-and-compute infrastructure with a 4+ GW NVIDIA capacity commitment and an 8 GW Ohio campus buildout through 2032, emphasizing vertical integration across power, data centers, and chips. The model access and routing API layer is becoming a competitive pricing battlefield, highlighted by the Stripe–OpenRouter deal and recent price cuts by OpenRouter and Vercel. Cursor launched Origin, an AI-native IDE aiming for full control over coding workflows, signaling a shift toward agentic coding platforms. Multi-agent orchestration is evolving from demos to operational patterns with specialized, persistent-context agents, as seen in projects by Hermes Desktop, Bot Mode, and Codex orchestration. Evaluation tools like Hamel Husain’s eval-skills plugin and Agent Arena are advancing harness-level measurement with data from over 1.7M sessions. Enterprise agent tooling is improving with sandboxed, permissioned execution environments from Vanta and LangChain. Open models like Qwen3.8-27B are compressing the capability frontier, reaching performance comparable to DeepSeek V4-Pro and GPT-5.6 Luna on the Artificial Analysis Intelligence Index, marking a milestone for local models.
not much happened today
gpt-5.6-luna gpt-5.6-terra gpt-5.6-sol arc-agi-3 inkling-small inkling gemini-robotics-2 openai thinking-machines lmsys modal unsloth artificial-analysis google price-optimization agent-systems memory-retention context-compaction multimodality mixture-of-experts model-compression benchmarking open-weights multimodal-models model-efficiency model-deployment embodied-ai robotics long-context sama fchollet kimmonismus gneubig scaling01 mervenoyann
OpenAI aggressively cut prices for GPT-5.6 Luna by 80% and Terra by 20%, introducing a faster Sol Fast tier with up to 2.5× lower latency at double the price, improving agent workflow costs by roughly 10×. The ARC-AGI-3 debate highlighted that the complete agent system, including memory retention and tool orchestration, is critical beyond just the base model. Thinking Machines released Inkling-Small, an open-weights, multimodal MoE model with 276B parameters (12B active), delivering performance comparable to the original Inkling at a quarter of the size, supporting audio, images, and Python-based image inspection. Benchmarks show Inkling-Small excels in coding and multimodality tasks, with 1M-context support and broad open inference stack adoption. The news also mentions Google's Gemini Robotics 2 advancing embodied AI from tabletop to full-body control.
not much happened today
gpt-5.6 codex bonsai-27b qwen-3.6-27b hy3-295b gemma-4 qwen3.5-122b-a10b glm-4.7-flash deepseek-v4-flash mimo-v2.5 glm-5.2-nvfp4 moss-vl-realtime openai jetbrains langchain prismml tencent-hunyuan miaai_lab openmoss agentic-ai model-quantization local-inference multimodality video-understanding model-compression evals observability long-context tool-use sama reach_vb kimmonismus swyx theo andykonwinski
OpenAI's agent products saw a 2.5x weekly usage growth driven by Codex + ChatGPT Work and demand for GPT-5.6 Sol. JetBrains adopted Codex as a recommended agent, while LangChain enhanced tracing and observability across multiple tools. PrismML released Bonsai 27B, a compressed variant of Qwen 3.6 27B enabling local multimodal agentic workflows on consumer devices. Tencent Hunyuan introduced 1-bit and 4-bit quantized Hy3 295B model deployable on a single GPU. Quantization advances like NVFP4 dynamic quants for Gemma-4 and others support serious local inference. OpenMOSS launched MOSS-VL-Realtime 11B for continuous video stream perception with a 256K context window. "Harness quality and observability are becoming a first-class differentiator" and local inference is now viable for agentic workflows.
not much happened today
glm-4.7-flash glm-4.7 glm-4.5 qwen3-vl qwen meta-ai-fair carnegie-mellon sakana-ai zhipu-ai transformer-memory model-architecture mixture-of-experts adaptive-position-encoding long-context model-compression inference-optimization local-inference model-deployment benchmarking coding agentic-ai
AI News for 1/16/2026-1/19/2026 covers new architectures for scaling Transformer memory and context, including STEM from Carnegie Mellon and Meta AI, which replaces part of the FFN with a token-indexed embedding lookup enabling CPU offload and asynchronous prefetch. RePo from Sakana AI introduces adaptive positional reordering to improve robustness on noisy and long-range contexts. Model releases highlight Zhipu AI's GLM-4.7-Flash, a 30B-class MLA + small MoE model optimized for coding and agentic tasks, noted for strong benchmark performance and a compression narrative from larger to smaller models. Inference and deployment updates include mlx-lm 0.30.3 supporting GLM-4.7-Flash with efficient 4-bit performance on laptops. The report emphasizes practical takeaways on static sparsity, adaptive ordering, and the resurgence of small, fast models for interactive tasks. "Sparse capacity doesn’t have to mean MoE routers + expert parallelism; static sparsity can be systems-friendly."
DeepSeek-OCR finds vision models can decode 10x more efficiently with ~97% accuracy of text-only, 33/200k pages/day/A100
deepseek-ocr deepseek3b-moe-a570m veo-3.1 deepseek-ai google-deepmind krea ocr vision multimodality model-compression long-context model-architecture video-generation autoregressive-models model-efficiency precision-editing karpathy teortaxestex reach_vb _akhaliq eliebakouch vikhyatk demishassabis
As ICCV 2025 begins, DeepSeek releases a novel DeepSeek-OCR 3B MoE vision-language model that compresses long text as visual context with high accuracy and efficiency, challenging traditional tokenization approaches. The model achieves ~97% decoding precision at <10× compression and processes up to ~33M pages/day on 20 A100-40G nodes, outperforming benchmarks like GOT-OCR2.0. Discussions highlight the potential for unlimited context windows and tokenization-free inputs, with contributions from @karpathy, @teortaxesTex, and others. In video generation, google-deepmind's Veo 3.1 leads community benchmarks with advanced precision editing and scene blending, while Krea open-sources a 14B autoregressive video model enabling realtime long-form generation at ~11 FPS on a single B200 GPU.
not much happened today
embeddinggemma qwen-2.5-coder minicpm-v-4.5 gpt-4o gemini-2.0-pro google-deepmind hugging-face jina-ai lighton microsoft stanford openai ollama weaviate langchain llamaindex embeddings retrieval-augmented-generation quantization multilingual-models on-device-ai semantic-search contrastive-learning dataset-release vision multimodality video-generation text-to-speech optimizer-benchmarking training-recipes model-compression video-token-compression fine-tuning osanseviero _philschmid tomaarsen ollama weaviate_io lusxvr andimarafioti thibaudfrere _akhaliq clementdelangue gordonwetzstein konstmish wen_kaiyue percyliang
Google DeepMind released EmbeddingGemma (308M), a small multilingual embedding model optimized for on-device retrieval-augmented generation and semantic search, supporting over 100 languages and running efficiently with quantization and EdgeTPU latency under 15ms. Jina AI introduced new code-focused embedding models (0.5B/1.5B) with GGUF quantization, achieving state-of-the-art retrieval across multiple languages and tasks. LightOn demonstrated large-scale retrieval training without distillation using contrastive training on billions of passages. Hugging Face released the FineVision dataset with 17.3M images and 9.5B answer tokens for vision-language model training, showing significant benchmark improvements. The MiniCPM-V 4.5 (8B) multimodal model reported surpassing GPT-4o and Gemini-2.0 Pro on OpenCompass benchmarks with innovative video token compression. Microsoft’s VibeVoice TTS and Stanford’s Mixture-of-Contexts video generation also featured. Additionally, a Stanford study benchmarked optimizers like Muon, Soap, Mars, and Sophia, finding diminishing speedups over AdamW at larger scales but advantages at smaller scales. The new ChatGPT branching feature was noted for its simplicity and popularity. "Everyone's a decacorn now."
Not much happened today
mistral-small-3.2 magenta-realtime afm-4.5b llama-3 openthinker3-7b deepseek-r1-distill-qwen-7b storm qwen2-vl gpt-4o dino-v2 sakana-ai mistral-ai google arcee-ai deepseek-ai openai amazon gdm reinforcement-learning chain-of-thought fine-tuning function-calling quantization music-generation foundation-models reasoning text-video model-compression image-classification evaluation-metrics sama
Sakana AI released Reinforcement-Learned Teachers (RLTs), a novel technique using smaller 7B parameter models trained via reinforcement learning to teach reasoning through step-by-step explanations, accelerating Chain-of-Thought learning. Mistral AI updated Mistral Small 3.2 improving instruction following and function calling with experimental FP8 quantization. Google Magenta RealTime, an 800M parameter open-weights model for real-time music generation, was released. Arcee AI launched AFM-4.5B, a sub-10B parameter foundation model extended from Llama 3. OpenThinker3-7B was introduced as a new state-of-the-art 7B reasoning model with a 33% improvement over DeepSeek-R1-Distill-Qwen-7B. The STORM text-video model compresses video input by 8x using Mamba layers and outperforms GPT-4o on MVBench with 70.6%. Discussions on reinforcement learning algorithms PPO vs. GRPO and insights on DINOv2's performance on ImageNet-1k were also highlighted. "A very quiet day" in AI news with valuable workshops from OpenAI, Amazon, and GDM.
Grok 3 & 3-mini now API Available
grok-3 grok-3-mini gemini-2.5-flash o3 o4-mini llama-4-maverick gemma-3-27b openai llamaindex google-deepmind epochairesearch goodfireai mechanize agent-development agent-communication cli-tools reinforcement-learning model-evaluation quantization-aware-training model-compression training-compute hybrid-reasoning model-benchmarking
Grok 3 API is now available, including a smaller version called Grok 3 mini, which offers competitive pricing and full reasoning traces. OpenAI released a practical guide for building AI agents, while LlamaIndex supports the Agent2Agent protocol for multi-agent communication. Codex CLI is gaining traction with new features and competition from Aider and Claude Code. GoogleDeepMind launched Gemini 2.5 Flash, a hybrid reasoning model topping the Chatbot Arena leaderboard. OpenAI's o3 and o4-mini models show emergent behaviors from large-scale reinforcement learning. EpochAIResearch updated its methodology, removing Maverick from high FLOP models as Llama 4 Maverick training compute drops. GoodfireAI announced a $50M Series A for its Ember neural programming platform. Mechanize was founded to build virtual work environments and automation benchmarks. GoogleDeepMind's Quantisation Aware Training for Gemma 3 models reduces model size significantly, with open source checkpoints available.
not much happened today
gpt-4.5 claude-3.7-sonnet deepseek-r1 smolagents-codeagent gpt-4o llama-3-8b tinyr1-32b-preview r1-searcher forgetting-transformer nanomoe openai deepseek hugging-face mixture-of-experts reinforcement-learning kv-cache-compression agentic-ai model-distillation attention-mechanisms model-compression minimax model-pretraining andrej-karpathy cwolferesearch aymericroucher teortaxestex jonathanross321 akhaliq
The AI news recap highlights several key developments: nanoMoE, a PyTorch implementation of a mid-sized Mixture-of-Experts (MoE) model inspired by Andrej Karpathy's nanoGPT, enables pretraining on commodity hardware within a week. An agentic leaderboard ranks LLMs powering smolagents CodeAgent, with GPT-4.5 leading, followed by Claude-3.7-Sonnet. Discussions around DeepSeek-R1 emphasize AI model commoditization, with DeepSeek dubbed the "OpenAI of China." Q-Filters offer a training-free method for KV cache compression in autoregressive models, achieving 32x compression with minimal perplexity loss. The PokéChamp minimax language agent, powered by GPT-4o and Llama-3-8b, demonstrates strong performance in Pokémon battles. Other notable models include TinyR1-32B-Preview with Branch-Merge Distillation, R1-Searcher incentivizing search capability via reinforcement learning, and the Forgetting Transformer using a Forget Gate in softmax attention. These advancements reflect ongoing innovation in model architectures, compression, reinforcement learning, and agentic AI.