All tags
Model: "glm-4.7-flash"
not much happened today
gpt-5.6 codex bonsai-27b qwen-3.6-27b hy3-295b gemma-4 qwen3.5-122b-a10b glm-4.7-flash deepseek-v4-flash mimo-v2.5 glm-5.2-nvfp4 moss-vl-realtime openai jetbrains langchain prismml tencent-hunyuan miaai_lab openmoss agentic-ai model-quantization local-inference multimodality video-understanding model-compression evals observability long-context tool-use sama reach_vb kimmonismus swyx theo andykonwinski
OpenAI's agent products saw a 2.5x weekly usage growth driven by Codex + ChatGPT Work and demand for GPT-5.6 Sol. JetBrains adopted Codex as a recommended agent, while LangChain enhanced tracing and observability across multiple tools. PrismML released Bonsai 27B, a compressed variant of Qwen 3.6 27B enabling local multimodal agentic workflows on consumer devices. Tencent Hunyuan introduced 1-bit and 4-bit quantized Hy3 295B model deployable on a single GPU. Quantization advances like NVFP4 dynamic quants for Gemma-4 and others support serious local inference. OpenMOSS launched MOSS-VL-Realtime 11B for continuous video stream perception with a 256K context window. "Harness quality and observability are becoming a first-class differentiator" and local inference is now viable for agentic workflows.
not much happened today
glm-4.7-flash grok deepseek-r1 qwq x-ai unsloth-ai google deepseek ollama transformer-architecture recommendation-systems local-inference kv-cache quantization tensor-parallelism reasoning model-optimization fine-tuning giffmana david_sholz yuchenj_uw nearcyan sam_paech teortaxes_tex danielhanchen alexocheema nopmobiel rohanpaul_ai
X Engineering open-sourced its new transformer-based recommender algorithm, sparking community debate on transparency and fairness. GLM-4.7-Flash (30B-A3B) gains momentum as a strong local inference model with efficient KV-cache management and quantization tuning strategies. Innovations include tensor parallelism on Mac Minis achieving ~100 tok/s throughput. Research highlights "Societies of Thought" as a reasoning mechanism improving model accuracy by 20%+.
not much happened today
glm-4.7-flash glm-4.7 glm-4.5 qwen3-vl qwen meta-ai-fair carnegie-mellon sakana-ai zhipu-ai transformer-memory model-architecture mixture-of-experts adaptive-position-encoding long-context model-compression inference-optimization local-inference model-deployment benchmarking coding agentic-ai
AI News for 1/16/2026-1/19/2026 covers new architectures for scaling Transformer memory and context, including STEM from Carnegie Mellon and Meta AI, which replaces part of the FFN with a token-indexed embedding lookup enabling CPU offload and asynchronous prefetch. RePo from Sakana AI introduces adaptive positional reordering to improve robustness on noisy and long-range contexts. Model releases highlight Zhipu AI's GLM-4.7-Flash, a 30B-class MLA + small MoE model optimized for coding and agentic tasks, noted for strong benchmark performance and a compression narrative from larger to smaller models. Inference and deployment updates include mlx-lm 0.30.3 supporting GLM-4.7-Flash with efficient 4-bit performance on laptops. The report emphasizes practical takeaways on static sparsity, adaptive ordering, and the resurgence of small, fast models for interactive tasks. "Sparse capacity doesnโt have to mean MoE routers + expert parallelism; static sparsity can be systems-friendly."