All tags
Person: "finkd"
not much happened today
muse-code deepseek-v4-flash-vision-exp glm-5.3-flash qwen3.8-flash-next hy4-preview meta-ai-fair deepseek google tencent ollama agent-benchmarks agent-infrastructure context-management multi-agent-systems model-releases plugin-systems model-performance finkd alexandr_wang teortaxestex zizhpan arena valsai zhihufrontier teknuim dair_ai
Meta's Muse Code has exited beta with an SDK and subscription plans, enabling embedding custom agents and tool integration. DeepSeek V4 Flash Vision weights were released openly, adding vision parity with other models. GLM-5.3 Flash showed strong agentic cost/performance in benchmarks, ranking #19 overall and #4 among open models with a $0.12 median cost per task. Qwen3.8-Flash-Next also competed but ranked lower. Tencent Hunyuan's Hy4 Preview is a 770B MoE model with 49B active parameters and over 1M context length, showing rapid improvements post Hy3. On infrastructure, Hermes Agent v0.21.0 introduced multi-agent workflow features and improved context efficiency. DeepSeek Harness v0.1.2-alpha updated with breaking changes, highlighting challenges in plugin-heavy agent platforms. Context management is emerging as a key research area with new papers like WikiSkill / SKILL.state from Google and collaborators.
not much happened today
muse-glimmer muse-spark-1.2 claude claude-3 meta-ai-fair anthropic openai together-ai hugging-face ollama quantization agentic-ai multimodality model-architecture model-optimization long-context local-deployment benchmarking theorem-proving proof-assistance ai-assisted-reasoning finkd alexandr_wang jarredsumner jdlichtman
Meta re-enters the open-weight frontier with the release of Muse Glimmer, a 30B dense, multimodal, agent-focused model under Apache 2.0, optimized for always-on local agents and consumer hardware. It features quantization to keep the model under 20GB, a lightweight DFlash drafter for faster on-device generation, and architectural innovations like Gemma 4-style hybrid attention and scale-free QK norm. Benchmarks place Muse Glimmer at 35 on the Intelligence Index, notable for local self-hosting with ~60GB BF16, ~18GB 4-bit, and 128K context. Immediate ecosystem support includes vLLM, llama.cpp, Ollama, Together AI, and Hugging Face transformers. Meanwhile, Anthropic's unreleased Claude variant improved a Riemann Hypothesis bound from 41.6% to 67.2% using over 31M output tokens, showcasing AI-assisted theorem search and proof iteration.