All tags
Person: "sebastian_raschka"
not much happened today
deepseek-v4.1-flash glm-5.3-flash deepseek baseten ollama causal-encoder-decoder inference-efficiency model-architecture multimodality model-optimization vision model-quantization model-compression context-windows sebastian_raschka
DeepSeek launched V4.1-Flash, a new open-weight flagship model focused on extreme inference efficiency and low cost, featuring a 763B total-parameter causal encoder-decoder architecture with 8B active input and 16B active output parameters and 1M-token context. It scored 40 on the Artificial Analysis Intelligence Index, outperforming its predecessor and ranking just below GLM-5.3-Flash. The model supports text and image input, is available under an MIT license, and is accessible via US/API. The architecture introduces a novel causal encoder-decoder design aimed at reducing active compute and KV/cache costs, with a hybrid sparse/local approach and a unique vision encoder differing from recent Chinese models. Early layers use a SWA-only pattern, and the model has an effective depth of about 40 layers with 20 decoder layers. Baseten and Ollama have begun supporting and rolling out the model to users.
not much happened today
muse-spark llama-4-maverick glm-5.1 deepseek-v3.2 meta-ai-fair zhipu-ai deepseek multimodality tool-use visual-chain-of-thought multi-agent-systems training-efficiency test-time-scaling parallel-inference image-to-code model-benchmarking model-architecture alexandr_wang shengjia_zhao jack_w_rae ananyaku _jasonwei artificialanlys valsai epochairesearch scale_ai matthuang omarsar0 skirano mattdeitke garrytan sebastian_raschka
Meta Superintelligence Labs launched Muse Spark, a natively multimodal reasoning model featuring tool use, visual chain of thought, and multi-agent orchestration. It is live on meta.ai and the Meta AI app with a private API preview and plans for open-sourcing future versions. Independent benchmarks rank Muse Spark highly, with strong performance on intelligence indices and efficiency, notably using over 10× less compute than Llama 4 Maverick. Key technical highlights include training efficiency, test-time scaling, and parallel multi-agent inference. Community testing shows strengths in image-to-code and one-shot game generation. Additionally, Zhipu AI's GLM-5.1 is recognized as a leading open-weight model with architecture similar to DeepSeek-V3.2.