All tags
Model: "deepseek-v4.1-flash"
not much happened today
deepseek-v4.1-flash glm-5.3-flash deepseek baseten ollama causal-encoder-decoder inference-efficiency model-architecture multimodality model-optimization vision model-quantization model-compression context-windows sebastian_raschka
DeepSeek launched V4.1-Flash, a new open-weight flagship model focused on extreme inference efficiency and low cost, featuring a 763B total-parameter causal encoder-decoder architecture with 8B active input and 16B active output parameters and 1M-token context. It scored 40 on the Artificial Analysis Intelligence Index, outperforming its predecessor and ranking just below GLM-5.3-Flash. The model supports text and image input, is available under an MIT license, and is accessible via US/API. The architecture introduces a novel causal encoder-decoder design aimed at reducing active compute and KV/cache costs, with a hybrid sparse/local approach and a unique vision encoder differing from recent Chinese models. Early layers use a SWA-only pattern, and the model has an effective depth of about 40 layers with 20 decoder layers. Baseten and Ollama have begun supporting and rolling out the model to users.