All tags
Model: "deepseek-v4-flash-vision-exp"
not much happened today
muse-code deepseek-v4-flash-vision-exp glm-5.3-flash qwen3.8-flash-next hy4-preview meta-ai-fair deepseek google tencent ollama agent-benchmarks agent-infrastructure context-management multi-agent-systems model-releases plugin-systems model-performance finkd alexandr_wang teortaxestex zizhpan arena valsai zhihufrontier teknuim dair_ai
Meta's Muse Code has exited beta with an SDK and subscription plans, enabling embedding custom agents and tool integration. DeepSeek V4 Flash Vision weights were released openly, adding vision parity with other models. GLM-5.3 Flash showed strong agentic cost/performance in benchmarks, ranking #19 overall and #4 among open models with a $0.12 median cost per task. Qwen3.8-Flash-Next also competed but ranked lower. Tencent Hunyuan's Hy4 Preview is a 770B MoE model with 49B active parameters and over 1M context length, showing rapid improvements post Hy3. On infrastructure, Hermes Agent v0.21.0 introduced multi-agent workflow features and improved context efficiency. DeepSeek Harness v0.1.2-alpha updated with breaking changes, highlighting challenges in plugin-heavy agent platforms. Context management is emerging as a key research area with new papers like WikiSkill / SKILL.state from Google and collaborators.
not much happened today
glm-5.3-vision glm-5.2 deepseek-v4-flash-vision-exp opus-4.8 gpt-5.6-sol codex zhipu-ai deepseek-ai openai multimodality post-training agentic-ai api pricing model-efficiency benchmarking inference spend-controls theo kimmonismus tim_dettmers scaling01 teortaxestex zhihufrontier
Ox Alpha emerged as a mystery model with strong coding and agentic performance, likely a Zhipu/GLM-family model such as GLM-5.3 Vision. Analysts suggest its gains come from post-training and infrastructure improvements rather than sheer size, based on the 743B base of GLM-5.2 with enhancements like SAO for long-horizon tasks. DeepSeek released DeepSeek-V4-Flash-Vision-Exp, adding multimodal support and mixed text+image API capabilities, with performance near Opus-4.8. Chinese AI labs are advancing on price/performance and multimodal agents, pressuring US labs. OpenAI cut GPT-5.6 Sol pricing by over 20% for three months and reported explosive Codex usage hitting 20M active users, while adding better spend controls for API usage.