All tags
Person: "zixuan_li"
not much happened today
glm-5.3-flash glm-5.2 claude-3-opus z.ai huggingface coreweave baseten multimodality context-window model-benchmarking model-performance coding vision open-source api model-distribution rasbt zixuan_li cline
Z.ai launched GLM-5.3-Flash, a natively multimodal model with a 1M-token context window, 320B total parameters / 18B active parameters, under the MIT License. It is positioned as a price-competitive successor to GLM-5.2 and claims performance on par with Claude Opus 4.8 on coding tasks. The model is available via weights on Hugging Face, API, chat, coding plan, and AutoClaw, and runs entirely on Chinese AI chips. Early third-party support includes CoreWeave and Baseten. Independent evaluation by Artificial Analysis scored GLM-5.3-Flash 57 on their Intelligence Index. Community reactions highlight its potential as a best intelligence-per-dollar option, though some critique its vision capabilities.
not much happened today
glm-5.1 gemini-3.1 gpt-5.4 claude-3-sonnet haiku opus sonnet qwen-3.6-plus qwen3-coder-next-80b z-ai anthropic berkeley langchain alibaba openai model-performance agent-frameworks orchestration model-routing fine-tuning agent-harness model-selection workflow-automation zixuan_li akshay_pachaar harrison_chase walden_yan yuchen_jin sentdex
GLM-5.1 has reached #3 on Code Arena, surpassing Gemini 3.1 and GPT-5.4, and matching Claude Sonnet 4.6 in coding performance. Z.ai now holds the #1 open model rank close to the top overall. The advisor pattern, combining a cheap executor with an expensive advisor, is gaining traction, improving performance and efficiency in models like Haiku + Opus and Sonnet + Opus. Alibaba's Qwen Code v0.14.x introduces orchestration features including remote control channels, cron tasks, and sub-agent model selection. Model routing is becoming a product-level concern due to specialization and spikiness in top models such as Opus and GPT-5.4. The Hermes Agent ecosystem shows strong momentum with a new workspace mobile app, FAST mode for OpenAI/GPT-5.4, and over 50k GitHub stars. Practitioners report Hermes as a reliable agent framework, with local Qwen3-Coder-Next 80B 4-bit replacing parts of workflows previously reliant on Claude Code. The harness layer is emerging as a key abstraction in agent frameworks.