All tags
Model: "hy4-preview"
not much happened today
muse-code deepseek-v4-flash-vision-exp glm-5.3-flash qwen3.8-flash-next hy4-preview meta-ai-fair deepseek google tencent ollama agent-benchmarks agent-infrastructure context-management multi-agent-systems model-releases plugin-systems model-performance finkd alexandr_wang teortaxestex zizhpan arena valsai zhihufrontier teknuim dair_ai
Meta's Muse Code has exited beta with an SDK and subscription plans, enabling embedding custom agents and tool integration. DeepSeek V4 Flash Vision weights were released openly, adding vision parity with other models. GLM-5.3 Flash showed strong agentic cost/performance in benchmarks, ranking #19 overall and #4 among open models with a $0.12 median cost per task. Qwen3.8-Flash-Next also competed but ranked lower. Tencent Hunyuan's Hy4 Preview is a 770B MoE model with 49B active parameters and over 1M context length, showing rapid improvements post Hy3. On infrastructure, Hermes Agent v0.21.0 introduced multi-agent workflow features and improved context efficiency. DeepSeek Harness v0.1.2-alpha updated with breaking changes, highlighting challenges in plugin-heavy agent platforms. Context management is emerging as a key research area with new papers like WikiSkill / SKILL.state from Google and collaborators.
not much happened today
glm-5.3 hy4-preview qwen3.8-flash z.ai tencent alibaba vllm_project perplexity-ai agentic-coding cyber-defense model-quantization speculative-decoding moe long-context multimodality benchmarking inference search kimmonismus zixuanli_ yuchenj_uw
Z.ai released the GLM-5.3 open-weight model family, optimized for agentic coding and cyber defense, with impressive specs like 744B total / 40B active parameters, 1M context window, and a 239GB 2-bit variant retaining 81% accuracy. Tencent launched Hy4-preview, a top-tier open-source MoE model with 770B total / 49B active parameters and 1M context, showing strong benchmark performance and innovative serving design. Alibaba introduced Qwen3.8-Flash, a cheaper, long-context MoE with 125B total / 6B active parameters and multimodality, though early user reports noted some stability issues resolved by switching KV cache to BF16. On the systems side, vLLM published a detailed speculative decoding benchmark across multiple models and hardware, emphasizing no one-size-fits-all solution. Additionally, search systems like Perplexity Search are gaining prominence as evaluated subsystems with strong economic and performance metrics. "There is no universal winner" in speculative decoding, highlighting the need for workload-specific tuning.