All tags
Topic: "cyber-defense"
not much happened today
glm-5.3 hy4-preview qwen3.8-flash z.ai tencent alibaba vllm_project perplexity-ai agentic-coding cyber-defense model-quantization speculative-decoding moe long-context multimodality benchmarking inference search kimmonismus zixuanli_ yuchenj_uw
Z.ai released the GLM-5.3 open-weight model family, optimized for agentic coding and cyber defense, with impressive specs like 744B total / 40B active parameters, 1M context window, and a 239GB 2-bit variant retaining 81% accuracy. Tencent launched Hy4-preview, a top-tier open-source MoE model with 770B total / 49B active parameters and 1M context, showing strong benchmark performance and innovative serving design. Alibaba introduced Qwen3.8-Flash, a cheaper, long-context MoE with 125B total / 6B active parameters and multimodality, though early user reports noted some stability issues resolved by switching KV cache to BF16. On the systems side, vLLM published a detailed speculative decoding benchmark across multiple models and hardware, emphasizing no one-size-fits-all solution. Additionally, search systems like Perplexity Search are gaining prominence as evaluated subsystems with strong economic and performance metrics. "There is no universal winner" in speculative decoding, highlighting the need for workload-specific tuning.