All tags
Topic: "agent-runtimes"
not much happened today
glm-5.3 qwen3.8-27b qwen3.8-2.4t-a95b deepseek-v4-pro dots3-note z-ai alibaba deepseek rednote vllm together-ai fireworks modal digitalocean deepinfra unsloth post-training reinforcement-learning agent-runtimes long-horizon-training multimodality model-infrastructure runtime-architecture model-benchmarking open-weight apache-2.0-license model-optimization multimodal-models mixture-of-experts context-windows
Z.ai launched GLM-5.3, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger base model. Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0 with a 262K native context extendable to 1M, designed for real-world coding and office workflows, with broad inference support from multiple platforms. DeepSeek V4-Pro and RedNote's dots3-note, a 280B multimodal MoE model with 16B active parameters and 512K context, continue the China open-model wave, introducing new RL methods like TEMPO for long-horizon self-evaluation. The ecosystem features multiple Chinese labs specializing in open models with different strengths. DeepSeek's harness is highlighted as a modular agent runtime infrastructure with replaceable components and lifecycle management via Cordis.
not much happened today
claude-code composer-2 cursor openai anthropic langchain cognition reinforcement-learning developer-tooling agent-systems agent-runtimes security credential-management multi-agent-systems model-training benchmarking software-engineering enterprise-ai kimmonismus mntruell theo ellev3n11 amanrsanger charliermarsh gdb yuchenj_uw neilhtennek simonw yuvalinthedeep lvwerra hrishioa
Cursor launched Composer 2, a frontier-class coding model with major cost reductions and strong benchmark scores like 61.3 on CursorBench and 73.7 on SWE-bench Multilingual. The model was improved via a first continued pretraining run feeding into reinforcement learning, trained across 3–4 clusters worldwide by a ~40-person team. OpenAI acquired Astral, the team behind Python tools uv, ruff, and ty, strengthening its developer platform. Anthropic expanded Claude Code with messaging app channels for persistent developer workflows. The focus in AI agents is shifting from single agents to managed fleets and runtimes, with LangChain launching LangSmith Fleet for enterprise agent management emphasizing agent identity, credential management, and auditability. Other launches include Cognition's teams of Devins, AgentUI by lvwerra, and discussions on agent runtimes with features like checkpointing and rollback. Security and permissions are emerging as critical constraints in agent system design.