All tags
Company: "fireworks"
not much happened today
glm-5.3 qwen3.8-27b qwen3.8-2.4t-a95b deepseek-v4-pro dots3-note z-ai alibaba deepseek rednote vllm together-ai fireworks modal digitalocean deepinfra unsloth post-training reinforcement-learning agent-runtimes long-horizon-training multimodality model-infrastructure runtime-architecture model-benchmarking open-weight apache-2.0-license model-optimization multimodal-models mixture-of-experts context-windows
Z.ai launched GLM-5.3, a coding- and cyber-focused model with significant gains on agentic and security benchmarks, achieved through scaled post-training rather than a larger base model. Alibaba released Qwen3.8-27B, a native multimodal dense model under Apache 2.0 with a 262K native context extendable to 1M, designed for real-world coding and office workflows, with broad inference support from multiple platforms. DeepSeek V4-Pro and RedNote's dots3-note, a 280B multimodal MoE model with 16B active parameters and 512K context, continue the China open-model wave, introducing new RL methods like TEMPO for long-horizon self-evaluation. The ecosystem features multiple Chinese labs specializing in open models with different strengths. DeepSeek's harness is highlighted as a modular agent runtime infrastructure with replaceable components and lifecycle management via Cordis.
GLM 5.2: the top Frontend Coding model in the world, IndexShare reduces costs
glm-5.2 z.ai lmsys deepseek cloudflare openrouter ollama baseten deepinfra fireworks notion coding agentic-ai long-context mixture-of-experts sparse-attention speculative-decoding multi-token-prediction model-benchmarking inference-optimization mervenoyann sentdex scaling01 omarsar0 teortaxestex
Z.ai released GLM-5.2, an MIT-licensed open-weight frontier model targeting coding and long-horizon agentic tasks with a 1M-token context window and two reasoning-effort modes. It features a 744B-parameter mixture-of-experts architecture with 40B active parameters per token, built on DeepSeek Sparse Attention extended by IndexShare, and supports improved multi-token prediction (MTP) for speculative decoding. The model achieved strong leaderboard placements, including #3 on FrontierSWE, #1 on Design Arena, and #1 open model on Agent Arena, with ecosystem support from platforms like Transformers, vLLM, SGLang, Cloudflare Workers AI, OpenRouter, Ollama Cloud, Baseten, DeepInfra, Fireworks, and Notion. Early testers praised its potential as a substitute for Opus/GPT-class workflows, though some called for further evaluation and long-horizon validation.
not much happened today
vllm-0.20.0 poolside-laguna-xs.2 ling-2.6-flash nemotron-3-nano-omni qwen-3.5 vllm poolside nvidia opensrouter lmstudio ollama unsloth fal fireworks deepinfra togethercompute baseten canonical memory-optimization mixture-of-experts model-optimization inference-speed quantization model-deployment multimodality hardware-optimization model-benchmarking open-models agentic-ai jeremyphoward maharshii teortaxestex aymericroucher piotrz
vLLM v0.20.0 introduces significant improvements in memory and MoE serving efficiency, including TurboQuant 2-bit KV cache for 4× KV capacity and a 2.1% latency improvement. The update supports multiple hardware platforms like DeepSeek V4 MegaMoE on Blackwell, Jetson Thor, ROCm, Intel XPU, and Grace-Blackwell setups. Early benchmarks show DeepSeek V4 Pro on B300 hardware can be up to 8× faster than H200. The ecosystem is rapidly adopting day-0 support for new open models such as Poolside Laguna XS.2, Ling-2.6-flash, and NVIDIA Nemotron 3 Nano Omni.
Poolside released Laguna XS.2, a 33B total / 3B active MoE coding model under Apache 2.0, capable of running on a single GPU, with hybrid attention and FP8 KV cache, performing near Qwen-3.5.
NVIDIA launched Nemotron 3 Nano Omni, a 30B / A3B multimodal MoE with 256K context, supporting text, image, video, audio, and documents, with immediate distribution across multiple platforms. Discussions highlighted tradeoffs in quantization methods and a shift away from CUDA lock-in towards heterogeneous accelerator support.
not much happened today
kimi-k2.5 claude-code cursor kimi fireworks anthropic langchain model-attribution fine-tuning reinforcement-learning open-source agent-products model-licensing software-integration product-differentiation clementdelangue leerob amanrsanger yuchenj_uw kimmonismus
Cursor's Composer 2, built on Kimi K2.5, sparked discussion over model attribution and licensing, highlighting a shift toward post-trained derivatives of open-source models with domain-specific fine-tuning and reinforcement learning. Claude Code is expanding into third-party tools like T3 Code and communication channels such as Telegram and Discord, while LangChain is evolving from orchestration to multi-agent products with offerings like Deep Agents/Open SWE and LangSmith Fleet. The discourse emphasizes the importance of clear base-model attribution, licensing compliance, and product differentiation through fine-tuning and user experience.
Llama 3.1: The Synthetic Data Model
llama-3-405b llama-3-1 llama-3 meta-ai-fair groq fireworks synthetic-data fine-tuning reinforcement-learning multilinguality long-context tool-use code-generation math model-licensing inference-speed model-deployment bindureddy thomas
Meta AI has released Llama 3.1, including a 405B parameter model that triggers regulatory considerations like the EU AI Act and SB 1047. The model incorporates extensive synthetic data techniques for code, math, multilinguality, long context, and tool use fine-tuning, with RLHF using synthetic preference data from Llama 2. The launch was coordinated across major inference providers, with Groq demonstrating 750 tokens per second inference speed and Fireworks leading in pricing. The updated license explicitly allows synthetic data generation, marking a significant step in open frontier-class LLMs and cost-efficiency improvements since March.