All tags
Model: "grok-4.5"
not much happened today
kimi-k3 grok-4.5 chatgpt codex moonshot baseten nvidia red-hat-ai perplexity-ai togethercompute cursor_ai mixture-of-experts model-architecture attention-mechanisms reinforcement-learning infrastructure model-deployment agentic-ai mobile-ai multimodality model-distillation gpu-optimization system-design zhihufrontier rasbt bhavinjawade danizeres amansanger
Moonshot released the Kimi K3, a 2.8T-parameter MoE model with 104B active parameters/token, featuring innovations like Kimi Delta Attention (KDA), Gated MLA, and LatentMoE. The release includes infrastructure components such as MoonEP, FlashKDA, and AgentEnv, emphasizing system-level design. Despite open weights, running K3 requires significant hardware investment (minimum 8Ć MI355X GPUs, production at 64+ GPUs) with costs reaching six figures USD or tens of millions RMB. Hosted access is available via Perplexity, Baseten, and Together. Additionally, agent-based workflows are advancing with mobile orchestration, highlighted by ChatGPT Voice + Codex, Cursor's Start in India powered by Grok 4.5, and Perplexity's Personal Computer local agent with multi-model comparison via Model Council. "If you ever want to feel dumb just read the Kimi K3 technical report" captures community reaction to the dense technical details.
not much happened today
gpt-5.6-sol grok-4.5 terra-max fable-5-max opus-4.8 100b-reasoning-model prime-intellect vllm langchain threepointone factory cognition arena artificial-analysis parlance-labs agentic-reinforcement-learning rollout-traces message-dags long-horizon-reinforcement-learning multimodality harness-design cost-per-task coding-agents benchmarks model-efficiency real-world-evaluation task-specialization johannes_hage willccbb mikasenghaas xeophon omarsar0 skirano imjaredz
Prime Intellect released verifiers v1, a redesigned environment stack for agentic reinforcement learning and evaluations, improving efficiency by storing rollout traces as message DAGs to reduce complexity from O(n²) to O(n). This enables practical long-horizon multimodal rollouts, demonstrated with a 100B reasoning model running 40-turn SWE agent tasks on 6 H200 nodes in under 2 days. The ecosystem support includes vLLM integration to avoid tokenization drift. Discussions highlight that harnesses are becoming critical as the product surface for coding agents, with task-specialized harnesses favored over generic wrappers. Benchmarks are shifting focus from token price to cost per task, with models like Terra Max, Fable 5 Max, and Opus 4.8 compared on efficiency and cost. Real-world agent benchmarks show GPT-5.6 Sol ranking #2 and Grok-4.5 jumping to #13 on Arena's leaderboard, emphasizing cost per task as a key metric for long-horizon knowledge work.
not much happened today
grok-4.5 opus-4.7 opus-4.8 gpt-5.6 xai cursor scaling01 coding agents model-scaling context-window model-pricing token-efficiency model-training model-performance elonmusk
xAI publicly launched Grok 4.5, a new coding-and-agents-focused frontier model emphasizing capability-per-dollar rather than benchmark supremacy. Elon Musk described it as "Opus-class" but faster, more token-efficient, and lower cost, with a 1.5 trillion parameter size, making it 3x larger than Grok 4.3. The model is priced at $2 per 1M input tokens and $6 per 1M output tokens, with discounts for cache hits and a context window expected to return to 1 million tokens soon. Cursor partnered in training Grok 4.5, highlighting it as their most powerful model yet and expanding beyond software engineering. Early ecosystem support includes Grok Build/API, Hermes Agent, Portal, and OpenRouter.