All tags
Topic: "mobile-ai"
not much happened today
kimi-k3 grok-4.5 chatgpt codex moonshot baseten nvidia red-hat-ai perplexity-ai togethercompute cursor_ai mixture-of-experts model-architecture attention-mechanisms reinforcement-learning infrastructure model-deployment agentic-ai mobile-ai multimodality model-distillation gpu-optimization system-design zhihufrontier rasbt bhavinjawade danizeres amansanger
Moonshot released the Kimi K3, a 2.8T-parameter MoE model with 104B active parameters/token, featuring innovations like Kimi Delta Attention (KDA), Gated MLA, and LatentMoE. The release includes infrastructure components such as MoonEP, FlashKDA, and AgentEnv, emphasizing system-level design. Despite open weights, running K3 requires significant hardware investment (minimum 8× MI355X GPUs, production at 64+ GPUs) with costs reaching six figures USD or tens of millions RMB. Hosted access is available via Perplexity, Baseten, and Together. Additionally, agent-based workflows are advancing with mobile orchestration, highlighted by ChatGPT Voice + Codex, Cursor's Start in India powered by Grok 4.5, and Perplexity's Personal Computer local agent with multi-model comparison via Model Council. "If you ever want to feel dumb just read the Kimi K3 technical report" captures community reaction to the dense technical details.
not much happened today
chatgpt-4o deepseek-r1 o3 o3-mini gemini-2-flash qwen-2.5 qwen-0.5b hugging-face openai perplexity-ai deepseek-ai gemini qwen metr_evals reasoning benchmarking model-performance prompt-engineering model-optimization model-deployment small-language-models mobile-ai ai-agents speed-optimization _akhaliq aravsrinivas lmarena_ai omarsar0 risingsayak
Smolagents library by Huggingface continues trending. ChatGPT-4o latest version
chatgpt-40-latest-20250129 released. DeepSeek R1 671B sets speed record at 198 t/s, fastest reasoning model, recommended with specific prompt settings. Perplexity Deep Research outperforms models like Gemini Thinking, o3-mini, and DeepSeek-R1 on Humanity's Last Exam benchmark with 21.1% score and 93.9% accuracy on SimpleQA. ChatGPT-4o ranks #1 on Arena leaderboard in multiple categories except math. OpenAI's o3 model powers Deep Research tool for ChatGPT Pro users. Gemini 2 Flash and Qwen 2.5 models support LLMGrading verifier. Qwen 2.5 models added to PocketPal app. MLX shows small LLMs like Qwen 0.5B generate tokens at high speed on M4 Max and iPhone 16 Pro. Gemini Flash 2.0 leads new AI agent leaderboard. DeepSeek R1 is most liked on Hugging Face with over 10 million downloads.