All tags
Person: "zizhpan"
not much happened today
muse-code deepseek-v4-flash-vision-exp glm-5.3-flash qwen3.8-flash-next hy4-preview meta-ai-fair deepseek google tencent ollama agent-benchmarks agent-infrastructure context-management multi-agent-systems model-releases plugin-systems model-performance finkd alexandr_wang teortaxestex zizhpan arena valsai zhihufrontier teknuim dair_ai
Meta's Muse Code has exited beta with an SDK and subscription plans, enabling embedding custom agents and tool integration. DeepSeek V4 Flash Vision weights were released openly, adding vision parity with other models. GLM-5.3 Flash showed strong agentic cost/performance in benchmarks, ranking #19 overall and #4 among open models with a $0.12 median cost per task. Qwen3.8-Flash-Next also competed but ranked lower. Tencent Hunyuan's Hy4 Preview is a 770B MoE model with 49B active parameters and over 1M context length, showing rapid improvements post Hy3. On infrastructure, Hermes Agent v0.21.0 introduced multi-agent workflow features and improved context efficiency. DeepSeek Harness v0.1.2-alpha updated with breaking changes, highlighting challenges in plugin-heavy agent platforms. Context management is emerging as a key research area with new papers like WikiSkill / SKILL.state from Google and collaborators.
DeepSeek-R1-0528 - Gemini 2.5 Pro-level model, SOTA Open Weights release
deepseek-r1-0528 gemini-2.5-pro qwen-3-8b qwen-3-235b deepseek-ai anthropic meta-ai-fair nvidia alibaba google-deepmind reinforcement-learning benchmarking model-performance open-weights reasoning quantization post-training model-comparison artificialanlys scaling01 cline reach_vb zizhpan andrewyng teortaxestex teknim1 lateinteraction abacaj cognitivecompai awnihannun
DeepSeek R1-0528 marks a significant upgrade, closing the gap with proprietary models like Gemini 2.5 Pro and surpassing benchmarks from Anthropic, Meta, NVIDIA, and Alibaba. This Chinese open-weights model leads in several AI benchmarks, driven by reinforcement learning post-training rather than architecture changes, and demonstrates increased reasoning token usage (23K tokens per question). The China-US AI race intensifies as Chinese labs accelerate innovation through transparency and open research culture. Key benchmarks include AIME 2024, LiveCodeBench, and GPQA Diamond.
not much happened today
deepseek-r1 qwen-2.5 qwen-2.5-max deepseek-v3 deepseek-janus-pro gpt-4 nvidia anthropic openai deepseek huawei vercel bespoke-labs model-merging multimodality reinforcement-learning chain-of-thought gpu-optimization compute-infrastructure compression crypto-api image-generation saranormous zizhpan victormustar omarsar0 markchen90 sakanaailabs reach_vb madiator dain_mclau francoisfleuret garygodchaux arankomatsuzaki id_aa_carmack lavanyasant virattt
Huawei chips are highlighted in a diverse AI news roundup covering NVIDIA's stock rebound, new open music foundation models like Local Suno, and competitive AI models such as Qwen 2.5 Max and Deepseek V3. The release of DeepSeek Janus Pro, a multimodal LLM with image generation capabilities, and advancements in reinforcement learning and chain-of-thought reasoning are noted. Discussions include GPU rebranding with NVIDIA's H6400 GPUs, data center innovations, and enterprise AI applications like crypto APIs in hedge funds. "Deepseek R1's capabilities" and "Qwen 2.5 models added to applications" are key highlights.