All tags
Person: "garrytan"
not much happened today
flux-3 flux-mimic qwen-audio-3.0-tts hugging-face black-forest-labs mimicrobotics alibaba open-datasets code-datasets distillation multimodality robotics video-modeling audio-generation tts model-training model-architecture anton_lozhkov loubnabenallal1 lvwerra eliebakouch gergelyorosz schmidhuberai suhail garrytan bfl_ai hila_chefer robrombach mimicrobotics generalistai alibaba_qwen
The Stack v3 is released as the largest open code dataset with 114 TB raw data, 224M repositories, and 5T deduplicated tokens, significantly expanding data for open code models and cyber-defense. The debate on distillation continues as a key ideological fault line, with calls for stronger investment in open-weight domestic models. Black Forest Labs launched FLUX 3, a unified multimodal model covering image, video, audio, and action prediction, with robotics transfer demonstrated by FLUX-mimic for general-purpose dexterity on a single GPU. Alibaba introduced Qwen-Audio-3.0-TTS supporting 16 languages and advanced control features, claiming the top spot on the Artificial Analysis TTS leaderboard.
not much happened today
muse-spark llama-4-maverick glm-5.1 deepseek-v3.2 meta-ai-fair zhipu-ai deepseek multimodality tool-use visual-chain-of-thought multi-agent-systems training-efficiency test-time-scaling parallel-inference image-to-code model-benchmarking model-architecture alexandr_wang shengjia_zhao jack_w_rae ananyaku _jasonwei artificialanlys valsai epochairesearch scale_ai matthuang omarsar0 skirano mattdeitke garrytan sebastian_raschka
Meta Superintelligence Labs launched Muse Spark, a natively multimodal reasoning model featuring tool use, visual chain of thought, and multi-agent orchestration. It is live on meta.ai and the Meta AI app with a private API preview and plans for open-sourcing future versions. Independent benchmarks rank Muse Spark highly, with strong performance on intelligence indices and efficiency, notably using over 10× less compute than Llama 4 Maverick. Key technical highlights include training efficiency, test-time scaling, and parallel multi-agent inference. Community testing shows strengths in image-to-code and one-shot game generation. Additionally, Zhipu AI's GLM-5.1 is recognized as a leading open-weight model with architecture similar to DeepSeek-V3.2.