All tags
Person: "teknuim"
not much happened today
claude-fable-5 muse-image muse-video audex anthropic langchain google meta-ai-fair nvidia cohere weaviate agent-design background-execution task-management human-in-the-loop agentic-generation reinforcement-learning model-scaling moe context-windows audio-processing video-generation image-generation open-source model-release mikeyk kimmonismus lilian_weng sakana _philschmid officiallogank dimillian reach_vb teknuim victorialslocum omarsar0 alexandr_wang _tim_brooks
Anthropic expanded the "background agent" UX with Claude Cowork for mobile and web, emphasizing task-running background teammates. They also extended access to Claude Fable 5 on paid plans. The concept of a harness in agent design gained traction, highlighted by Lilian Weng and echoed by LangChain with a new Deep Agents course and open-source project. Google's Gemini API Managed Agents introduced features like background execution and custom function calling. Operator-facing agent infrastructure saw updates from Codex Mobile iOS, Hermes Agent with 1Password integration, and Weaviate 1.38 enabling runtime-gated write access. Experimentation with human-in-the-loop control via phone/SMS was noted. In model releases, Meta AI launched Muse Image and previewed Muse Video, featuring an agentic generation loop with planning, web search, and self-refinement, achieving top ranks on Image and Video Arena. NVIDIA released Audex, a 30B parameter MoE model with 1M context for unified text and audio tasks.
not much happened today
gpt-5.2-codex gpt-5.3-codex openai langchain baseten ollama openrouter agent-orchestration context-pipelines coding-agents pricing-models multi-agent-systems workflow-optimization model-agnostic-orchestration prompt-engineering memory-optimization anthony_maio mason_drxy hwchase17 sydneyrunkle naroh teknuim vtrivedy dbreunig zachtratar theo petergostev cheatyyyy
AI Twitter Recap highlights the shift from model-centric AI to context pipelines and agent orchestration as key performance drivers. Notably, gpt-5.2-codex and gpt-5.3-codex showed significant benchmark improvements through prompt and middleware tuning. The ecosystem around open harnesses like Hermes, deepagents, and Flue is rapidly evolving, with innovations in multi-agent coordination and model-agnostic orchestration. Developer workflows are adapting to coding agents such as Codex and Claude Code, with emerging challenges in pricing models due to high token usage in agentic workloads. The practical takeaway is that agent performance depends on the synergy of model × harness × memory/context strategy, not just model weights alone.
not much happened today
opus-4.6 glm-5 anthropic ibm perplexity-ai llamaindex deepseek google-chrome persistent-memory agent-infrastructure cross-device-synchronization long-context sparse-attention inference-optimization computer-architecture task-completion systems-performance pamelafox tadasayy llama_index bromann dair_ai omarsar0 abxxai teknuim bcherny kimmonismus _catwu alexalbert__ realyushibai
MCP tools remain relevant for deterministic APIs despite ergonomic criticisms, with new web MCP support in Chrome v146 enabling continuous browsing agents. Persistent memory is emerging as a key differentiator for agents, with IBM improving task completion rates and multi-agent memory framed as a computer architecture challenge. Agent UX is evolving towards always-on, cross-device operation, exemplified by Perplexity Computer on iOS and Claude Code session management. Anthropic released Opus 4.6 1M context as default with no extra long-context API charges, achieving 78.3% on MRCR v2 at 1M tokens. Sparse attention optimizations like IndexCache in DeepSeek Sparse Attention yield significant speedups on large models with minimal code changes.