All tags
Topic: "persistent-agents"
not much happened today
qwen3.8-27b carnice-v3-27b claude-melon-eap claude-marshmallow-eap qwen-4 gpt-astra nvidia anthropic agent-harness persistent-agents self-modifying-agents enterprise-infrastructure skill-lift open-source model-leaks pre-release-access model-benchmarking long-running-workloads rollback durability self-debugging fine-tuning omarsar0 dair_ai andykonwinski claudedevs _philschmid kaiostephens lentils80 kimmonismus eliebakouch
Agent harnesses are becoming a key optimization focus, with NVIDIA research showing traditional skill checks poorly predict agent usefulness and proposing a new metric called "Skill Lift". Open-source implementations of persistent and self-modifying agents like Headlong and exo emphasize durability features such as rollback and continuous operation. Anthropic advances enterprise infrastructure with MCP connectors featuring managed auth and support for long-running workloads. In model releases, Qwen3.8-27B ranks highly in Code Arena: WebDev, and open-source derivatives like Carnice-V3-27B target consumer GPUs. Rumors swirl around unreleased frontier models including claude-melon-eap, claude-marshmallow-eap, Ox Alpha, Qwen 4, and GPT Astra, highlighting pre-release access asymmetry in the ecosystem.
GDM leadership reset
gemini muse-spark-1.2 muse-code claude-code codex google-deepmind alphabet discovery-loop radical-ventures khosla-ventures lightspeed kleiner-perkins doerr-capital meta-ai-fair artificial-analysis automated-discovery machine-learning coding-agents model-harness-co-design benchmarking public-benefit-corporation venture-capital long-context parallel-computing persistent-agents demis-hassabis koray-kavukcuoglu jeff-dean sanjay-ghemawat oriol-vinyals quoc-le nat-friedman nathan-lambert andrew-ng alexandr-wang fink
Google DeepMind undergoes a leadership reshuffle with Demis Hassabis moving to Chair and Chief Scientist roles, while Koray Kavukcuoglu takes operational control focusing on Gemini and product execution. The launch of Discovery Loop by founders including Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, and Quoc Le targets automated machine learning and scientific discovery, backed by major venture firms. Meta AI releases Muse Spark 1.2 and Muse Code (beta), co-trained model and harness for coding agents, achieving strong benchmark scores and emphasizing harness-model co-design, entering the coding-agent competition alongside systems like Claude Code and Codex. The market views these moves as pivotal for AI-for-science and coding agent development.
not much happened today
nemotron-3-super gpt-oss-120b qwen3.5-122b-a10b nvidia perplexity replit base44 vllm llama.cpp ollama togethercompute baseten wandb langchain unsloth model-architecture model-optimization inference-speed kv-cache multi-token-prediction agent-infrastructure orchestration persistent-agents model-serving product-launches karpathy ctnzr bnjmn_marie artificialanlys
NVIDIA’s Nemotron 3 Super is a 120B parameter / ~12B active open model featuring a hybrid Mamba-Transformer / SSM Latent MoE architecture and 1M context window, delivering up to 2.2x faster inference than GPT-OSS-120B in FP4 with strong throughput gains. It supports agentic workloads and is unusually open with weights, data, and infrastructure details released. The model scored 36 on the AA Intelligence Index, outperforming GPT-OSS-120B but behind Qwen3.5-122B-A10B. Community and infrastructure support from projects like vLLM, llama.cpp, Ollama, Together, Baseten, W&B Inference, LangChain, and Unsloth GGUFs was immediate. Key technical innovations include native multi-token prediction (MTP) and a significant KV-cache efficiency advantage.
On the product side, a shift towards persistent agent runtimes and orchestration layers is highlighted, with Andrej Karpathy advocating for a "bigger IDE" concept where agents replace files as the unit of work, enabling legible, forkable agentic organizations with real-time control. New launches fitting this vision include Perplexity’s Personal Computer, an always-on local/cloud hybrid running on Mac mini, and Computer for Enterprise orchestrating 20 specialized models and 400+ apps. Replit Agent 4 offers a collaborative, canvas-like workflow with parallel agents, while Base44 Superagents provide integrated solutions for nontechnical users. The engineering focus is increasingly on the orchestration harness rather than just the model.