All tags
Person: "andykonwinski"
not much happened today
qwen3.8-27b carnice-v3-27b claude-melon-eap claude-marshmallow-eap qwen-4 gpt-astra nvidia anthropic agent-harness persistent-agents self-modifying-agents enterprise-infrastructure skill-lift open-source model-leaks pre-release-access model-benchmarking long-running-workloads rollback durability self-debugging fine-tuning omarsar0 dair_ai andykonwinski claudedevs _philschmid kaiostephens lentils80 kimmonismus eliebakouch
Agent harnesses are becoming a key optimization focus, with NVIDIA research showing traditional skill checks poorly predict agent usefulness and proposing a new metric called "Skill Lift". Open-source implementations of persistent and self-modifying agents like Headlong and exo emphasize durability features such as rollback and continuous operation. Anthropic advances enterprise infrastructure with MCP connectors featuring managed auth and support for long-running workloads. In model releases, Qwen3.8-27B ranks highly in Code Arena: WebDev, and open-source derivatives like Carnice-V3-27B target consumer GPUs. Rumors swirl around unreleased frontier models including claude-melon-eap, claude-marshmallow-eap, Ox Alpha, Qwen 4, and GPT Astra, highlighting pre-release access asymmetry in the ecosystem.
not much happened today
gpt-5.6 codex bonsai-27b qwen-3.6-27b hy3-295b gemma-4 qwen3.5-122b-a10b glm-4.7-flash deepseek-v4-flash mimo-v2.5 glm-5.2-nvfp4 moss-vl-realtime openai jetbrains langchain prismml tencent-hunyuan miaai_lab openmoss agentic-ai model-quantization local-inference multimodality video-understanding model-compression evals observability long-context tool-use sama reach_vb kimmonismus swyx theo andykonwinski
OpenAI's agent products saw a 2.5x weekly usage growth driven by Codex + ChatGPT Work and demand for GPT-5.6 Sol. JetBrains adopted Codex as a recommended agent, while LangChain enhanced tracing and observability across multiple tools. PrismML released Bonsai 27B, a compressed variant of Qwen 3.6 27B enabling local multimodal agentic workflows on consumer devices. Tencent Hunyuan introduced 1-bit and 4-bit quantized Hy3 295B model deployable on a single GPU. Quantization advances like NVFP4 dynamic quants for Gemma-4 and others support serious local inference. OpenMOSS launched MOSS-VL-Realtime 11B for continuous video stream perception with a 256K context window. "Harness quality and observability are becoming a first-class differentiator" and local inference is now viable for agentic workflows.
not much happened today
arc-agi-3 claude-code anthropic langchain arcprize primeintellect agentic-reasoning interactive-environments benchmarking efficiency-metrics zero-preparation-generalization agent-infrastructure trainable-agents classifier-approval fchollet mikeknoop scaling01 _rockt mark_k andykonwinski bradenjhancock jeremyphoward togelius bracesproul hwchase17 caspar_br _catwu
ARC-AGI-3 benchmark introduced by @arcprize and François Chollet resets the frontier for general agentic reasoning with humans solving 100% of tasks versus under 1% for current models, focusing on zero-preparation generalization and human-like learning efficiency. The scoring protocol sparked debate over its harsh efficiency-based metric compared to prior ARC versions and other benchmarks like NetHack. The community acknowledges the benchmark highlights weaknesses in current LLM agents in interactive, sparse-feedback environments. Concurrently, agent infrastructure advances with LangChain launching Fleet shareable skills for reusable domain knowledge, and Anthropic revealing Claude Code auto mode for classifier-mediated approval balancing autonomy and manual confirmation. Browser and coding agents are evolving into trainable systems beyond prompt wrappers, exemplified by BrowserBase and Prime Intellect collaboration.