All tags
Topic: "plugin-architecture"
not much happened today
ornith-1.5 qwen3.8-27b claude-opus-5 kimi-k3 glm-5.2 grok-4.5 gpt-5.6-luna grok-4.6 glm-5.3 trueforge ornith vllm ollama unsloth qwen arena valsai deepseek truefoundry claude model-compression quantization reinforcement-learning agent-evaluation plugin-architecture open-agent-runtime cost-efficiency session-management tooling benchmarking ornith_ unslothai danielhanchen arena valsai zhihufrontier theturingpost truefoundry omarsar0 kimmonismus bradenjhancock dbreunig rseroter claudedevs
Ornith-1.5 launches as a new open-weight model family with 9B dense, 35B MoE, and 397B MoE variants under MIT license, featuring quantized formats like FP8, GGUF, MLX, and NVFP4 and showcasing end-to-end self-improvement capabilities. Compression techniques improve accuracy and efficiency, with Qwen3.8-27B GGUFs using Dynamic V3 achieving 10% higher accuracy and 1-bit quantization retaining 77% BF16 accuracy on 8GB RAM. Agent evaluation boards highlight models like Claude Opus 5 (High), Kimi K3, GLM 5.2, Grok 4.5, and GPT-5.6 Luna leading in quality and value. DeepSeek Harness (DSH) introduces a plugin-based open agent runtime architecture optimized for extensibility and tooling. TrueFoundry open-sources TrueForge, a self-hostable, vendor-neutral agent harness that reduces token usage by 30% and cuts costs by 75% while maintaining accuracy, emphasizing the growing importance of session, environment, memory, and tools layers in agent platforms.
not much happened today
gpt-5.4-mini gpt-5.4-nano gpt-5.4 codex openai langchain stripe ramp coinbase nous-research hermes-agent coding multimodality subagents context-window model-performance pricing behavior-tuning secure-execution plugin-architecture attention-mechanisms agent-infrastructure hwchase17 michpokrass
OpenAI released GPT-5.4 mini and GPT-5.4 nano, their most capable small models optimized for coding, multimodal understanding, and subagents, featuring a 400k context window and over 2x speed compared to GPT-5 mini. The mini model approaches larger GPT-5.4 performance while using only 30% of Codex quota, becoming the default for many coding workflows. Pricing concerns and truthfulness tradeoffs were noted, with mixed third-party evaluations on reasoning and resistance to false premises. OpenAI also addressed behavior tuning issues in a recent update. Meanwhile, agent infrastructure is evolving with secure code execution and orchestration tools like LangChain's LangSmith Sandboxes and Open SWE, inspired by internal systems at Stripe, Ramp, and Coinbase. Subagents and secure execution are now key product features, with releases like Hermes Agent v0.3.0 showcasing plugin architectures, live Chrome control, and voice mode. Research on attention mechanisms, including Attention Residuals and vertical attention, is gaining traction.