All tags
Person: "liam_fedus"
not much happened today
jalapeno gpt-astra codex openai microsoft nvidia inference-optimization hardware-efficiency agent-systems benchmarking software-engineering kernel-optimization model-assistance latency-reduction power-efficiency sama liam_fedus omarsar0 kimmonismus eliebakouch
OpenAI announced benchmark results for its custom inference chip Jalapeรฑo, showing 1.5โ1.9ร better efficiency and 1.7โ3.6ร lower latency compared to NVIDIA GB200/GB300. Deployment starts by year-end with Gen 2 and Gen 3 in development. The chip runs at 700W but stayed below 550W in tests. Model-assisted kernel optimization using GPT-Astra + Codex improved performance by 1.5โ1.8ร. This signals a shift in inference stack economics, potentially reducing NVIDIA's dominance. Additionally, research on agent harnesses like AutoSaddler shows system-level improvements can surpass model changes, with significant gains on benchmarks like GAIA2 and SWE-Bench Pro. A new Harness Card standard is proposed to disclose harness variance, highlighting the importance of software engineering in AI agent performance.