Tag: benchmarks

39 posts tagged with "benchmarks" — Page 2 of 2

Preview image for AGENTS.md Best Practices

Most root-level AGENTS.md files deliver negligible or negative returns for AI coding tools, per 2026 ETH Zurich research. Curated minimal files with only non-inferable rules cut task time by 28% and reduce agent-generated bugs by 35-55%. Avoid bloat, redundant overviews, and stale content to boost performance and lower inference costs.

Preview image for The Velvet Rope Apocalypse: Why GPT-5.6’s 'Trusted Partner' Program is a Declaration of War on the Open Source Community

OpenAI's GPT-5.6 launch restricts frontier model access to a small group of U.S. government-vetted 'trusted partners' under a new dual-track release system. This structure creates a hard barrier for the open source community, blocking independent research, transparent benchmarking, and competitive development of open source AI alternatives.

Preview image for AI Coding Agent Benchmarks: Why Harness Matters Over Model

This guide explains why AI coding agent benchmark scores are often misleading, as the agent harness and scaffolding can shift scores by 10–20 percentage points without changing the underlying model. It provides a critical framework for evaluating benchmark claims, noting that real-world coding agent performance is roughly half of reported leaderboard scores. Engineering teams should prioritize production-representative internal evaluations over vendor-reported benchmark claims when selecting AI.