Agent Evals Are the New Unit Tests — And Yours Are Flaky
Evals are the only thing standing between your agent and production, and most of them fail for reasons that have nothing to do with the agent. Here is how I stopped trusting a green run.
Diego Vallejo
Analysis and reflections on technology, language, and thought
RSSEvals are the only thing standing between your agent and production, and most of them fail for reasons that have nothing to do with the agent. Here is how I stopped trusting a green run.
How Mexico’s real-time interbank payment rail went from back-office plumbing to the default way millions pay — and why cards are now the underdog.
The Model Context Protocol buys you a socket, not a clean integration. Here is how the promise of plug-and-play agent tooling quietly becomes the same old adapter problem.
Why every abstraction that saves you once will eventually cost you, and how to know when to rewrite it.
Version 1.1 replaces SWE-Bench-Pro-Hard-AA with DeepSWE and averages pass@1 across three complementary benchmarks. Here is what each component measures, what the headline score hides, and how to use it when choosing a coding agent.
El rediseño de la app de BBVA México y las caÃdas continuas en 2025 y 2026 exponen qué sucede cuando la experiencia de usuario y la estabilidad del backend desacoplan sus prioridades: una lección sobre arquitectura financiera.
August 2026 developer radar: stateless MCP, universal memory layers, open-source meta-harnesses, and the new model tier reshaping agentic coding.
A layer-by-layer tour of modern accelerator design: SMs vs CUs, tensor cores, reticle limits, prefill vs decode, the 1000W power wall, HBM4, and why the compiler is the final boss.
An interactive policy-debate demo built on the POLARIS multi-agent framework: five biased agents, four layers, and a quality gate that decides when the answer is good enough.
Seven days without AI autocomplete. What got slower, what got sharper, and what I turned back on when the week was over.
Why stuffing the prompt is a losing strategy, and how to design a tiered memory architecture with deterministic eviction and retrieval budgets.
An empirical analysis of the mid-2026 development stack: from AI-native IDEs and autonomous orchestration to agentic DevSecOps.
How Policy Optimization via Layered Agent Recursive Inference Search enforces determinism through agent isolation.
Reflections after orchestrating a dozen minor side projects. Here is what we missed in the tutorial: the paradigm shift from reactive assistance to proactive predictive engineering.
Superando el 80%+ de fracaso en proyectos de IA: Una arquitectura estructural para transformar la experimentación corporativa en rendimiento empÃrico, contable y defendible en el P&L.
An empirical analysis of the mid-2026 development stack: from AI-native IDEs and autonomous orchestration to agentic DevSecOps.
Sustituyendo la rumiación psicológica por un modelo matemático: TeorÃa de Prospectos, entropÃa de Shannon y optimización no lineal para entender la incertidumbre.
Eradicating opaque abstractions by aligning software design with empirical system realities.
A chronological tour of every major React feature — from JSX to Server Components — with live demos you can play with.
A reference for conditional types, inference, branded primitives, and template literal magic. Strict mode only.
Why `while(true)` is technical debt, and how to implement O(1) event dispatching for scalable agent systems in Node.js.
Ditch the monolithic form for a state-driven, accessible multi-step architecture that scales.
Moving beyond naive useState hooks to handle complex data flows and atomic dirty-state validation.
Why natural language persuasion fails in production AI pipelines and how to implement strict logical boundaries via schema enforcement and logit bias.
Architectural Redefinition: From Synaptocentric to Hybrid Models.
Why A11y compliance must be a development axiom, not a post-production patch.
No articles match your search.