Using Execution Traces to Evaluate AI Agent Behavior

Execution traces reveal how agents reason through developer workflows and provide observable evidence for measuring their reliability and behavior.

The word trace formed from white halftone dots

Latest Articles

Articles on AI evaluation workflows, benchmarking methods, applied AI systems, and Quantiles product updates.

How to Evaluate AI Agents

Evaluate agent performance using representative tasks, controlled environments, observable execution traces, and repeated trials.

Run and Customize Public AI Benchmarks

Use Quantiles to run ready-to-use public benchmarks on your models, test smaller samples, customize configurations, and compare benchmark results.

How to build Custom No-code AI Evaluations

Create and run custom AI evaluations using your own datasets and prompts, without writing code.

Quantiles: Open-Source, Local-First Infrastructure for AI Evaluation Workflows

Run lightweight checks or full-scale evaluations locally with the open-source Quantiles CLI, manually and with coding agents.

Evaluating Emergency Recognition in Patient-Facing AI Systems

In AI supported patient portals, emergency recognition should be evaluated using local message data, slice analysis, and post deployment monitoring.

Recovery, Escalation, and Workflow Reliability in Clinical AI Agents

Strong recovery and escalation enable clinical AI agents to surface and contain failures early, improving handoff quality and workflow reliability.

Evaluating Agentic Clinical AI Systems

Evaluating clinical AI agents requires a broader, multi-axis approach that captures sequential behavior, tool dependencies, and governance risk in evolving care environments.

Post-deployment Monitoring of Healthcare AI

Healthcare AI evaluation is moving beyond one-time validation toward a structured post-deployment lifecycle that supports continuous learning and adapts to evolving healthcare environments.

Open and Proprietary Benchmarks

Rigorous healthcare AI evaluation requires combining open and proprietary benchmarks to balance transparency, comparability, and real-world clinical utility.

LLM-as-a-judge Evaluations

LLM-as-a-judge makes clinical judgement scalable and auditable, shifting healthcare AI evaluation toward how models behave in clinical context.

Evaluating Healthcare AI with OpenAI's HealthBench

HealthBench is a rubric-based healthcare AI benchmark that evaluates model behavior across safety, reliability, and communication dimensions in realistic clinical scenarios.

How to Analyze and Interpret Evaluations in Healthcare AI

Interpret benchmark signals for healthcare AI by linking deterministic metrics, calibration, agreement, generation metrics, and statistical validity to clinical risk and production readiness.

AI Benchmarks for Healthcare

How industry-standard and domain-specific benchmarks, plus continuous monitoring, keep healthcare AI safe, reliable, and compliant.