Benchmark Hub

The Quantiles Benchmark Hub is a reference library for built-in benchmarks, evaluation suites, and metrics used to test AI systems. It highlights benchmarks that can be run directly in Quantiles while also helping teams choose broader measurement approaches for accuracy, hallucinations, reasoning, calibration, robustness, and safety.

What the Hub is for

The Benchmark Hub serves as a centralized reference for:

  • Built-in benchmarks that are ready to run with qt run, without setting up datasets, scoring logic, or custom evaluation code.
  • Curated descriptions of widely used and emerging AI evaluation benchmarks
  • A shared reference point for teams building and evaluating AI systems

Built-in Benchmarks

Ready-to-run benchmarks with predefined datasets, scoring methodologies, and metrics.