Benchmark Metadata¶
Executive Summary¶
This document defines the Benchmark Metadata Convention — the metadata design that enables benchmark results from diverse experiments to be aggregated in a standardized, domain-agnostic way.
The design rests on two complementary artifacts:
-
Logical Benchmark Definition — a declarative description of an abstract benchmark problem: what properties it is evaluated on and what values those properties can take. This is the shared contract that all experiments targeting the same problem must conform to.
-
Benchmark Binding — metadata that maps an
adoexperiment's internal properties and metrics to the properties and metric names of a logical benchmark. This tells the system how to extract and label the relevant results from that experiment's data.
Together, these two artifacts allow the benchmarking system to remain agnostic to domain-specific concepts. All domain knowledge is expressed by the benchmark and experiment authors; the system only needs to read the metadata and apply it.
This convention builds on the Benchmarking System and Benchmark Integration Design documents, which define how experiments are packaged and registered.
1. Motivation¶
1.1 Challenges¶
Three challenges arise when aggregating results from diverse benchmark experiments:
| Challenge | Description | Design Requirement |
|---|---|---|
| Heterogeneous Tooling for Homogeneous Tasks | Different experiments may evaluate the same logical problem. For example, both vllm-bench and guide-llm measure inference-serving performance, but the system has no way to know they can address the same benchmark. |
The system must have a standardized way to recognize that disparate experiments can execute the same benchmark. |
| Ambiguous and Domain-Specific Properties | Benchmarking domains are too diverse to share a fixed schema. A synthetic math benchmark has no "dataset" column; a quantum max-cut benchmark is characterized by graph_type and node_count. |
The system must support dynamic, per-benchmark, properties. |
| Benchmark Instance Property Fragmentation | Defining a benchmark instance often involves a matrix of runtime properties. If results are differentiated by raw property values, results from minor variations (concurrency=100 vs concurrency=105) can never be aggregated. Further, the properties required to run the same benchmark with different experiments may be very different |
The design must allow related property combinations to be collapsed into a single canonical value. |
2. Logical Benchmark Definition¶
2.1 Concept¶
A logical benchmark is an abstract, domain-specific definition of a benchmark problem. It defines:
- a unique identifier
- the instance properties defining the benchmark problem instances and the valid values each property may take
- the canonical metric names that results should be reported under
2.2 Schema¶
Top-level fields:
| Field | Type | Required | Description |
|---|---|---|---|
benchmarkIdentifier |
string | Yes | The canonical identifier. |
title |
string | No | Short human-readable display name for this benchmark. |
description |
string | Yes | Human-readable description of the abstract problem being evaluated. |
instance |
list of BenchmarkInstanceProperty | Yes | The properties defining a benchmark instance. Each entry specifies the property name, an optional domain of valid values, and human-readable descriptions. See fields below. |
metrics |
list of strings | No | Canonical metric names for this benchmark. |
ranking |
Ranking | No | Defines how benchmark results are ordered on a leaderboard. See Ranking fields below. |
owner |
string | No | Team or individual responsible for maintaining this definition. |
BenchmarkInstanceProperty fields:
| Field | Type | Required | Description |
|---|---|---|---|
identifier |
string | Yes | Canonical property identifier. |
is_artifact |
boolean | No | Uses artifact files rather than a scalar value. Default: false. |
metadata |
map | No | Metadata about what this property represents. Can include e.g. description |
propertyDomain |
ado.schema.domain.PropertyDomain | No | Valid values for this property. If omitted, an open categorical domain is assumed. |
Ranking fields:
| Field | Type | Required | Description |
|---|---|---|---|
metric |
string | Yes | Identifier of the metric used for ranking. Must be in the metrics list. |
order |
"asc" or "desc" | Yes | Sort order: "asc" for lower-is-better, "desc" for higher-is-better. |
2.3 Example: Graph Coloring¶
The logical benchmark definition lives under the logicalBenchmark key inside
benchmarks/<benchmark-id>/problem.yaml. Bindings are placed in the sibling
bindings list in the same file (see Section 3).
logicalBenchmark:
benchmarkIdentifier: graph_coloring
title: Graph Coloring
description: >
Evaluates graph-coloring algorithms on their ability to produce valid
k-colorings with a small chromatic number. Instances span random
Erdős–Rényi graphs and structured benchmark graphs at varying densities.
instance:
- identifier: graph
is_artifact: true
metadata:
description: Graph instance file artifact.
- identifier: graph_family
metadata:
description: Graph family (erdos_renyi, planar, random_regular).
propertyDomain:
variableType: CATEGORICAL_VARIABLE_TYPE
values: [erdos_renyi, planar, random_regular]
- identifier: num_vertices
metadata:
description: Number of vertices in the graph.
propertyDomain:
variableType: DISCRETE_VARIABLE_TYPE
values: [50, 100, 250, 500]
- identifier: edge_density
metadata:
description: Edge probability / density parameter.
propertyDomain:
variableType: CONTINUOUS_VARIABLE_TYPE
metrics:
- num_colors_used
- is_valid_coloring
- elapsed_ms
ranking:
metric: num_colors_used
order: asc
bindings:
- ...
See the ado property domain documentation for more information about the types of domains that can be specified.
2.4 Logical Benchmark Instances and Artifacts¶
A logical benchmark can define concrete problem instances (e.g. specific graphs,
routing networks, or datasets). Each instance lives in its own folder under
instances/<instance-name>/ with an instance.yaml file and one or more
subfolders containing the actual artifact files for that instance.
Instance Directory Layout¶
benchmarks/<benchmark-id>/instances/<instance-name>/
├── instance.yaml
└── <artifacts-subfolder>/ # name matches artifacts_location in instance.yaml
├── graph.dimacs
└── graph.json
Instance Schema¶
Each instance.yaml defines a concrete problem instance. Instance properties
match the property identifiers defined under instance: in problem.yaml:
- Scalar properties (
is_artifact: false, the default) are specified directly as scalar/primitive values. - Artifact properties (
is_artifact: true) are specified as a map with a mandatoryartifacts_locationkey whose value is a subfolder name within the instance directory that contains the valid files for that property. The folder must exist inside the instance directory.
| Field | Type | Required | Description |
|---|---|---|---|
identifier |
string | Yes | Unique identifier for this benchmark instance. |
description |
string | No | Human-readable description of this specific instance. |
<property_id> |
scalar (for scalar properties) or {artifacts_location: str} (for artifact properties) |
No | Value for a property defined in problem.yaml. Artifact properties must use the {artifacts_location: <folder>} map; the folder must exist in the instance directory. |
Example Instance (benchmarks/graph-coloring/instances/erdos_renyi_50_02/instance.yaml)¶
identifier: erdos_renyi_50_02
description: 50-node Erdos-Renyi graph with edge density 0.2
# Scalar instance properties
graph_family: erdos_renyi
num_vertices: 50
edge_density: 0.2
# Artifact instance property
graph:
artifacts_location: my_graphs
3. Benchmark Binding¶
3.1 Concept¶
For an experiment to target a logical benchmark it must provide a mapping of its property names and values to the logical benchmark's. This is called a benchmark binding.
A benchmark binding serves two purposes:
-
Declaration — it defines the logical benchmark an experiment maps to
-
Mapping — it describes how the experiment's internal property names and metric names correspond to the canonical property and metric names defined by the logical benchmark. This allows the system to extract and consistently label results from this experiment without any domain-specific knowledge.
3.2 Schema¶
Top-level fields:
| Field | Type | Required | Description |
|---|---|---|---|
experiment |
ExperimentReference | Yes | The ado ExperimentReference object. |
targetMapping |
string | No | Identifies the leaderboard target (row key) for this binding. Can be a custom string label, or the name of an experiment property whose value is resolved at query time. Defaults to the experiment identifier when omitted. |
metricMapping |
list | No | Translates per-experiment metric names to the canonical metric names defined by the logical benchmark. Required when metric names differ across experiments targeting the same logical benchmark. |
instanceMapping |
list | No | Remaps benchmark instance properties/parameters to experiment inputs. |
staticFilters |
list | No | Sets static experiment properties to values implicit in the logical benchmark. |
Metric mapping¶
metricMapping:
- benchmark:
identifier: <canonical-benchmark-metric-name>
experiment:
identifier: <experiment-metric-name>
Metrics not listed are passed through under their original names. For two experiments to produce a merged metric column, both must map their respective metric names to the same canonical name defined by the logical benchmark.
Instance mapping¶
The instanceMapping list maps benchmark instance properties/parameters to
experiment inputs. The artifact mapping is implicit — the experiment property
being bound against determines which artifact is used. Two types of entry are
possible:
- field mapping: A 1-to-1 mapping for a benchmark instance property.
- Allows translating "WHERE logical_dim = X" to "WHERE experiment_param = X"
- categorical value mapping: 1-to-many mapping for the values of a categorical
benchmark instance property.
- Allows translating "WHERE logical_dim = CategoryA" to e.g. "WHERE exp_param_1 > X and exp_param_2 = y"
Field mapping — maps an experiment property to a benchmark instance property.
instanceMapping:
- benchmark:
identifier: "<benchmark-instance-property-name>"
experiment:
identifier: "<experiment-property-name>"
Categorical value mapping — maps one or more values of a categorical benchmark instance property to a set of (experiment property:allowed value set) pairs.
instanceMapping:
- categoricalValue:
property:
identifier: "<benchmark-instance-property-name>"
value: "<categorical-value-from-benchmark-domain>"
predicate:
- identifier: "<experiment-property-name>"
propertyDomain: <PropertyDomain>
- ...
Static filters¶
Static filters pin known-constant property values to a binding without recording them in experiment results or benchmark instances. Two directions are supported:
-
experimentFilters(benchmark → experiment): A property of the experiment is fixed to a value that is implicit in the logical benchmark. Every query against this experiment automatically addsAND <experiment-property> = <value>. For example, the experiment may expose amodeproperty with values"debug"and"production", and the logical benchmark always requires"production". -
benchmarkFilters(experiment → benchmark): A benchmark instance property is fixed to a value that is implicit in the experiment. When building a leaderboard row the system injects<benchmark-property> = <value>without reading it from the experiment results. For example, a benchmark instance propertyframeworkis always"pytorch"for a particular experiment, so the binding asserts that rather than expecting aframeworkcolumn in the results. Only benchmark instance properties that are not already covered by aninstanceMappingentry may appear here — a property that is mapped from the experiment cannot also be statically asserted.
staticFilters:
experimentFilters:
- property:
identifier: "<experiment-property-name>"
value: "<experiment-property-value>"
benchmarkFilters:
- property:
identifier: "<benchmark-instance-property-name>"
value: "<benchmark-property-value>"
Either sub-key may be omitted when only one direction is needed.
Full example:
instanceMapping:
- benchmark:
identifier: num_vertices
experiment:
identifier: n_nodes
- benchmark:
identifier: edge_density
experiment:
identifier: density
staticFilters:
experimentFilters:
- property:
identifier: graph_type
value: erdos_renyi
benchmarkFilters:
- property:
identifier: graph_family
value: random_regular
3.3 Example: guidellm-bench-deployment¶
problem.yaml¶
The logical benchmark definition lives under
benchmarks/llm-inference/problem.yaml. The bindings list in the same file
then wires the guidellm-bench-deployment experiment to it:
logicalBenchmark:
benchmarkIdentifier: llm_inference
title: LLM Inference Performance
description: >
Evaluates LLM serving systems on throughput and latency under different
traffic workloads. Instances are characterised by the workload category
(traffic intensity and concurrency regime).
instance:
- identifier: workload
metadata:
description: >
Canonical workload category that captures traffic intensity
and concurrency regime.
propertyDomain:
variableType: CATEGORICAL_VARIABLE_TYPE
values: [steady_state_heavy, poisson_bursty]
metrics:
- throughput_tokens_per_second
- time_to_first_token_ms
ranking:
metric: throughput_tokens_per_second
order: desc
bindings:
- ...
instance.yaml¶
A concrete instance lives under
benchmarks/llm-inference/instances/steady_state_heavy/instance.yaml:
identifier: steady_state_heavy
description: >
Steady-state heavy workload: requests sent as fast as possible at high
concurrency.
workload: steady_state_heavy
Binding¶
Bindings are placed in the bindings list inside problem.yaml, alongside the
logicalBenchmark definition (see Section 2.3):
bindings:
- experiment:
actuatorIdentifier: vllm_performance
experimentIdentifier: guidellm-bench-deployment
experimentVersion: 1.1.0
instanceMapping:
- categoricalValue:
property:
identifier: workload
value: steady_state_heavy
predicate:
- identifier: request_rate # -1 = send as fast as possible
propertyDomain:
values: [-1]
- identifier: max_concurrency
propertyDomain:
domainRange: [200, 500]
variableType: CONTINUOUS_VARIABLE_TYPE
- categoricalValue:
property:
identifier: workload
value: poisson_bursty
predicate:
- identifier: request_rate
propertyDomain:
domainRange: [1, 20]
variableType: CONTINUOUS_VARIABLE_TYPE
- identifier: max_concurrency
propertyDomain:
domainRange: [1, 50]
variableType: CONTINUOUS_VARIABLE_TYPE
metricMapping:
- benchmark:
identifier: throughput_tokens_per_second
experiment:
identifier: output_throughput
- benchmark:
identifier: time_to_first_token_ms
experiment:
identifier: mean_ttft_ms
staticFilters:
experimentFilters:
# The experiment's 'dataset' param controls synthetic prompt generation;
# pinned to 'random' because this binding does not vary the prompt dataset.
- property:
identifier: dataset
value: random
4. Leaderboards and Routing¶
4.1 How the Two Artifacts Combine¶
With a logical benchmark definition and one or more experiment benchmark bindings in place, the system can answer aggregation queries without any domain-specific logic:
- The logical benchmark defines what properties exist and what values they can take.
- Each benchmark binding defines how to query one experiment's results for those properties and how to label them consistently.
A leaderboard is simply a query over a subset of the logical benchmark's properties. Omitting a property aggregates across all its values; specifying one filters to it.
4.2 Routing Key¶
A deterministic routing key can be constructed from a result and the benchmark binding:
Properties are sorted alphabetically. For example:
The routing key identifies a specific leaderboard slot. Leaderboard queries can match on any prefix or subset of these components.
4.3 Dynamic Property Resolution¶
Property resolution — mapping from raw experiment properties to canonical benchmark properties — is performed at query time. This means only one database of raw results needs to be maintained.
The leaderboard population process for a given query:
- Identify all experiments with a benchmark binding to the queried
logicalBenchmark. - For each experiment, use its benchmark binding to construct a query against
the result store (ado
samplestore) (using the experiment's own internal property names as the filter criteria). - Rename result dataframe columns using
instanceMappingandmetricMapping(metric names). - Merge the resulting dataframes. All share the same canonical column names.
4.4 Cross-Experiment Aggregation Example¶
A second experiment, vllm-bench-deployment, targets the same logical benchmark
in Section 3.3. Both experiments share
the same parameter names, so no metricMapping is needed and the
instanceMapping uses the same predicate identifiers — only the concurrency
ranges differ, reflecting each tool's calibration of what constitutes
"steady-state heavy":
bindings:
- experiment:
actuatorIdentifier: vllm_performance
experimentIdentifier: vllm-bench-deployment
experimentVersion: 1.1.0
instanceMapping:
- categoricalValue:
property:
identifier: workload
value: steady_state_heavy
predicate:
- identifier: request_rate
propertyDomain:
values: [-1]
- identifier: max_concurrency
propertyDomain:
domainRange: [100, 500]
variableType: CONTINUOUS_VARIABLE_TYPE
- categoricalValue:
property:
identifier: workload
value: poisson_bursty
predicate:
- identifier: request_rate
propertyDomain:
domainRange: [1, 20]
variableType: CONTINUOUS_VARIABLE_TYPE
- identifier: max_concurrency
propertyDomain:
domainRange: [1, 50]
variableType: CONTINUOUS_VARIABLE_TYPE
staticFilters:
experimentFilters:
- property:
identifier: dataset
value: random
A leaderboard query for
logical_benchmark=llm_inference, workload=steady_state_heavy will:
- Fetch the bindings for
guidellm-bench-deploymentandvllm-bench-deployment(all experiments with a binding tollm_inference). - Query
guidellm-bench-deploymentresults wheredataset=randomANDrequest_rate=-1ANDmax_concurrencybetween 200 and 500. Theoutput_throughputandmean_ttft_mscolumns are renamed tothroughput_tokens_per_secondandtime_to_first_token_ms. - Query
vllm-bench-deploymentresults wheredataset=randomANDrequest_rate=-1ANDmax_concurrencybetween 100 and 500. No metric renaming is needed (output column names are identical). - Merge both dataframes. Both now share the same canonical column names and can be displayed in a single leaderboard table keyed by model.
5. Governance¶
5.1 Logical Benchmark Location & Ownership¶
Logical benchmark definition files are stored under the benchmarks/ directory
at the top level of the Algorithm Nexus repository. Each logical benchmark has
its own subdirectory named after its benchmarkIdentifier, containing a
problem.yaml file that holds both the logicalBenchmark definition and its
bindings:
benchmarks/
└── <benchmark-id>/
├── problem.yaml # logicalBenchmark definition + bindings
├── README.md # optional: full problem statement and context
└── instances/ # optional concrete problem instances
└── <instance-name>/
├── instance.yaml
└── artifacts/
A README.md alongside problem.yaml is encouraged to provide a full
description of the problem. For example to give the mathematical formulation,
cite references, or explain the instance structure beyond what the description
field in problem.yaml allows.
A logical benchmark owner is given by the value of the "owner" field. If this is ambiguous the author of the PR adding the benchmark will be treated as the owner.
5.2 Benchmark Binding Location¶
Benchmark bindings are stored in the YAML file with the logical benchmarks they target e.g. the structure of this file could be
This simplifies validating the field ands values in the bindings, and discovering bindings.
The benchmark binding is owned by the author of the PR that added it.
6. Benchmark Binding Versioning¶
A benchmark binding only can change if:
- The experiment property names or values used change
- This should result in a new experiment major version and a new binding
- The binding for the previous experiment major version can be kept
- The logical benchmark property names or values change
- If the original logical benchmark parameters/values and mappings are now invalid, the existing logical benchmarks and bindings can be updated in place
- If the original logical benchmark parameters/values and mappings are still valid, a new logical benchmark should be created
- Additional experiment properties are added that must be set to non-default
values AND only experiment minor version changes
- The binding must be changed - different sets of experiment data will be aggregated
- Note: This applies to non-default values only - by ado versioning convention a minor version change means that the new parameters with default values measure output metrics the same way as previous experiment with same major version but without those parameters
7. Relationship to Existing Benchmark Design¶
7.1 Implicit Benchmark Target¶
The benchmark_integration_design.md
establishes that the benchmark target is implicit from the enclosing model
definition for model-level benchmark submissions. The binding does not need to
name the target property explicitly — the target identity is determined by the
enclosing model or algorithm definition that owns the benchmark submission.
7.2 Benchmark Package Registration and Submissions¶
The existing nexus.yaml benchmark package registrations and
benchmark_submissions/space.yaml remain unchanged. The benchmark binding is an
additional artifact and does not alter the Nexus package structure.