Beyond the basics: ado with real data¶
New to ado?
This walkthrough introduces the key concepts — sample stores, discovery spaces, and operations — as we go. If you'd prefer a full overview first, the Concepts page has you covered.
This walkthrough takes you end-to-end through a real-data ado workflow using an imported cloud benchmark dataset, so you can focus on how ado works with pre-existing measurements rather than synthetic examples.
Once imported, ado's replay mechanism transparently serves the benchmark results as the operation runs — no re-execution needed.
We will:
- Load pre-existing benchmark data into a sample store
- Define a discovery space over the cloud hardware configurations
- Run a random walk operation to sample the space
Prerequisites
ado-coreinstalled (pip install ado-core)- The repository cloned from GitHub:
git clone https://github.com/ibm/ado.git
cd ado
All commands are run from the repository root (ado/).
TL;DR
ado create store -f examples/ml-multi-cloud/ml_multicloud_sample_store_from_root.yaml
ado create space -f examples/ml-multi-cloud/ml_multicloud_space.yaml --use-latest samplestore
ado create operation -f examples/ml-multi-cloud/randomwalk_ml_multicloud_operation.yaml --use-latest space
ado show measurements operation --use-latest
Step 1 — Load the benchmark data¶
This example uses pre-existing benchmark data from ml_export.csv — results of running a workload on different cloud hardware configurations across three providers.
In ado, configurations are called entities and are stored, together with measurement results, in a sample store.
The file examples/ml-multi-cloud/ml_multicloud_sample_store_from_root.yaml tells ado how to import the CSV:
# Copyright IBM Corporation 2025, 2026
# SPDX-License-Identifier: MIT
copyFrom:
- module:
moduleClass: CSVSampleStore
moduleName: ado.core.samplestore.csv
parameters:
experiments:
- actuatorIdentifier: replay
constitutivePropertyMap:
- cpu_family
- vcpu_size
- nodes
- provider
experimentIdentifier: benchmark_performance
observedPropertyMap:
- wallClockRuntime
- status
generatorIdentifier: multi-cloud-ml
identifierColumn: config
storageLocation:
path: examples/ml-multi-cloud/ml_export.csv
metadata:
description: samplestore initialised with ml_multi_cloud data
name: ml_multi_cloud
specification:
module:
moduleClass: SQLSampleStore
moduleName: ado.core.samplestore.sql
Create the sample store:
ado create store -f examples/ml-multi-cloud/ml_multicloud_sample_store_from_root.yaml
Success! Created sample store with identifier $SAMPLE_STORE_IDENTIFIER
Tip
The sample store only needs to be created once. It can be reused across multiple discovery spaces and examples that need the ml_export.csv data.
Step 2 — Define a discovery space¶
A discovery space brings together:
- an entity space — the set of points to measure
- a measurement space — the experiments to apply to them
The file examples/ml-multi-cloud/ml_multicloud_space.yaml describes the space of cloud configurations and maps them to the replay.benchmark_performance experiment:
# Copyright IBM Corporation 2025, 2026
# SPDX-License-Identifier: MIT
entitySpace:
- identifier: "provider"
propertyDomain:
values: ['A', 'B', 'C']
- identifier: "cpu_family"
propertyDomain:
values: [0, 1]
- identifier: "vcpu_size"
propertyDomain:
values: [0, 1]
- identifier: "nodes"
propertyDomain:
values: [2, 3, 4, 5]
experiments:
- experimentIdentifier: 'benchmark_performance'
actuatorIdentifier: 'replay'
metadata:
name: "ml_multicloud_basic"
Create the discovery space, linking it to the sample store you just created:
ado create space -f examples/ml-multi-cloud/ml_multicloud_space.yaml --use-latest samplestore
Success! Created space with identifier: $DISCOVERY_SPACE_IDENTIFIER
Inspect it to confirm the entity space dimensions and experiment interface:
ado describe space --use-latest
Entity Space:
Number of entities: 48
Categorical properties:
name ┃ values
━━━━━━━━━━╋━━━━━━━━━━━━━━━━━
provider ┃ ['A', 'B', 'C']
Discrete properties:
name ┃ range ┃ interval ┃ values
━━━━━━━━━━━━╋━━━━━━━╋━━━━━━━━━━╋━━━━━━━━━━━━━━
cpu_family ┃ None ┃ None ┃ [0, 1]
vcpu_size ┃ None ┃ None ┃ [0, 1]
nodes ┃ None ┃ None ┃ [2, 3, 4, 5]
Measurement Space:
Experiments:
base identifier ┃ required major version ┃ parameterization
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╋━━━━━━━━━━━━━━━━━━━━━━━━╋━━━━━━━━━━━━━━━━━━
replay.benchmark_performance ┃ None ┃ None
Have a look
The entity space has 48 entities — the Cartesian product of provider (3), cpu_family (2), vcpu_size (2), and nodes (4). Compare this to the number of rows in ml_export.csv.
Step 3 — Run a random walk¶
An operation drives an operator over the discovery space. Here we use random_walk, which samples entities at random.
The file examples/ml-multi-cloud/randomwalk_ml_multicloud_operation.yaml configures the walk:
# Copyright IBM Corporation 2025, 2026
# SPDX-License-Identifier: MIT
metadata:
name: 'randomwalk-all'
description: 'Perform a random walk on all points in a space'
spaces:
- 'space-630588-bfebfe'
operation:
module:
operatorName: "random_walk"
operationType: "explore"
parameters:
numberEntities: 48
batchSize: 1
singleMeasurement: True
samplerConfig:
samplerType: generator
mode: random
Run the operation:
ado create operation -f examples/ml-multi-cloud/randomwalk_ml_multicloud_operation.yaml --use-latest space
Tip
The spaces: field of the operation contains a placeholder identifier. The --use-latest space flag replaces it with the discovery space you just created.
ado prints a line for each entity as it is measured. When finished, you will see a summary like:
identifier: random_walk@2.0.0-31d4c6
kind: operation
metadata:
entities_submitted: 48
experiments_requested: 74
status:
- event: finished
exit_state: success
recorded_at: "2026-07-09T09:26:52.072109Z"
Memoization
ado transparently reuses any measurement already in the sample store rather than re-running it. The operator does not need to know whether a result will come from live execution or replay — ado handles it automatically.
Step 4 — View the results¶
ado show measurements operation --use-latest
┌───────────────┬──────────────┬────────────────────────┬──────────────────────────────┬────────────┬───────┬──────────┬───────────┬────────────────────┬──────────────┬───────┐
│ request_index │ result_index │ identifier │ experiment_id │ cpu_family │ nodes │ provider │ vcpu_size │ wallClockRuntime │ status │ valid │
├───────────────┼──────────────┼────────────────────────┼──────────────────────────────┼────────────┼───────┼──────────┼───────────┼────────────────────┼──────────────┼───────┤
│ 0 │ 0 │ A_f1.0-c0.0-n4 │ replay.benchmark_performance │ 1.0 │ 4 │ A │ 0.0 │ 158.70639538764954 │ ok │ True │
│ 1 │ 0 │ A_f1.0-c0.0-n2 │ replay.benchmark_performance │ 1.0 │ 2 │ A │ 0.0 │ 378.31657004356384 │ ok │ True │
│ 2 │ 0 │ B_f0.0-c0.0-n3 │ replay.benchmark_performance │ 0.0 │ 3 │ B │ 0.0 │ 153.51639366149902 │ ok │ True │
│ ... │ ... │ ... │ ... │ ... │ ... │ ... │ ... │ ... │ ... │ ... │
└───────────────┴──────────────┴────────────────────────┴──────────────────────────────┴────────────┴───────┴──────────┴───────────┴────────────────────┴──────────────┴───────┘
Warning
Some entities have valid: False — these are points in the discovery space for which no matching data existed in ml_export.csv. The status column shows not_measured for those rows.
To see results aggregated across the full space (not just this operation):
ado show measurements space --use-latest
Exploring Further¶
Here are a variety of commands you can try after executing the example above:
Viewing entities¶
There are multiple ways to view the entities related to a discoveryspace. Try:
ado show measurements space --use-latest
ado show measurements space --use-latest --aggregate mean
ado show measurements space --use-latest --include unmeasured
ado show measurements space --use-latest --property-format target
Also, the following command will give you summary statistics of what has been measured:
ado show stats discoveryspace --use-latest
Note
If you want to run these commands against the most recent space in the current context, use the --use-latest flag as above.
Resource provenance¶
The related sub-command shows resource provenance:
ado show related operation --use-latest
Operation timeseries¶
The following commands give more details of the operation timeseries:
ado show trace operation --use-latest --unroll-entities
ado show trace operation --use-latest
Resource templates¶
Another helpful command is template which will output a default example of a resource YAML along with an (optional) description of its fields. Try:
ado template operation --include-schema --operator-name random_walk --output-file random_walk_template.yaml
Summary¶
| Step | What you did | ado concept |
|---|---|---|
| 1 | Imported benchmark data into a sample store | Sample store / replay |
| 2 | Described the cloud configuration space in space.yaml | Discovery space |
| 3 | Ran a random walk to sample 48 entities | Operation / operator |
| 4 | Retrieved measured results | ado show measurements |
What's next¶
-
Search using an optimizer
Use RayTune to perform an optimization instead of a random walk.
-
Discover important dimensions
Identify which hardware parameters most influence runtime using
ado's importance analysis.