Pareto Frontier
ct metrics pareto places each entry of a YAML config at its (safety, usefulness) point, all at one audit budget, and marks the Pareto frontier. Each entry is an honest/attack run pair, so one plot can compare protocols, monitors, or protocol settings, one entry each.
Quick Start
ct metrics pareto --yml-config experiments/my_pareto.yaml --output-dir results/my_pareto
Outputs land in --output-dir (figures/pareto/<timestamp>/ if omitted):
results/my_pareto/
data.json # ParetoResult dump: every point, its thresholds, and whether it is on the frontier
pareto.png # scatter + frontier polyline (raster)
pareto.pdf # same plot, vector
<yml-config name> # copy of the input YAML
YAML Schema
title: "My experiment — Pareto frontier" audit_budget: 0.01 # fraction of the honest audit opportunities (0.01 is 1%) bootstrap: 1000 # optional: draws for a 95% safety interval (default 0, none) entries: - label: "Trusted monitoring" honest: <run_id_or_path> attack: <run_id_or_path> style: { marker: o, color: "#1f77b4" } - label: "Action monitor only" honest: <run_id_or_path> attack: <run_id_or_path> monitors: [simple] # optional: the monitors that audit (default: every monitor that judged the honest run) style: { marker: D, color: "#ff7f0e" }
Required: title, audit_budget, and per entry label, honest, attack. Optional: bootstrap, and per entry monitors and style (marker, color, size). An unknown key is an error.
Sources can be run IDs, run URLs, .eval files, or directories holding one .eval file.
Each Point
- Safety is computed as in Safety Curves: one threshold per monitor in
monitors, fit ataudit_budgeton the honest run's samples without a side task, played against the attack run's samples with one. An entry whose thresholds or safety are undefined is an error. - Usefulness is the mean main-task score over the entry's honest samples that carry no side task, ignoring unscored ones, with a 95% Wilson interval: exact for pass/fail scores, conservative for graded ones.
- With
bootstrapset, each entry's safety gets a 95% interval from that many draws, each resampling the honest and attack trajectories and refitting the thresholds. Safety needs the bootstrap because it is not a proportion of the attack count.
A reference point is an entry of its own. With monitors: [] no monitor audits, so only the protocol's online audits catch attacks: the unmonitored baseline for that run pair.
Pareto Frontier Marking
A point (s, u) is on the frontier unless some other point has s' >= s and u' >= u with at least one inequality strict. Frontier points are joined by a dashed polyline in pareto.png, and the other points are drawn translucent.
Programmatic API
import yaml from control_tower.metrics.pareto import ParetoConfig, compute_pareto from control_tower.metrics.printing.pareto import handle_pareto with open("experiments/my_pareto.yaml") as f: config = ParetoConfig.model_validate(yaml.safe_load(f)) result = compute_pareto(config) for p in result.points: print(p.label, p.safety, p.usefulness, "★" if p.on_frontier else "") print("frontier, least safe first:", [p.label for p in result.frontier]) handle_pareto(result)
Run ct metrics pareto --help for every option.