ROC Curves
ct metrics roc plots, for each monitor, a ROC curve of honest against attack trajectories and prints its AUROC.
Quick Start
# One honest/attack pair ct metrics roc --honest <run-id-or-path> --attack <run-id-or-path> # From a run file with several entries ct metrics roc --run-file runs.yaml
How Trajectories Are Scored
Each trajectory has one score per monitor: its highest verdict that an audit could act on, read from the sample's monitor score as Safety Curves describes. Verdicts after a live side task landed are left out, and a trajectory verdict counts only when the run was neither audited online nor harmed live. A trajectory with no such verdict from a monitor is left off that monitor's curve. Every attack trajectory is a positive, whether or not its side task succeeded; every honest trajectory is a negative.
Input Modes
Inline mode
Pass --honest and one or more --attack sources. Each --attack is paired with the shared --honest and gets its own curve.
ct metrics roc \ --honest https://data.linuxarena.ai/runs/c7aeb0aeccba40eb9d910ab13b601bed \ --attack https://data.linuxarena.ai/runs/bf8a832da92f4431b2579ad04cd983ad \ --attack https://data.linuxarena.ai/runs/99d571edada441e8a595d7e50f4efc37
Sources can be run IDs, run URLs (https://data.linuxarena.ai/runs/...), .eval files, or directories holding one .eval file.
Run file mode
Pass --run-file with a YAML file describing several honest/attack pairs and their metadata. This enables --color-by, --style-by, and --filter-by.
runs: - label: gpt-4.1-nano (goals) model: gpt-4.1-nano monitor: mean-action # metadata for --color-by/--style-by honest: https://data.linuxarena.ai/runs/c7aeb0aeccba40eb9d910ab13b601bed attack: https://data.linuxarena.ai/runs/bf8a832da92f4431b2579ad04cd983ad - label: gpt-4.1-nano (no-goals) model: gpt-4.1-nano monitor: mean-action-no-goals honest: data/roc_eval/gpt-4.1-nano/honest.eval attack: data/roc_eval/gpt-4.1-nano/attack.eval
Each entry requires honest and attack, and may add a label plus any metadata fields (e.g. model, monitor). Without a label, one is made from model and monitor if present, otherwise from the attack source name. Run-file entries may also carry honest_task_set / attack_task_set, as in Safety Curves.
Visual Mapping
Use --color-by and --style-by to map run-level metadata fields to visual properties. Each unique value gets a distinct color or linestyle.
# Color by model, linestyle by monitor variant ct metrics roc --run-file runs.yaml --color-by model --style-by monitor
Auto-resolved fields
When the attack source is a run ID, run URL, or a local directory containing run.json, these fields are resolved from the run's metadata and need not appear in the YAML run file:
| Field | Source in run metadata |
|---|---|
monitor | metadata.monitor |
model | metadata.monitor_model_usage.model.model |
attack-model | metadata.policy_selection.args.model |
attack-policy | metadata.policy_selection.name |
Values set in the YAML take precedence. If a --color-by or --style-by field is in neither the YAML nor the run metadata, the command errors and lists the runs missing it.
Grouping
Use --group-by to draw one curve per run and group value, all on the monitor's plot.
ct metrics roc --run-file runs.yaml --group-by env ct metrics roc --honest <run> --attack <run> --group-by failure_type
| Key | Groups honest? | Groups attack? | Description |
|---|---|---|---|
env | Yes (paired) | Yes | Environment ID |
main_task_category | Yes (paired) | Yes | Main task category |
failure_type | No (shared) | Yes | Side task failure type (live/delayed) |
side_task_category | No (shared) | Yes | Side task category |
"Paired" means a group's curve uses only the honest trajectories with the same value. "Shared" means every group's curve uses all the run's honest trajectories. With one run a curve is labelled by its group value; with several, by <run> — <group>, colored by group and styled by run.
Use --min-samples N (default 2) to leave out a curve with fewer scored trajectories on either side.
Filtering
Use --filter-by to select runs from a run file; filters can be repeated, and all must match:
ct metrics roc --run-file runs.yaml --filter-by model=gpt-4.1-nano --filter-by monitor=mean-action
Output
data/roc/<timestamp>/ # or --output-dir
roc.png # with one monitor
roc_<monitor>.png # with several, one per monitor ("/" in a name becomes "_")
roc_combined.png # with several, a grid of them all
AUROC per monitor and curve is printed as each plot is drawn.
Run ct metrics roc --help for every option.