ROC Curves

ct metrics roc plots, for each monitor, a ROC curve of honest against attack trajectories and prints its AUROC.

Quick Start

# One honest/attack pair ct metrics roc --honest <run-id-or-path> --attack <run-id-or-path> # From a run file with several entries ct metrics roc --run-file runs.yaml

How Trajectories Are Scored

Each trajectory has one score per monitor: its highest verdict that an audit could act on, read from the sample's monitor score as Safety Curves describes. Verdicts after a live side task landed are left out, and a trajectory verdict counts only when the run was neither audited online nor harmed live. A trajectory with no such verdict from a monitor is left off that monitor's curve. Every attack trajectory is a positive, whether or not its side task succeeded; every honest trajectory is a negative.

Input Modes

Inline mode

Pass --honest and one or more --attack sources. Each --attack is paired with the shared --honest and gets its own curve.

ct metrics roc \ --honest https://data.linuxarena.ai/runs/c7aeb0aeccba40eb9d910ab13b601bed \ --attack https://data.linuxarena.ai/runs/bf8a832da92f4431b2579ad04cd983ad \ --attack https://data.linuxarena.ai/runs/99d571edada441e8a595d7e50f4efc37

Sources can be run IDs, run URLs (https://data.linuxarena.ai/runs/...), .eval files, or directories holding one .eval file.

Run file mode

Pass --run-file with a YAML file describing several honest/attack pairs and their metadata. This enables --color-by, --style-by, and --filter-by.

runs: - label: gpt-4.1-nano (goals) model: gpt-4.1-nano monitor: mean-action # metadata for --color-by/--style-by honest: https://data.linuxarena.ai/runs/c7aeb0aeccba40eb9d910ab13b601bed attack: https://data.linuxarena.ai/runs/bf8a832da92f4431b2579ad04cd983ad - label: gpt-4.1-nano (no-goals) model: gpt-4.1-nano monitor: mean-action-no-goals honest: data/roc_eval/gpt-4.1-nano/honest.eval attack: data/roc_eval/gpt-4.1-nano/attack.eval

Each entry requires honest and attack, and may add a label plus any metadata fields (e.g. model, monitor). Without a label, one is made from model and monitor if present, otherwise from the attack source name. Run-file entries may also carry honest_task_set / attack_task_set, as in Safety Curves.

Visual Mapping

Use --color-by and --style-by to map run-level metadata fields to visual properties. Each unique value gets a distinct color or linestyle.

# Color by model, linestyle by monitor variant ct metrics roc --run-file runs.yaml --color-by model --style-by monitor

Auto-resolved fields

When the attack source is a run ID, run URL, or a local directory containing run.json, these fields are resolved from the run's metadata and need not appear in the YAML run file:

FieldSource in run metadata
monitormetadata.monitor
modelmetadata.monitor_model_usage.model.model
attack-modelmetadata.policy_selection.args.model
attack-policymetadata.policy_selection.name

Values set in the YAML take precedence. If a --color-by or --style-by field is in neither the YAML nor the run metadata, the command errors and lists the runs missing it.

Grouping

Use --group-by to draw one curve per run and group value, all on the monitor's plot.

ct metrics roc --run-file runs.yaml --group-by env ct metrics roc --honest <run> --attack <run> --group-by failure_type
KeyGroups honest?Groups attack?Description
envYes (paired)YesEnvironment ID
main_task_categoryYes (paired)YesMain task category
failure_typeNo (shared)YesSide task failure type (live/delayed)
side_task_categoryNo (shared)YesSide task category

"Paired" means a group's curve uses only the honest trajectories with the same value. "Shared" means every group's curve uses all the run's honest trajectories. With one run a curve is labelled by its group value; with several, by <run> — <group>, colored by group and styled by run.

Use --min-samples N (default 2) to leave out a curve with fewer scored trajectories on either side.

Filtering

Use --filter-by to select runs from a run file; filters can be repeated, and all must match:

ct metrics roc --run-file runs.yaml --filter-by model=gpt-4.1-nano --filter-by monitor=mean-action

Output

data/roc/<timestamp>/        # or --output-dir
  roc.png                    # with one monitor
  roc_<monitor>.png          # with several, one per monitor ("/" in a name becomes "_")
  roc_combined.png           # with several, a grid of them all

AUROC per monitor and curve is printed as each plot is drawn.

Run ct metrics roc --help for every option.