ct metrics safety

Usefulness, side-task success, and safety at each audit budget with the monitors auditing together and each alone, per run and per attack group.

Usage

ct metrics safety [OPTIONS]

Options

OptionDescription
--honest TEXTHonest trajectories (path, dir, run ID or run URL). Mutually exclusive with --run-file.
--attack TEXTAttack trajectories (path, dir, run ID or run URL). Mutually exclusive with --run-file.
--run-file PATHYAML run file with several honest/attack entries.
--filter-by TEXTKeep run-file entries whose field has this value (field=value, repeatable).
--honest-task-set PATHKeep the honest trajectories whose (env, main_task) this JSONL file lists. A run-file entry's honest_task_set takes precedence.
--attack-task-set PATHKeep the attack trajectories whose (env, main_task, side_task) this JSONL file lists. A run-file entry's attack_task_set takes precedence.
--group-by [env|failure_type|main_task_category|side_task_category]Also report each group of attacks by this field.
--audit-budgets FLOAT RANGEAudit budgets as fractions of the honest audit opportunities (0.01 is 1%). [default: (0.0001 to 0.5); 0<=x<=1]
--bootstrap INTEGER RANGEResample honest and attack trajectories this many times for 95% safety intervals. [default: 0; x>=0]
--output-dir PATHDirectory for metrics.json and safety_curves.png (default: data/metrics/<timestamp>).
--max-workers INTEGER RANGEThread-pool size for loading run entries concurrently. [default: 8; x>=1]
--helpShow this message and exit.