The question
A security team drowns in alerts, most of them noise. This study asks how much of the triage an AI can carry, and, because capability keeps moving, how fast that share is growing.
The method
Lab A takes a labelled corpus of real-shape alerts and runs it across a ladder of models spanning today's capability range, against a human-analyst baseline on the same corpus. We fit the result against a capability axis and report the slope with its uncertainty, so the finding is a measured trend rather than a single snapshot. A well-resourced defender can use frontier models our self-hosted ladder cannot reach, and we state that truncation as a limit on the defensive read.
Publication transparency
Triage accuracy across the capability range against the human baseline, the growth slope with uncertainty, and the protocol.