How We Measure
Goals
Each of our studies aims to answer a specific question within two related, but separate, registers: given a security task, (1) how much does AI change what an attacker can do, and (2) how much does AI change what a defender can do?
Delta Measurement
We run a task to set a baseline, then run it again with AI, while striving to maintain parity of conditions (authentication, access, and so on). What matters is the difference the AI makes when it is cleanly separated from any advantage of access or resources.
Environment Security
Before a study runs, we write down its protocol, its endpoints, and its analysis plan, and we freeze them cryptographically. The freeze is enforced by our harness: an altered protocol is refused, a widened study is refused, and the question therefore stays fixed before the answer is known.
Target Selection and Limitations
We run against synthetic environments we build and own, so the true state of the target is known by construction. Each study carries a guardrail against its own most likely way of being wrong, and states its limits in the protocol.
For more about how we work, take a look at our Labs.