Lure quality in a scaling capability environment
The question
The cheapest attackers send the worst lures. This study asks whether AI closes that gap: whether a hobbyist-grade model now writes business-email and phishing lures that used to take a professional operation to produce.
The method
Lab A runs a ladder of self-hosted models across today's capability range, entirely inside an isolated range we own, and measures how lure quality scales as capability climbs. A human panel rates the output under our documented human-subjects standard. We measure the effect. We publish the measurement and withhold the method: no lure, no generation recipe, and no reusable artifact leaves the study. This is offense-side research run to help the defense, and the line between those is where we keep it.
Publication transparency
How lure quality scales with capability, where a hobbyist-tier model crosses the professional bar if it does, and the protocol. No generated content.