AI flag accuracy
Statistically indistinguishable from the 89.0% expert-human baseline.
Open research + software
We tested five AI models on customer screening for synthetic DNA orders. The best matched expert-human accuracy on four flag criteria at about one-tenth the cost. We published the data, prompts, and a tool you can try.
Results
Gemini 2.5 Pro with web, bibliographic, and sanctions tools was the strongest system in our 41-profile comparison across five information-gathering tasks.
Statistically indistinguishable from the 89.0% expert-human baseline.
Per customer, compared with $14.04 for manual screening.
Average AI cost before human review, about 50 times cheaper than manual screening.
The study evaluated information gathering, source quality, source fidelity, and flags—not fulfillment decisions. Human reviewers retained authority over follow-up and fulfillment.
Try it
Enter a public, fictional, or authorized customer profile. This demo adapts the paper's prompts, searches public records, and returns cited evidence for review. The deployed demo itself was not evaluated.
Open the toolOpen by default
Developed by
Hanna Pálya is a PhD student in mathematical epidemiology at the University of Warwick who has researched regulatory routes for DNA synthesis screening.
Alejandro Acelas is a data scientist and developer who has researched risks at the intersection of AI and biosecurity.
Run a screening