This is not the latest version of this item. The latest version can be found here.
A machine learning toolkit for genetic engineering attribution to facilitate biosecurity
Name
s41467-020-19612-0.pdf
Description
Published version
Size
2.19 MB
Format
Adobe PDF
Checksum (MD5)
a89e132ffa3a0dcd8605ebdba06f8061
Author(s) • • • • • • • •
Alley, Ethan C
Turpin, Miles
Liu, Andrew Bo
Kulp-McDowall, Taylor
Swett, Jacob
Edison, Rey
Von Stetina, Stephen E
Church, George M
Esvelt, Kevin M
Date Issued
2020
Journal
Nature Communications
Publisher
Springer Science and Business Media LLC
Version
Final published version
Abstract
© 2020, The Author(s). The promise of biotechnology is tempered by its potential for accidental or deliberate misuse. Reliably identifying telltale signatures characteristic to different genetic designers, termed ‘genetic engineering attribution’, would deter misuse, yet is still considered unsolved. Here, we show that recurrent neural networks trained on DNA motifs and basic phenotype data can reach 70% attribution accuracy in distinguishing between over 1,300 labs. To make these models usable in practice, we introduce a framework for weighing predictions against other investigative evidence using calibration, and bring our model to within 1.6% of perfect calibration. Additionally, we demonstrate that simple models can accurately predict both the nation-state-of-origin and ancestor labs, forming the foundation of an integrated attribution toolkit which should promote responsible innovation and international security alike.
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
10.1038/s41467-020-19612-0