<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T19:51:46Z</responseDate><request verb="GetRecord" identifier="oai:dspace.mit.edu:1721.1/140985" metadataPrefix="dim">https://dspace.mit.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:dspace.mit.edu:1721.1/140985</identifier><datestamp>2022-03-04T03:13:05Z</datestamp><setSpec>com_1721.1_7582</setSpec><setSpec>com_1721.1_7581</setSpec><setSpec>col_1721.1_131023</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Esvelt, Kevin Michael</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author">Ethan Chase Alley</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department">Program in Media Arts and Sciences (Massachusetts Institute of Technology)</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2022-03-03T19:28:35Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2022-03-03T19:28:35Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2021-06</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="submitted">2022-02-27T16:47:23.914Z</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1721.1/140985</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="orcid">0000-0002-8219-7382</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract">The promise of biotechnology is tempered by its potential for accidental or deliberate misuse. Reliably identifying provenance by examining telltale signatures characteristic to different genetic designers, termed genetic engineering attribution, would deter misuse, yet is still considered unsolved. In this work, we present analysis of the biosecurity implications of improved tools for attribution, arguing that the technology has robust co-benefits for deterring misuse and promoting responsible innovation. Then, we demonstrate that recurrent neural networks trained on DNA motifs and basic phenotype data can reach 70% attribution accuracy distinguishing between over 1,300 labs. To make these models usable in practice, we introduce a framework for weighing predictions against other investigative evidence using calibration, and bring our model to within 1.6% of perfect calibration. Additionally, we demonstrate that simple models can accurately predict both the nation-state-of-origin and ancestor labs, forming the foundation of an integrated attribution toolkit which should promote responsible innovation and international security alike. Finally, we discuss ongoing work to crowdsource improved attribution tools via an open data science challenge.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="degree">S.M.</dim:field>
   <dim:field mdschema="dc" element="publisher">Massachusetts Institute of Technology</dim:field>
   <dim:field mdschema="dc" element="rights">In Copyright - Educational Use Permitted</dim:field>
   <dim:field mdschema="dc" element="rights">Copyright MIT</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri">http://rightsstatements.org/page/InC-EDU/1.0/</dim:field>
   <dim:field mdschema="dc" element="title">Machine learning to promote transparent provenance of genetic engineering</dim:field>
   <dim:field mdschema="dc" element="type">Thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="mit" element="thesis" qualifier="degree">Master</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="name">Master of Science</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="b563fd00-2b7c-4004-a9b6-ef0f461f8f2c">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
   	&lt;Title>Machine learning to promote transparent provenance of genetic engineering&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2021-06&lt;/PublicationDate>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Ethan Chase Alley&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;DisplayName>Massachusetts Institute of Technology&lt;/DisplayName>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>http://rightsstatements.org/page/InC-EDU/1.0/&lt;/License>
   	&lt;Abstract>The promise of biotechnology is tempered by its potential for accidental or deliberate misuse. Reliably identifying provenance by examining telltale signatures characteristic to different genetic designers, termed genetic engineering attribution, would deter misuse, yet is still considered unsolved. In this work, we present analysis of the biosecurity implications of improved tools for attribution, arguing that the technology has robust co-benefits for deterring misuse and promoting responsible innovation. Then, we demonstrate that recurrent neural networks trained on DNA motifs and basic phenotype data can reach 70% attribution accuracy distinguishing between over 1,300 labs. To make these models usable in practice, we introduce a framework for weighing predictions against other investigative evidence using calibration, and bring our model to within 1.6% of perfect calibration. Additionally, we demonstrate that simple models can accurately predict both the nation-state-of-origin and ancestor labs, forming the foundation of an integrated attribution toolkit which should promote responsible innovation and international security alike. Finally, we discuss ongoing work to crowdsource improved attribution tools via an open data science challenge.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>