<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T21:58:19Z</responseDate><request verb="GetRecord" identifier="oai:dspace.mit.edu:1721.1/147435" metadataPrefix="dim">https://dspace.mit.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:dspace.mit.edu:1721.1/147435</identifier><datestamp>2023-01-20T03:37:06Z</datestamp><setSpec>com_1721.1_7582</setSpec><setSpec>com_1721.1_7581</setSpec><setSpec>col_1721.1_131023</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Kim, Sangbae</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author">Ackerman, Liam J.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2023-01-19T19:50:12Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2023-01-19T19:50:12Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2022-09</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="submitted">2022-09-16T20:23:57.251Z</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1721.1/147435</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract">Deep reinforcement learning has been used to craft robust and performant control policies for legged robotics. However, the engineering processes to create these policies are often plagued by long training times that slow down engineering iteration. This thesis suggests that model-based controllers offer a wealth of successful computation that may be used within reinforcement learning control pipelines to improve learning efficiency. Two ideas incorporate this engineering expertise to increase reinforcement learning efficiency. First, successful model-based computations are pre-processed and incorporated directly into network observations. Introducing these terms into the reinforcement learning architecture is shown to increase learning speeds and policy performance dramatically. Next, inspired by model-based task hierarchies, more structure is added to the reinforcement learning objective function to activate and deactivate reward terms based on an agent’s state. This structure is intended to avoid local minima which impede learning. This reward restructure is shown to avoid local minima during training but degrades final policy performance at edge-cases.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="degree">M.Eng.</dim:field>
   <dim:field mdschema="dc" element="publisher">Massachusetts Institute of Technology</dim:field>
   <dim:field mdschema="dc" element="rights">In Copyright - Educational Use Permitted</dim:field>
   <dim:field mdschema="dc" element="rights">Copyright MIT</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri">http://rightsstatements.org/page/InC-EDU/1.0/</dim:field>
   <dim:field mdschema="dc" element="title">Leveraging Engineering Expertise in Deep Reinforcement Learning</dim:field>
   <dim:field mdschema="dc" element="type">Thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="mit" element="thesis" qualifier="degree">Master</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="name">Master of Engineering in Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="6dc4d3cf-f037-45cd-9eee-b06dc763f6d0">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
   	&lt;Title>Leveraging Engineering Expertise in Deep Reinforcement Learning&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2022-09&lt;/PublicationDate>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Ackerman, Liam J.&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;DisplayName>Massachusetts Institute of Technology&lt;/DisplayName>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>http://rightsstatements.org/page/InC-EDU/1.0/&lt;/License>
   	&lt;Abstract>Deep reinforcement learning has been used to craft robust and performant control policies for legged robotics. However, the engineering processes to create these policies are often plagued by long training times that slow down engineering iteration. This thesis suggests that model-based controllers offer a wealth of successful computation that may be used within reinforcement learning control pipelines to improve learning efficiency. Two ideas incorporate this engineering expertise to increase reinforcement learning efficiency. First, successful model-based computations are pre-processed and incorporated directly into network observations. Introducing these terms into the reinforcement learning architecture is shown to increase learning speeds and policy performance dramatically. Next, inspired by model-based task hierarchies, more structure is added to the reinforcement learning objective function to activate and deactivate reward terms based on an agent’s state. This structure is intended to avoid local minima which impede learning. This reward restructure is shown to avoid local minima during training but degrades final policy performance at edge-cases.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>