<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T05:15:59Z</responseDate><request verb="GetRecord" identifier="oai:dspace.mit.edu:1721.1/124251" metadataPrefix="dim">https://dspace.mit.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:dspace.mit.edu:1721.1/124251</identifier><datestamp>2026-06-06T00:49:00Z</datestamp><setSpec>com_1721.1_7582</setSpec><setSpec>com_1721.1_7581</setSpec><setSpec>col_1721.1_131023</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor" lang="en_US">Leslie Pack Kaelbling.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author" lang="en_US">LaGrassa, Alex Licari.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="other" lang="en_US">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department" lang="en_US">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2020-03-24T15:36:27Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2020-03-24T15:36:27Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="copyright" lang="en_US">2019</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued" lang="en_US">2019</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1721.1/124251</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="oclc" lang="en_US">1145122826</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">This electronic version was submitted by the student author. The certified thesis is available in the Institute Archives and Special Collections.</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Thesis: M. Eng., Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2019</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Cataloged from student-submitted PDF version of thesis.</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Includes bibliographical references (pages 73-78).</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en_US">Engineering reinforcement learning agents for application on a particular target domain requires making decisions such as the learning algorithm and state representation. We empirically study the performance of three reference implementations of model-free reinforcement learning algorithms: Covariance Matrix Adaptation Evolution Strategy, Deep Deterministic Policy Gradients, and Proximal Policy Optimization. We compare their performance on various target domains to measure quantitatively their dependence on varied features of the environment. We study the effect of actuation noise, observation noise, reward sparsity and task horizon. Then, we explore automatically generated state encodings for learning using a lower-dimensional encoding from high dimensional sensor data. A proof-of- concept end-to-end system for scooping beads of different sizes in the real world generates, uses, then follows force traces along with a positional controller to execute a scoop.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="statementofresponsibility" lang="en_US">by Alex Licari LaGrassa.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="degree" lang="en_US">M.Eng.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="collection" lang="en_US">M.Eng. Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="extent" lang="en_US">78 pages</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso" lang="en_US">eng</dim:field>
   <dim:field mdschema="dc" element="publisher" lang="en_US">Massachusetts Institute of Technology</dim:field>
   <dim:field mdschema="dc" element="rights" lang="en_US">MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri" lang="en_US">http://dspace.mit.edu/handle/1721.1/7582</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Electrical Engineering and Computer Science.</dim:field>
   <dim:field mdschema="dc" element="title" lang="en_US">Selecting appropriate reinforcement-learning algorithms for robot manipulation domains</dim:field>
   <dim:field mdschema="dc" element="type" lang="en_US">Thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="dspace" element="imported" lang="en_US">2020-03-24T15:36:26Z</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="mit" element="thesis" qualifier="degree" lang="en_US">Master</dim:field>
   <dim:field mdschema="mit" element="thesis" qualifier="department" lang="en_US">EECS</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="24173d7e-67a5-46ac-b5d7-9dcacea3570a">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
	&lt;Language>eng&lt;/Language>
   	&lt;Title>Selecting appropriate reinforcement-learning algorithms for robot manipulation domains&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2019&lt;/PublicationDate>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>LaGrassa, Alex Licari.&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;DisplayName>Massachusetts Institute of Technology&lt;/DisplayName>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>http://dspace.mit.edu/handle/1721.1/7582&lt;/License>
    &lt;Keyword>Electrical Engineering and Computer Science.&lt;/Keyword>
   	&lt;Abstract>Engineering reinforcement learning agents for application on a particular target domain requires making decisions such as the learning algorithm and state representation. We empirically study the performance of three reference implementations of model-free reinforcement learning algorithms: Covariance Matrix Adaptation Evolution Strategy, Deep Deterministic Policy Gradients, and Proximal Policy Optimization. We compare their performance on various target domains to measure quantitatively their dependence on varied features of the environment. We study the effect of actuation noise, observation noise, reward sparsity and task horizon. Then, we explore automatically generated state encodings for learning using a lower-dimensional encoding from high dimensional sensor data. A proof-of- concept end-to-end system for scooping beads of different sizes in the real world generates, uses, then follows force traces along with a positional controller to execute a scoop.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>