<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-18T21:46:42Z</responseDate><request verb="GetRecord" identifier="oai:dspace.mit.edu:1721.1/115632" metadataPrefix="dim">https://dspace.mit.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:dspace.mit.edu:1721.1/115632</identifier><datestamp>2026-06-16T18:55:16Z</datestamp><setSpec>com_1721.1_7582</setSpec><setSpec>com_1721.1_7581</setSpec><setSpec>col_1721.1_131022</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor" lang="en_US">Patrick H. Winston.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author" lang="en_US">Kraft, Adam Davis</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="other" lang="en_US">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science.</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2018-05-23T15:05:36Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2018-05-23T15:05:36Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="copyright" lang="en_US">2018</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued" lang="en_US">2018</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">http://hdl.handle.net/1721.1/115632</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="oclc" lang="en_US">1036987419</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Thesis: Ph. D., Massachusetts Institute of Technology, Department of Electrical Engineering and Computer Science, 2018.</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">This electronic version was submitted by the student author.  The certified thesis is available in the Institute Archives and Special Collections.</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Cataloged from student-submitted PDF version of thesis.</dim:field>
   <dim:field mdschema="dc" element="description" lang="en_US">Includes bibliographical references (pages 127-134).</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract" lang="en_US">Human visual intelligence is robust. Vision is versatile in its variety of tasks and operating conditions, it is flexible, adapting facilely to new tasks, and it is introspective, providing compositional explanations for its findings. Vision is fundamentally underdetermined, but it exists in a world that abounds with constraints and regularities perceived not only through vision but through other senses as well. These observations suggest that the imperative of vision is to exploit all sources of information to resolve ambiguity. I propose an alignment model for vision, in which computational specialists eagerly share state with their neighbors during ongoing computations, availing themselves of neighbors' partial results in order to ll gaps in evolving descriptions. Connections between specialists extend across sensory modalities, so that the computational machinery of many senses may be brought to bear on problems with strictly-visual inputs. I anticipate that this alignment process accounts for vision's robust attributes, and I call this prediction the alignment hypothesis. In this document I lay the groundwork for evaluating the hypothesis. I then demonstrate progress toward that goal, by way of the following contributions: -- I performed an experiment to investigate and characterize the ways that high-performing computer-vision models fall short of robust perception, and evaluated whether alignment models can address the shortcomings. The experiment, which relied on a procedure to remove signal energy from natural images while preserving high classication condence by a neural network, revealed that the type of object depicted in the original image is a strong predictor of whether humans recognize the reduced-energy image. -- I implemented an alignment model based on a network of propagators. The model can use constraints to infer locations and heights of pedestrians and locations of occluding objects in an outdoor urban scene. I used the results of the effort to refine the requirements of mechanisms to use in building alignment models. -- I implemented an alignment model based on neural networks. Alignment-motivated design empowers the model, trained to estimate depth maps from single images, to perform the additional task of depth super-resolution without retraining. The design thus demonstrates flexibility, a property of robust vision systems.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="statementofresponsibility" lang="en_US">by Adam Davis Kraft.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="degree" lang="en_US">Ph.D.</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="extent" lang="en_US">134 pages</dim:field>
   <dim:field mdschema="dc" element="language" qualifier="iso" lang="en_US">eng</dim:field>
   <dim:field mdschema="dc" element="publisher" lang="en_US">Massachusetts Institute of Technology</dim:field>
   <dim:field mdschema="dc" element="rights" lang="en_US">MIT theses are protected by copyright. They may be viewed, downloaded, or printed from this source but further reproduction or distribution in any format is prohibited without written permission.</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri" lang="en_US">http://dspace.mit.edu/handle/1721.1/7582</dim:field>
   <dim:field mdschema="dc" element="subject" lang="en_US">Electrical Engineering and Computer Science.</dim:field>
   <dim:field mdschema="dc" element="title" lang="en_US">Vision by alignment</dim:field>
   <dim:field mdschema="dc" element="type" lang="en_US">Thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="dspace" element="authorsordered">false</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="4d0dffc9-7048-4831-96b4-8ae995e0254c">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
	&lt;Language>eng&lt;/Language>
   	&lt;Title>Vision by alignment&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2018&lt;/PublicationDate>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Kraft, Adam Davis&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;DisplayName>Massachusetts Institute of Technology&lt;/DisplayName>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>http://dspace.mit.edu/handle/1721.1/7582&lt;/License>
    &lt;Keyword>Electrical Engineering and Computer Science.&lt;/Keyword>
   	&lt;Abstract>Human visual intelligence is robust. Vision is versatile in its variety of tasks and operating conditions, it is flexible, adapting facilely to new tasks, and it is introspective, providing compositional explanations for its findings. Vision is fundamentally underdetermined, but it exists in a world that abounds with constraints and regularities perceived not only through vision but through other senses as well. These observations suggest that the imperative of vision is to exploit all sources of information to resolve ambiguity. I propose an alignment model for vision, in which computational specialists eagerly share state with their neighbors during ongoing computations, availing themselves of neighbors&amp;apos; partial results in order to ll gaps in evolving descriptions. Connections between specialists extend across sensory modalities, so that the computational machinery of many senses may be brought to bear on problems with strictly-visual inputs. I anticipate that this alignment process accounts for vision&amp;apos;s robust attributes, and I call this prediction the alignment hypothesis. In this document I lay the groundwork for evaluating the hypothesis. I then demonstrate progress toward that goal, by way of the following contributions: -- I performed an experiment to investigate and characterize the ways that high-performing computer-vision models fall short of robust perception, and evaluated whether alignment models can address the shortcomings. The experiment, which relied on a procedure to remove signal energy from natural images while preserving high classication condence by a neural network, revealed that the type of object depicted in the original image is a strong predictor of whether humans recognize the reduced-energy image. -- I implemented an alignment model based on a network of propagators. The model can use constraints to infer locations and heights of pedestrians and locations of occluding objects in an outdoor urban scene. I used the results of the effort to refine the requirements of mechanisms to use in building alignment models. -- I implemented an alignment model based on neural networks. Alignment-motivated design empowers the model, trained to estimate depth maps from single images, to perform the additional task of depth super-resolution without retraining. The design thus demonstrates flexibility, a property of robust vision systems.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>