Neural Feature Fields for Language-Guided Robot Manipulation
Name
shen-willshen-sm-eecs-2023-thesis.pdf
Description
Thesis PDF
Size
12.14 MB
Format
Adobe PDF
Checksum (MD5)
ad63283203ea246484d33fd075692000
Author(s)
Shen, William
Advisor(s)
Kaelbling, Leslie P.
Lozano-Pérez, Tomás
Date Issued
September 2023
Publisher
Massachusetts Institute of Technology
Abstract
Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed understanding of 3D geometry, which is often lacking in 2D image features. This work bridges this 2D-to-3D gap for robotic manipulation by leveraging distilled feature fields to combine accurate 3D geometry with rich semantics from 2D foundation models. We present a few-shot learning method for 6-DOF grasping and placing that harnesses these strong spatial and semantic priors to achieve in-the-wild generalization to unseen objects. Using features distilled from a vision-language model, CLIP, we present a way to designate novel objects for manipulation via free-text natural language, and demonstrate its ability to generalize to unseen expressions and novel categories of objects.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Terms of Use
In Copyright - Educational Use Permitted
Copyright retained by author(s)
Persistent DSpace Link