<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/style.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-19T20:59:51Z</responseDate><request verb="GetRecord" identifier="oai:dspace.mit.edu:1721.1/151608" metadataPrefix="dim">https://dspace.mit.edu/server/oai/request</request><GetRecord><record><header><identifier>oai:dspace.mit.edu:1721.1/151608</identifier><datestamp>2023-08-01T03:19:54Z</datestamp><setSpec>com_1721.1_7582</setSpec><setSpec>com_1721.1_7581</setSpec><setSpec>col_1721.1_131023</setSpec></header><metadata><dim:dim xmlns:dim="http://www.dspace.org/xmlns/dspace/dim" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:doc="http://www.lyncode.com/xoai" xsi:schemaLocation="http://www.dspace.org/xmlns/dspace/dim http://www.dspace.org/schema/dim.xsd">
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Oliva, Aude</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="advisor">Feris, Rogerio</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="author">Purohit, Sonia</dim:field>
   <dim:field mdschema="dc" element="contributor" qualifier="department">Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="accessioned">2023-07-31T19:52:20Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="available">2023-07-31T19:52:20Z</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="issued">2023-06</dim:field>
   <dim:field mdschema="dc" element="date" qualifier="submitted">2023-06-06T16:35:17.730Z</dim:field>
   <dim:field mdschema="dc" element="identifier" qualifier="uri">https://hdl.handle.net/1721.1/151608</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="abstract">Automated visual understanding is an essential part of the sports industry, particularly in the context of major sports tournaments. The scale of generated video footage necessitates the use of automated systems to generate insights and enhance fan experiences. One area where this is particularly challenging is commentary, which requires detailed information about play-by-play action, a task that cannot be efficiently carried out by human commentators at scale.&#xd;
&#xd;
We tackle this problem for grand-slam tennis through an IBM partnership with the Championships, Wimbledon. This thesis introduces a novel system that utilizes computer vision to extract play-by-play metadata and convert it into fluent commentary using large language models. Our computer vision module utilizes a single camera feed to understand every detail of the game – court and net detection, player and ball tracking, player poses, and fine-grained shot classification, all in near-real-time. This metadata is then combined with additional information from other modalities, such as crowd audio and radar-measured ball speed, and fed into a "data2text" large language model to generate commentary in natural language.&#xd;
&#xd;
Our system not only supports the narration of match content at scale, but also powers the collection of additional metadata to facilitate additional match insights in the future.</dim:field>
   <dim:field mdschema="dc" element="description" qualifier="degree">M.Eng.</dim:field>
   <dim:field mdschema="dc" element="publisher">Massachusetts Institute of Technology</dim:field>
   <dim:field mdschema="dc" element="rights">In Copyright - Educational Use Permitted</dim:field>
   <dim:field mdschema="dc" element="rights">Copyright retained by author(s)</dim:field>
   <dim:field mdschema="dc" element="rights" qualifier="uri">https://rightsstatements.org/page/InC-EDU/1.0/</dim:field>
   <dim:field mdschema="dc" element="title">AI Commentator: Narrating Sports Games through&#xd;
Multimodal Perception and Large Language Models</dim:field>
   <dim:field mdschema="dc" element="type">Thesis</dim:field>
   <dim:field mdschema="dc" element="format" qualifier="mimetype">application/pdf</dim:field>
   <dim:field mdschema="mit" element="thesis" qualifier="degree">Master</dim:field>
   <dim:field mdschema="thesis" element="degree" qualifier="name">Master of Engineering in Electrical Engineering and Computer Science</dim:field>
   <dim:field mdschema="dspace" element="entity" qualifier="type">Publication</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="others" element="access-status">unknown</dim:field>
   <dim:field mdschema="cerif" element="openaire" authority="" confidence="-1">&lt;Publication xmlns="https://www.openaire.eu/cerif-profile/1.1/" id="be75a659-2cbb-4c17-bd8e-e72bb99f91c6">
	&lt;Type xmlns="https://www.openaire.eu/cerif-profile/vocab/COAR_Publication_Types">http://purl.org/coar/resource_type/c_1843&lt;/Type>
   	&lt;Title>AI Commentator: Narrating Sports Games through&#xd;
Multimodal Perception and Large Language Models&lt;/Title>
   	&lt;PublishedIn>
    	&lt;Publication>
      	&lt;/Publication>
   	&lt;/PublishedIn>
   	&lt;PublicationDate>2023-06&lt;/PublicationDate>
   	&lt;Authors>
      	&lt;Author>
        	&lt;DisplayName>Purohit, Sonia&lt;/DisplayName>
         	&lt;Affiliation>
         		&lt;OrgUnit>
         		&lt;/OrgUnit>
         	&lt;/Affiliation>
      	&lt;/Author>
	&lt;/Authors>
   	&lt;Editors>
	&lt;/Editors>
    &lt;Publishers>
        &lt;Publisher>
            &lt;DisplayName>Massachusetts Institute of Technology&lt;/DisplayName>
            &lt;OrgUnit />
        &lt;/Publisher>
    &lt;/Publishers>
    &lt;License>https://rightsstatements.org/page/InC-EDU/1.0/&lt;/License>
   	&lt;Abstract>Automated visual understanding is an essential part of the sports industry, particularly in the context of major sports tournaments. The scale of generated video footage necessitates the use of automated systems to generate insights and enhance fan experiences. One area where this is particularly challenging is commentary, which requires detailed information about play-by-play action, a task that cannot be efficiently carried out by human commentators at scale.&#xd;
&#xd;
We tackle this problem for grand-slam tennis through an IBM partnership with the Championships, Wimbledon. This thesis introduces a novel system that utilizes computer vision to extract play-by-play metadata and convert it into fluent commentary using large language models. Our computer vision module utilizes a single camera feed to understand every detail of the game – court and net detection, player and ball tracking, player poses, and fine-grained shot classification, all in near-real-time. This metadata is then combined with additional information from other modalities, such as crowd audio and radar-measured ball speed, and fed into a &amp;quot;data2text&amp;quot; large language model to generate commentary in natural language.&#xd;
&#xd;
Our system not only supports the narration of match content at scale, but also powers the collection of additional metadata to facilitate additional match insights in the future.&lt;/Abstract>
	&lt;Access xmlns="http://purl.org/coar/access_right" 
    >
    &lt;/Access>
&lt;/Publication>
</dim:field>
</dim:dim>
</metadata></record></GetRecord></OAI-PMH>