The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures
Author(s)
Mysore, Sheshera; Jensen, Zach; Kim, Edward; Huang, Kevin Joon-Ming; Chang, Haw-Shiuan; Strubell, Emma; Flanigan, Jeffrey; McCallum, Andrew; Olivetti, Elsa A.; ... Show more Show less
DownloadAccepted version (349.5Kb)
Open Access Policy
Open Access Policy
Creative Commons Attribution-Noncommercial-Share Alike
Terms of use
Metadata
Show full item recordAbstract
© 2019 Association for Computational Linguistics Materials science literature contains millions of materials synthesis procedures described in unstructured natural language text. Large-scale analysis of these synthesis procedures would facilitate deeper scientific understanding of materials synthesis and enable automated synthesis planning. Such analysis requires extracting structured representations of synthesis procedures from the raw text as a first step. To facilitate the training and evaluation of synthesis extraction models, we introduce a dataset of 230 synthesis procedures annotated by domain experts with labeled graphs that express the semantics of the synthesis sentences. The nodes in this graph are synthesis operations and their typed arguments, and labeled edges specify relations between the nodes. We describe this new resource in detail and highlight some specific challenges to annotating scientific text with shallow semantic structure. We make the corpus available to the community to promote further research and development of scientific information extraction systems.
Date issued
2019-07Department
Massachusetts Institute of Technology. Department of Materials Science and Engineering; Massachusetts Institute of Technology. Institute for Data, Systems, and SocietyJournal
LAW 2019 - 13th Linguistic Annotation Workshop, Proceedings of the Workshop
Citation
2019. "The materials science procedural text corpus: Annotating materials synthesis procedures with shallow semantic structures." LAW 2019 - 13th Linguistic Annotation Workshop, Proceedings of the Workshop.
Version: Author's final manuscript