Discovery and recognition of formula concepts using machine learning

2023 | journal article. A publication with affiliation to the University of Göttingen.

Jump to: Cite & Linked | Documents & Media | Details | Version history

Cite this publication

​Discovery and recognition of formula concepts using machine learning​
Scharpf, P.; Schubotz, M. ; Cohl, H. S.; Breitinger, C.   & Gipp, B. ​ (2023) 
Scientometrics128(9) pp. 4971​-5025​.​ DOI: https://doi.org/10.1007/s11192-023-04667-9 

Documents & Media

License

GRO License GRO License

Details

Authors
Scharpf, Philipp; Schubotz, Moritz ; Cohl, Howard S.; Breitinger, Corinna ; Gipp, Béla 
Abstract
Abstract Citation-based Information Retrieval (IR) methods for scientific documents have proven effective for IR applications, such as Plagiarism Detection or Literature Recommender Systems in academic disciplines that use many references. In science, technology, engineering, and mathematics, researchers often employ mathematical concepts through formula notation to refer to prior knowledge. Our long-term goal is to generalize citation-based IR methods and apply this generalized method to both classical references and mathematical concepts. In this paper, we suggest how mathematical formulas could be cited and define a Formula Concept Retrieval task with two subtasks: Formula Concept Discovery (FCD) and Formula Concept Recognition (FCR). While FCD aims at the definition and exploration of a ‘Formula Concept’ that names bundled equivalent representations of a formula, FCR is designed to match a given formula to a prior assigned unique mathematical concept identifier. We present machine learning-based approaches to address the FCD and FCR tasks. We then evaluate these approaches on a standardized test collection (NTCIR arXiv dataset). Our FCD approach yields a precision of 68% for retrieving equivalent representations of frequent formulas and a recall of 72% for extracting the formula name from the surrounding text. FCD and FCR enable the citation of formulas within mathematical documents and facilitate semantic search and question answering, as well as document similarity assessments for plagiarism detection or recommender systems.
Issue Date
2023
Journal
Scientometrics 
Organization
Institut für Informatik ; Niedersächsische Staats- und Universitätsbibliothek Göttingen 
ISSN
0138-9130
eISSN
1588-2861
Language
English
Sponsor
Deutsche Forschungsgemeinschaft http://dx.doi.org/10.13039/501100001659
Deutsche Forschungsgemeinschaft http://dx.doi.org/10.13039/501100001659
Niedersächsisches Ministerium für Wissenschaft und Kultur http://dx.doi.org/10.13039/501100010570
Volkswagen Foundation http://dx.doi.org/10.13039/501100001663
Georg-August-Universität Göttingen http://dx.doi.org/10.13039/501100003385

Reference

Citations


Social Media