A Joint Probabilistic Classification Model of Relevant and Irrelevant Sentences in Mathematical Word Problems

Autor/inn/en	Cetintas, Suleyman; Si, Luo; Xin, Yan Ping; Zhang, Dake; Park, Joo Young; Tzur, Ron
Titel	A Joint Probabilistic Classification Model of Relevant and Irrelevant Sentences in Mathematical Word Problems
Quelle	In: Journal of Educational Data Mining, 2 (2010) 1, S.83-101 (19 Seiten) PDF als Volltext Verfügbarkeit
Sprache	englisch
Dokumenttyp	gedruckt; online; Zeitschriftenaufsatz
ISSN	2157-2100
Schlagwörter	Probability; Word Problems (Mathematics); Classification; Difficulty Level; Identification; Sentences; Models; Correlation; Intelligent Tutoring Systems; Mathematics Education; Computational Linguistics; Statistical Analysis; Mathematical Formulas; Textbook Content + Suchen Sie Ihr Suchwort? Wahrscheinlichkeitsrechnung; Wahrscheinlichkeitstheorie; Textaufgabe; Classification system; Klassifikation; Klassifikationssystem; Schwierigkeitsgrad; Identifikation; Identifizierung; Sentence analysis; Satzanalyse; Analogiemodell; Korrelation; Intelligentes Tutorsystem; Mathematische Bildung; Linguistics; Computerlinguistik; Statistische Analyse; Mathematische Formel; Lehrbuchtext
Abstract	Estimating the difficulty level of math word problems is an important task for many educational applications. Identification of relevant and irrelevant sentences in math word problems is an important step for calculating the difficulty levels of such problems. This paper addresses a novel application of text categorization to identify two types of sentences in "mathematical word problems", namely "relevant" and "irrelevant" sentences. A novel joint probabilistic classification model is proposed to estimate the joint probability of classification decisions for all sentences of a math word problem by utilizing the correlation among all sentences along with the correlation between the question sentence and other sentences, and sentence text. The proposed model is compared with (i) a SVM classifier which makes independent classification decisions for individual sentences by only using the sentence text, and (ii) a novel SVM classifier that considers the correlation between the question sentence and other sentences along with the sentence text. An extensive set of experiments demonstrates the effectiveness of the joint probabilistic classification model for identifying relevant and irrelevant sentences, as well as the novel SVM classifier that utilizes the correlation between the question sentence and other sentences. Furthermore, empirical results and analysis show that (i) it is highly beneficial not to remove stopwords, and (ii) utilizing part of speech tagging does not make a significant improvement although it has been shown to be effective for the related task of math word problem type classification. (As Provided).
Anmerkungen	International Working Group on Educational Data Mining. e-mail: jedm.editor@gmail.com; Web site: http://www.educationaldatamining.org/JEDM/index.php/JEDM/index
Erfasst von	ERIC (Education Resources Information Center), Washington, DC
Update	2020/1/01