6533b833fe1ef96bd129ba26

RESEARCH PRODUCT

DeepEva: A deep neural network architecture for assessing sentence complexity in Italian and English languages

Daniele SchicchiGiovanni PilatoGiosué Lo Bosco

subject

Artificial intelligenceComputer engineering. Computer hardwareText simplificationComputer scienceText simplificationcomputer.software_genreLexiconAutomatic-text-complexity-evaluationDeep-learningField (computer science)TK7885-7895Automatic text copmplexity evaluationText-complexity-assessmentText complexity assessmentStructure (mathematical logic)Settore INF/01 - InformaticaText-simplificationbusiness.industryDeep learningNatural language processingNatural-language-processingDeep learningGeneral MedicineQA75.5-76.95Artificial-intelligenceSupport vector machineElectronic computers. Computer scienceGradient boostingArtificial intelligencebusinesscomputerSentenceNatural language processing

description

Abstract Automatic Text Complexity Evaluation (ATE) is a research field that aims at creating new methodologies to make autonomous the process of the text complexity evaluation, that is the study of the text-linguistic features (e.g., lexical, syntactical, morphological) to measure the grade of comprehensibility of a text. ATE can affect positively several different contexts such as Finance, Health, and Education. Moreover, it can support the research on Automatic Text Simplification (ATS), a research area that deals with the study of new methods for transforming a text by changing its lexicon and structure to meet specific reader needs. In this paper, we illustrate an ATE approach named DeepEva, a Deep Learning based system capable of classifying both Italian and English sentences on the basis of their complexity. The system exploits the Treetagger annotation tool, two Long Short Term Memory (LSTM) neural unit layers, and a fully connected one. The last layer outputs the probability of a sentence belonging to the easy or complex class. The experimental results show the effectiveness of the approach for both languages, compared with several baselines such as Support Vector Machine, Gradient Boosting, and Random Forest.

10.1016/j.array.2021.100097http://www.sciencedirect.com/science/article/pii/S2590005621000424