BERT Embeddings for Automatic Readability Assessment

Joseph Marvin Imperial

2021-06-15RANLP 2021 9Text Classification

Abstract

Automatic readability assessment (ARA) is the task of evaluating the level of ease or difficulty of text documents for a target audience. For researchers, one of the many open problems in the field is to make such models trained for the task show efficacy even for low-resource languages. In this study, we propose an alternative way of utilizing the information-rich embeddings of BERT models with handcrafted linguistic features through a combined method for readability assessment. Results show that the proposed method outperforms classical approaches in readability assessment using English and Filipino datasets, obtaining as high as 12.4% increase in F1 performance. We also show that the general information encoded in BERT embeddings can be used as a substitute feature set for low-resource languages like Filipino with limited semantic and syntactic NLP tools to explicitly extract feature values for the task.

Results

Task	Dataset	Metric	Value	Model
Text Classification	OneStopEnglish (Readability Assessment)	Accuracy (5-fold)	0.744	Logistic Regression
Classification	OneStopEnglish (Readability Assessment)	Accuracy (5-fold)	0.744	Logistic Regression

Related Papers

Making Language Model a Hierarchical Classifier and Generator2025-07-17 GNN-CNN: An Efficient Hybrid Model of Convolutional and Graph Neural Networks for Text Representation2025-07-10 The Trilemma of Truth in Large Language Models2025-06-30 Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack2025-06-30 Perspectives in Play: A Multi-Perspective Approach for More Inclusive NLP Systems2025-06-25 Can Generated Images Serve as a Viable Modality for Text-Centric Multimodal Learning?2025-06-21 SHREC and PHEONA: Using Large Language Models to Advance Next-Generation Computational Phenotyping2025-06-19 Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages2025-06-12