Textual patterns and machine learning classification in academic writing: a linguistic analysis of theses and dissertations from a Brazilian graduate program
This study investigates linguistic patterns in academic texts produced within the Graduate Program in Linguistic Studies (PosLin) at the Federal University of Minas Gerais. A corpus comprising 1,270 documents, 730 master's dissertations and 540 doctoral theses was compiled and analyzed using co...
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Online |
| Language: | English |
| Published: |
Universidade Estadual de Campinas
2025
|
| Subjects: | |
| Online Access: | https://econtents.sbu.unicamp.br/inpec/index.php/joss/article/view/20586 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| Summary: | This study investigates linguistic patterns in academic texts produced within the Graduate Program in Linguistic Studies (PosLin) at the Federal University of Minas Gerais. A corpus comprising 1,270 documents, 730 master's dissertations and 540 doctoral theses was compiled and analyzed using computational linguistic techniques. Exploratory analyses included the extraction of unigrams, bigrams, trigrams, and the classification of the most frequent tokens into morphological categories (nouns, verbs, adjectives and adverbs). Despite the shared institutional context and research tracks, subtle differences in lexical and structural features were observed between the two academic levels. To evaluate whether these differences could support automated classification, machine learning models were trained on bag-of-words representations of the texts. Gradient Boosting emerged as the most effective algorithm, achieving an AUC of 0.989 with only the 1,000 most frequent tokens, demonstrating that high classification accuracy can be reached without extensive computational overhead. The results show that textual analysis combined with supervised learning can effectively distinguish academic genres within a single graduate program. Furthermore, the approach holds potential for broader applications in genre classification, fake news detection, and discourse analysis. This study also reinforces the importance of continued research in computational linguistics for underrepresented languages such as Brazilian Portuguese, especially in the context of formal and academic writing. |
|---|---|
| ISSN: | 2236-9740 |