The study proposes a multilevel methodological approach for analyzing Piani Triennali dell'Offerta Formativa (PTOF), the official documents through which Italian schools describe their educational projects. The framework combines two levels of analysis: Measurement level: texts from PTOF documents are represented through sentence embeddings generated by a BERT model. These representations are then processed using UMAP (dimensionality reduction) and HDBSCAN (clustering) to identify thematically coherent groups in a fully replicable manner. Interpretative level: text mining and lexicometric techniques are used to characterize the emerging clusters. Specifically, keyness analysis (based on chi-square statistics) is applied to reconstruct the lexicon specific to each cluster, while co-occurrence network analysis is proposed as a subsequent analytical step for exploring the internal structure of selected semantic fields. The integration of semantic clustering and lexical analysis provides a structured representation of the corpus and lays the groundwork for future analytical outputs, such as interactive dashboards and knowledge graphs, combining distributional semantics and classical text statistics to deliver scalable, interpretable, and methodologically transparent analyses of large institutional text corpora.
A Multilevel Analysis of Italy’s Three-Year School Planning Documents Using a Lexical and Distributional-Semantic Framework
Pasquale Pavone
;
2026-01-01
Abstract
The study proposes a multilevel methodological approach for analyzing Piani Triennali dell'Offerta Formativa (PTOF), the official documents through which Italian schools describe their educational projects. The framework combines two levels of analysis: Measurement level: texts from PTOF documents are represented through sentence embeddings generated by a BERT model. These representations are then processed using UMAP (dimensionality reduction) and HDBSCAN (clustering) to identify thematically coherent groups in a fully replicable manner. Interpretative level: text mining and lexicometric techniques are used to characterize the emerging clusters. Specifically, keyness analysis (based on chi-square statistics) is applied to reconstruct the lexicon specific to each cluster, while co-occurrence network analysis is proposed as a subsequent analytical step for exploring the internal structure of selected semantic fields. The integration of semantic clustering and lexical analysis provides a structured representation of the corpus and lays the groundwork for future analytical outputs, such as interactive dashboards and knowledge graphs, combining distributional semantics and classical text statistics to deliver scalable, interpretable, and methodologically transparent analyses of large institutional text corpora.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
