This study investigates how actions for respect education and gender equali-ty are articulated in the Three-Year Educational Plans (PTOFs) of Italian schools, using a national corpus with explicit document identifiers, rich cat-egorical metadata, and internal section markers. The objective is twofold: (i) to identify where gender-related discourse appears within PTOF macro-sections, and (ii) to reconstruct the main thematic strands through a scala-ble, unsupervised workflow. We adopt a corpus-driven approach starting from a curated set of high-precision multiword expressions (MWEs) strictly related to gender issues. MWEs are used as anchors to quantify the presence and distribution of the theme across sections and to extract a targeted sub-corpus at chunk level, balancing interpretability and contextual coherence. Each chunk is represented through BERT-based contextual embeddings, subsequently projected with UMAP to support structure discovery and visu-alization. Topic-like groupings are derived via agglomerative hierarchical clustering (Ward), enabling multi-resolution exploration of the discourse. Cluster quality and partition choices are assessed through internal validation indices, including Calinski–Harabasz and silhouette measures. While the workflow generates multiple complementary outputs, we report here the core result: a twelve-cluster semantic typology obtained via charac-teristic-language profiling, which captures recurring institutional framings through which schools translate the policy mandate into planning language.
Profiling Gender Equality Discourse in Italian Three-Year School Plans (PTOFs) with Contextual Embeddings
Pasquale Pavone
;
2026-01-01
Abstract
This study investigates how actions for respect education and gender equali-ty are articulated in the Three-Year Educational Plans (PTOFs) of Italian schools, using a national corpus with explicit document identifiers, rich cat-egorical metadata, and internal section markers. The objective is twofold: (i) to identify where gender-related discourse appears within PTOF macro-sections, and (ii) to reconstruct the main thematic strands through a scala-ble, unsupervised workflow. We adopt a corpus-driven approach starting from a curated set of high-precision multiword expressions (MWEs) strictly related to gender issues. MWEs are used as anchors to quantify the presence and distribution of the theme across sections and to extract a targeted sub-corpus at chunk level, balancing interpretability and contextual coherence. Each chunk is represented through BERT-based contextual embeddings, subsequently projected with UMAP to support structure discovery and visu-alization. Topic-like groupings are derived via agglomerative hierarchical clustering (Ward), enabling multi-resolution exploration of the discourse. Cluster quality and partition choices are assessed through internal validation indices, including Calinski–Harabasz and silhouette measures. While the workflow generates multiple complementary outputs, we report here the core result: a twelve-cluster semantic typology obtained via charac-teristic-language profiling, which captures recurring institutional framings through which schools translate the policy mandate into planning language.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
