A survey of 212 studies (2020–August 2026) mapping AI and machine learning across seven ILSA programs, with a five-category synthesis and a policy-actionability framework (mean 3.70 / 5).
"A structured survey of AI applications in ILSAs across 212 studies, pairing a five-category synthesis with a policy-actionability rubric (mean 3.70 / 5; 85.8% high actionability)."
International Large-Scale Assessments (ILSAs) are a primary evidence base for comparing education systems. AI and data-mining applications to ILSA data have grown rapidly, yet the field remains prediction-oriented, relies on aggregate performance indicators, and only partially converts computational advances into educationally interpretable insight.
This study synthesizes AI and data-driven research on ILSA datasets through five categories: (i) predictive modelling of achievement, (ii) process data and learning analytics, (iii) socio-emotional and behavioral modelling, (iv) assessment engineering, and (v) computational psychometrics. A schema-constrained LLM-assisted extraction pipeline with full human verification produced the open machine-readable dataset released with the survey.
Practical usefulness is evaluated with a three-dimensional policy-actionability framework (inferential warrant, effect specification, population boundedness). Across 212 studies (2020–August 2026), mean actionability was 3.70 (SD = 0.80), with 85.8% reaching Score 4–5. Study type, not algorithmic family, was the primary differentiator: predictive modelling families scored uniformly high, while review/methodology studies scored lower (mean 3.23).
| Sheet / Table | Records | Description | Key fields |
|---|---|---|---|
| Articles | 212 |
One record per study containing publication metadata, AI/ML methods, study design, survey methodology, and quality indicators. | 31 columns · One record per study |
| Findings | 382 |
One record per reported finding, including the ILSA program, assessment cycle, outcome, key predictors, model performance, and standardized evidence labels. | 12 columns · One record per finding |
| Predictors | 3,088 |
One record per predictor–study pair with standardized variable names, educational level, and controlled taxonomy labels. | 7 columns · One record per predictor |
Shares among predictive modelling families in the survey taxonomy (Ensemble n = 77, Classical n = 22, Deep Learning n = 16, Hybrid/Penalised n = 14; total 129). Review/Methodology studies (n = 83; mean actionability 3.23) are excluded from this breakdown and scored separately.
Domain shares are finding-level counts from the survey synthesis (n = 345 coded outcomes). The open dataset contains 382 extracted finding records before methodological-output exclusions.
Three-dimensional rubric: inferential warrant × effect specification × population boundedness. Predictive families (Ensemble, Classical, Deep Learning, Hybrid) all mean 4.00; Review/Methodology mean 3.23.
Beyond mapping methods, the survey translates ILSA–AI evidence into guidance that teachers, counselors, and education systems can use—while keeping clear the line between prediction, explanation, and intervention.
Across the corpus, self-efficacy, metacognitive strategies, and school belonging repeatedly emerge as stronger, more transferable predictors of achievement than demographic or resource factors alone—pointing educators toward supports they can actually change.
Process data reveal distinct behavioral routes among students with similar outcomes. Engagement and strategy traces can flag potential disengagement earlier than achievement alone, supporting targeted guidance rather than ranking.
XAI clarifies which features drive a model’s predictions, but those attributions are not causal mechanisms. Instructional choices should combine model output with classroom context and domain knowledge—not treat SHAP rankings as interventions.
Predictor importance shifts across education systems. Findings should be replicated and piloted locally before cross-national scale-up—the actionability rubric makes that boundedness explicit for decision-makers.
“The central challenge is no longer computational capacity but connecting prediction to explanation, explanation to intervention, and intervention to context-sensitive educational action.”
Extraction used one model version (gpt-5.4-nano, snapshot 2026-03-17) with a fixed schema. Reproducibility of derived fields applies to Layer-2 rule-based scripts after human verification.
The D₁–D₃ rubric has not been externally validated. It is a structured analytical lens rather than a consensus psychometric instrument; a follow-up expert re-rating is planned.
Independent human scoring was checked on n = 37 (exact agreement 86.5%; κ = 0.76, κw = 0.83). Full-corpus verification relied on author consensus rather than dual coding of every study.
Review/methodology papers receive lower D₁ by design (family mean 3.23). Most empirical studies cluster at Score 4 (180/212), limiting discrimination among predictive papers.
English-only search in Web of Science (991) and Google Scholar (225); Scopus was excluded. Scopus-only and some IEEE proceedings may be underrepresented. Cutoff: August 2026.
Move beyond single-country models with measurement invariance and contextual heterogeneity so generalizable effects can be separated from system-specific ones.
Pair predictive models with explainability methods and substantive educational theory so feature attributions are not read as causal mechanisms.
Go beyond response time and action counts to sequential and multilevel models of decision transitions in CBA log data.
Expand BART/BCF and related designs where ILSA structure supports stronger inferential warrants for policy claims.
Only 32.5% of extracted findings address mathematics, reading, and science jointly as core cognitive domains; denser coverage would strengthen synthesis.