Abstract Methodology Dataset Findings Contributions Education Limitations Cite
Survey Paper  ·  Open Dataset

Artificial Intelligence Applications in International Large-Scale Assessments:
A Survey with LLM-Assisted Evidence Synthesis

A survey of 212 studies (2020–August 2026) mapping AI and machine learning across seven ILSA programs, with a five-category synthesis and a policy-actionability framework (mean 3.70 / 5).

PISATIMSSPIRLS TALISICCSICILS PIAAC
212
Studies
382
Findings
3,088
Predictors

Abstract

What this survey does

"A structured survey of AI applications in ILSAs across 212 studies, pairing a five-category synthesis with a policy-actionability rubric (mean 3.70 / 5; 85.8% high actionability)."

International Large-Scale Assessments (ILSAs) are a primary evidence base for comparing education systems. AI and data-mining applications to ILSA data have grown rapidly, yet the field remains prediction-oriented, relies on aggregate performance indicators, and only partially converts computational advances into educationally interpretable insight.

This study synthesizes AI and data-driven research on ILSA datasets through five categories: (i) predictive modelling of achievement, (ii) process data and learning analytics, (iii) socio-emotional and behavioral modelling, (iv) assessment engineering, and (v) computational psychometrics. A schema-constrained LLM-assisted extraction pipeline with full human verification produced the open machine-readable dataset released with the survey.

Practical usefulness is evaluated with a three-dimensional policy-actionability framework (inferential warrant, effect specification, population boundedness). Across 212 studies (2020–August 2026), mean actionability was 3.70 (SD = 0.80), with 85.8% reaching Score 4–5. Study type, not algorithmic family, was the primary differentiator: predictive modelling families scored uniformly high, while review/methodology studies scored lower (mean 3.23).

Motivation

Computational capability in ILSA research has expanded faster than theoretically grounded, policy-defensible insight. The survey maps that gap across seven programs.

gap analysis
Approach

PRISMA-guided search, schema-constrained LLM extraction with full human verification, five-category synthesis, and a three-dimensional policy-actionability rubric.

survey paper
Output

Open evidence resource: 212 studies, 382 finding records, 3,088 predictors, actionability scores, and a CC BY 4.0 dataset for reuse and secondary analysis.

open dataset

Methodology

Study selection & screening process

RQ 1
How is the methodological landscape of AI-driven ILSA research distributed across predictive modelling, process-data analytics, socio-emotional modelling, assessment engineering, and computational psychometrics (2020–August 2026)?
RQ 2
How actionable are the field’s findings for educational policy when scored on inferential warrant, effect specification, and population boundedness, and does actionability vary by study type or algorithmic family?
RQ 3
Under what conditions do computational advances in ILSA research translate into educationally defensible evidence, and which methodological gaps still limit classroom- and system-level use?
1,216
Step 01 · Identification
Literature search
991 records were identified in Web of Science (Topic field) using search strings combining ILSA program names with AI, machine learning, explainability, educational data mining, and computational psychometric terminology. An additional 225 records were identified in Google Scholar using the allintitle operator. Both searches were restricted to English-language publications from 2020 to August 2026.
293
Step 02 · Screening
Title and abstract screening
Title-and-abstract screening excluded 815 Web of Science and 119 Google Scholar records that did not use large-scale assessment data, leaving 183 and 110 records, respectively.
212
Step 03 · Eligibility
Full-text assessment
After merging sources and removing 70 duplicates, 212 unique full-text articles were assessed for eligibility. Both authors read every article in full against the predefined inclusion criteria.
212
Step 04 · Final Corpus
Evidence synthesis
All 212 studies were retained in the final corpus. Each record was processed through a schema-constrained LLM-assisted extraction pipeline and fully human-verified, then coded into a five-category synthesis framework and a three-dimensional policy-actionability rubric.
Inclusion Criteria
  • English-language publications, 2020–August 2026
  • Uses ILSA / large-scale assessment data (PISA, TIMSS, PIRLS, TALIS, PIAAC, ICCS, ICILS)
  • AI, ML, EDM, LA, XAI, or computational-psychometric methods
  • Peer-reviewed research, reviews, and methodology papers
  • Full text available for human-verified extraction
Exclusion Criteria
  • No large-scale assessment data used
  • Non-ILSA national assessments only (unless ILSA-linked)
  • Duplicate records across WoS and Google Scholar
  • Non-English publications
  • Records outside the 2020–August 2026 window
Five-Category Synthesis Framework
  • Predictive modelling of achievement
  • Process data and learning analytics
  • Socio-emotional and behavioral modelling
  • Assessment engineering (scoring / items)
  • Computational psychometrics & methods
Policy Actionability Rubric
  • D₁ Inferential warrant (descriptive → correlational → causal)
  • D₂ Effect specification (none → directional → quantified)
  • D₃ Population boundedness (global/unspecified → bounded)
  • Composite ordinal score 1–5
  • Human inter-rater check: n = 37, κ = 0.76, κw = 0.83
  • Mean 3.70 (SD 0.80); 85.8% Score 4–5
Dataset

Open structured research dataset

Sheet / Table Records Description Key fields
Articles
212
One record per study containing publication metadata, AI/ML methods, study design, survey methodology, and quality indicators.
DOIML techniquesML familyPV handlingSampling weightsStudy type
31 columns · One record per study
Findings
382
One record per reported finding, including the ILSA program, assessment cycle, outcome, key predictors, model performance, and standardized evidence labels.
DOITarget variableTop predictorsPerformance metricsOutcome domain
12 columns · One record per finding
Predictors
3,088
One record per predictor–study pair with standardized variable names, educational level, and controlled taxonomy labels.
DOIVariable nameCategoryPredictor levelPredictor category
7 columns · One record per predictor
// sample extraction record
"metadata": { "title": "ML to predict science achievement TIMSS 2019", "year": 2024, "open_access": true }, "data": { "ml_techniques": { "primary": "Random Forest", "all": ["Random Forest", "XGBoost"] }, "plausible_values_handling": "not_reported", "survey_design": { "student_weights_used": false }, "main_findings": [{ "target_variable": "Science (TIMSS 2019)", "performance_metrics": "R² = 0.71" }] }
// controlled vocabulary taxonomy
source_category
Peer-Reviewed Review Article Methodology Paper
ml_family
Ensemble Deep Learning Classical Hybrid Review/Methods
target_domain
Mathematics Science Reading Non-cognitive
predictor_level
Student School System
pv_filter_label
Not Applicable Rubin Rules Single PV Average PVs WLE/IRT All PVs Not Reported
weights_filter
True False Unknown
Key Findings

ML landscape & methodological rigor

Ensemble Learning
Random Forest · XGBoost · Gradient Boosting · LightGBM · SHAP
60%
Classical / Statistical
Regression · HLM · Logistic · Discriminant Analysis
17%
Deep Learning
Neural Networks · CNN · LSTM · Autoencoders
12%
Hybrid / Penalised
LASSO · Ridge · Elastic Net · Stacking · Causal ML
11%

Shares among predictive modelling families in the survey taxonomy (Ensemble n = 77, Classical n = 22, Deep Learning n = 16, Hybrid/Penalised n = 14; total 129). Review/Methodology studies (n = 83; mean actionability 3.23) are excluded from this breakdown and scored separately.

Outcome domains targeted
25%
Other / Unspecified
24%
Composite / Multi-Domain
18%
Mathematics
10%
Non-Cognitive
8%
Reading
7%
Problem Solving
7%
Science
1%
Civic Education

Domain shares are finding-level counts from the survey synthesis (n = 345 coded outcomes). The open dataset contains 382 extracted finding records before methodological-output exclusions.

Policy Actionability Framework
Mean actionability score (1–5 scale)
3.70
SD across 212 studies
0.80
High actionability — Score 4–5 (182 studies)
85.8%

Three-dimensional rubric: inferential warrant × effect specification × population boundedness. Predictive families (Ensemble, Classical, Deep Learning, Hybrid) all mean 4.00; Review/Methodology mean 3.23.

Methodological rigor indicators
Performance metrics reported
87%
Sampling weights applied
20%
Plausible values correctly handled
24%
Sample size explicitly reported
75%
Countries / economies specified
89%
Cross-validation or test-set reported
45%
Contributions

What This Survey Contributes

01
First Comprehensive Survey of AI in ILSAs
The first survey explicitly mapping AI methods across seven ILSA programs, synthesizing 212 studies (2020–August 2026) through a five-category framework spanning prediction, process data, socio-emotional modelling, assessment engineering, and computational psychometrics.
7 ILSA programs212 studiesFive-category taxonomy
02
Open Structured Evidence Repository
A schema-constrained LLM-assisted extraction pipeline with full human verification produced a reusable open dataset of study metadata, findings, predictors, and actionability scores under CC BY 4.0.
3 relational tables3,682 structured recordsCC BY 4.0
03
Policy Actionability Framework
A domain-adapted three-dimensional rubric (inferential warrant × effect specification × population boundedness) scores what ILSA–AI findings can support for policy. Mean score 3.70 (SD 0.80); 85.8% reach Score 4–5, with study type as the main differentiator.
D₁–D₃ rubricκ = 0.76Score 4–5: 85.8%
04
Research Gaps and Future Directions
Identifies under-mined process sequences, limited causal designs, fragmented outcome domains (only 32.5% of findings in core cognitive domains), and the need for multi-country invariance-aware analyses.
Evidence gapsCausal MLProcess data
Educational Impact

What this work offers education

Beyond mapping methods, the survey translates ILSA–AI evidence into guidance that teachers, counselors, and education systems can use—while keeping clear the line between prediction, explanation, and intervention.

01 Classroom & school practice
Prioritize modifiable psychosocial levers

Across the corpus, self-efficacy, metacognitive strategies, and school belonging repeatedly emerge as stronger, more transferable predictors of achievement than demographic or resource factors alone—pointing educators toward supports they can actually change.

self-efficacy metacognition belonging
02 Counseling & learning support
See learning pathways, not only scores

Process data reveal distinct behavioral routes among students with similar outcomes. Engagement and strategy traces can flag potential disengagement earlier than achievement alone, supporting targeted guidance rather than ranking.

process data engagement strategies
03 Instructional decisions
Use explainable AI with caution

XAI clarifies which features drive a model’s predictions, but those attributions are not causal mechanisms. Instructional choices should combine model output with classroom context and domain knowledge—not treat SHAP rankings as interventions.

XAI not causal context-aware
04 Policy & system transfer
Localize before scaling

Predictor importance shifts across education systems. Findings should be replicated and piloted locally before cross-national scale-up—the actionability rubric makes that boundedness explicit for decision-makers.

cross-cultural pilot first actionability

“The central challenge is no longer computational capacity but connecting prediction to explanation, explanation to intervention, and intervention to context-sensitive educational action.”

Limitations

Known constraints

Single extraction model

Extraction used one model version (gpt-5.4-nano, snapshot 2026-03-17) with a fixed schema. Reproducibility of derived fields applies to Layer-2 rule-based scripts after human verification.

Internally developed actionability rubric

The D₁–D₃ rubric has not been externally validated. It is a structured analytical lens rather than a consensus psychometric instrument; a follow-up expert re-rating is planned.

Partial reliability subsample

Independent human scoring was checked on n = 37 (exact agreement 86.5%; κ = 0.76, κw = 0.83). Full-corpus verification relied on author consensus rather than dual coding of every study.

Rubric structure and score concentration

Review/methodology papers receive lower D₁ by design (family mean 3.23). Most empirical studies cluster at Score 4 (180/212), limiting discrimination among predictive papers.

Search coverage

English-only search in Web of Science (991) and Google Scholar (225); Scopus was excluded. Scopus-only and some IEEE proceedings may be underrepresented. Cutoff: August 2026.

Future Directions

Multi-country comparative designs

Move beyond single-country models with measurement invariance and contextual heterogeneity so generalizable effects can be separated from system-specific ones.

Routine XAI with domain grounding

Pair predictive models with explainability methods and substantive educational theory so feature attributions are not read as causal mechanisms.

Micro-level process sequences

Go beyond response time and action counts to sequential and multilevel models of decision transitions in CBA log data.

Causal and counterfactual ML

Expand BART/BCF and related designs where ILSA structure supports stronger inferential warrants for policy claims.

Core cognitive outcome coverage

Only 32.5% of extracted findings address mathematics, reading, and science jointly as core cognitive domains; denser coverage would strengthen synthesis.

Citation

Cite this work

BibTeX
@article{dede_cetinkaya2026ilsa_survey, title = {Artificial Intelligence Applications in International Large-Scale Assessments: A Survey with LLM-Assisted Evidence Synthesis}, author = {Dede, Merve and Çetinkaya, Ekrem}, year = {2026}, note = {Open dataset: HuggingFace Datasets}, url = {https://huggingface.co/datasets/dedemerve/ILSA-Survey-Dataset} }
Access
Open Dataset on HuggingFace View on GitHub