Classroom observation is widely used to measure teaching practices but is costly at scale, limiting its use as a formative tool for teacher improvement in low- and middle-income countries (LMICs). We test whether a large language model (LLM) can agree with a certified human observer as closely as the World Bank's Teach reliability protocol requires of trained raters. Across 199 Peruvian primary-classroom transcripts, we evaluate eight LLM configurations (Gemini 2.0 and 2.5 Flash) varying rubric quality, prompt format, and reasoning depth, and apply the Teach certification rule to the LLM-observer agreement via Monte Carlo simulation. Our best configuration meets the within-one-point certification criterion against the certified observer in 92.9% of simulated exams (95% CI [91.3, 94.5]), with element-level agreement of 81-99% across Teach elements—in the range trained human raters reach, though our gold standard is a single certified field observer rather than the program's master coders. A regression that isolates each design lever shows that rubric content and distributional priors are the two dominant drivers of agreement, while chain-of-thought reasoning and snapshot tagging are not; a minimal-rubric baseline almost never clears the criterion. A reliable LLM scorer can help ministries deliver formative feedback within teacher-support programs more frequently and at lower marginal cost.
| Repository name | URI |
|---|---|
| Reproducible Research Repository (World Bank) | https://reproducibility.worldbank.org |
Paper exhibits were reproduced on a computer with the following specifications:
• OS: Windows 11 Enterprise
• Processor: Intel(R) Xeon(R) Gold 5218 CPU @ 2.30GHz (2.30 GHz) (2 processors)
• Memory available: 16.0 GB
• Software version: Python 3.11.1
Run time: ~5 minutes
To reproduce the findings in this paper, a replicator must:
requirements.txt, set the environment variable AITEACH_USE_CACHED_LLM=1 to run the cached version of the pipeline using the shipped de-identified intermediate data, and run 00_master.py.Since not all the data is included, the package includes the results produced by replicators. These files can be used to review the results presented in the paper.
Some data is restricted and has not been included in the reproducibility package. For more details, please refer to the README file. Intermediate data is forthcoming in the World Bank Microdata Library under a restricted license and can be used to fully reproduce the results in the paper.
| Author | Affiliation | |
|---|---|---|
| Carolina Lopez | World Bank | carolina_lopez@worldbank.org |
| Ezequiel Molina | World Bank | molina@worldbank.org |
| Matthew D. Krasnow | Harvard University | matt@swiftscore.org |
| Jostin Kitmang | Harvard University | jkitmang@g.harvard.edu |
| Carolina Moreira Vasquez | World Bank | cmm9mn@virginia.edu |
| William J. Krasnow | Swiftscore | will@swiftscore.org |
2026-07-29
| Location | Code |
|---|---|
| Peru | PER |
The materials in the reproducibility packages are distributed as they were prepared by the staff of the International Bank for Reconstruction and Development/The World Bank. The findings, interpretations, and conclusions expressed in this event do not necessarily reflect the views of the World Bank, the Executive Directors of the World Bank, or the governments they represent. The World Bank does not guarantee the accuracy of the materials included in the reproducibility package.
| Name | URI |
|---|---|
| MIT License | https://opensource.org/license/mit |
| World Bank IGO Rider | https://github.com/worldbank/metadata-editor/blob/main/WB-IGO-RIDER.md |
| Name | Affiliation | |
|---|---|---|
| Carolina Lopez | World Bank | carolina_lopez@worldbank.org |
| Reproducibility WBG | World Bank | reproducibility@worldbank.org |
| Name | Abbreviation | Affiliation | Role |
|---|---|---|---|
| Reproducibility WBG | DECDI | World Bank - Development Impact Department | Verification and preparation of metadata |
2026-07-29
1