Advancing Data Extraction Workflows for Single-Case Design Meta-Analysis Using Large Language Models
ORCID
https://orcid.org/0000-0001-8879-8376
Date of Award
Summer 2026
Language
English
Embargo Period
7-29-2026
Document Type
Dissertation
Degree Name
Doctor of Philosophy (PhD)
College/School/Department
Department of Educational and Counseling Psychology
Program
Educational Psychology and Methodology
First Advisor
Mariola Moeyaert
Committee Members
Benjamin G. Solomon, Xin Li
Keywords
Large language models, meta-analysis, automated data extraction, single-case designs, prompt engineering, evidence synthesis
Subject Categories
Educational Assessment, Evaluation, and Research | Educational Methods | Educational Psychology
Abstract
Single-case experimental designs (SCEDs) are a form of longitudinal research that is particularly valuable for studying marginalized and underrepresented populations. Meta-analysis plays an essential role in synthesizing evidence across studies and enhancing the generalizability of findings. However, as interest in SCEDs grows and the evidence base expands, conducting meta-analyses has become increasingly labor-intensive and time-consuming. This dissertation addresses these challenges by examining the applicability of large language models (LLMs) for data extraction and meta-analytic dataset compilation in the context of SCED meta-analysis. Because LLM performance is influenced by model selection, prompt design, and the specific materials being extracted, this dissertation has three overarching objectives: first, to develop prompts that enable LLMs to extract SCED data effectively and efficiently and produce meta-analytic datasets that are readily usable for analysis; second, to identify which LLMs perform best for specific types of data extraction tasks; and third, to determine the conditions under which LLM-based data extraction is appropriate and yields strong performance. These objectives were addressed through two separate studies. The first study focuses on textual data extraction and examines whether LLM performance varies across target variables, including author name, number of participants, physical setting, and outcome domain. The second study focuses on graphical data extraction and investigates whether LLM performance differs as a function of graph characteristics, including the number of data points, number of outcomes, trend, and variability. Three LLM families were evaluated using the most advanced model versions available at the time of each study. Across both studies, Gemini 3.0 Pro demonstrated strong potential for extracting data from both textual and graphical sources. The findings of this dissertation provide empirical evidence regarding the usability of LLMs in the data extraction and dataset compilation process for SCED meta-analysis. By identifying the strengths and limitations of LLM-based data extraction, this dissertation contributes methodological guidance for using LLMs to improve the efficiency and scalability of SCED meta-analytic workflows. Practical recommendations for prompt development, model selection, and human verification are also provided for researchers conducting SCED evidence synthesis and meta-analysis.
License
This work is licensed under the University at Albany Standard Author Agreement.
Recommended Citation
Lou, Yaosheng, "Advancing Data Extraction Workflows for Single-Case Design Meta-Analysis Using Large Language Models" (2026). Electronic Theses & Dissertations (2024 - present). 519.
https://scholarsarchive.library.albany.edu/etd/519
Included in
Educational Assessment, Evaluation, and Research Commons, Educational Methods Commons, Educational Psychology Commons