ORCID

https://orcid.org/0000-0001-8879-8376

Date of Award

Summer 2026

Language

English

Embargo Period

7-29-2026

Document Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

College/School/Department

Department of Educational and Counseling Psychology

Program

Educational Psychology and Methodology

First Advisor

Mariola Moeyaert

Committee Members

Benjamin G. Solomon, Xin Li

Keywords

Large language models, meta-analysis, automated data extraction, single-case designs, prompt engineering, evidence synthesis

Subject Categories

Educational Assessment, Evaluation, and Research | Educational Methods | Educational Psychology

Abstract

Single-case experimental designs (SCEDs) are a form of longitudinal research that is particularly valuable for studying marginalized and underrepresented populations. Meta-analysis plays an essential role in synthesizing evidence across studies and enhancing the generalizability of findings. However, as interest in SCEDs grows and the evidence base expands, conducting meta-analyses has become increasingly labor-intensive and time-consuming. This dissertation addresses these challenges by examining the applicability of large language models (LLMs) for data extraction and meta-analytic dataset compilation in the context of SCED meta-analysis. Because LLM performance is influenced by model selection, prompt design, and the specific materials being extracted, this dissertation has three overarching objectives: first, to develop prompts that enable LLMs to extract SCED data effectively and efficiently and produce meta-analytic datasets that are readily usable for analysis; second, to identify which LLMs perform best for specific types of data extraction tasks; and third, to determine the conditions under which LLM-based data extraction is appropriate and yields strong performance. These objectives were addressed through two separate studies. The first study focuses on textual data extraction and examines whether LLM performance varies across target variables, including author name, number of participants, physical setting, and outcome domain. The second study focuses on graphical data extraction and investigates whether LLM performance differs as a function of graph characteristics, including the number of data points, number of outcomes, trend, and variability. Three LLM families were evaluated using the most advanced model versions available at the time of each study. Across both studies, Gemini 3.0 Pro demonstrated strong potential for extracting data from both textual and graphical sources. The findings of this dissertation provide empirical evidence regarding the usability of LLMs in the data extraction and dataset compilation process for SCED meta-analysis. By identifying the strengths and limitations of LLM-based data extraction, this dissertation contributes methodological guidance for using LLMs to improve the efficiency and scalability of SCED meta-analytic workflows. Practical recommendations for prompt development, model selection, and human verification are also provided for researchers conducting SCED evidence synthesis and meta-analysis.

License

This work is licensed under the University at Albany Standard Author Agreement.

Share

COinS