Honors & Awards


  • Stanford HAI Graduate Fellow, Stanford University HAI (2025)
  • Graduate Student Best Poster Award, National Council on Measurement in Education (2025)
  • Stanford Interdisciplinary Graduate Fellowship, Stanford University (2023)
  • Distinguished Poster Award, Psychometrics Society (2023)
  • Best Paper Nomination, The 13th Annual International Conference for Computer-Supported Collaborative Learning (2019)
  • Letha Hurd Morgan Award, New York University (2018)

Professional Education


  • Bachelor of Science, New York University, Computer Science; Education (2018)
  • Bachelor of Science, Boston University, Computer Science and Education (2018)
  • Master of Science, University of Pennsylvania, Learning Sciences and Techolog (2019)
  • M.S.Ed, University of Pennsylvania, Learning Sciences and Technologies (2019)
  • B.S., New York University, Computer Science, Teaching Chemistry 7–12 (with Honors) (2018)

Work Experience


  • Chemistry Subject Expert Teacher, BASIS Independent Brooklyn (2019 - 2021)

    Location

    Brooklyn, NY

All Publications


  • A comparison of the predictive performance of continuous and class-based latent trait models. Psychological methods Ma, W. A., Liu, Y., Kanopka, K., Ma, W., Domingue, B. W. 2026

    Abstract

    The ability of a student can be conceptualized as either a continuously varying entity (e.g., conventional analysis using dichotomous item response theory [IRT] models; Lord & Novick, 1968) or a bundle of latent classes (e.g., cognitive diagnostic models [CDMs]; Rupp et al., 2010; von Davier & Lee, 2019). This article builds on recent efforts to focus on predictive differences between measurement models in an attempt to examine the degree to which such approaches, which utilize quite distinctive notions regarding the nature of ability, produce different predictions of response behavior in the real world. We first present two simulation studies in which data are generated from variants of CDMs that differ in sample size, attribute hierarchical structures, and attribute estimation methods. We illustrate that, as we would expect given that they were used to generate data, CDM-based predictions uniformly outperform those of IRT models. We then compare the performance of CDM- and IRT-based approaches across 11 empirical data sets previously analyzed using CDMs. Our findings indicate that overfitting is a pervasive issue across CDM-based predictions, particularly with the generalized deterministic inputs, noisy "and" gate model. Furthermore, in no case does the CDM show superior performance when using the maximum a posteriori estimator, and only six out of 11 data sets show improved model fit for CDMs over the two-parameter logistic model when using the marginal mastery probabilities estimator. Researchers and practitioners may need to balance the diagnostic appeal of CDMs with the fact that their complexity can come at the cost of predictive accuracy. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

    View details for DOI 10.1037/met0000860

    View details for PubMedID 42475042

  • A fair lexical decision task for monolingual and multilingual Spanish-speakers. Frontiers in psychology Siebert, J. M., Jimenez, M., Ma, W. A., Townley-Flores, C., Saavedra, A., Yeatman, J. D. 2026; 17: 1826045

    Abstract

    This study describes the development and validation of ROAR Palabra, a novel Spanish lexical decision task designed for use with both Spanish-speaking children and Spanish-English bilinguals. This self-administered task requires students to decide whether a string of letters presented on the screen is a real word in Spanish. While there is evidence that scores on English lexical decision tasks are highly predictive of performance on conventional (time- and resource-intensive) word reading assessments in English, we explore whether this holds in Spanish, which has a much more transparent orthography. The specific goals are (i) to create a linguistically fair task using item-response theory and (ii) to evaluate whether such task can serve as a reliable proxy for conventional word reading measures, offering a quick and easy-to-administer tool for assessing reading skills across linguistic and cultural contexts. Results demonstrated strong correlations between performance on ROAR Palabra and standardized word reading assessments such as the Woodcock-Muñoz Batería IV, suggesting its effectiveness as a substitute measure. Notably, the task was sensitive to differences in language proficiency across both monolingual and multilingual groups, reflecting expected developmental and environmental influences. While not specifically designed for the comparisons between monolingual and multilingual populations, the findings underscore the potential of this task as a versatile and culturally adaptable tool for reading assessments in different Spanish-speaking and bilingual contexts.

    View details for DOI 10.3389/fpsyg.2026.1826045

    View details for PubMedID 42416169

    View details for PubMedCentralID PMC13337784

  • Improving validity and efficiency of digital dyslexia screening through trial-by-trial feedback. Psychological assessment Ma, W. A., Jimenez, M., Siebert, J. M., Saavedra, A., Townley-Flores, C., Richie-Halford, A., Domingue, B. W., Yeatman, J. D. 2026

    Abstract

    Trial-by-trial feedback is a powerful tool across scientific disciplines, offering real-time information that shapes learning, attention, and behavioral adjustment. However, its role in large-scale educational and psychological assessments remains underexplored. This study examines whether providing trial-by-trial performance feedback during a digital dyslexia screener enhances engagement without compromising score validity or interpretability. Engagement is a critical threat to validity in assessments, particularly among young children in digital universal screening contexts. We conducted two large-scale experiments involving 6,332 students from Grades 1 to 12 in Colombia and the United States in which students were randomly assigned to one of two conditions: (a) informative feedback (auditory cues indicating correct/incorrect) or (b) non-informative feedback (neutral sound). The assessment was administered in two formats: an adaptively ordered English version and a randomly ordered Spanish version. We evaluated differences in response behaviors, item response theory model fit, and concurrent validity between conditions. Results indicate that informative feedback significantly increased compliance and reduced disengagement. Assessment scores and item response theory model fit were comparable across conditions. Notably, informative feedback improved the concurrent validity of scores and was associated with shorter completion times. These findings suggest that trial-by-trial feedback can enhance engagement while preserving score interpretability, leading to more efficient and valid digital universal screening assessments. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

    View details for DOI 10.1037/pas0001474

    View details for PubMedID 42189560

  • Similarities (and Differences) in the Learning Patterns of Single-Word Reading of an Alphabetic Orthography in Monolingual and Bilingual Primary School Children: A Cross-Sectional Study. Brain sciences Smith, G., Bassoli, E., Ozturk, Y., Arteaga-Garcia, E., Ma, W. A., , , Yeatman, J. D., Mastrogiuseppe, M., Caffarra, S. 2026; 16 (4)

    Abstract

    Background/Objectives: With growing waves of migration, children speaking a home language different from the language of school literacy have become increasingly common in Western education systems. In this context, understanding and monitoring bilinguals' reading development is crucial to inform both educational and clinical practices and ensure equitable services. The present study contributes to the literature by investigating learning patterns in single-word reading across primary school grades. Monolingual and bilingual children learning to read in an alphabetic orthography were examined. Methods: The sample consisted of 565 typically developing monolingual and bilingual primary school children from grades 1-5 (bilinguals = 162). Participants completed a computerised Lexical Decision task (LDT) recording accuracy and response times, and standardised tests of reading and cognition. A parental questionnaire was used to gather socio-demographic and linguistic information. Results: Response bias-corrected accuracy rates in the LDT revealed an increase in sensitivity across school years after correcting for potential confounds (SES, vocabulary, nonverbal intelligence). No significant effect of bilingualism was observed. Response times for correct responses also decreased consistently across grades after controlling for the same confounds. Although no significant main effect of bilingualism emerged, an interaction with grade revealed a greater decrease in response times for second-grade bilinguals compared to monolingual peers. Conclusions: Monolingual and bilingual children showed comparable sensitivity rates and reading times, suggesting similar decoding skill acquisition. However, an earlier decrease in response times for bilinguals points to a facilitatory effect in the early stages of reading development, consistent with a bilingual advantage during skill learning.

    View details for DOI 10.3390/brainsci16040356

    View details for PubMedID 42041767

    View details for PubMedCentralID PMC13115128

  • ROAR-CAT: Rapid Online Assessment of Reading ability with Computerized Adaptive Testing. Behavior research methods Ma, W. A., Richie-Halford, A., Burkhardt, A. K., Kanopka, K., Chou, C., Domingue, B. W., Yeatman, J. D. 2025; 57 (1): 56

    Abstract

    The Rapid Online Assessment of Reading (ROAR) is a web-based lexical decision task that measures single-word reading abilities in children and adults without a proctor. Here we study whether item response theory (IRT) and computerized adaptive testing (CAT) can be used to create a more efficient online measure of word recognition. To construct an item bank, we first analyzed data taken from four groups of students (N = 1960) who differed in age, socioeconomic status, and language-based learning disabilities. The majority of item parameters were highly consistent across groups (r = .78-.94), and six items that functioned differently across groups were removed. Next, we implemented a JavaScript CAT algorithm and conducted a validation experiment with 485 students in grades 1-8 who were randomly assigned to complete trials of all items in the item bank in either (a) a random order or (b) a CAT order. We found that, to achieve reliability of 0.9, CAT improved test efficiency by 40%: 75 CAT items produced the same standard error of measurement as 125 items in a random order. Subsequent validation in 32 public school classrooms showed that an approximately 3-min ROAR-CAT can achieve high correlations (r = .89 for first grade, r = .73 for second grade) with alternative 5-15-min individually proctored oral reading assessments. Our findings suggest that ROAR-CAT is a promising tool for efficiently and accurately measuring single-word reading ability. Furthermore, our development process serves as a model for creating adaptive online assessments that bridge research and practice.

    View details for DOI 10.3758/s13428-024-02578-y

    View details for PubMedID 39810042

    View details for PubMedCentralID PMC11732908

  • Development and validation of a rapid and precise online sentence reading efficiency assessment FRONTIERS IN EDUCATION Yeatman, J. D., Tran, J. E., Burkhardt, A. K., Ma, W., Mitchell, J. L., Yablonski, M., Gijbels, L., Townley-Flores, C., Richie-Halford, A. 2024; 9
  • Rapid online assessment of reading and phonological awareness (ROAR-PA). Scientific reports Gijbels, L., Burkhardt, A., Ma, W. A., Yeatman, J. D. 2024; 14 (1): 10249

    Abstract

    Phonological awareness (PA) is at the foundation of reading development: PA is introduced before formal reading instruction, predicts reading development, is a target for early intervention, and is a core mechanism in dyslexia. Conventional approaches to assessing PA are time-consuming and resource intensive: assessments are individually administered and scoring verbal responses is challenging and subjective. Therefore, we introduce a rapid, automated, online measure of PA-The Rapid Online Assessment of Reading-Phonological Awareness-that can be implemented at scale without a test administrator. We explored whether this gamified, online task is an accurate and reliable measure of PA and predicts reading development. We found high correlations with standardized measures of PA (CTOPP-2, r = .80) for children from Pre-K through fourth grade and exceptional reliability (α = .96). Validation in 50 first and second grade classrooms showed reliable implementation in a public school setting with predictive value of future reading development.

    View details for DOI 10.1038/s41598-024-60834-9

    View details for PubMedID 38704429

    View details for PubMedCentralID 2752890