Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License
ORCID
https://orcid.org/0009-0002-2408-1397
Date of Graduation
5-14-2026
Semester of Graduation
Spring
Degree Name
Master of Arts (MA)
Department
Department of Graduate Psychology
First Advisor
Brian Leventhal
Abstract
In this study, I evaluated the feasibility of using an automatic speech recognition (ASR) system and large language model pipeline to transcribe and classify student audio into predefined thematic categories as a proof-of-concept step toward scalable qualitative analysis for culturally responsive assessment development. I collected audio recordings of college students reading short stories representing themes of risk and needs and experimentally manipulated the recordings by overlaying background noise at varying signal-to-noise ratios (SNR). I transcribed the audio using Whisper Turbo ASR and evaluated transcription quality using word error rate (WER). I then examined downstream thematic classification performance across multiple modeling approaches, including fine-tuned and frozen RoBERTa models and zero-shot and few-shot prompting with Copilot Chat. Across analyses, acoustic conditions influenced transcription performance more strongly than the speaker characteristics represented in this sample. Classification performance declined as transcription quality worsened, although this pattern varied across models. Encoder-based approaches, especially the fine-tuned RoBERTa model, produced stronger and more stable classification performance than decoder-based Copilot Chat approaches under the conditions examined. These findings suggest that automated ASR-LLM pipelines show promise for supporting large-scale analysis of spoken data. However, their usefulness depends on transcription quality, modeling approach, and evaluation under realistic acoustic conditions before they can be applied to authentic educational interview contexts.
