Scoring Employment Interviews With Large Language Models: Evaluation Design Components, Validity Investigations, and Best Practice Recommendations
成果类型:
Article; Early Access
署名作者:
Stockdale, Kayden; Hickman, Louis; Liu, Siyi
署名单位:
Virginia Polytechnic Institute & State University
刊物名称:
JOURNAL OF APPLIED PSYCHOLOGY
ISSN/ISSBN:
0021-9010
DOI:
10.1037/apl0001396
发表日期:
2026
关键词:
assimilation
reliability
accuracy
ratings
摘要:
Recently, supervised machine learning models served as alternatives to humans for scoring open-ended assessments, including employment interviews. More recently, large language models (LLMs) emerged as alternative raters that do not require task-specific training data. However, interview evaluation design considerations have predominantly been conceptualized for human raters, and thus, it is unclear whether the psychometric benefits of interview evaluation best practices generalize to LLM raters. Additionally, LLM raters introduce novel evaluation design decisions that could be psychometrically consequential. We investigated the effects of several LLM rater evaluation practices-specifically, prompt design, model selection, hyperparameters, and number of interviewers-on the LLM interview scores' psychometric properties in two interview data sets. In the first, interviewee Big Five personality traits were evaluated (N = 954). In the second, interviewees' question responses were evaluated on the targeted construct on Behaviorally Anchored Rating Scales (N = 144). We then investigated the LLM scores' intrarater reliabilities, test-retest correlations, convergent, discriminant, and criterion evidence of validity, group differences, and measurement bias. We compared this evidence, when possible, to the same evidence for human raters and supervised machine learning models. The results suggest that ensembles of larger, newer LLMs using prompts with detailed construct information hold potential for scoring employment interviews with psychometric properties comparable to or superior to supervised machine learning models and single human raters. We detail the reasons that organizations may want to be cautious in adopting LLMs for scoring high-stakes open-ended assessments, but since organizations have already begun adopting them, we also offer best practice recommendations.