AI evaluation · Assurance · Linguistics
Brett Reynolds is an AI evaluation and assurance researcher grounded in linguistics and philosophy of measurement. He is a linguist at Humber Polytechnic and an adjunct professor of linguistics at the University of Toronto. His work asks what labels, scores, traces, and judge results license us to infer, where those inferences project, and where they break. In linguistics, this means English grammar and grammatical categories: syntactic annotation, usage-based and constructionist analysis, the philosophy of linguistics, and what grammaticality judgments can – and can't – tell us. In AI evaluation, it means benchmark validity, evaluator reliability, language-mediated control, delegated authority, and audit evidence. Recent work includes an adversarial-pragmatics seed benchmark and evaluation pipeline for instruction conflict, embedded commands, policy ambiguity, refusal calibration, and LLM-judge validation; a delegation-assurance framework for tool-using AI systems; and an evidentiary-assurance framework for audit, challenge, and remediation. He co-authored the second edition of A Student's Introduction to English Grammar (2021) and co-edited a new edition of Negation in English and Other Languages (2025). Language Landscapes (a TESL textbook) was published by Language Science Press in 2026. His current monograph is Words That Won't Hold Still: How Linguistic Categories Work, a completed manuscript under peer review at Cambridge University Press.
Research focus: AI evaluation and assurance; benchmark validity; evaluator and gold-label validity; language-mediated control; delegated authority; audit evidence and contestability; category stability and inference; English grammar and grammatical categories; philosophy of linguistics.