Current work
Public preprint · arXiv:2607.01153 (2026)
Safety evaluations turn on judgments about ambiguous language: whether a model followed an instruction, refused well, resisted an embedded command, or misreported progress. Pass/fail labels hide which of those failed.
- Scope18-item seed benchmark with validator-enforced metadata, plus a 54-row seed pilot.
- CoversInstruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, multi-turn agent transcripts.
- FindingA rubric-aided LLM judge, grading its own outputs with the expected-behaviour fields visible, still missed the safety-relevant minority classes.
- LimitA seed benchmark and a protocol. It does not establish that these categories hold at scale, across model families, or under other judges. Testing that is what the metrics are for.
The benchmark, protocol, and code are open to inspection because a headline score doesn’t show where its inferences will travel.
In development
Two companion frameworks: delegation assurance, for tool-using systems that act under delegated authority, and evidentiary assurance, for audit, challenge, and remediation. Two completed manuscripts, being prepared for public release.
What does a label license you to infer?
Labels don’t need essences to be useful. They need a record of supporting the right inferences for a stated field and purpose. I study how that warrant is produced, how far it projects, and how it fails, in grammatical categories, benchmark scores, and evaluator judgments.
Each row is one form of that question. Each column is a field that has to settle it on different evidence. Some work answers a row in more than one field at once, and appears in more than one cell; that overlap is the point rather than an error in the filing.
Trace one field: AI evaluation · philosophy of science · English grammar · show all
| In AI evaluation | In philosophy of science | In English grammar |
|---|---|---|
| What does membership let you predict? | ||
| In AI evaluation Adversarial Pragmatics Public preprint — Adversarial Pragmatics · What a safety-benchmark score licenses | In English grammar Interjection as a lexical category Public preprint — Interjection as a lexical category · What membership lets us predict | |
| Who is competent to judge, and in what role? | ||
| In AI evaluation The adjudication protocol Public preprint — The adjudication protocol · LLM-judge validation, open to inspection | In philosophy of science Truth-tracking profiles Public preprint — Truth-tracking profiles · What large language models participate in | In English grammar Expert grammaticality judges Public preprint — Expert grammaticality judges · Evaluators, not participants. The same argument constrains who may adjudicate an AI evaluation |
| Where does the inference stop? | ||
| In AI evaluation Delegation assurance In development · Tool-using systems acting under delegated authority | In philosophy of science Effective without warrant Completed manuscript — Effective without warrant · Status that is effective without being warranted | In English grammar The homeostatic maintenance of English countability Completed manuscript — The homeostatic maintenance of English countability · The cluster dissociates in a constrained order |
| What keeps a category standing when nothing essential holds it together? | ||
| In AI evaluation Evidentiary assurance In development · Audit, challenge, and remediation | In philosophy of science Not every stable cluster is homeostatic Public preprint — Not every stable cluster is homeostatic · Separating achievements the homeostatic component conflates | In English grammar Grammaticality de-idealized Public preprint — Grammaticality de-idealized · Grammaticality as conditioned stability |
| General framework | ||
| Words That Won’t Hold Still: How Linguistic Categories WorkCompleted manuscript · Book-length statement. | ||
| Kinds as projectibility profiles: support grades and demotion rulesPublic preprint — Kinds as projectibility profiles: support grades and demotion rules · How a category earns, keeps, or loses its standing | ||
Open questions
Four things I don’t know, and what would make me give up each position.
- When do individually warranted benchmark inferences fail to compose? Each step of an evaluation chain can be licensed on its own while the chain is not. The seed benchmark defines metrics for whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. Give it up if the categories project cleanly across all four conditions, since then the diagnostic apparatus is answering a question nobody needed asked.
- Is grammaticality conditioned stability, or just frequency wearing a new name? The operator–value model predicts that opportunity structure, preemption, community licensing, and exposure interact, and that some rejected constructions soften under exposure while others resist. Give it up if a hierarchical study with partial pooling finds exposure effects that do not interact with opportunity and preemption, and the model fits no better than frequency, surprisal, or processing accounts.
- Why does the English count cluster dissociate in a constrained order? Count properties cluster tightly enough to be projectible, but quasi-count nouns (cattle, police, clergy) come apart in a specific sequence. An account has to explain the projectibility and its limits together. Give it up if the dissociation order turns out to vary arbitrarily across speakers and varieties, which would make it noise rather than structure.
- Do stability, network order, maintenance, and corrective control really come apart? Homeostatic property cluster theory can be read as covering several achievements that are not the same thing. The claim is that they dissociate, and that only one of them earns the word homeostatic. Give it up if a survey of candidate kinds finds the four always co-occur, which would make the distinction bookkeeping rather than a finding.
If you work on any of these, I’d like to hear from you: brett.reynolds@humber.ca
Selected publications
- Huddleston, R., Pullum, G. K., & Reynolds, B. (2022). A Student’s Introduction to English Grammar (2nd ed.). Cambridge University Press.
- Reynolds, B. (2026). Language Landscapes: The ESL teacher’s guide to how English works. Language Science Press.
- Jespersen, O. (2025). Negation in English and Other Languages. New edition co-edited with Peter Evans. Language Science Press.
- Reynolds, B. (2026). The lexicon–syntax boundary in English numerals. English Language and Linguistics, 1–19.
- Reynolds, B. (2024). Why more and less are never adverbs. Journal of Linguistics.
- Arora, A., Schneider, N., & Reynolds, B. (2023). Unified syntactic annotation of English in the CGEL framework. Proceedings of the 17th Linguistic Annotation Workshop.
Full publication list, 1998 to present.