Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

← Back to publications

A linguistically controlled benchmark and annotation protocol for evaluating language-model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. Designed to extend to multi-turn agent transcripts, though the seed set represents that family with a single-turn tool-result contrast.

Status: arXiv preprint.

The Markdown file is an author-manuscript mirror provided for accessibility, search, and machine readability. Use the linked public record as the canonical citation target unless a later publisher version supersedes it.

Project Pages