Agent systems · deployment · evaluation

The agent is only part of the system.

I'm a forward-deployed AI engineer working on conversational agents, deployment tooling, and evaluation methods. I'm interested in both sides of the work: building systems people can use and understanding what those systems are actually doing. My background in linguistics helps with the second part, and often the first.

Selected work

All work →

Evaluation methodology · 2025

Evaluability taxonomy for summary faithfulness

For this multi-author study, I designed a taxonomy of ambiguity, vagueness, and other conditions that make factuality judgments ill-posed. I also improved annotation quality through data analysis, annotator interviews, revised guidelines, and a Socratic review tool.

Read the paper ↗

Meta-evaluation benchmark · 2025

MDSEval

A benchmark for testing whether automatic evaluation methods agree with human assessment of multimodal dialogue summaries.

Read the paper ↗

Method in practice

Before Five9, I spent four years at AWS across Lex, Transcribe, and Contact Lens. I designed evaluation methods, built annotation and data tooling, and developed agentic systems for large-scale annotation and synthetic conversation generation.

Linguistics enters my work both directly and through the way I approach evidence. In synthetic-dialogue systems, I draw on pragmatics, discourse structure, and language variation to model how conversations unfold and what realistic behavior looks like. My dissertation used controlled acceptability experiments to test whether English shows the same attenuation of relative-clause island effects observed in other languages. The results suggested that it does, raising broader questions about grammar, processing, and comprehension. That work also trained me to isolate the signal a test is meant to measure, account for what else could produce it, and be clear about what the result does and does not show. I bring that same discipline to agent design and evaluation now.