Conversational AI · Meta AI · 2025
Rebuilding AI evaluation around the conversation, not just the turn

Overview
The team's existing evals judged either a single model turn in isolation, or a turn plus its preceding composition history — but scored it against static per-turn metrics (should this turn be verbose, should it use cognitive empathy, etc.).
That didn't account for the fact that different conversations — and different moments within a conversation — call for different things. The dynamic of a conversation varies; a "good" turn in one context is wrong in another.
Approach
Built a better model of conversational dynamics through discourse analysis: reviewed conversation-analysis and dialectic literature.
Developed a new set of rubrics for evaluating conversations as a whole, not just single turns.
Experimented with different eval protocols to measure conversation-level quality, and explicitly tested whether the new measures added benefit over the existing per-turn measures — rather than assuming they would.
Impact
Partnered with cross-functional leaders to formalize the new rubrics and methodology for assessing conversation-level and multi-turn attributes at scale, with precision.
Used the same triangulation approach to build a strategic point of view on top AI chatbot use cases for the 1–3 year roadmap — including defining Education as a priority investment area.
Refined prompts and workflows for AI-assisted grounded theory, scaling reliable qualitative coding to datasets too large for manual analysis.
"What stood out most about working with Yann was his enthusiasm and can-do attitude — he was always ready to move projects forward and explore different approaches to complex problems. In his work on AI conversation evaluations, Yann demonstrated genuine curiosity about the various ways to tackle research challenges and showed great initiative driving progress for a complex project."
Frank J Kanayet, UXR Manager AI, Meta