I translate human values and intent into behavioral specifications engineers can implement and evaluations can grade — on shipping agentic products at consumer scale, across voice, text, and wearable surfaces where the right behavior differs by modality.
That isn't a copy task or a prompt tweak. It's a specification problem: define the boundary precisely enough that it can be trained toward, graded against, and enforced at runtime. It's the seam most teams leave unowned — where model behavior, evaluation, permissions, and audit all touch — and it's the one craft I've carried across three eras.
MFA, Writing for Performance (CalArts). Years directing how meaning lands live — timing, intent, the gap between what's said and what's understood.
Directed dialogue, IVR, and NLU at regulated scale — intents, entities, and recovery for tens of millions of interactions where a wrong turn had real cost.
Behavioral boundaries for agentic products — the taxonomies, gold sets, grading guidance, and safeguards that make those boundaries trainable and enforceable.
I author the behavioral layer for AI on Meta's consumer hardware — the rules for when the assistant acts, confirms, defers, or recovers, on surfaces where behavior is felt in the body, not read on a screen.
Authored the labeling guidance and gold set defining which utterances count as a hands-free "stats" request, plus the success / false-activation grading rules the classifier is measured against. The boundary held false activations under 10% during physical exertion.
Designed fallback logic and authored the Voice Behavior Layer across an array of wearables — how the assistant recovers, defers, and hands off when a request can't be completed cleanly.
Authored output and recourse / recovery rules for pre-release VR and robotics assistants: when the system acts autonomously, when it must confirm or defer, and how it recovers when it fails.
Defined the AI behavior layer for a Deaf and hard-of-hearing assistant on wearables — how the system adapts when audio-first behavior can't be assumed.
A governed multi-agent orchestration system for high-stakes contract analysis, where I treat separation of authority as safety architecture — each agent operates under a written behavioral contract and an explicit list of actions it will not take.
Designed how the system decides when to answer, ask, defer, or refuse — grounded in the known unreliability of self-reported model confidence. The specific mechanism is proprietary and withheld.
Authored the taxonomy the system uses to flag concealment and error — each entry pairing recognition criteria with the required response and, deliberately, what the system must not do, including dual-use and boundary cases.
An independent audit layer distinct from the components doing the work, permission controls that bound each agent, and citation verification kept separate from the generative path — integrity checked by something that didn't produce the output.
The system distinguishes a true capability failure from a question that belongs to a human professional — and routes accordingly instead of guessing.
Designed and tuned enterprise AI telephony handling 2M+ calls per week for prescription refills and Rx coverage — where precision, compliance, and trustworthy error handling weren't features, they were the requirement.
Defined the KPI framework for conversational success — instrumenting where each journey began and ended, and the criteria separating a completed task from a drop-out.
Ran the intent / entity taxonomy as a live loop: analyzed real calls, authored intents, re-sampled each cycle to close coverage gaps and correct mislabeled entities — improving precision and recall release over release.
Built the platform's first Spanish-language intent libraries, improving recognition and reducing fallback for bilingual callers.
Implemented error-recovery and override flows to keep continuity during misrecognition.
Chat-to-IVR expansion; preserved intent across modalities and reduced fallback through improved intent handling.
Enterprise scope from the start: interaction models and system behaviors for contact-center software serving 28,000 employees and 6,000 physicians; bilingual IVR for AutoZone, Virgin Money, Globe Life, and GAP.
Training data and structured input patterns to improve multilingual NLU coverage and inclusivity.
Bilingual ASR / IVR consultancy for U.S. / LATAM; Spanish-native voice assistants and cross-border CxD talent matching.
For twenty years I've worked the same problem in different clothing: how meaning is made, and who controls it. On stage it was performance. In the enterprise it was millions of calls a week where a misheard word had a cost. Now it's models — where the behavior a system exhibits is the product, and the specification of that behavior is the leverage.
I don't think the performer and the systems designer are two people. The instinct is the same: know exactly what should be said, what should be withheld, and when the silence is the more honest choice. That's dramaturgy, and it's also model behavior. The best people at this are creative people — which is why the field pays for taste, not just rigor.
What I keep finding is a seam no one wants to own, and it's where the trust actually lives. So I write the specs — the taxonomies, the gold sets, the grading guidance, the contracts — that turn "the model should behave well" into something an engineer can build and an eval can grade. And I build my own governed systems to keep the craft honest.
If you're building agentic or multimodal systems and the question of when a model should help, ask, or refuse is still unowned — that's my seam.