How non-verbal behavior signals enhanced multimodal AI interaction
Multimodal LLMs are empowering the next generation of voice agents
With the rapid rise of LLMs and wearable smart devices like AR glasses, users now have unprecedented opportunities to interact with on-device assistants through both voice and gesture. While Voice User Interfaces (VUIs) are well-understood, the potential for full-body gestures that contextualize a user’s surroundings remains largely unexplored.
At Meta, we are heavily investing in immersive hardware, from VR headsets like the Oculus Quest 3 to AR glasses such as Meta Ray-Ban, which all support rich multi-modal inputs including facial expressions and body gestures alongside voice. It only makes sense to push further by deepening our understanding of how these smart devices can interpret user intention through the integration of this new data with traditional voice inputs.
Rather than jumping straight into designing specific interaction techniques for our voice agents, I took a step back and advocated to establish the fundamentals first. Before we can innovate on what users interact with, we must understand who they are and how they naturally communicate when given multi-modal tools. Just as we use body language daily to convey intent without words, we needed to de-risk future development by validating how users naturally gesture with voice assistants. This foundational work is essential not only for technical feasibility but also for proving real consumer desirability before scaling our product roadmap.
What we needed to learn
To truly understand the user landscape, I structured this research around three key pillars. The goal was not just to gather data, but to uncover users' natural behaviors before any innovation is introduced and identify the underlying motivations and rationales driving those actions.
- 1. Opinions: What are users' general attitudes towards interacting with multi-modal VUIs?
- 2. Preferences: How do users perceive the use of gestures in multimodal VUIs, and what specific preferences guide their choices?
- 3. Gestures: Which implicit (unconscious) and explicit (conscious) gestures are commonly used when interacting with multi-modal VUIs, and what functions or intents do these gestures serve? :::
::: sec-label "Methods & Process"
A Two-Phase Wizard-of-Oz Study
To investigate future technology without the heavy resource cost of full implementation, I used the Wizard-of-Oz (WoZ) methodology. This approach allowed us to simulate advanced capabilities and answer our core research questions efficiently with timeline and resource constraints. To ensure these findings translated into actionable product insights, I organized internal critique sessions with senior research leadership and cross-functional teams. Additionally, we conducted iterative design workshops where designers and developers tested the generated design space, providing rapid feedback on the co-design experience to accelerate feature iteration.
::: methods-grid
Phase 1: Discovery Wizard-of-Oz Study
I recruited 6 participants in a simulated environment to mock up six daily voice assistant usage scenarios based on common tasks identified through my literature review. This initial phase yielded essential categories of gesture functions and uncovered four key factors influencing how users interact with voice assistants via gestures.
Phase 2: Validation Wizard-of-Oz Study
To validate these findings and refine the initial design space, I conducted a second WoZ study with 12 participants in a fully staged environment designed to mimic real-life situations. I confirmed the four influential factors and built a robust design space featuring four validated categories of gestures specifically for conveying user attention.
Stakeholder Alignment
Throughout both phases, I held weekly sessions with senior research leadership and multifunctional stakeholders. These meetings ensured the study results were not only methodologically sound but also directly applicable to our business objectives.
Internal Design Workshops
I facilitated workshops where cross-functional teams co-designed interaction flows. These sessions focused on rapid prototyping and feedback integration, allowing us to iterate quickly before moving into engineering implementation.
Three findings, three product implications
Each key finding was traced directly to a product implication for future development, ensuring the research had measurable impact.
Users recognize the benefits of using gestures with voice assistance.
Implications: To promote future products using voice and gestures together, priority should shift from mere feature implementation to helping users build an intuitive mental model of how gestures interact with voice. Simultaneously, providing the public with more information can increase social acceptance for using these combinations in everyday scenarios.
Users can learn and retain gestures with voice assistants.
Implications: Since gesture usage can be learned and maintained over time, it is crucial to provide a smooth onboarding experience and actively encourage users to adopt gestures from the very beginning of their interaction with a new product. Without this early reinforcement, users are likely to revert to voice-only interactions.
The most intuitive gestures mimic real-life object manipulation.
Implications: The system should be designed to accommodate users who will naturally use customized gestures mimicking daily physical interactions. When creating interaction techniques for the new product, designers must translate these everyday physical actions into effective gesture-and-voice design.
Actionable design framework and insights driving future development.
Delivered within a strict six-month timeline, this project successfully transitioned from zero to one by investigating user intentions for gesture-voice integration, establishing a foundational understanding of motivation and rationale. The research synthesizes an actionable design framework to accelerate feature iteration while culminating in a peer-reviewed publication accepted at a top-tier HCI venue.
- Empirical Insights User preference and attitudes toward future products.
- Categorized gestures Explicit gestures organized into 4 functional categories.
- Product Implications 3 essential implications for future product development.
- Design Framework Streamlined, structured workflow for rapid iteration. :::
::: sec-label "Reflection"
Honest tradeoffs and what I'd do differently
Every project involves compromises in scope, depth, and timing. Here's what I would change if revisiting this work:
::: reflection-list