MENU
02 — 翻譯所見
2024 – 2025 | MULTIMODAL AI DESIGN LEAD ON DISPLAY GLASSES

Connect people with the world,
so they look forward, not down.

CONTEXT

Meta's mission was building the future of human connection. We needed compelling use cases to find product-market fit on display glasses.

ROLE

I led the end-to-end multimodal AI design leveraging photo input and OCR to help people understand and take actions on what they see, including standardizing the human-AI interaction model and gesture vocabulary, contributing to design system, and creative direction for onboarding, education, and commercial artwork.

OUTCOME

Launched at Meta Connect 2025. The interaction model my engineer and I designed was filed as a patent, with me as first inventor.

A restaurant abroad — a menu in a language you can't read
IMAGINE・身臨其境

A restaurant abroad. A menu you can't read.

COMPLEX CHALLENGES

How can wearers consume translations on a transparent display the size of a postage stamp?

If the wearer can't discover the gesture, can't navigate, and can't read the translation — the AI feels broken regardless of whether the translation model is accurate.

Consumer

PAIN POINT 01

Discovery

Wearers didn't know zoom and pan were possible on the 600 × 600px display.

PAIN POINT 02

Gesture friction

Captouch swipe to zoom was easily mistaken as taps. Head-IMU panning felt jittery.

PAIN POINT 03

Legibility

The additive display, contrast, and image resolution can make it hard to read texts.

Technical

COMPLEXITY 01

Natural invocation

Corresponding visual response templates only fired on rigid commands — wearers had to remember exact trigger phrases for the AI to show rich visual responses.

COMPLEXITY 02

Image resolution ↔ latency

Sharper text on the 600 × 600px display needs higher-resolution renders and deeper zoom — heavier processing, slower response.

COMPLEXITY 03

System understanding ≠ human understanding

OCR groups real-world text by adjacency — orientation, line breaks, and layout work against it.

COMPLEXITY 04

What you see ≠ what you hear

1 OCR input feeds 2 separate models. The rendered translation includes over-detected noise the LLM's audio response cuts — and even on clean input, 2 models means 2 different wordings.

OPPORTUNITIES BREAKDOWN

A translation you can
read, move through, and act on —
with nothing to touch.

Every design decision held to 3 principles: intuitive, easy, quick.

FOCUS 01

Result format

How might we show the translated result on the display and in the AI's audio response — what's the relationship between the two?

FOCUS 02

Navigation

How might we enable wearers to move through the translated content — with no direct-manipulation surface like a touchscreen?

FOCUS 03

Follow-up actions

How might we offer the right actions once the translation lands to help wearers reach their goal?

RELENTLESS ITERATION

The devil is in the details.

01 · THE FORMAT CALL

Understanding was never the goal —
acting on it was.

A translation is only useful if you can map it back to the world in front of you. I made the call on intuition and competitor analysis, we built it, then validated with UXR.

The generic overlay flow: translation overlaid on OCR-detected text, auto-zoom and centring on the target when the gesture is released, then a follow-up rail to toggle between original and translation.
The finger-pointing overlay flow: the pointed line item is highlighted, auto-zoomed and centred while the rest stays translated, then the same follow-up rail toggles original and translation.
HOW WE SHOW TRANSLATIONS?

Two ways to put a translation on a display you can’t touch. I took the one that keeps your eyes in one place.

Four display screens with the translation drawn directly over the captured menu, so the words sit where the original text is.
PRO
  • Can transfer attention to the display entirely, with no extra eye movement to scan and compare between capture preview and translation
CON
  • Can’t scale to include image
  • Might be challenging to read depending on the color of the documents and the background environment users are in
02 · THE NAVIGATION STRUGGLE

Machine picks by proximity;
human follows a path.

Next, how do you move through it — on a screen this small, with nothing to touch? Constant pinch-pan-zoom is exhausting, so I bet on removing the manual work with Smart text selection.

Text selection may sound simple. But in reality it's much messier than you can imagine. Real-world text has various orientation, linebreak, and layout. Jiaqian and I relentlessly refined the calculation so the system identifies orientation, proximity, and chooses the next text block aligned with human reading order. The calculation also reduces panning to the minimum to avoid dizziness.

HOW DO YOU MOVE THROUGH IT?

Two ways to get from one line to the next with nothing to touch. Both auto-zoom; they disagree about what moves — the image, or the selection.

The snapping flow: the capture pans under a fixed viewport, and on release the nearest text box snaps to centre and auto-zooms.
PRO
  • Same mental model as most of the image panning mechanism: fixed viewport + moving image
CON
  • Less confident about which line item is going to be zoomed in
  • When text boxes are really close to each other, how does the bounding box work
03 · THE DISAGREEMENT

Flexibility was the design —
not the fallback.

Not everyone agreed on leaning this hard on automation. Leadership pushed to keep manual free pan-and-zoom. We didn't settle it by winning; we made the modes complement each other: smart selection for the common path, free pan-and-zoom whenever you want it.

The end-to-end Visual Translation flow
DAY-0 KEY FEATURES

How it works.

1
2 WAYS TO ASK

Ask generically by natural invocation or specify with your finger.

2
2 WAYS TO INTERACT

Pinch and drag to pan where you want. Wrist roll to zoom in on the section. Swipe to consume each line.

3
THE RIGHT CLICK

Reveal the original language or shift long-form text into system template tuned for 600 × 600 legibility. Other follow up actions are designed and scoped in future milestones.

A restaurant abroad — a menu in a language you can't read
DAY-90 REFINEMENT · PARTNERED WITH INTERACTION MODEL & EDU TEAM TO DEFINE THE PLATFORM-WIDE SOLUTION

Resolve gesture conflict, provide contextual education

IMPACT

Shipped, patented, recognized.

Multimodal AI design systemConnect 2025 launchPatent filed · first inventorTIME Best Inventions 2025UX Design Award 2026
Multimodal AI design systemConnect 2025 launchPatent filed · first inventorTIME Best Inventions 2025UX Design Award 2026
REFLECTION

I’d go upstream sooner.

If I did it again, I'd build a tighter feedback loop with the ML, OCR, and Camera teams earlier. Some of my design decisions were band-aids for upstream model and camera limitations — earlier collaboration would have let me shape UX-relevant model behavior directly, instead of designing around it downstream.

NEXT — HANDS-FREE PAYMENTS