text · image · video
Multimodal content understanding
At Meta, I work on LLMs that understand content across language, images, and video—connecting model behavior to real product experiences.
Meta · Multimodal LLMs
Huiyu (Yvette) ChenMachine Learning Engineer at Meta
I’m a machine learning engineer at Meta, based in Singapore. My current work focuses on multimodal content understanding across language, images, and video—plus the training, evaluation, and production systems that make it useful.
research brain, product hands

Now · 现在
Building multimodal content-understanding LLMs at Meta—across language, images, and video—while staying interested in the whole path from how a model learns to how a person experiences it.
Full profile ↗︎What I do · 研究方向
I’m happiest in the messy middle—where a promising idea has to become a reliable, measurable experience.
text · image · video
At Meta, I work on LLMs that understand content across language, images, and video—connecting model behavior to real product experiences.
train · evaluate · ship
I care about training, alignment, evaluation, and production systems that stay measurable, efficient, and understandable at real scale.
read · question · explain
I translate papers, experiments, and hard-won debugging lessons into notes that another engineer can actually use.
A short timeline · 简历
details belong on LinkedIn 〰
Now
Building multimodal content-understanding LLM systems across text, images, and video.
Previously
Built production AI assistants for e-commerce across multiple markets. That chapter taught me how to carry an LLM idea from training to real users—without turning this homepage into a quarterly report.
Before that
Research training in NLP and machine learning, with a lasting habit of reading the appendix.
Off screen
Still learning, still moving, still evolving.
Speaking · 演讲
A practical talk about taking chatbot assistants from architecture decisions and model training to alignment and production scale.
Open 27 slides ↗︎APAC Data Innovation Summit 2026
I shared the system and model choices behind production-grade conversational AI—from retrieval and fine-tuning to alignment, evaluation, and deployment.
Writing · 写作
01
Today we skip paper lists and do one deep dive.
02
I’ve recently been diving into memory management for dialog-based AI, especially how to construct and retrieve memories in long-term conversations. During my exploration I came across an eye-opening ICLR 2025 paper—**”Se…
03
DeepSeek-R1、OpenAI o3-mini 和 Google Gemini 2.0 Flash Thinking 是如何通过“推理”框架将 LLM(大型语言模型, Large Language Models) 扩展到新高度的典型示例。
Elsewhere · 生活支线
A spinning planet of places I’ve been, plus the less polished thoughts that escape onto Xiaohongshu.