Meta · Multimodal LLMs
PATH 04 · 4 CHAPTERS
HSTU, from actions to generative recommendation
Trace HSTU tensor by tensor, train a tiny sequential recommender, then examine ragged execution, candidate scoring and evaluation.
- 01
HSTU starts with a question: which sequence are we predicting?
Separate content, actions, targets and candidates; understand sequential transduction and leakage boundaries.
- 02
HSTU tensor by tensor: SiLU aggregation, temporal bias and gating
Compare softmax attention with pointwise aggregation and trace U/Q/K/V, normalization and residual paths.
- 03
Train a tiny HSTU-inspired recommender: data, loss and top-k
Train executable PyTorch teaching code and verify causality, padding, gradients and recommendation outputs.
- 04
Scaling HSTU: ragged execution, candidate amortization and content features
Separate variable-length waste, candidate-dependent work and content embeddings to connect recommendation with multimodal understanding.