All learning paths阅读中文版 →

PATH 04 · 4 CHAPTERS

HSTU, from actions to generative recommendation

Trace HSTU tensor by tensor, train a tiny sequential recommender, then examine ragged execution, candidate scoring and evaluation.

  1. 01

    HSTU starts with a question: which sequence are we predicting?

    Separate content, actions, targets and candidates; understand sequential transduction and leakage boundaries.

  2. 02

    HSTU tensor by tensor: SiLU aggregation, temporal bias and gating

    Compare softmax attention with pointwise aggregation and trace U/Q/K/V, normalization and residual paths.

  3. 03

    Train a tiny HSTU-inspired recommender: data, loss and top-k

    Train executable PyTorch teaching code and verify causality, padding, gradients and recommendation outputs.

  4. 04

    Scaling HSTU: ragged execution, candidate amortization and content features

    Separate variable-length waste, candidate-dependent work and content embeddings to connect recommendation with multimodal understanding.