← All writingHSTU, from actions to generative recommendation · 02

HSTU tensor by tensor: SiLU aggregation, temporal bias and gating

Compare softmax attention with pointwise aggregation and trace U/Q/K/V, normalization and residual paths.

阅读中文版 →

HSTU tensor by tensor: SiLU aggregation, temporal bias and gating

A complete explanation from first principles

Begin with the block's function

HSTU means Hierarchical Sequential Transduction Unit. Treat it first as a function that transforms behavior-sequence representations, then inspect projections, aggregation, normalization and gating. Replacing Transformer softmax with another function alone does not reproduce the full design.

Efficient implementations may use ragged storage rather than dense [B,T,d]. This chapter uses a fixed-length sequence for inspection. Logical tensor relationships and physical storage are separate layers of understanding.

Follow the U, V, Q and K branches

Learned projections produce branches with different roles: Q/K form position-matching scores, V supplies aggregated content and U participates in gating. An implementation may obtain them through one larger projection followed by slicing, while retaining distinct parameter subspaces.

A schematic flow forms biased QK scores, applies a SiLU-like transformation and valid-position constraints, aggregates V, normalizes, multiplies by a U-related branch, projects and adds a residual. Exact scales, normalization and layouts belong to the paper and implementation; the browser example explicitly omits components.

SiLU aggregation coefficients are not probabilities

SiLU is x*sigmoid(x). At -1 it is approximately -0.269, at zero it is zero, and at one it is approximately 0.731. Unlike softmax, it does not produce nonnegative row-normalized coefficients.

With values [1,0] and [0,1], those coefficients yield [-0.269,0.731], outside their convex-combination line segment. Negative coefficients can subtract feature directions, but do not directly mean dislike for an item. Gating and output projections further change the representation.

Do not copy softmax masking mechanically

Softmax commonly masks illegal scores with negative infinity before normalization. Directly applying SiLU to negative infinity can involve infinity multiplied by zero and produce NaN. A valid implementation can zero illegal aggregation coefficients after the nonlinearity or use another verified treatment.

Distinguish history, padding, prediction and candidate positions. A triangular-looking mask is not a full correctness test. Change future inputs and verify that earlier outputs remain unchanged.

Relative time adds information content similarity lacks

Two interactions with the same item can have different relevance when one happened a minute ago and the other a month ago. Relative-time or position biases make such distinctions expressible without declaring all old behavior useless.

Units, bucketing and truncation matter. Training in seconds but serving milliseconds can place events in different bias regions. Input-contract errors may be less visible than formula errors.

The experiments check signed coefficients, the effect of changing temporal bias on actual aggregates, and causal/padding boundaries. The browser notebook fixes embeddings and parts of the aggregator, then trains a scoring head. The full PyTorch notebook offers a more complete but still educational HSTU-inspired block with end-to-end training. Neither claims to reproduce an entire production recommendation system. Explicit omissions make each experiment's evidence easier to interpret.

Similar shape, different aggregation

Standard attention softmax-normalizes scores across each row. HSTU includes pointwise activated aggregation: transform biased scores with a function such as SiLU, aggregate V, then normalize and gate. Without sequence-wise softmax, the weight matrix is not a probability distribution.

Our single-head teaching equation is , followed by . The demo uses fixed sequence length N. Check official code for actual scaling, head layout, bias and normalization.

Trace four projections

Normalized [B,T,D] inputs produce U/V/Q/K projections. Q/K match positions, V supplies aggregated values and U gates the output. Q/K head dimensions can differ from U/V dimensions. Fusing projections can reduce operator overhead, but the split must match weight layout.

import torch
import torch.nn.functional as F
q, k, v = [torch.randn(1,4,8) for _ in range(3)]
mask = torch.ones(4,4,dtype=torch.bool).tril()
a = F.silu(q @ k.transpose(-2,-1)) / 4
a = a.masked_fill(~mask, 0)
y = a @ v
print(a.sum(-1))  # Not required to equal one.

SiLU can produce negative coefficients. Probability entropy cannot be applied directly to this matrix without defining a different meaningful measure.

Temporal bias is not simply fixed forgetting

Relative positions and elapsed time distinguish three consecutive events from three events spread over months. Real bias parameterization includes learned implementation-specific details. The interactive bias is deliberately simple and is not the paper's complete construction.

Increasing α can make old-position weights negative. Combined with normalization and gating, effects on outputs need not be monotonic. Inspect outputs and ablations rather than guessing recommendation behavior from heatmap colors.

Check your understanding

Question: Does replacing softmax with SiLU fully implement HSTU?

AnswerNo. Data formulation, projections, gating, normalization, temporal/position bias and system optimizations also matter. The notebook is explicitly HSTU-inspired teaching code, not a complete reproduction.

Primary sources and further reading

Sources checked on 2026-09-10. Teaching examples are not production benchmarks; confirm APIs and model support against the linked version.

INTERACTIVE LAB · LIVE NUMERICAL COMPUTATION

HSTU-style aggregation need not sum to one

Fixed 2D Q/K, a causal mask and teaching bias −α(i−j). Displaying SiLU(QKᵀ+bias)/4; full HSTU also includes U/V projections, normalization, gating and a residual.

Weights may be negative; these are not probabilities
Q / Kt0t1t2t3
t00.1830.0000.0000.000
t1-0.0470.1830.0000.000
t20.0000.0780.4400.000
t3-0.047-0.047-0.0670.243

Row sums: 0.183 / 0.136 / 0.518 / 0.081

Increasing α can make earlier positions negative, rather than simply less probable. Softmax entropy analysis does not apply to this matrix.

NOTEBOOK · LIVE PYTHON

Run the notebook inside this article

Edit and run cells, or run all. Variables persist between cells until you leave the page or restart the kernel. After editing an upstream cell, rerun the cells below it.

This edition runs inspectable small computations and training in real Python without a GPU. This kernel does not include PyTorch/CUDA. The other tab contains the full pretrained-model notebook; the foundations notebook also has a separate editable cell for real MiniLM inference.

Kernel not loaded; click Run to start.

Loading notebook…