YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
HSTU Generative Recommendation Model Test Configuration
config.json in this directory is the test configuration for the HSTU (Hierarchical Sequential Transduction Unit) generative recommendation model. It describes a test model with an intended size of approximately 0.2B parameters and declares HSTUForCausalLM as its model entry point.
Note:
use_random_modelis set totrue. The configuration therefore uses randomly initialized weights for inference-pipeline, operator, and performance tests; it does not indicate that a trained model checkpoint has been loaded.
1. Configuration Summary
| Configuration | Value | Description |
|---|---|---|
model_type |
hstu |
HSTU model type |
architectures |
HSTUForCausalLM |
Model loading entry point. CausalLM is a framework-compatible interface name; the actual task is candidate scoring for recommendation. |
| Intended model size | Approximately 0.2B |
Target size of the test model represented by this configuration |
num_hidden_layers / hstu_config.num_layers |
12 |
Number of HSTU layers |
hidden_size |
1024 |
Hidden-state dimension at each token or feature position |
num_attention_heads |
8 |
Number of attention heads per layer |
head_dim |
128 |
Dimension of each attention head; 8 Γ 128 = 1024 |
max_seq_len |
8832 |
Maximum sequence length allowed for one request. The exact number of history and candidate positions depends on the input-packing scheme. |
torch_dtype / hstu_config.dtype |
float32 |
Data type used by this test configuration |
is_causal |
true |
Enables causal masking, so a position cannot read future positions |
residual |
true |
Enables residual connections in HSTU layers |
has_ffn |
true |
Includes a feed-forward network (FFN) in each HSTU layer |
dropout_ratio |
0 |
Dropout is disabled for inference |
norm_epsilon |
1e-5 |
Numerical-stability epsilon used by normalization layers |
2. Input Features and Embedding Tables
task_config.embedding_configs defines the recommendation features used by the model:
| Feature | Table | Vocabulary Size | Embedding Dimension | Dynamic Embedding | Purpose |
|---|---|---|---|---|---|
action_weights |
act |
1024 |
1024 |
No | User behavior/action type or behavior weight |
video_id |
item |
100000 |
1024 |
Yes | Video/item ID used as a candidate-item feature |
The configuration also declares:
item_feature_name = "video_id"action_feature_name = "action_weights"- Both input embedding dimensions match
hidden_size = 1024.
In a typical HSTU inference implementation, input features first go through embedding lookup. Multiple features at the same position are combined through concatenation and/or a linear projection to form a 1024-dimensional input representation. Positional information is then added before the representation is passed to the HSTU backbone.
3. HSTU Model Architecture
The logical path from model inputs to recommendation scores is:
User-history features + candidate-item features
β
ββ action_weights embedding: 1024 β 1024
ββ video_id embedding: 1024 β 1024
β
ββ Feature combination / linear projection + positional embedding
β
βΌ
ββββββββββββββββββββββββ
β HSTU Layer 1 β
β β¦ β
β HSTU Layer 12 β
ββββββββββββββββββββββββ
β
βΌ
Hidden states at candidate positions
β
βΌ
Dense/MLP score head
β
βΌ
One scalar score for each candidate
A single HSTU layer can be summarized as follows:
Input x
β
ββ Input normalization (epsilon = 1e-5)
ββ Fused UVQK linear projection
β ββ Split into U, V, Q, and K representations
ββ SiLU activation
ββ Causal attention (8 heads, head_dim = 128)
ββ Attention-output normalization
ββ U gating / element-wise modulation of the attention output
ββ Linear projection back to 1024 dimensions
ββ SwiGLU FFN (has_ffn = true)
ββ Residual connection (residual = true)
β
Output x'
This is not a conventional language model that generates text tokens. It uses a ForCausalLM-style interface to apply causal sequence modeling to the joint representation of user behavior history and candidate items, and ultimately produces a ranking score for each candidate. The current config.json does not explicitly specify the intermediate architecture of the prediction head under task_config; the concrete head configuration is determined by the model implementation or loader defaults. Its task semantics are to output one score per candidate.
4. Inference Paradigm: Generative Recommendation vs. LLM
Both model types may expose a CausalLM-style interface, but their decode-stage objectives are fundamentally different.
4.1 Generative Recommendation: Parallel Candidate Scoring
Given user history (h) and a candidate set (C = {c_1, ldots, c_M}), HSTU computes a score for every candidate:
[ s_i = f_\theta(h, c_i), \qquad i = 1, \ldots, M ]
During decode, multiple candidates are arranged as positions or packed sequences that can be evaluated in parallel. A forward pass produces the score vector:
[ \mathbf{s} = [s_1, s_2, \ldots, s_M] ]
The candidates are then sorted by score and truncated to Top-K. Thus, the primary parallel dimension during decode is the candidate dimension. As the candidate count grows, candidate batching, candidate parallelism, and memory usage become important. The model does not need to generate (c_1) first and then generate (c_2) conditioned on (c_1).
4.2 LLM: Autoregressive Token-by-Token Generation
Given a prompt prefix (x_{<t}), an LLM computes the distribution of the next token at step (t):
[ p(x_t \mid x_{<t}) = \operatorname{softmax}(W h_t) ]
After selecting or sampling a token, the token is appended to the sequence and the model proceeds to step (t+1), continuing until a complete sequence is generated or an EOS token is reached. The logits for the whole vocabulary can be computed in parallel at each step, but the time dimension remains autoregressive and sequential. The KV cache is primarily used to reuse the already generated prefix states.
4.3 Comparison
| Dimension | HSTU Generative Recommendation | Conventional LLM |
|---|---|---|
| Decode objective | Evaluate a set of candidate items simultaneously | Generate the next token |
| Computation unit | Candidate / candidate position | Token / time step |
| Typical output | M candidate scores, followed by sorting and Top-K selection |
One token per step, repeated to form a complete sequence |
| Dependency pattern | Candidates generally do not require autoregressive dependencies on one another | The current token depends on previously generated tokens |
| Main parallel dimension | Candidate dimension; candidates can be batched and evaluated in parallel | Vocabulary logits are parallel within a step, but time steps are sequential |
| Stopping condition | Candidate scoring completes, followed by ranking/truncation | EOS, maximum generation length, or another stopping rule |
| Cache purpose | Reuse the user-history representation to reduce candidate-scoring cost | Reuse historical token K/V states to reduce incremental-generation cost |
In simplified form:
HSTU decode: history + [candidate_1 ... candidate_M]
βββββββββ one parallel forward pass βββββββββ
β [score_1 ... score_M] β Top-K
LLM decode: prompt β token_1 β token_2 β β¦ β token_T
Each step determines or samples only the next token
5. Configuration Consistency Check
The top-level vocab_size is 1,001,024, while the two explicitly configured embedding-table vocabulary sizes sum to 1024 + 100000 = 101,024. If vocab_size represents the actual number of rows in a merged embedding table, verify whether the additional 900,000 IDs correspond to reserved IDs, dynamic-embedding capacity, or another offset. If they have no additional purpose, this field should be checked before loading real weights and tokenizer/ID mappings to avoid embedding out-of-range errors or unnecessary memory allocation.
- Downloads last month
- 3