《LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation》
我们提出了 LLaTTE(LLM-Style Latent Transformers for Temporal Events),一种用于 production ads recommendation 的 scalable transformer 架构。通过系统性实验,我们证明了推荐系统中的 sequence modeling 遵循与 LLM 类似的 predictable power-law scaling。至关重要的是,我们发现 semantic features 改变了 scaling curve:它们是 scaling 的一个前提条件(prerequisite),使模型能够有效利用 deeper and longer architectures 的容量。为了在严格的latency constraints 下实现 continued scaling 的收益,我们引入了一种两阶段架构,将 large, long-context models 的heavy computation 卸载到异步的 upstream user model。我们证明了 upstream improvements 可预测地迁移到 downstream ranking tasks。作为 Meta 部署的 largest user model,这种 multi-stage framework 在 Facebook Feed 和 Reels 上带来了 4.3% 的 conversion 提升,且 serving 开销极小,为在工业推荐系统中利用 scaling laws 建立了一个实用的蓝图。
sequence modeling 和 recommendation systems的融合为人工智能开辟了新的前沿。现代推荐系统,特别是在广告和内容平台中,每天处理数十亿次 user interactions,因为每个用户在产品、广告和内容中的旅程形成了一个丰富的时间序列,编码了 user preferences 和 user intent。虽然大型语言模型(Large Language Models: LLMs)已经证明 deep sequence models 可以通过 predictable scaling laws(《Scaling Laws for Neural Language Models》)实现卓越的性能(能力随着模型大小、数据量和计算量的 power-law functions 而提高),但 production 推荐系统在很大程度上仍然局限于较浅的 sequence architectures (《Self-Attentive Sequential Recommendation》、《BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer》、《Behavior Sequence Transformer for E-commerce Recommendation in Alibaba》)。
尽管 user behavior的顺序性(sequential nature)和 the potential for similar scaling benefits,这种 gap 依然存在。挑战是双重的:
(i):将 deep sequence models 的计算需求与 production serving 的严格延迟要求相协调,在这些系统中必须在毫秒内对数百个 candidates 进行排序。
(ii):弥合传统的基于 Factorization Machin: FM 的模型(《Wukong: Towards a Scaling Law for Large-Scale Recommendation》)和基于 Transformer 的序列模型之间的架构鸿沟。
FM-based models 擅长从大量 sparse and dense feature spaces(user IDs, item IDs, request and contextual features)中学习,但缺乏 sequential modeling 能力。
而 sequence models 捕获时间动态(temporal dynamics),但难以有效整合对 production 性能至关重要的高维 non-sequence features。
如何在 production 约束下有效 scale sequence models 同时保留两种范式的优势,并在这个 hybrid regime 中利用 predictable scaling laws,仍然是一个连接 research 和实际部署的挑战。
为了解决这些挑战,我们将研究围绕三个研究问题展开:
RQ1:我们如何设计集成了 FM-based architectures 的 sequence models,以同时利用 sparse collaborative signals 和 temporal dynamics?
RQ2:推荐系统是否表现出与 LLM 类似的 predictable scaling behaviors,具有可识别的 trade-offs、capacity bottlenecks 和 data quality impacts ?
RQ3:我们如何在满足严格的 production 延迟要求的同时利用 scaling benefits ?
我们引入了 LLaTTE(LLM-Style Latent Transformers for Temporal Events),一种在我们大规模 production 系统中部署的架构范式,它解决了所有三个问题。LLaTTE 通过两个关键创新实现了高效的 scaling,这两个创新将 deep sequence modeling 与我们的 production constraints 相协调:
Target-Aware Adaptive Transformer:我们的 sequence module 通过利用 Multihead Latent Attention: MLA(《DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model》)将 non-sequence sparse features 和 candidate information 整合到 extended query tokens 中。该架构支持 adaptive pyramidal output extraction,这可以选择性地用于渐进式地降低计算复杂度。这种设计与 FM-style modules 无缝集成,使我们能够在严格的 production inference budgets 内利用 deep sequence processing 的 scaling gains。
Multi-Stage Architecture:为了在严格的延迟约束下获取 scaling benefits,我们引入了一个 upstream 阶段,该阶段运行大型 LLaTTE encoders,由 high-value user events 来触发,从而异步生成和缓存 compressed user embeddings (这个 compressed user embeddings 不受 request-time latency 的限制)。online ranking 阶段将这些 cached embeddings 与由 a lightweight LLaTTE counterpart 处理的 fresh, short-horizon sequential signals 相结合。两个阶段共享共同的 sequence architecture,但在截然不同的规模上运行。具体来说,upstream model 消耗 FLOPs 大于 45 倍的 the sequence FLOPs of its online counterpart 。这种不对称性使我们能够将绝大部分 sequence computation 卸载,而不影响 online serving latency。
通过涵盖模型 depth、模型 width 、序列长度、data enrichment via content understanding model features、以及 cross-stage transfer dynamics 的系统性实验,我们对推荐系统环境中的 foundational scaling behavior questions 提供了全面的答案,揭示了几个关键见解:
(i):性能遵循 predictable log-linear scaling laws,sequence length 作为一个主要的杠杆(a primary lever)。
(ii):data enrichment(如 semantic content embeddings)不仅仅是 additive 的,而是更陡峭 scaling curves 的先决条件,有效地将 scaling laws 弯曲到 sparse ID signals 所允许的范围之外。
(iii):upstream modeling 的改进可预测地迁移到 online ranking 阶段,具有高迁移率(约 50%),即使在严格的 information bottleneck 下也验证了我们 multi-stage design 的有效性。
(iv):模型 width 扮演了容量瓶颈(capacity bottleneck);在 depth scaling 变得有效之前,必须建立足够的 width 。
这些发现表明,当架构约束(architectural constraints)得到适当解决时,推荐系统可以从类似于 LLM 的 scaling laws 中受益。
我们的 production 部署验证了这些发现,证明两阶段架构在我们主要的 revenue-generating models 上实现了 0.25% 的 Normalized Entropy 改进,且 serving 开销极小。总之,我们的架构创新和 empirical scaling laws 为在大规模 production 推荐系统中 scaling sequence models 提供了一个系统框架。
我们的工作涉及三个研究领域:推荐系统中 transformer-based sequence modeling、神经网络的 scaling laws 、以及用于 production 部署的多阶段架构。我们将回顾每个领域并定位我们的贡献。
Sequence Modeling in Recommendation:
早期的序列推荐方法使用循环神经网络(RNN)来建模 sequential behavior。
GRU4Rec(《Session-based Recommendations with Recurrent Neural Networks》)开创了 RNN 用于 session-based recommendation。
随后是各种 LSTM 变体(《Sequential User-based Recurrent Neural Network Recommendations》)。
self-attention 机制的引入标志着该领域的重大进步。
SASRec(《Self-Attentive Sequential Recommendation》)和 HLLM(《HLLM: Enhancing Sequential Recommendations via Hierarchical Large Language Models for Item and User Modeling》)将 self-attention 应用于 next-item prediction,实现了并行处理和更好的 long-range modeling。
BERT4Rec(《BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer》)采用 masked item prediction 适配了双向 Transformer,而 S3-Rec(《S3-Rec: Self-Supervised Learning for Sequential Recommendation with Mutual Information Maximization》)为序列推荐引入了 self-supervised pretraining。
Target-aware attention 机制对于 production CTR prediction 也至关重要。
DIN(《Deep Interest Network for Click-Through Rate Prediction》)引入了由 target item 加权的 attention.
随后由 DIEN(《Deep Interest Evolution Network for Click-Through Rate Prediction》)通过 interest evolution modeling 进行了扩展。
近期工作探索了整合 sequential features 和 non-sequential features 的不同架构选择。
HSTU(《Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations》)、OneTrans(《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》)和 LONGER(《LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders》)主张采用 pure sequence-based architectures 的架构。
而 InterFormer(《InterFormer: Effective Heterogeneous Interaction Learning for Click-Through Rate Prediction》)则提出在每一层交错 sequence computation 和 non-sequence computation。
然而,这些方法需要与特定架构紧密耦合,限制了 independent scaling of sequence modules 的灵活性,并阻碍了跨在线和离线阶段轻松分离 model inference 。这种架构刚性使得 production 系统只能使用相对较小的 sequence models 处理有限的上下文。
Scaling Laws:《Scaling Laws for Neural Language Models》 证明了 language model 性能与模型大小、数据集大小和计算预算遵循 power-law 关系。虽然近期工作已在推荐系统中探索了 scaling(《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》、《Wukong: Towards a Scaling Law for Large-Scale Recommendation》),但这些研究通常要么单独考察各个 scaling 维度,要么考察它们的复合效应,而没有系统研究不同 scaling axes 如何交互。此外,它们忽视了 data richness 和 data composition 等关键维度,这些对于大规模 production 推荐系统尤为重要。我们通过建模 depth 、 width 和 feature richness 的联合交互(joint interaction)来解决这个问题,识别 scaling behavior 发生根本性转变的关键阈值。
Multi-Stage Architectures in Production RecSys:Multi-stage architectures 将 representation learning 与 task-specific prediction 分离,已成为 production 推荐系统中建模的一种实用范式。这些方法旨在平衡模型表达力与 online serving 的严格延迟要求。
Embedding-based 的方法预先计算 user representations,供 online ranking models 使用。
Pinnerformer(《PinnerFormer: Sequence Modeling for User Representation at Pinterest》)从 sequential engagement data 中学习 user embeddings。这个 user embedding 是使用一个 transformer,在 engagement windows 上的 dense all-action loss 训练而来。
SUM(《Scaling User Modeling: Large-scale Online User Representations for Ads Personalization in Meta》)提供了一个框架,通过异步 inference 在数百个 production models 之间共享 user representations。
这两种系统都将上游计算与 online serving 解耦,使得能够使用比实时 latency budgets 所允许的更大的模型。
TransAct(《TransAct: Transformer-based Realtime User Action Model for Recommendation at Pinterest》)引入了一种混合架构,将 precomputed embeddings 与 real-time processing of recent user actions 相结合。该系统通过 Transformer 处理 recent actions(precomputed embeddings 来捕获长期兴趣。
TransAct V2(《TransAct V2: Lifelong User Action Sequence Modeling on Pinterest Recommendation》)使用 candidate-anchored nearest neighbor search 和 custom GPU kernels,将 real-time component 扩展到更长的序列(
另一种方法使用 retrieval-based designs ,这种方法在每个 request 内同步执行。
SIM(《Search-based User Interest Modeling with Lifelong Sequential Behavior Data for Click-Through Rate Prediction》)引入了 GSU-ESU 范式,其中 General Search Unit: GSU 首先检索 relevant behaviors,然后 Exact Search Unit: ESU 单元应用 target-aware attention。
TWIN(《TWIN: TWo-stage Interest Network for Lifelong User Behavior Modeling in CTR Prediction at Kuaishou》)通过确保两个阶段使用相同的 relevance metrics 来改进 stages 之间的一致性,在 Kuaishou 扩展到 behaviors 。
TWIN V2(《TWIN V2: Scaling Ultra-Long User Behavior Sequence Modeling for Enhanced CTR Prediction at Kuaishou》)通过 hierarchical clustering compression 扩展到 lifecycle-scale sequences。
尽管有这些进展,先前的 multi-stage 工作尚未探索遵循 predictable scaling laws 的系统性 model scaling。这些系统通常沿单一维度 scaling(例如 TransAct V2 中的 sequence length),同时整体保持模型相对较小。在这项工作中,我们展示了一个统一的 scaling 框架同时适用于 real-time and asynchronous upstream ranking settings。我们系统研究了 increasing model capacity across multiple dimensions 、以及 allocating capacity between online and upstream stages 是如何影响 production 推荐系统的性能。据我们所知,这是第一个为这种 hybrid upstream-downstream architectures 建立 empirical scaling laws 的工作。此外,我们严格量化了 stages 之间的 transfer ratio,提供了一个度量标准来衡量 scaling laws 如何有效地穿越 inference bottleneck。
我们考虑一个标准的 multi-task ads ranking problem,其中模型预测 engagement 概率,例如点击率(click-through rate: CTR)和转化率(conversion rate: CVR)。这些 calibrated probabilities 用于下游 ranking 和 auction pricing。
Task:给定用户 candidate ad request context a vector of engagement probabilities:
其中:prediction heads 集合。
Feature Structure:我们沿两个轴对 input features 进行分类。
按特征来源,我们区分:
(1):用户特征
(2):广告特征
(3):user-ad 交叉特征 historical interactions)。
(4):上下文特征 surface)。
按特征类型,特征被处理为:
Sparse IDs categorical entities 被映射到 embeddings 。
Dense features precomputed embeddings(例如,来自 content encoders)。
Float attributes continuous variables)。
Sequences ordered lists of temporal events)。
User Sequences:User behavioral sequence list of actions。每个 action action type、item ID、surface type、可选的 content embeddings 和其他 metadata。在我们的实验中,序列长度 500 到 5000 个 actions 不等。
Objective and Evaluation:我们使用加权的 multi-task binary cross-entropy loss 进行训练。我们的主要评估指标是归一化熵(Normalized Entropy: NE)的相对改进。NE 定义为 average log loss 除以 the entropy of the empirical CTR:
其中: ground truth 观察到的经验正例率(empirical positive rate)。
NE 是工业推荐任务中的标准指标。
为了系统研究 recommendation 中的 scaling laws,我们需要一个既代表现代 production 系统、又足够高效以 scale 到 massive contexts 的架构。我们采用模块化设计,称为 LLaTTE(LLM-Style Latent Transformers for Temporal Events),它综合了 efficient sequence modeling 的最佳实践。
该架构(Figure 1)由三个组件组成:
(i):编码 user behavioral histories 的 Sequence Module。
(ii):将 sequence summaries 与 static features 整合的 Non-Sequence Module。
(iii):用于预测的Task Heads 。
注意:
Upstream User Model生成的user representation,被用作于Online Ranking Model的Dense Features。

DHEN 的架构如下:

Compute Allocation:在我们的 production setting 中,sequence module 在 large-scale upstream model 中占 total inference FLOPs 的 90% 以上,而 lightweight sequence module counterpart 约占 online ranking model FLOPs 的 30 %。在后续 scaling 实验中,我们默认将研究重点放在 large-scale upstream model 上,其中改变 Sequence Module 的容量 lightweight online ranking model 可能实现的更广泛的 compute FLOP range 内探索 scaling behaviors。 Non-Sequence backbone和 Task Heads 在整个过程中保持固定,以隔离 Sequence Module scaling 的影响。
为模型 layer num,代表模型depth;为向量维度,代表模型 width;为时间序列长度。
non-sequence module 实现了标准的 DHEN-style 架构(《DHEN: A Deep and Hierarchical Ensemble Network for Large-Scale Click-Through Rate Prediction》),这是一种结合 heterogeneous feature interaction modules 的 deep ensemble network。它处理所有 non-sequential features 、以及由 Transformer 产生的 sequence summaries Sparse categorical features 被嵌入并与 dense/float features 拼接起来。concatenated features 通过 feature interaction networks,产生 unified representation Prediction heads 是 shallow MLPs,产生概率:
这里的
是 Transformer Layer最后一层的输出。在最后一层,它输出( query token num)个representations。
sequence module 是我们 scaling study 的焦点。为了在 production 约束下 scaling 到长上下文(lengths > 1000),我们采用了一种 Target-Aware transformer,具有两个关键的 efficiency optimizations:多头潜在注意力(Multi-head Latent Attention: MLA)和自适应金字塔形输出(Adaptive Pyramidal Output)。
该模块分五步操作:
1) Tokenization:将每个 action a token vector
2) Fusion:append query tokens 以形成
我们为 online model 编码 the query tokens with user-ad candidate context
我们为 upstream model 编码 the query tokens with user-only context ad candidate context 尚不可用。
这种灵活的编码使得同一架构能够桥接 offline 阶段和 online 阶段。
3) Latent Attention:使用 MLA(《DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model》)应用到 Transformer 以减少内存占用。
参考附录部分。
4) Pyramidal Reduction:在 deeper layers 选择性地移除 older tokens(《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》)。这将计算集中在recent events 和 query tokens上,允许模型在 context length 和 FLOPs 之间进行权衡。
5) Readout:从 the final query-token representations 中抽取 fixed-size summaries。
这种自适应设计(adaptive design)允许我们在整个技术栈中部署不同风格的 LLaTTE,同时保持统一的建模范式。
在异步的 upstream阶段,我们部署 high-capacity 的变体,其消耗的 sequence FLOPs 是 main ranker 的 90% 集中在 Sequence Module),利用 full self-attention 最大化 representation quality。
相反,online ranking models 采用激进的金字塔修剪(pyramidal trimming)以满足严格的延迟预算。
该架构由标准化组件组成;完整规格见附录 A。
为了严格量化推荐中 scaling 的好处,我们超越了临时的 architecture search,采用系统性的 scaling 框架。与 LLM(其中 input text distribution 通常被视为固定)不同,推荐系统允许我们沿 architectural dimensions 和 input information densities 两方面进行 scale。
我们将 scaling behavior 表述 compute budget sequence module 的 compute budget 并以 FLOPs 来衡量。当 information density 保持恒定时,我们假设归一化熵(Normalized Entropy: NE)的改进(相对于我们强大的 production baseline)与计算量遵循幂律关系(power-law relationship),建模为对数线性(log-linearly):
其中: scaling coefficient(效率)。
我们研究四个主要的 scaling 的轴:
1) Model Capacity Transformer depth embedding width objective 是:确定最大化 parameter efficiency 的纵横比 sparse-feature regime 中,在 depth scaling 变得有效之前是否存在关键的 width thresholds。
2) Temporal Horizon information horizon)的直接代理。我们考察延长 history events
3) Information Density:标准 scaling laws 假设同质的 token distribution。我们通过改变 tokens semantic richness)来挑战这一点。我们将 sparse ID-based tokens 与 tokens enriched with dense semantic embeddings from foundational content encoders 进行比较,将 signal quality 视为缩放系数
4) Cross-Stage Transfer Efficiency:为了量化跨阶段效率(cross-stage efficiency ),我们引入了迁移率(Transfer Ratio) large up-stream model 的改进如何转化为 production online ranking loss 的改善。该指标考虑了所有 production 干扰因素,如异步 inference latency 惩罚。
其中:
offline user modeling task 中的 loss reduction。
online ranking task 中的 loss reduction。
该指标使我们能够确定哪些 scaling axes 最能通过多阶段架构中固有的 information bottleneck 而保留 gains。
Training infrastructure:所有 scaling 实验都在同一个大规模 production recommendation task 上进行,使用专有的工业训练平台。如前所述,我们的实验在大型 upstream model 上进行,以研究最大可达到的 FLOP horizon,超越 online model 的计算极限。模型使用混合精度在 128 NVIDIA H100 GPUs 上训练,并使用 FlashAttention(《FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness》)以提高内存和吞吐量效率。除非另有说明,每个 configuration 在采样自 production traffic distribution 的 30 billion examples 上训练 229K steps。
Evaluation metric:我们使用归一化熵(Normalized Entropy: NE)报告所有结果。我们报告相对于 strong production baseline 的 NE 相对降低百分比(%),较低的 NE 值表示更好的性能。在我们的内部数据集上,NE 降低 0.02% 被认为具有统计显著性,并足以产生可衡量的收入影响。
我们首先研究如何在模型 depth width a fixed compute budget。我们在固定序列长度 grid search。
结果揭示了 depth 和 width 之间的非线性关系:
当模型较窄时(depth 会产生递减的收益;representation capacity 受限于 embedding 维度。
然而,一旦 width 达到临界阈值 depth scaling 变得显著更有效。
相反,将计算纯粹分配给 width (固定 NE gains。
为了说明这些 regimes,Table 1 比较了四种代表性配置。我们包含了一个 "Deep-Balanced" 变体以展示当 depth 和 width 都足够时的 scaling behavior。
极端或不平衡地分配给 width 或 depth 都会导致递减收益。
足够的 width (depth scaling 有效的前提,而过大的 width 而 depth 不足(例如
Balanced configurations 最大化性能和计算效率。
受这些观察的启发,我们在研究其他 scaling axes(如 sequence length 和 content features)时,默认采用 balanced configurations(

我们现在研究 LLaTTE 性能如何随时间视界(temporal horizon)(它与 events 数量 scale、以及如何随着每个 event 上的 semantic content embeddings 而 scale。
Extending the temporal horizon:我们在固定 width 的条件下,对 depth sequence horizon Figure 2 所示,在所有 depth 上,NE 随 user histories 持续提高 ranking quality。
additional context 的 incremental benefit 在很大程度上取决于模型容量。Deeper models 在 NE gains,这反映在 Figure 2 中

Attention-score probability analyses 也揭示了 long contexts 的效用。The cumulative attention probability distributions over positions (Figure 3)表明,a non-trivial fraction of attention probability 分配在整个 200-1600 token range 内,而不是狭隘地聚焦于 most recent events,这表明中期历史和长期历史对 predictions 具备有意义的贡献。这项研究的细节见附录(见 A.4 节)。

Balancing freshness and signal strength:我们接下来研究固定长度下序列的 composition。在这种设置中,user histories 结合了 high-frequency, lower per-event signal actions(ad views)和 low-frequency, high-value actions(conversions)。我们固定 sources 之间的分配(Table 2,上半部分)。
views and conversions 的 balanced mixture 表现最佳。
两种极端情况都明显次优。
Pure-view sequences表现不佳,因为它们包含许多 low-signal events。
low-signal event:即这些事件包含了很少的关于用户兴趣的信号,例如ad views。
Pure-conversion sequences 虽然由非常高价值的 events 组成,但性能也会下降,因为 conversions 在时间上是稀疏的: filling a length-1000 sequence with only conversions 迫使模型依赖 older events,而这些 older events 可能不再反映 current intent 。相比之下,Views 提供了 dense and recent coverage。
表现最佳的 allocations 结合了 the high per-event signal of conversions 和 the temporal freshness provided by frequent views。

Content-aware scaling: strong content features bend the scaling curve:最后,我们研究 feature richness 如何与 architectural scaling 进行交互。除了 sparse IDs 外,我们还引入了由 fine-tuned LLaMA models 以及处理 heterogeneous multimodal signals(如文本和图像)的 Content Understanding models 所生成的 dense content embeddings,并在不同 depths 比较了 models with and without these features(Table 2,下半部分)。
有两个观察结果:
首先,在没有 content features 的情况下,将 depth 从一层增加到四层仅产生微小的改进:the 4-layer ID-only model 仅略好于 1-layer baseline。这对应于非常平缓的 scaling 斜率,表明 additional capacity 主要用于 memorizing ID patterns。
其次,当启用 LLaMA content embeddings 时,相同的 4-layer configuration 实现了显著更大的增益,并且 the relative benefit of content 对于更 deeper 的模型明显强于较浅的模型。
这些结果表明,strong semantic content features 不是可以忽略的附加物(a marginal add-on),而是 effective scaling 的前提条件。The presence of semantic features 实质上改变了 compute–NE curve 的斜率:
with ID-only inputs,depth and sequence-length scaling 很快出现递减收益。
而 with content-enriched sequences ,相同 increases 会转化为显著更大的 NE gains。
从这个意义上说,LLaTTE 的 scaling law 本质上是 content-aware 的;孤立地分析模型大小和计算量不足以预测性能。
最后,我们汇总上述实验以考察 global compute scaling。Figure 4 在对数尺度上绘制了 relative NE gain 与 sequence compute budget sequence FLOPs 衡量)的关系。
结果表明,recommendation 性能与计算量遵循可预测的 power-law 关系。我们对经验数据拟合了一个对数线性模型(log-linear model):
其中 scaling 系数 scaling 策略决定。我们观察到每个轴的不同行为:
Sequence Length:该维度表现出最陡的 scaling 斜率。延长 temporal horizon 始终产生每单位计算量(per unit of compute)最大的性能改进,表明模型有效利用了 long-term history。然而,这伴随着 logging and materializing longer user engagement histories 的额外成本。
Model Depth:增加 depth 提供了稳健的 scaling behavior。一旦模型 width 被充分 scale up,增加层数为提高 ranking quality 提供了可靠路径。
Model Width:width 在架构中扮演基础性角色(foundational role)。虽然 scaling 斜率相对于 depth 或 length 较平缓,但足够的 width 对于最大化其他维度的容量至关重要。
Content Quality:如 Figure 4 所示,semantic features 的存在显著增强了其他所有维度的 scaling laws。 Content-enriched models 相比 their ID-only counterparts(黑色线)表现出更陡的斜率,表明 high-quality semantic features 是模型从 increased compute 中 extract value 所必需的。
这些发现确立了 sequence-based recommendation 是一种 scalable 的范式,能将计算转化为可预测的 performance gains。我们在下一个章节中利用这个 scaling hierarchy,在 production 约束下优化我们的 multi-stage architecture。

第 1.5.4 节表明,随着我们 scale depth,特别是 scale sequence length,ranking quality 持续提高。然而,production ranker 每天在严格的 latency budget 下服务数万亿次 requests,这限制了 deployed sequence module 只有少数几层,并且每个 source(即 engagement type)的 horizon 大约为 events/tokens。为了在不违反 latency 约束的情况下从有利的 scaling curves 中受益,我们采用了一个 multistage 架构。
The latency-constrained downstream model 保持紧凑,并联合关注 user sequences features 和 ad/context features。
An asynchronous upstream user model 使用相同的 LLaTTE 架构处理更长的 histories,但仅使用 user-side features,并发布 cached user embeddings,ranker 在 request 时刻消费这些 embeddings。
这种 separation 引入了一个严格的 information bottleneck,因为 upstream encoder 必须将数千个 historical events 总结为 a single fixed-size vector(在我们的生产部署中,fixed-bandwidth bottleneck 如何影响 scaling laws。
我们首先单独描述 the upstream user model 的 scaling behavior。使用与第 1.5.4 节相同的 log-linear 拟合程序,我们将估计到的斜率作为 NE improvement,它是 upstream sequence FLOPs 的函数。Table 3 (top)总结了 upstream 斜率和 downstream 斜率之间的比较。
Table 3a 报告了针对 total model FLOPs 的 scaling slopes。downstream model 在所有维度上都表现出显著更陡的斜率。然而,这种差异反映了一种结构不对称:在 downstream ranker 中,sequence module 仅占 total FLOPs 的约 30%;而在 upstream model 中,sequence module 占 total FLOPs 的约 90%。因此,在 downstream setting 中,Scaling the sequence module 产生更大的 relative improvement per total FLOP。
为了隔离 sequence module 的内在 scaling efficiency,Table 3b 报告了针对 sequence-only FLOPs的斜率。这种比较揭示了 depth and width scaling efficiency 在两个设置中是一致的。关键区别在于 sequence length:upstream scaling 仅保留了约 50% 的 downstream efficiency。这种差距产生的原因是:upstream encoder 缺少 candidate context,迫使upstream encoder 将 user history 压缩为 a generic representation,而不是执行 candidate-aware attention。因此,只有与 candidates 相关的 long-range signals 能够穿过瓶颈。

为了量化我们多阶段架构的效率,回顾迁移率 upstream representation quality 转化为 downstream ranking accuracy 的转化率。
例如,Table 3c 中的 Seq-heavy configuration 展示了这种关系:升级 upstream encoder 产生 0.14% 的 upstream improvement,这转化为 downstream ranker 0.07% 的改进,对应迁移率
Benchmarking Transfer Efficiency:将收益从 asynchronous upstream models 迁移到 online rankers 本质上是损耗性的(lossy)。几个结构性因素通常将迁移效率限制在 25% - 30% 范围内:
massive upstream encoders 与 compact downstream rankers 之间的容量差距(capacity gap)。
asynchronous inference 引入的陈旧性(staleness)。
以及将 user histories 压缩为 fixed-size user embeddings 所带来的信息瓶颈。
在此背景下,我们实验中观察到的约 50% 的迁移率是显著的。它表明通过 scaling the sequence encoder 所捕获的 high-level intent signals 对 compression 和 temporal delays 具有鲁棒性,尽管存在 aggressive compression 和 separation from the online scoring context,仍保留了实质性的 predictive value。
Architecture Robustness:为了确定 specific architectural choices 是否影响这种效率,我们比较了两种具有 matched compute budgets(约 12 GFLOPs/sample)的 upstream configurations,代表不同的 scaling 优先级:
Sequence-heavy:
Model-heavy:
如 Table 3 (bottom) 所示,两种模型实现了类似的 upstream gains(-0.14% vs -0.13%),并且关键的是,相同的 downstream gains(-0.07%)。这表明在固定的 compute budget 下,架构对于 depth 和 sequence length 之间的分配是鲁棒的。upstream model 的总计算量决定了 downstream 性能,验证了我们可以在 history length 和 model capacity 之间灵活权衡,而不降低迁移率。
结合 downstream and upstream observations ,这产生一个简单策略:
Downstream online ranker:将 latency budget 分配给 short sequence lengths and shallow depths,辅以足够的 width 和 rich content features( sequence lengths 最高约
Asynchronous upstream user model:利用宽松的 latency budget 来 scale total sequence-modeling compute 。结果表明,downstream 性能主要由 total upstream compute budget 驱动,而非 specific allocation between depth and length。因此,我们部署 upstream models,其中 upstream models 的 sequence compute budget 是 online ranker 的
这种 two-stage asymmetric strategy 使我们在 latency-constrained ranker 用尽 local budget 后,仍能继续遵循第 1.5.4 节中确定的 scaling frontier,将 additional offline compute 转化为 predictable online gains。
在生产中,我们部署了 two-stage LLaTTE 架构。
downstream ranking model 继续在线运行 compact LLaTTE module,在严格的 per-request latency budget 下,在万亿级请求规模上关注 short user histories(per-source horizon 上限约为 ad and context features。这个 online LLaTTE model 捕获 fresh user intent signals 并利用 ad-specific interactions,但有意比第 1.5.4 节 scaling study 中考察的模型小得多。
为了在不影响延迟的情况下利用更大的 sequence models,我们引入了一个 asynchronous embedding service,该 service 托管 upstream user understanding LLaTTE model。Embedding updates 不是 computed per request;而是由 high-value user events(主要是 user conversions)触发的。Triggered events 在 dedicated cluster of H100 GPU hosts上处理,使用针对 high-throughput transformer inference 优化好的 deeper and longer LLaTTE models。作为 Meta 部署的最大 user model,这些 upstream encoders 仅使用 user-side features,生成 user representations (embeddings),并将其 compressed representations 写入 a feature store。在 request 时刻, downstream ranking models 将这些 embeddings 作为 dense features 读取,并在 online downstream LLaTTE ranking model 中使用。
将 heavy sequence modeling 卸载到 asynchronous embedding service,这使 ranking path上的 additional cost 保持为a single feature lookup。online downstream LLaTTE model 在设计上保持保守( depth 和 sequence length)。与没有 upstream LLaTTE user embeddings 的 baseline 相比,我们观察到 P99 ranking latency 没有 measurable change。来自 larger upstream LLaTTE models 的 extra FLOPs 由 H100 cluster 吸收,该 cluster 以低得多的 QPS 运行且 heavily batched。
在多个大规模 A/B tests 中,the combination of the compact online LLaTTE and the upstream LLaTTE user embeddings 在我们的旗舰 ads ranker 上实现了约 0.25% 的 NE reduction,对应于 Facebook Feed 和 Reels 上 4.3% 的 conversion 提升,年收入的业务影响达数亿美元。这验证了所提出的 two-stage scaling 策略是一种在严格 serving 约束下将 additional sequence-modeling compute 转化为 production value 的有效方式。
在这项工作中,我们提出了 LLaTTE 范式,显著 scale up 了推荐系统中的 sequence learning。通过将大部分计算资源分配给 sequence modeling,我们实现了可预测的且稳健的 scaling behavior。我们系统地分析了跨 standard dimensions(model depth and width)以及 less explored factors(sequence length, composition and richness)的 scaling,证明 optimal scaling 需要同时平衡所有这些因素。
为了最大化 production impact,我们开发了一种 multistage modeling and deployment strategy:一个在 user-only features 上训练好的计算昂贵的模型,离线且不频繁地运行,同时有效地将这些 gains 迁移到轻量级的 online production model。
我们的工作与推荐系统 sequence modeling 方面其他显著的行业进展(《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》、《OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment》)是并行开发的。这些努力共同表明了 sequence modeling 在大规模推荐系统中的兴起。展望未来,我们计划探索高效 long-context kernels、reinforcement learning、scalable infrastructure solutions 以及 the upper bounds of scaling laws ——因为我们正站在推荐系统 LLM-scale 时代的起点。
令 embedding lookup for sparse categorical features。我们通过 concatenating embedded sparse features, dense features, float features, and sequence summaries 来构建 initial non-sequence representation:
注意,这里
是拼接起来的,而不是经过池化的。
然后我们应用 a feature-interaction network:
得到:
在 production,DCN/DeepFM/DLRM 风格架构来实例化。
Action Embeddings:令 ordered sequence of actions,最早的在前。每个 action token:
形成:
我们对 a subset of hidden dimensions 应用 additive timestamp encodings。
在
Transformer中,每个token(这里是一个用户历史动作)会被表示为一个向量。向量的每一个分量就是一个hidden dimension。在这里,a subset of hidden dimensions指的是:不把时间戳编码加到整个维向量上,而是只选择其中一部分维度来承载时间信息。 论文采用这种设计可能有以下原因:
保留原始语义:用户行为
token中已经包含了item ID、内容嵌入等丰富语义,如果对整个向量加时间编码,可能会干扰这些语义。只对子集维度加时间信息,可以让其余维度专注于内容表示。参数效率:时间编码只需要覆盖子集维度,减少了参数量和计算量。
解耦时间与内容:让模型在部分维度上专门感知时间,便于学习时间模式(如周期性、衰减等)。
Query Tokens and Fusion:我们引入 query tokens summarize the user–ad request。根据阶段不同,这些 tokens 可以编码 candidate ad features, request context, user-level features, and learned seed tokens(《Perceiver: General Perception with Iterative Attention》)。我们将 sequence 和 query tokens 拼接起来:
并将 Transformer。
Transformer Layers and Pyramidal Schedule:我们应用具有 Multi-head Latent Attention: MLA (《DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model》)和 RMS 归一化(《Root Mean Square Layer Normalization》)的 L-layer causal transformer 。令
遵循 《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》,我们在时间维度上使用自适应金字塔调度(adaptive pyramidal schedule)。如果
其中 older tokens,减少 attention cost FFN cost
我们使用三种模式:
Full self-attention (all sequence and query tokens。
Pyramidal attention(query tokens 加上 most recent actions。
Cross-attention (query tokens;在最后一层使用。
Offline (upstream) models 通常使用更多层,这些层采用 full self-attention(除了 final cross-attention layer)。 Online ranking models 在 final cross-attention layer 之前,在 early layers 使用更激进的 pyramidal trimming。
Sequence Summaries:final layer 在 query-token positions 输出 fixed-size sequence summaries :
其中每个 low-rank adapted MLP 。
为完整起见,我们包含 MLA parameterization。对于 heads,给定输入
一个关键观察是:连续的线性投影可以在代数上吸收到单一矩阵中。具体来说:
这种简化揭示 MLA 在数学上等价于 latent space 上的 Multi-Query Attention: MQA (Shazeer,2019),其中 queries 被投影到多个 heads ,而 key-value representations 在 compressed latent space 中保持共享。重写的方程如下:
.
我们研究了 attention weight distribution
为了增强 visualization of attention weights,我们在 a single-layer transformer LLaTTE backbone 中同时减少了 the number of query tokens 和 the number of attention heads。我们提取了每个 query token 相对于整个 user sequence 的 attention probabilities,并在来自不同 user requests 的足够多样本上聚合了这些概率。
我们的分析表明,模型将更大比例的注意力分配给 most recent events。然而,它继续对 longer-term history 分配 non-negligible probabilities,表明模型不会完全忽略 earlier events(Figure 3)。有趣的是,我们还考察了通过 selecting topK event tokens 而不是 most recent k tokens 的 cumulative attention weight,这代表了通过 selecting k tokens 可达到的理论最大 cumulative attention weight。

此外,我们在 daily level 识别出显著的季节模式。如 Figure 5 和 Figure 6 所示,attention weight 峰值出现在 24-hour intervals 的尖锐处。这表明个体用户倾向于在一天中的相似时间表现出重复的 interests 和 behaviors。
虽然这些发现与我们的直觉一致,但我们认识到对 user behavior 与 behavioral sequences 之间更细粒度的理解仍然是一个开放的探索领域。我们呈现这些初步结果,以鼓励更广泛社区的进一步研究和贡献。

