一、 TokenMinds [2026]

《TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems》

  1. 工业推荐系统中的 user modeling 通常生成 dense embeddings,这些 embeddings 受限于 representational constraints:它们固有地采用固定维度的向量。一种新兴的 discrete user representation 替代方法——使用 LLM 生成 text-based user tokens ——捕获的是主题共现(topical co-occurrences)而非 deep sequential behavior dynamics,并且产生的 outputs 难以与 item attributes 关联。同时,基于 Semantic ID: SID 的 item tokenization 已被证明在 generative recommendation 中能有效提升泛化能力,然而 discrete SID-based 的 user representations 在很大程度上仍未得到探索。我们提出 TokenMinds,一个工业规模的系统,它将 PLUM 框架(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)从 item retrieval 扩展到 user modeling,通过从 pre-trained LLMs 改编的一个 encoder-decoder 架构,同时生成 discrete SID-based user tokens 和 dense user embeddings。这种 dual-output 的设计提供了离散的、语义上有据可依的 user representations 的互补优势,同时保持了与现有下游模型(这些模型依赖 dense embeddings)的兼容性。此外,shared SID vocabulary 自然地扩展到跨场景建模(cross-scenario modeling):通过将长视频行为和短视频行为统一到单个模型中,我们显著降低了 training 成本和 serving 成本。我们通过广泛的离线实验和在多个 YouTube surfaces 的线上 launches 来验证 TokenMinds,通过一种异步 infrastructure 来服务全量用户流量(数十亿用户),该 infrastructure 将 representation generation 与下游 scoring 进行解耦。以 ranking 作为主要的下游用例,我们的结果证实了 SID-based user tokens 在工业规模上的实际可行性,并证明 tokens 和 dense embeddings 在不同 production ranking systems 中提供了互补价值。

  2. 推荐系统已经从早期的因子分解机(factorization machines)发展到现代深度学习架构,并且越来越多地从多级级联架构(multi-stage cascaded architectures)转向 end-to-end 方法。sequential user modeling 侧重于利用用户交互过的 history,长期以来在实现推荐系统个性化方面发挥了关键作用。Large Embedding Models: LEMs 作为主导的工业范式,依赖大规模的 embedding tables 来表示 high-cardinality item IDs。虽然 Large Embedding Models 在记忆 user-item interactions 方面有效,但它们在复杂网络上的泛化能力常常受限(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。

    相反,Large Language Models: LLMs 将其参数预算(parameter budgets)主要分配给神经网络而非大规模 embedding tables,在 recommendation 方面已显示出巨大潜力。近期工作利用它们的 reasoning 和 contextual understanding 能力进行 feature extraction 、Large Embedding Model integration 和 end-to-end sequence modeling。然而,将 LLMs 应用于 domain-specific recommendations 揭示了两个关键瓶颈。

    • 首先,models pretrained on natural language 在处理大规模、非文本的 ID spaces(例如,数十亿视频)时存在根本性的模态鸿沟(modality gap)。

    • 其次,通过将 LLMs 与 traditional large embedding tables 耦合来弥合这一鸿沟,继承了 LLMs 本应超越的那些局限性:scaling 约束、vocabulary 变动和有限的表达能力。

    为了克服这些瓶颈,Semantic IDs tokenize item 为 hierarchical discrete codewords (这些 codewords 源自 content semantics),消除了对 embedding table 的依赖,并通过有意义的碰撞(collisions)提升了泛化能力。在 SID 解决了 traditional representations 的 scaling 和 expressiveness 的局限性之后,PLUM 框架通过 Continued Pre-Training: CPT 和 task-specific post-training,将 new SID vocabulary 与 LLM’s pre-existing knowledge 对齐,从而弥合了 remaining modality gap。

    虽然 PLUM 框架已被证明对 item retrieval 有效,但其在 sequential user modeling: SUM 方面的潜力尚未被探索。在本文中,我们提出了 TokenMinds,它扩展了 PLUM 从而生成 SID-based user tokens 作为 user representations。SID-based representation 与 using dense user embeddings 的常见方法形成对比,后者将 a user’s full spectrum of interests 压缩到一个或几个固定维度的向量中,可能会丢失细粒度的信号。同时,using LLMs to generate text-based user profiles 也日益受到关注。虽然 text profiles 可以被视为 discrete user representation 的另一种形式,但它们面临关键局限性: LLMs without pre-training on user behavior signals 倾向于捕获主题共现(topical co-occurrences)而 deep sequential behavior dynamics(《Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?》),它们的 text outputs 难以与 item attributes in any target domain 关联,并且当集成到 non-textual downstream systems 时会引入 modality gaps (《Llm4msr: An llm-enhanced paradigm for multi-scenario recommendation》)。在 TokenMinds 中,我们通过利用 PLUM 中的 CPT 方法,并将 model outputs 锚定到语义上有意义、且基于 recommendation 语料库的 SID tokens 来克服这些挑战。

    具体来说,TokenMinds 采用 encoder-decoder 架构——这是 generative retrieval 中广泛使用的范式——但将其重新用于一种新颖的 dual-output 设计:decoder 生成 SID-based user tokens,而 encoder 处理 sequential user features 并同时产生 dense user embeddings,确保与使用 dense embeddings 的现有下游模型兼容。以 ranking 作为主要下游用例,我们研究如何将 SID-based user tokens 正确地集成到 production ranking models 中。为了在严格的 latency 约束下为数十亿用户服务该模型,我们在异步 serving infrastructure 之上部署 TokenMinds,该 infrastructure 将繁重的 representation generation 与 real-time scoring 解耦。

    除了单场景 user modeling 之外,由于 TokenMinds 在从 pre-trained LLM 所继承的 shared token space 中运行,它自然地支持两个进一步的扩展。

    • 首先,诸如 textual search queries 之类的 heterogeneous signals 可以轻松地与 watch histories 在 input sequence 中交织起来。

    • 其次,shared SID vocabulary 为长视频(long-form video: LFV)和短视频(short-form video: SFV)之间的跨场景建模提供了天然桥梁,这两个场景具有不相交的 video ID spaces 和不同的消费模式。我们在每个 watch 前添加 scenario-specific condition tokens (LFV/SFV),使单个模型能够在按时间交错排列的 cross-scenario sequences 上进行训练。在推理时,我们引入了 multi-context decoding:encoder 在单次前向传递中产生 a shared user representation,decoder 通过以不同的 scenario prefixes 为条件生成 scenario-specific user tokens ——有效地实现了 scenario-aware output 而无需多模型开销。

    我们总结贡献如下:

    • Dual-Output 架构:我们提出 TokenMinds,一个为大规模工业推荐系统设计的统一的 user modeling 架构。其新颖的 dual-output 机制联合生成标准的 dense continuous embeddings 和 discrete SID-based user tokens。

    • Tokenized User Representations & Adaptation:我们提出了一种 SID-based discrete tokens 的新颖 application 从而用于 user representation,并确定了在下游模型中适配它们的最优策略。虽然 TokenMinds 支持各种下游任务,但本文侧重于 ranking integration,我们在其中展示了 tokenized representations 在工业规模上的实际可行性,并证明了 dense embeddings 和 discrete tokens 可以捕获互补信号,放大了性能提升。

    • Cross-Scenario Modeling :我们引入了一种高效的 modeling 范式,原生地整合了 cross-scenario content training and serving。通过在单个框架内联合建模根本不同的场景( LFV 和 SFV),与维护 separate models 相比,我们将上游训练计算量减少了 50%,serving 计算量减少了 31%,同时保持了 core engagement 质量,并在两个场景中显著改善了 fresh content metrics。

    • Industrial-Scale Deployment:我们通过广泛的离线评估和在 production ranking systems 上的线上 A/B testing 验证了 TokenMinds,在 core user metrics 上实现了高达 +0.11% 的统计显著提升,在 core engagement metrics 上实现了高达 +0.62% 的提升。TokenMinds 已在多个主要 YouTube surfaces 的生产环境中部署,在全量用户流量上服务 LFV 和SFV 推荐。

1.1 相关工作

  1. Sequential User Modeling:

    • sequential recommendation 已经从 RNN(《Session-based Recommendations with Recurrent Neural Networks》)和 CNN(《Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding》)演变为强大的 Transformer-based 的模型,如 SASRec(《Self-Attentive Sequential Recommendation》)和万亿参数 transducers 如 HSTU(《Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations》)。尽管架构复杂,这些模型从根本上依赖于 atomic item IDs 的大规模的 continuous embedding tables。

      最近,Semantic IDs: SIDs 已成为 item representation 的一种有吸引力的替代方案,它能增强泛化能力(《Recommender systems with generative retrieval》、《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)并提高样本效率(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)。

    • 在用户侧,dense continuous embeddings 是工业界 user representation 的主流标准(《Longer: Scaling up long sequence modeling in industrial recommenders》、《Pinnerformer: Sequence modeling for user representation at pinterest》、《Empowering General-purpose User Representation with Full-life Cycle Behavior Modeling》)。

      最近 LLM-based 的方法试图将 preferences 总结为 natural language profiles(《Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5)》、《Idgenrec: Llm-recsys alignment with textual id learning》、《Palr: Personalization aware llms for recommendation》),但这些方法倾向于捕获 topical co-occurrences 而非 user’s sequential dynamics(《Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?》),并且在 non-textual downstream systems 中面临 modality gaps 的挑战(《Llm4msr: An llm-enhanced paradigm for multi-scenario recommendation》)。 discrete, SID-based token 的 user representations 在很大程度上仍未得到探索。

      我们的工作通过将 SID 范式从 items 扩展到 users 来填补这一空白,生成紧凑的 discrete user tokens,同时继续输出 dense embeddings 以与现有下游系统兼容。

  2. Generative Recommendations:

    • generative recommendation 将 traditional embedding-table-based paradigm 重新构建为 sequence-to-sequence 任务,例如 PLUM(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)、GenRank(《Towards large-scale generative ranking》)和 GPR(《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》)。

      我们的工作将 PLUM 框架从 retrieval 扩展到 user modeling。retrieval models 用于预测下一个即时 items,然而 user modeling 在更长的时间窗口内捕获更广泛的 spectrum of intents。它利用更粗粒度的 semantic granularity 来识别 distinct areas of interest,避免了严格映射回specific individual items 的需要。

      与 GPR 不同,后者将 user representation 与下游任务指标和 policy optimization 对齐,我们将 learning of user representations 与 specific downstream training objectives 解耦,以提供对用户兴趣的通用理解。

    • 与此同时,LIGER(《Unifying generative and dense retrieval for sequential recommendation》)和 COBRA(《Sparse meets dense: Unified generative recommendations with cascaded sparse-dense representations》)证明了 sparse semantic IDs 和 dense embeddings 为 retrieval 提供了互补的 item representations。然而,两者都仅在 item level 操作。

      TokenMinds 将这种 sparse-dense 互补性从 items 扩展到 users,通过统一的 encoder-decoder 架构同时生成 discrete SID-based user tokens 和 continuous embeddings。

    • 在线上部署 heavy generative models 面临着现实世界系统的 scaling 瓶颈。近期工作通过优化线上架构本身来解决这一问题,例如 HSTU(《Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations》)的高效 transducers、LONGER(《Longer: Scaling up long sequence modeling in industrial recommenders》)的 optimized representations 、或 parallel generation 方法(《Generating long semantic ids in parallel for recommendation》)。由于这些模型在 real-time scoring 期间执行大量计算,严格的 computational budgets 通常将它们限制在广告(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》、《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》)等子领域,而非核心自然流量(core organic traffic)。

      为了支持数十亿用户,TokenMinds采 用了一种不同的 infrastructural 方法,利用异步的 User Behavior Service: UBS 框架(《Short-form Video Needs Long-term Interests: An Industrial Solution for Serving Large User Sequence Models》),该框架将 representation generation 与 scoring 解耦。

  3. Cross-Scenario Modeling:在不同场景之间建模 user behaviors 仍然是现实世界推荐系统中的关键挑战(《A survey on cross-domain recommendation: taxonomies, methods, and future directions》)。

    • 早期方法利用 multi-source user histories 来提升特定场景(如 CTR prediction)的性能(《Mixed information flow for cross-domain sequential recommendations》、《Mlora: Multi-domain low-rank adaptive network for ctr prediction》)。

    • 最近,采用 foundation model 范式的工业框架在统一 multi-scenario behaviors 方面显示出有希望的结果(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》、《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》。

      虽然这些方法通常将不同场景中的 user behaviors 投影到 shared continuous space 中,但 TokenMinds 更进一步,同时生成统一的 intent embedding 和 scenario-specific discrete token representations,使下游模型能够利用 scenario-aware user signals。

1.2 TokenMinds Framework

  1. 我们首先介绍 TokenMinds 的概述,包括设计背后的直觉。然后我们详细描述如何同时生成 SID-based discrete user tokens 和 dense user embeddings。之后我们考虑具有不同格式数据的跨场景情况。然后我们讨论 SID-based user token 在下游模型中的适配,并以 serving system 的介绍结束本节。

1.2.1 系统概述

  1. 我们采用 sequence-to-sequence: seq2seq 框架用于 TokenMinds ,如 Figure 1 所示。 User behavior signals 被 tokenized 为 sequential input,并被 a pre-trained foundation model 处理,以同时产生 dense user embeddings 和 discrete SID-based user tokens 。由于 foundation models 先前未接触过我们的视频语料库,我们遵循 PLUM(《Plum: Adapting pre-trained language models for industrial-scale generative recommendation》)进行 Continued Pre-Training: CPT,以使 new SID modality 与模型的现有知识对齐(在实验章节中进行了消融实验),然后进行 Supervised Fine-Tuning: SFT 从而用于 user modeling 。下面我们描述 inputs、resentations 和架构。

  2. User Data:我们整合了不同的 user behavior signals,如 textual search queries 、以及跨多个 recommender applications(例如,从首页直接推荐,由 what to watch next 等特色功能建议)的 watch histories。这些 behavioral data 包括 user histories 中的时间信号和 engagement 信号,如 interaction timestamps 和 likes or dislikes。

  3. Video Representation:我们不用随机分配的 video IDs: VIDs 来表示每个视频,而是用 Semantic ID: SID(《Recommender systems with generative retrieval》)来表示:通过一个 Residual-Quantized Variational AutoEncoder: RQ-VAE (《Plum: Adapting pre-trained language models for industrial-scale generative recommendation》)从视频的 content features导出的 a hierarchical sequence of discrete codewords。SID 为 user modeling 提供了两个关键优势:

    • 从头部到尾部视频上更好的泛化能力(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。

    • 以及更好的时间稳定性(temporal stability),因为 VID 随着语料库的演变而遭受 vocabulary 变动,这在建模长期 user histories 时问题会被放大。

  4. Foundation Model Architecture:虽然 decoder-only models 在 generative recommendation 中很流行,但我们采用 encoder-decoder 架构,原因有二。

    • 首先,众所周知,encoders 能在完整的 user history 上捕获更全面的 sequential patterns(《Generative Recommendation with Semantic IDs: A Practitioner’s Handbook》),可以自然地从此类 contextualized representations 中抽取 dense user embeddings,而 decoder 自回归地生成 discrete SID-based user tokens。

    • 其次,该架构提供了部署灵活性:encoder 和 decoder 可以解耦从而用于 serving (《Extremely Long User History Modeling at Instagram》、《OneRec Technical Report》),将 low-frequency heavy encoder 用于 history compression,将 high-frequency lightweight decoder 用于 more recent behaviors。

      一个 cutoff timestamp T 将用户历史行为分为 history window [W1,⋯,Wt](T 之前的 watches)和 future window {Wt+1,⋯,Wn}([T,T+24h] 内的 watches)。

      • encoder 负责处理 T 时刻之前的长历史、压缩成稳定用户画像。长期画像变化慢,所以低频异步刷新即可。

      • decoder 负责处理 T 时刻之后的短的近期兴趣。近期兴趣变化快,所以需要高频刷新;而 decoder 又不需要重读全历史,因而可以轻量。

    在训练期间,decoder 通过 cross-attention 关注 full encoder output,并学习为 targeted watches 生成 SID tokens。在 serving 时,我们提取 dual-outputs :

    • encoder outputs 通过 pooling(例如,last-token 或 mean pooling)得到 dense user embedding。

    • 通过对 decoder 应用 beam search 从而生成 multiple SID sequences 来产生 discrete user tokens,每个 SID sequence 被截断为粗粒度前缀(coarse-grained prefix) 从而作为 a single user token。

      粗粒度前缀就是:一个完整 Semantic ID 序列的前 L 个 codewords 组成的前缀,而不是完整 SID。在 TokenMinds 里,视频的 SID 由 RQ-VAE 生成,完整长度是 Lfull=8。TokenMinds 实际只保留前 L=4 级,表示一个较粗的语义兴趣簇,而不是具体视频。

      实验表明,使用完整的 Lfull 级会降低性能。

    我们也可以通过限制 decoder cross-attention 仅关注少量的 output embeddings from encoder (例如,learnt or attention pooled)。然而,这限制了模型的表达能力,可能降低 decoding 性能。为了缓解这一问题,可以引入具有 constrained cross-attention 的 auxiliary decoding 任务。

1.2.2 Model Training

  1. 在 Figure 2 中,我们可视化了 the training for one user example,其中用户的 watches 按时间顺序呈现并馈入模型以预测multiple near-future watches。为清晰起见,这描绘了基础的单场景情况;跨场景 condition tokens 和 search query integration 的扩展在后续章节中描述。例如,input sequence 'A12 B278 C23 D77' 表示 a video’s hierarchical SID 的truncated L-token prefix(L=4),而后续 tokens 如 '100.0s'、'20%' 和 'IOS' 编码了数值特征和文本特征(例如,watch time ratio 和 device platform)。

  2. 符号:

    • 设 W1,⋯,Wn 表示用户的时间顺序的 watch sequence。

    • 一个 cutoff timestamp T 将其分为 history window [W1,⋯,Wt](T 之前的 watches)和 future window {Wt+1,⋯,Wn}([T,T+24h] 内的 watches)。

    • 每个 watch Wk 由SID 来表示:给定 watch Wk 的 content embeddings,一个具有 Lfull=8个 codebook levels 的 RQ-VAE 产生 a full codeword sequence [SIDk,1,⋯,SIDk,Lfull];我们仅保留 prefix-L codes(L<Lfull),利用 SID hierarchy 以更粗粒度来表达视频,以鼓励 diversity 并缓解 memorization 问题(《OneRec Technical Report》)(参考消融实验)。

  3. Input Tokenization:每个 watch Wk 将其 prefix-L SID 作为 hard tokens 与 non-SID features(编码为 hard tokens 或 soft tokens)相结合。

    • hard tokens 通过对 dense features 进行分桶、或将 categorical/text features 映射到 vocabulary 来获得。

    • soft tokens 通过独立地 embedding 每个 feature ,拼接 resulting representations,并通过 MLP 将 concatenation 投影到 M 个 embeddings 来产生。

    transformer layer 的输入都是 embeddings。

    • 对于 hard tokens,它们被嵌入,然后将 output embeddings 馈入 transformer layer。

    • 对于 soft tokens,它们直接被馈入 transformer layers。注意,它不会被作为 target,因此不需要计算它的 loss。

    论文主实验里取 M=1,即每个 watch 只用一个 soft token 来表示所有non-SID 特征。论文也提到,用 1 个 soft token 会牺牲约 5% 离线 recall,但能大幅缩短序列长度。

  4. Training Objective:我们采用两个与 standard next-watch prediction的关键的偏离,两者都在实验章节中通过消融实验验证。

    • 首先,我们采用前瞻采样(look-ahead sampling)(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》):不是预测 immediate next watch Wt+1,而是从 future window {Wt+1,…,Wn} 中随机选择最多 N 个 targets,防止过拟合到 immediate watches,并通过近似 near-future interests 来改善泛化。

    • 其次,同时预测 multiple targets 比 single-target training 提高了训练效率。

    形式上,TokenMinds 最小化:

    L=−∑i=1Nr(Wi)×∑j=1Llog⁡P(SIDi,j∣W1,⋯,Wt,W<i,SIDi,<j)

    其中:

    • I 是 N 个 sampled targets 的索引。

    • L 仅在 prefix-L SID tokens 上计算。

    • r(Wi) 为 engagement reward,它鼓励多样化的、高价值的 consumption。它被表述为 multiple user signals 的组合。按根据 rewards 的比例来采样 training examples (《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)并在上述公式中平等地加权它们,在计算上比应用不同的加权权重更高效。

    注意,上述公式仅适用于 decoder outputs;encoder 仅通过 decoder 的 cross-attention layers 来接收梯度,这隐式地监督 encoder 产生 support accurate SID generation 的 representations。

1.2.3 Cross-Scenario Modeling

  1. 上述架构为每个内容场景(content scenario)独立地训练和服务。为了将它们整合到单个模型中,我们需要考虑长视频(long-form video: LFV)和短视频(short-form video: SFV)消费之间的关键差异。

    • SFV 以连续浏览的方式消费,无需显式的 click 发起,因此与 LFV 相比产生更强的反馈循环(feedback loop)。

    • 此外,用户兴趣并非按视频格式隔离;近一半的用户同时参与 SFV 和 LFV,并且它们的 SID 共享 significant vocabulary overlap(例如,the first two prefixes 大约有 40% 重叠)。这表明用户在 LFV 和 SFV 中的兴趣存在重叠。开发一个统一模型可以提高计算效率,促进 transfer learning,并捕获更全面的 user preferences 视图。

  2. Unified Training:为了构建 the unified input sequence,我们在每个 watch 之前添加 discrete condition tokens (例如,<LFV>、<SFV>)以区分 content scenarios。由于 TokenMinds 在从 a pre-trained LLM 继承来的 a shared token space 中运行,textual signals 可以自然地与 SID-based watch tokens 交错。我们通过对 S 个最近的 search queries 的每个 search query 之前添加 a <Search> token 来利用这一点,将它们按照时间顺序与 watch sequence 交织在一起(在实验章节进行了消融实验)。

    示例:

    这样,在每个 search 开头都有一个 <Search> token;每个长视频开头都有一个 <LFV> token;每个短视频开头都有一个 <SFV> token 。所有的 search、长视频、短视频都按照发生的时刻进行排序,交织在一起。

    对于 training objective,我们将 the random sampling of future target watches(见前面的章节)扩展为 uniformly include both SFVs and LFVs。Condition tokens 被排除在 loss 之外,因为预测它们过于简单,且若纳入其中会降低 SID prediction quality。

    • 在跨场景统一训练中,未来 target watches 的采样要均衡覆盖长视频和短视频,避免场景偏置。

    • 而 <LFV>、<SFV> 这类场景条件 tokens 只作为输入条件,不计算损失。因为预测它们太简单、信息量低,纳入 loss 会分散模型容量,最终损害 SID user token 的预测质量。

  3. Multi-Context Decoding:A unified model 应该在 serving 时为每个 context 生成 scenario-specific user tokens。一种朴素的方法需要为每个 context 单独进行 inference runs,冗余地处理 user history,抵消了 a unified model 的效率提升。

    这里的意思是,分别以 <LFV> token 和 <SFV> token 作为 condition token,来解码生成结果。用户历史是同一份,但 encoder 被跑了两次。

    我们提出 multi-context decoding 来消除这种冗余。给定 a shared encoder pass over the user history ,我们引入 a context stage,将 decoding 划分为并行的 sub-batches,每个 target context 一个 sub-batch。每个 sub-batch 以其各自的 condition token 来初始化,并通过 beam search 来独立地解码,而所有 sub-batches 完全地共享来自 a single encoder pass 的 encoder hidden states。如 Figure 3 所示,这使得 a single prefill pass 能够进行并发的、context-specific 的解码——从相同的 cached user representation 同时生成 LFV user tokens 和 SFV user tokens。

    这里的解决方案是:

    • 先做一次共享的 encoder pass。

    • 然后在 encoder 输出之后、decoder 解码之前,加一个 context stage,把解码请求按目标场景分成多个并行的 sub-batch:

      • sub-batch 1:目标场景 LFV,用 <LFV> 初始化。

      • sub-batch 2:目标场景 SFV,用 <SFV> 初始化。

    • 每个 sub-batch 内部独立做 beam search,生成该场景的多个 SID 序列。

1.2.4 Downstream Adaptation

  1. 将 SID-based user tokens 集成到下游模型——特别是需要 dense vectors 的传统 LEMs ——需要将 discrete tokens 投影到 a continuous embedding space 中。我们评估三种 token-to-embedding 的方法(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》):

    • (1):Prefix Embedding Mapping:将每个 predicted L-prefix SID 映射回 original content embeddings (这个 content embeddings 就是 SID 的 embedding)。由于多个视频可能 collapse 为相同的 coarse-grained prefix,我们对sharing the predicted prefix 的所有视频的 original embeddings 进行均值池化,以获得 SID embedding。

    • (2):N-gram Embedding:将每个 predicted SID sequence 分割为固定长度的 N-gram sub-words(例如,对于 N=2 表示连续的 codeword pairs),每个 N-gram sub-word 被映射到 a learned embedding。SID representation 是 the sum of its sub-word embeddings。

    • (3):SPM Embedding(《Subword regularization: Improving neural network translation models with multiple subword candidates》):类似于 N-gram,但使用 SentencePiece 从 item distributions 中学习 variable-length sub-words。

    我们将 (2) 和 (3) 统称为 Learnable Embeddings: LEs,因为两者都使用随机初始化的 embedding tables,与下游模型一起被端到端地训练,与 (1) 的 static mapping 形成对比。

    N-gram Embedding 和 SPM Embedding 的效果要比 Prefix Embedding Mapping 更好。而N-gram Embedding 和 SPM Embedding 之间效果差不多。

    在 serving 时,beam search 为每个用户产生 B 个 SID sequences,每个 SID sequence 代表一个 predicted future interest。为了形成 a single user vector,我们通过池化(例如,attention-weighted 池化、均值池化、最大值池化、或top-k concatenation)聚合对应的 B 个 SID embeddings。所有策略产生相差无几的下游性能,表明收益主要来自 tokens 本身包含的信息,而非 aggregation choice。

    这 B 个 SID embeddings 代表了 B 个兴趣。这里采用均值池化即可用于下游模型。

    上述 encoder-derived dense embeddings 和 token-derived SID embeddings 都可以作为下游模型的 direct input features,或作为 cross-attention layers 中的 key-value pairs,其中 candidate items 对 user representations 进行关注(attend)。

1.2.5 Serving System

  1. 为了克服十亿用户规模下的 high serving costs 和 latency,我们将 user representation generation 与 real-time down-stream inference 解耦,如 Figure 4 所示。基于 User Behavior Service: UBS 框架(《Short-form Video Needs Long-term Interests: An Industrial Solution for Serving Large User Sequence Models》)构建,TokenMinds 异步生成 user embeddings 和 user tokens,并将其缓存在快速的 key-value store 中。real-time models 在 scoring 期间检索这些 cached representations,无论 TokenMinds 的底层复杂性如何,都保持恒定的延迟和成本。当 a scoring request 到达时,客户端检查有效的、未过期的 representations:

    • 如果可用,则获取它们并直接馈入下游模型。

    • 否则,立即触发后台刷新:Representation Refresh Service 读取用户最新的 watch history,通过 exported TokenMinds model 运行 inference,并将 updated representations 写回缓存。

1.3 实验

  1. 我们首先详细说明 TokenMinds 的实验设置,然后从离线指标和 downstream integration 的线上评估考察 recommendation quality。最后,我们分享来自 additional studies 的经验。我们尝试回答以下问题:

    • RQ1 (Token Adaptation & Viability):如何将 SID-based discrete user tokens 最优地适配用于下游 continuous models;并且当部署在工业规模的 recommendation surfaces 上时,它们是否能产生 measurable improvements ?

    • RQ2 (Complementary Values):TokenMinds 的 dual-output 性质是否通过同时生成 continuous embeddings 和 discrete tokens 来提供互补的价值?

    • RQ3 (Cross-Scenario Modeling Impact):统一 cross-scenario modeling 范式(联合训练 LFV 和 SFV)能否在不影响下游 recommendation quality 的情况下实现计算效率提升?

1.3.1 Experiment Set-up

  1. Model:TokenMinds 采用基于 Gemini V1.5 的 encoder-decoder 架构,包含一个 370M-parameter Mixture-of-Experts (MoE) encoder 和一个 370M-parameter dense decoder。两者都从 Continued Pre-Training (CPT) checkpoints(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)初始化。

  2. Training:TokenMinds 在 YouTube LFV and SFV watch histories interleaved with textual search queries 的数据上训练,使用为 Gemini 生态系统开发的内部 JAX 和 Pathways 框架。我们采用 continuous training,在最新的 daily data 上,从而捕获 fresh engagement signals,每天处理数百万样本,这比通常需要 billions of interactions 的 traditional LEMs 更具样本效率。

    对于每个 user sample,我们使用 most recent 1,200 watches interleaved with S=10 search queries,最大 input sequence length 为 1,024 tokens(在后续章节中进一步分析)。每个 watch 由一个 condition token、 prefix-L=4 SID tokens、和用于 non-SID features 的 M=1 soft token组成;using a single soft token 以大约 5% 的 offline recall 为代价,显著减少了 sequence length。Target selection 从 24-hour look-ahead window 中采样最多N=15个 watches。通过对学习率在 [10−6,10−3] 范围内进行 grid search,我们发现 linear warmup with cosine decay and cyclic restarts(《Language models are few-shot learners》、《Sgdr: Stochastic gradient descent with warm restarts》)达到了 peak same-day performance,但在 continuous training 下,a constant learning rate 对 day-to-day distribution shifts 更为鲁棒。

  3. Serving:TokenMinds 通过异步的 processing infra 具备 24-hour refresh cadence。对于每个用户:

    • 我们从 encoder 提取 a 1,152-dimensional dense embedding。

    • 并通过 beam search 来解码 B=40 SID sequences(20 sequences 用于 LFV,20 sequences 用于 SFV),所有 sequence 都使用 prefix length L 来形成 user’s discrete token representation。

1.3.2 Representation Quality

  1. 本节中的所有模型都在相同数据上训练至收敛,以确保公平比较。

  2. Token Accuracy and Training Ablations:我们评估 TokenMinds generated tokens 的预测准确性,并对关键 training design decisions 进行消融。我们在两种协议下评估 token accuracy:

    • Session Recall:使用近乎完整的 history [W1,⋯,Wn−1] 作为 input 并预测 final watch Wn。

    • Cold-Start Recall:使用 truncated history [W1,⋯,Wt] 作为 input 并预测从 future window {Wt+1,⋯,Wn} 中随机采样的观 watch。

    所有变体都在两种协议下进行评估以确保可比性。在两种情况下,我们生成 top-10 SID sequences(prefix length L)并计算Recall@10。

    Table 1 报告了结果以及三个消融实验,每个消融实验移除一项 training innovation,同时保持 evaluation 固定:

    • (1):w/o Multiple Targets:将 sampled targets 从 N=15 减少到 1。

    • (2):w/o Look-ahead Window:用标准 next-watch prediction( input [W1,⋯,Wt],target Wt+1)替换 look-ahead sampling。

    • (3):w/o SID Truncation: trains and encodes inputs with full-length SIDs (Lfull)而非 prefix-L ;在评估时,predictions 仍在 prefix-L granularity 上进行比较以确保指标可比性。

    所有三项都降低了准确性,其中 look-ahead removal、以及 full-length SIDs 对冷启动性能尤其有害(分别相对下降-10% 和-17%),确认了 temporal look-ahead sampling 和 coarse-grained prefixes 对于泛化到 immediate session 之外是必不可少的。

  3. Initialization Strategy:为了验证 the choice of initializing from a CPT checkpoint(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》),我们与两个基线进行比较:a general-purpose Pre-Trained Gemini checkpoint (without SID-specific grounding)、以及随机初始化。

    Table 2 显示了这些基线在 with and without search queries 的配置下的比较。我们观察到 CPT 始终优于 Pre-Trained Gemini,Pre-Trained Gemini 又优于随机初始化。这证实了在 CPT 期间 the semantic grounding of SIDs 对下游 task-specific fine-tuning 有益,超出了 base LLM 本身的 general sequence modeling 能力。

  4. Impact of Search Queries:为了量化 search queries 的贡献,我们将 our full model 与 a variant trained without search signals进行比较。如 Table 2 所示:

    • 添加 S=10 个交错的 queries 在所有初始化策略下持续改善 recall,确认了 explicit textual intent 提供了超出 watch history 之外的补充信号。

      既然 search queries 有效,那么为什么只选择 S=10 个而不是更多个 queries?两个原因:

      • 最大输入长度只有 1,024 tokens,watches 已经占了大头。所以 S=10​ 是在有限 tokens 预算下,给 search queries 分配的一个合理配额。如果再增加 search queries,每个 query 本身会 tokenize 成多个文本 tokens,还会加 <Search> token,就会挤掉 watches。

      • 只取 “最近的” search queries,更早的搜索时效性差。最近的 10 searches 已经能覆盖短期显式意图,再往前翻,边际信息价值下降。

    • 此外,search 的收益通过更强的 initialization 被放大,表明 an LLM backbone with aligned SID representations 更能够利用 textual search signals 与 behavioral watch histories 相结合。

  5. Token Diversity:在确立了预测准确性并确定了一些关键的驱动因素之后,我们接下来验证这种性能不是通过 redundant beams 来实现的。由于每个 decoded SID 作为下游模型的 a distinct interest signal,beam collapse ——多个 beams 收敛到近乎相同的 outputs ——将严重限制 representational capacity。

    我们使用两个指标评估多样性(diversity):

    • SID Token Collision Rate:测量位置 x 处相同 token 的频率,high collision 表示 semantic beam collapse。

    • SID Prefix Duplication Rate:统计直到位置 x 为止,相同 prefixes 的数量,high repetition 表示 isolated exploration of a single branch。

    我们在 5K users 的样本上评估这些指标,将 generated tokens 与 ground-truth watches from the same look-ahead 24-hour window 进行比较。为确保尽管规模差异仍能公平比较,我们对较大集合(ground-truth watches or generated SIDs)进行降采样以匹配 smaller set 的大小。为减少 subsampling 的方差,每次比较时,我们对 larger set 重复降采样 10 次,并报告平均 similarity 分数。different token indices 下的结果 Cumulative Distribution Function: CDF 绘制在 Figure 5 中,表明 TokenMinds 达到了与 ground-truth distribution 相当的 prediction diversity,确认了 beam search 产生 diverse interest signals 而不会坍缩为 redundant outputs。

    实线是 ground-truth,虚线是 ours。

  6. Embedding Quality:我们在 2K random sampled users 上评估 embedding consistency。对于每个用户,我们从其 full history 生成 an embedding EA,并通过随机丢弃 watches 生成一个扰动的变体 EA∗,然后与 a random user’s embedding EB 进行比较。平均余弦相似度 Sim(EA,EA∗)=0.993 远超 Sim(EA,EB)=0.761,表明 embeddings 在轻微 history 扰动下保持稳定,同时在用户间保持区分性。

1.3.3 Online Performance

  1. 我们通过线上 A/B 实验评估 TokenMinds 的现实世界影响,将用 user embeddings 和 user tokens 集成到多个 production ranking models 中。虽然 TokenMinds 已部署在 YouTube 的 ranking system 、retrieval system 和 LLM-based production systems 中,但本文侧重于 ranking integration。所有实验均在 a seven-day period 内进行。我们在 LFV 和SFV 中报告两个关键 quality metrics 上的性能:Engaged Users 和 Satisfied Engagement。在本小节,表格内的粗体数字表示在 95% 置信度下具有统计学显著性。

  2. Token Adaptation Pivot Study:在投入大规模计算到 full-scale dual-output architecture 之前,我们使用一个轻量级的 110M model 在 SFV surface 上进行了线上转折研究,以确定最优的 token adaptation 策略。如 Table 3 所示,我们比较了静态的 Prefix Embedding Mapping: EM 与 Learnable Embeddings: LE ——具体来说是 Unigram (N = 1) 用于 SFV,匹配该 surface 上现有的 item tokenization。LE 始终优于 EM,确认了 allowing downstream models to learn a specialized embedding space 能产生更好的 task adaptation(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。在 LFV platform 上使用 SPM-based LE 进行的 auxiliary evaluations 显示出相同的趋势。因此,我们在所有后续 full-scale evaluations 中专用地使用 LE,回答了RQ1 的前半部分。

  3. Downstream Quality:为了评估 TokenMinds 的各自部分的影响和互补的影响,我们比较了单独添加 continuous embeddings、单独添加 discrete SID-based tokens(通过 LE)、以及同时添加两者。Table 4 报告了 primary SFV surfaces 和 primary LFV surfaces 的结果。我们得出两个关键发现。

    • 首先,SID-based user tokens 在 production systems 中提供了 additive value,直接回答了 RQ1 的后半部分。除了 Table 4 中的 primary surfaces,仅在另外两个 LFV surfaces上的 token-only integration 在 Engaged Users 方面实现了 +0.04% / +0.16% 的统计显著提升,在 Satisfied Engagement 方面实现了 +0.07% / +0.11% 的提升,进一步确认了泛化性。

    • 其次,同时部署 embeddings 和 tokens 产生了放大的增益,验证了 RQ2 关于这两种模态的互补性质、以及我们的 dual-output 架构。

  4. Downstream Cost:对于 joint token and embedding generation,TokenMinds 大约需要 339ms per user,完全在后台处理中被吸收。每个用户的 discrete token representation 仅需 1,280 bytes——而 dense embedding 需要 4,608 bytes ——存储减少了 72%。在 serving 时,系统在来自多个 production surfaces 的 1.44M read requests per second 中达到了96.4% 的 cache hit rate。将 pre-computed embeddings 作为 features 所引入带来的计算成本可忽略不计;token integration 开销详见 Table 5。

  5. Cross-Scenario Modeling:我们评估 the unified cross-scenario model 与两个基线:

    • a LFV-only model(仅在 LFV data 上训练)用于 quality comparison。

    • separately trained LFV and SFV models 用于 efficiency comparison。

    从 Table 6(Panel B)来看,将 LFV 和 SFV 整合到单个框架中产生了可观的成本节约。

    • 训练计算减少了 50%,因为一个模型取代了两个模型。

    • 上游 serving 成本减少了 31%,因为 multi-context decoding 在场景间共享 a single encoder pass(481 chips),而不是 running two full model passes(总共 698 chips)。

    • 下游 integration cost 保持中性,因为 the unified model 产生的 outputs 形状与 LFV-only baseline 相同。

    在质量方面,Panel A 显示用 unified representations 替换 LFV-only baseline 持续改善了 freshness 指标(Fresh Engagement 衡量用户与新视频的 interaction),而没有降低 core engagement 指标。值得注意的是,the unified model 在相同的 fixed-length input sequence 中处理 LFV watches 和 SFV watches,几乎将可用的 LFV history 减半。尽管有这种减少,core metrics 仍保持中性这一事实表明,来自 SFV interactions 的 cross-scenario signals 补偿了 reduced LFV context ——这确认了 the unified paradigm 在实现显著效率提升的同时不损害质量。

    LFV-only 模型中,inputs 仅仅包含 LFV 视频,因此在相同 input length 下,LFV 视频数量相比 Cross-Scenario Model 翻倍。

1.3.4 Scaling Studies

  1. 为了优化计算效率,这些研究采用 an accelerated offline evaluation protocol(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》):模型在 7 randomly shuffled days 上训练,并在连续的第 8 天进行评估。

  2. Model Variants & Capacity Allocation:我们探讨 architectural design 和 capacity allocation 对模型质量的影响。受限于可用的 pre-trained Gemini checkpoints,我们研究三种 warm-started 变体:

    • a Balanced MoE model:370M Enc / 370M MoE Dec。

    • a Balanced Dense model:370M Enc / 370M Dense Dec。

    • an Unbalanced Dense model :420M Enc / 110M Dense Dec。

    我们围绕 Balanced Dense anchor 对学习率网格搜索 [0.1×,10×] (即,对 Balanced Dense anchor 的学习率除以 10 到乘以 10)。

    Figure 6 显示了 LFV 的 Training 和 8th-Day Recall 与 Iso-FLOPS 的关系( SFV 因趋势相似而省略)。

    • 虽然 the Unbalanced model 显示的 Training Recall 低于 the Balanced Dense baseline,但它达到了相差无几的 8th-Day Recall。这为异步 serving 策略打开了大门——将较重的 encoder 以较低频率刷新以获取稳定的 user embeddings,与较轻的 decoder 以较高频率刷新从而捕获 recent behaviors 相结合。

    • 此外,在相差无几的 FLOPS 下,the Balanced MoE decoder 在 8th-Day Recall 上优于 Dense 对应物,确认了 sparse expert routing 在泛化到 future user behaviors 方面提供了切实的优势。

  3. Input Length Scaling:虽然现代 user modeling 通常受益于 long-term watch histories(《DV365: Extremely Long User History Modeling at Instagram》),但开发 compute-optimal models 需要理解 scaling dynamics。我们分析了用户历史长度(history length: HL)的影响。如 Figure 7 所示:

    • LFV 和SFV 的 8th-Day Recall@10 在大约 1K watches 处开始饱和。

    • 扩展到 2K watches 在 SFV 上产生相差无几或略微下降的性能。

  4. Batch Size Scaling:我们评估了 4K、8K 和 16K 的 batch sizes,使用 8K 作为 anchor configuration 并进行了广泛调优好的学习率,并通过标准的N 规则(《Language models are few-shot learners》)扩展到其他 batch sizes。相对于 4K baseline,8th-Day Recall@10 在 8K 时提高了 +2.5% / +5.5%(SFV/LFV),在 16K 时提高了 +7.6% / +13.7%。与 PLUM (《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)中的发现一致,较大的 batches 加速了收敛,并在相同的 training window 内产生了更强的 representations。这突显了在硬件容量允许的情况下最大化 batch size 对于 compute-optimal training 的重要性。

1.4 结论

  1. 我们介绍了 TokenMinds ,一个通过从 pre-trained LLMs 改编的 a unified encoder-decoder architecture 来生成 discrete SID-based user tokens 和 dense user embeddings 的框架。Cross-scenario modeling 进一步将 LFV 和 SFV 统一在 a shared SID vocabulary 下,采用 multi-context decoding,大幅减少了上游计算。

    作为 YouTube 上已部署的系统,TokenMinds 证明了 SID-based user tokens 是 production ranking systems 的可行的 representation。将这些 tokens 与 dense embeddings 相结合产生了放大的性能增益,确认了 dual-output design 的互补的价值。我们的消融实验进一步验证了 CPT initialization 的重要性、以及 search query signals 对 SID-based user modeling 的协同效益。

    这些结果表明,离散的、语义上有据可依的 user representations 为 scaling user modeling 到超越 dense embeddings alone 的 representational constraints 之外提供了一条有希望的方向。