《TokenMinds: Pretrained User Tokens and Embeddings for User Understanding in Large Recommender Systems》
工业推荐系统中的 user modeling 通常生成 dense embeddings,这些 embeddings 受限于 representational constraints:它们固有地采用固定维度的向量。一种新兴的 discrete user representation 替代方法——使用 LLM 生成 text-based user tokens ——捕获的是主题共现(topical co-occurrences)而非 deep sequential behavior dynamics,并且产生的 outputs 难以与 item attributes 关联。同时,基于 Semantic ID: SID 的 item tokenization 已被证明在 generative recommendation 中能有效提升泛化能力,然而 discrete SID-based 的 user representations 在很大程度上仍未得到探索。我们提出 TokenMinds,一个工业规模的系统,它将 PLUM 框架(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)从 item retrieval 扩展到 user modeling,通过从 pre-trained LLMs 改编的一个 encoder-decoder 架构,同时生成 discrete SID-based user tokens 和 dense user embeddings。这种 dual-output 的设计提供了离散的、语义上有据可依的 user representations 的互补优势,同时保持了与现有下游模型(这些模型依赖 dense embeddings)的兼容性。此外,shared SID vocabulary 自然地扩展到跨场景建模(cross-scenario modeling):通过将长视频行为和短视频行为统一到单个模型中,我们显著降低了 training 成本和 serving 成本。我们通过广泛的离线实验和在多个 YouTube surfaces 的线上 launches 来验证 TokenMinds,通过一种异步 infrastructure 来服务全量用户流量(数十亿用户),该 infrastructure 将 representation generation 与下游 scoring 进行解耦。以 ranking 作为主要的下游用例,我们的结果证实了 SID-based user tokens 在工业规模上的实际可行性,并证明 tokens 和 dense embeddings 在不同 production ranking systems 中提供了互补价值。
推荐系统已经从早期的因子分解机(factorization machines)发展到现代深度学习架构,并且越来越多地从多级级联架构(multi-stage cascaded architectures)转向 end-to-end 方法。sequential user modeling 侧重于利用用户交互过的 history,长期以来在实现推荐系统个性化方面发挥了关键作用。Large Embedding Models: LEMs 作为主导的工业范式,依赖大规模的 embedding tables 来表示 high-cardinality item IDs。虽然 Large Embedding Models 在记忆 user-item interactions 方面有效,但它们在复杂网络上的泛化能力常常受限(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。
相反,Large Language Models: LLMs 将其参数预算(parameter budgets)主要分配给神经网络而非大规模 embedding tables,在 recommendation 方面已显示出巨大潜力。近期工作利用它们的 reasoning 和 contextual understanding 能力进行 feature extraction 、Large Embedding Model integration 和 end-to-end sequence modeling。然而,将 LLMs 应用于 domain-specific recommendations 揭示了两个关键瓶颈。
首先,models pretrained on natural language 在处理大规模、非文本的 ID spaces(例如,数十亿视频)时存在根本性的模态鸿沟(modality gap)。
其次,通过将 LLMs 与 traditional large embedding tables 耦合来弥合这一鸿沟,继承了 LLMs 本应超越的那些局限性:scaling 约束、vocabulary 变动和有限的表达能力。
为了克服这些瓶颈,Semantic IDs tokenize item 为 hierarchical discrete codewords (这些 codewords 源自 content semantics),消除了对 embedding table 的依赖,并通过有意义的碰撞(collisions)提升了泛化能力。在 SID 解决了 traditional representations 的 scaling 和 expressiveness 的局限性之后,PLUM 框架通过 Continued Pre-Training: CPT 和 task-specific post-training,将 new SID vocabulary 与 LLM’s pre-existing knowledge 对齐,从而弥合了 remaining modality gap。
虽然 PLUM 框架已被证明对 item retrieval 有效,但其在 sequential user modeling: SUM 方面的潜力尚未被探索。在本文中,我们提出了 TokenMinds,它扩展了 PLUM 从而生成 SID-based user tokens 作为 user representations。SID-based representation 与 using dense user embeddings 的常见方法形成对比,后者将 a user’s full spectrum of interests 压缩到一个或几个固定维度的向量中,可能会丢失细粒度的信号。同时,using LLMs to generate text-based user profiles 也日益受到关注。虽然 text profiles 可以被视为 discrete user representation 的另一种形式,但它们面临关键局限性: LLMs without pre-training on user behavior signals 倾向于捕获主题共现(topical co-occurrences)而 deep sequential behavior dynamics(《Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?》),它们的 text outputs 难以与 item attributes in any target domain 关联,并且当集成到 non-textual downstream systems 时会引入 modality gaps (《Llm4msr: An llm-enhanced paradigm for multi-scenario recommendation》)。在 TokenMinds 中,我们通过利用 PLUM 中的 CPT 方法,并将 model outputs 锚定到语义上有意义、且基于 recommendation 语料库的 SID tokens 来克服这些挑战。
具体来说,TokenMinds 采用 encoder-decoder 架构——这是 generative retrieval 中广泛使用的范式——但将其重新用于一种新颖的 dual-output 设计:decoder 生成 SID-based user tokens,而 encoder 处理 sequential user features 并同时产生 dense user embeddings,确保与使用 dense embeddings 的现有下游模型兼容。以 ranking 作为主要下游用例,我们研究如何将 SID-based user tokens 正确地集成到 production ranking models 中。为了在严格的 latency 约束下为数十亿用户服务该模型,我们在异步 serving infrastructure 之上部署 TokenMinds,该 infrastructure 将繁重的 representation generation 与 real-time scoring 解耦。
除了单场景 user modeling 之外,由于 TokenMinds 在从 pre-trained LLM 所继承的 shared token space 中运行,它自然地支持两个进一步的扩展。
首先,诸如 textual search queries 之类的 heterogeneous signals 可以轻松地与 watch histories 在 input sequence 中交织起来。
其次,shared SID vocabulary 为长视频(long-form video: LFV)和短视频(short-form video: SFV)之间的跨场景建模提供了天然桥梁,这两个场景具有不相交的 video ID spaces 和不同的消费模式。我们在每个 watch 前添加 scenario-specific condition tokens (LFV/SFV),使单个模型能够在按时间交错排列的 cross-scenario sequences 上进行训练。在推理时,我们引入了 multi-context decoding:encoder 在单次前向传递中产生 a shared user representation,decoder 通过以不同的 scenario prefixes 为条件生成 scenario-specific user tokens ——有效地实现了 scenario-aware output 而无需多模型开销。
我们总结贡献如下:
Dual-Output 架构:我们提出 TokenMinds,一个为大规模工业推荐系统设计的统一的 user modeling 架构。其新颖的 dual-output 机制联合生成标准的 dense continuous embeddings 和 discrete SID-based user tokens。
Tokenized User Representations & Adaptation:我们提出了一种 SID-based discrete tokens 的新颖 application 从而用于 user representation,并确定了在下游模型中适配它们的最优策略。虽然 TokenMinds 支持各种下游任务,但本文侧重于 ranking integration,我们在其中展示了 tokenized representations 在工业规模上的实际可行性,并证明了 dense embeddings 和 discrete tokens 可以捕获互补信号,放大了性能提升。
Cross-Scenario Modeling :我们引入了一种高效的 modeling 范式,原生地整合了 cross-scenario content training and serving。通过在单个框架内联合建模根本不同的场景( LFV 和 SFV),与维护 separate models 相比,我们将上游训练计算量减少了 50%,serving 计算量减少了 31%,同时保持了 core engagement 质量,并在两个场景中显著改善了 fresh content metrics。
Industrial-Scale Deployment:我们通过广泛的离线评估和在 production ranking systems 上的线上 A/B testing 验证了 TokenMinds,在 core user metrics 上实现了高达 +0.11% 的统计显著提升,在 core engagement metrics 上实现了高达 +0.62% 的提升。TokenMinds 已在多个主要 YouTube surfaces 的生产环境中部署,在全量用户流量上服务 LFV 和SFV 推荐。
Sequential User Modeling:
sequential recommendation 已经从 RNN(《Session-based Recommendations with Recurrent Neural Networks》)和 CNN(《Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding》)演变为强大的 Transformer-based 的模型,如 SASRec(《Self-Attentive Sequential Recommendation》)和万亿参数 transducers 如 HSTU(《Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations》)。尽管架构复杂,这些模型从根本上依赖于 atomic item IDs 的大规模的 continuous embedding tables。
最近,Semantic IDs: SIDs 已成为 item representation 的一种有吸引力的替代方案,它能增强泛化能力(《Recommender systems with generative retrieval》、《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)并提高样本效率(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)。
在用户侧,dense continuous embeddings 是工业界 user representation 的主流标准(《Longer: Scaling up long sequence modeling in industrial recommenders》、《Pinnerformer: Sequence modeling for user representation at pinterest》、《Empowering General-purpose User Representation with Full-life Cycle Behavior Modeling》)。
最近 LLM-based 的方法试图将 preferences 总结为 natural language profiles(《Recommendation as language processing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5)》、《Idgenrec: Llm-recsys alignment with textual id learning》、《Palr: Personalization aware llms for recommendation》),但这些方法倾向于捕获 topical co-occurrences 而非 user’s sequential dynamics(《Lost in Sequence: Do Large Language Models Understand Sequential Recommendation?》),并且在 non-textual downstream systems 中面临 modality gaps 的挑战(《Llm4msr: An llm-enhanced paradigm for multi-scenario recommendation》)。 discrete, SID-based token 的 user representations 在很大程度上仍未得到探索。
我们的工作通过将 SID 范式从 items 扩展到 users 来填补这一空白,生成紧凑的 discrete user tokens,同时继续输出 dense embeddings 以与现有下游系统兼容。
Generative Recommendations:
generative recommendation 将 traditional embedding-table-based paradigm 重新构建为 sequence-to-sequence 任务,例如 PLUM(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)、GenRank(《Towards large-scale generative ranking》)和 GPR(《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》)。
我们的工作将 PLUM 框架从 retrieval 扩展到 user modeling。retrieval models 用于预测下一个即时 items,然而 user modeling 在更长的时间窗口内捕获更广泛的 spectrum of intents。它利用更粗粒度的 semantic granularity 来识别 distinct areas of interest,避免了严格映射回specific individual items 的需要。
与 GPR 不同,后者将 user representation 与下游任务指标和 policy optimization 对齐,我们将 learning of user representations 与 specific downstream training objectives 解耦,以提供对用户兴趣的通用理解。
与此同时,LIGER(《Unifying generative and dense retrieval for sequential recommendation》)和 COBRA(《Sparse meets dense: Unified generative recommendations with cascaded sparse-dense representations》)证明了 sparse semantic IDs 和 dense embeddings 为 retrieval 提供了互补的 item representations。然而,两者都仅在 item level 操作。
TokenMinds 将这种 sparse-dense 互补性从 items 扩展到 users,通过统一的 encoder-decoder 架构同时生成 discrete SID-based user tokens 和 continuous embeddings。
在线上部署 heavy generative models 面临着现实世界系统的 scaling 瓶颈。近期工作通过优化线上架构本身来解决这一问题,例如 HSTU(《Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations》)的高效 transducers、LONGER(《Longer: Scaling up long sequence modeling in industrial recommenders》)的 optimized representations 、或 parallel generation 方法(《Generating long semantic ids in parallel for recommendation》)。由于这些模型在 real-time scoring 期间执行大量计算,严格的 computational budgets 通常将它们限制在广告(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》、《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》)等子领域,而非核心自然流量(core organic traffic)。
为了支持数十亿用户,TokenMinds采 用了一种不同的 infrastructural 方法,利用异步的 User Behavior Service: UBS 框架(《Short-form Video Needs Long-term Interests: An Industrial Solution for Serving Large User Sequence Models》),该框架将 representation generation 与 scoring 解耦。
Cross-Scenario Modeling:在不同场景之间建模 user behaviors 仍然是现实世界推荐系统中的关键挑战(《A survey on cross-domain recommendation: taxonomies, methods, and future directions》)。
早期方法利用 multi-source user histories 来提升特定场景(如 CTR prediction)的性能(《Mixed information flow for cross-domain sequential recommendations》、《Mlora: Multi-domain low-rank adaptive network for ctr prediction》)。
最近,采用 foundation model 范式的工业框架在统一 multi-scenario behaviors 方面显示出有希望的结果(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》、《GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation》。
虽然这些方法通常将不同场景中的 user behaviors 投影到 shared continuous space 中,但 TokenMinds 更进一步,同时生成统一的 intent embedding 和 scenario-specific discrete token representations,使下游模型能够利用 scenario-aware user signals。
我们首先介绍 TokenMinds 的概述,包括设计背后的直觉。然后我们详细描述如何同时生成 SID-based discrete user tokens 和 dense user embeddings。之后我们考虑具有不同格式数据的跨场景情况。然后我们讨论 SID-based user token 在下游模型中的适配,并以 serving system 的介绍结束本节。
我们采用 sequence-to-sequence: seq2seq 框架用于 TokenMinds ,如 Figure 1 所示。 User behavior signals 被 tokenized 为 sequential input,并被 a pre-trained foundation model 处理,以同时产生 dense user embeddings 和 discrete SID-based user tokens 。由于 foundation models 先前未接触过我们的视频语料库,我们遵循 PLUM(《Plum: Adapting pre-trained language models for industrial-scale generative recommendation》)进行 Continued Pre-Training: CPT,以使 new SID modality 与模型的现有知识对齐(在实验章节中进行了消融实验),然后进行 Supervised Fine-Tuning: SFT 从而用于 user modeling 。下面我们描述 inputs、resentations 和架构。

User Data:我们整合了不同的 user behavior signals,如 textual search queries 、以及跨多个 recommender applications(例如,从首页直接推荐,由 what to watch next 等特色功能建议)的 watch histories。这些 behavioral data 包括 user histories 中的时间信号和 engagement 信号,如 interaction timestamps 和 likes or dislikes。
Video Representation:我们不用随机分配的 video IDs: VIDs 来表示每个视频,而是用 Semantic ID: SID(《Recommender systems with generative retrieval》)来表示:通过一个 Residual-Quantized Variational AutoEncoder: RQ-VAE (《Plum: Adapting pre-trained language models for industrial-scale generative recommendation》)从视频的 content features导出的 a hierarchical sequence of discrete codewords。SID 为 user modeling 提供了两个关键优势:
从头部到尾部视频上更好的泛化能力(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。
以及更好的时间稳定性(temporal stability),因为 VID 随着语料库的演变而遭受 vocabulary 变动,这在建模长期 user histories 时问题会被放大。
Foundation Model Architecture:虽然 decoder-only models 在 generative recommendation 中很流行,但我们采用 encoder-decoder 架构,原因有二。
首先,众所周知,encoders 能在完整的 user history 上捕获更全面的 sequential patterns(《Generative Recommendation with Semantic IDs: A Practitioner’s Handbook》),可以自然地从此类 contextualized representations 中抽取 dense user embeddings,而 decoder 自回归地生成 discrete SID-based user tokens。
其次,该架构提供了部署灵活性:encoder 和 decoder 可以解耦从而用于 serving (《Extremely Long User History Modeling at Instagram》、《OneRec Technical Report》),将 low-frequency heavy encoder 用于 history compression,将 high-frequency lightweight decoder 用于 more recent behaviors。
一个
cutoff timestamp将用户历史行为分为 history window( 之前的 watches)和future window( 内的 watches)。
encoder负责处理时刻之前的长历史、压缩成稳定用户画像。长期画像变化慢,所以低频异步刷新即可。
decoder负责处理时刻之后的短的近期兴趣。近期兴趣变化快,所以需要高频刷新;而 decoder又不需要重读全历史,因而可以轻量。
在训练期间,decoder 通过 cross-attention 关注 full encoder output,并学习为 targeted watches 生成 SID tokens。在 serving 时,我们提取 dual-outputs :
encoder outputs 通过 pooling(例如,last-token 或 mean pooling)得到 dense user embedding。
通过对 decoder 应用 beam search 从而生成 multiple SID sequences 来产生 discrete user tokens,每个 SID sequence 被截断为粗粒度前缀(coarse-grained prefix) 从而作为 a single user token。
粗粒度前缀就是:一个完整
Semantic ID序列的前个 codewords组成的前缀,而不是完整SID。在TokenMinds里,视频的SID由RQ-VAE生成,完整长度是。 TokenMinds实际只保留前级,表示一个较粗的语义兴趣簇,而不是具体视频。 实验表明,使用完整的
级会降低性能。
我们也可以通过限制 decoder cross-attention 仅关注少量的 output embeddings from encoder (例如,learnt or attention pooled)。然而,这限制了模型的表达能力,可能降低 decoding 性能。为了缓解这一问题,可以引入具有 constrained cross-attention 的 auxiliary decoding 任务。
在 Figure 2 中,我们可视化了 the training for one user example,其中用户的 watches 按时间顺序呈现并馈入模型以预测multiple near-future watches。为清晰起见,这描绘了基础的单场景情况;跨场景 condition tokens 和 search query integration 的扩展在后续章节中描述。例如,input sequence 'A12 B278 C23 D77' 表示 a video’s hierarchical SID 的truncated L-token prefix(tokens 如 '100.0s'、'20%' 和 'IOS' 编码了数值特征和文本特征(例如,watch time ratio 和 device platform)。

符号:
设 watch sequence。
一个 cutoff timestamp history window watches)和 future window watches)。
每个 watch SID 来表示:给定 watch content embeddings,一个具有 codebook levels 的 RQ-VAE 产生 a full codeword sequence prefix-L codes(SID hierarchy 以更粗粒度来表达视频,以鼓励 diversity 并缓解 memorization 问题(《OneRec Technical Report》)(参考消融实验)。
Input Tokenization:每个 watch prefix-L SID 作为 hard tokens 与 non-SID features(编码为 hard tokens 或 soft tokens)相结合。
hard tokens 通过对 dense features 进行分桶、或将 categorical/text features 映射到 vocabulary 来获得。
soft tokens 通过独立地 embedding 每个 feature ,拼接 resulting representations,并通过 MLP 将 concatenation 投影到 embeddings 来产生。
transformer layer的输入都是embeddings。
对于
hard tokens,它们被嵌入,然后将output embeddings馈入transformer layer。对于
soft tokens,它们直接被馈入transformer layers。注意,它不会被作为target,因此不需要计算它的loss。论文主实验里取
,即每个 watch只用一个soft token来表示所有non-SID特征。论文也提到,用1个soft token会牺牲约5%离线recall,但能大幅缩短序列长度。
Training Objective:我们采用两个与 standard next-watch prediction的关键的偏离,两者都在实验章节中通过消融实验验证。
首先,我们采用前瞻采样(look-ahead sampling)(《Cross-Scenario Unified Modeling of User Interests at Billion Scale》):不是预测 immediate next watch future window targets,防止过拟合到 immediate watches,并通过近似 near-future interests 来改善泛化。
其次,同时预测 multiple targets 比 single-target training 提高了训练效率。
形式上,TokenMinds 最小化:
其中:
sampled targets 的索引。
prefix-L SID tokens 上计算。
engagement reward,它鼓励多样化的、高价值的 consumption。它被表述为 multiple user signals 的组合。按根据 rewards 的比例来采样 training examples (《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)并在上述公式中平等地加权它们,在计算上比应用不同的加权权重更高效。
注意,上述公式仅适用于 decoder outputs;encoder 仅通过 decoder 的 cross-attention layers 来接收梯度,这隐式地监督 encoder 产生 support accurate SID generation 的 representations。
上述架构为每个内容场景(content scenario)独立地训练和服务。为了将它们整合到单个模型中,我们需要考虑长视频(long-form video: LFV)和短视频(short-form video: SFV)消费之间的关键差异。
SFV 以连续浏览的方式消费,无需显式的 click 发起,因此与 LFV 相比产生更强的反馈循环(feedback loop)。
此外,用户兴趣并非按视频格式隔离;近一半的用户同时参与 SFV 和 LFV,并且它们的 SID 共享 significant vocabulary overlap(例如,the first two prefixes 大约有 40% 重叠)。这表明用户在 LFV 和 SFV 中的兴趣存在重叠。开发一个统一模型可以提高计算效率,促进 transfer learning,并捕获更全面的 user preferences 视图。
Unified Training:为了构建 the unified input sequence,我们在每个 watch 之前添加 discrete condition tokens (例如,<LFV>、<SFV>)以区分 content scenarios。由于 TokenMinds 在从 a pre-trained LLM 继承来的 a shared token space 中运行,textual signals 可以自然地与 SID-based watch tokens 交错。我们通过对 search queries 的每个 search query 之前添加 a <Search> token 来利用这一点,将它们按照时间顺序与 watch sequence 交织在一起(在实验章节进行了消融实验)。
示例:
<Search> NBA highlights<LFV> SID_A1 SID_A2 SID_A3 SID_A4 soft_A<SFV> SID_B1 SID_B2 SID_B3 SID_B4 soft_B<Search> cooking recipe<LFV> SID_C1 SID_C2 SID_C3 SID_C4 soft_C这样,在每个
search开头都有一个<Search> token;每个长视频开头都有一个<LFV> token;每个短视频开头都有一个<SFV> token。所有的search、长视频、短视频都按照发生的时刻进行排序,交织在一起。
对于 training objective,我们将 the random sampling of future target watches(见前面的章节)扩展为 uniformly include both SFVs and LFVs。Condition tokens 被排除在 loss 之外,因为预测它们过于简单,且若纳入其中会降低 SID prediction quality。
在跨场景统一训练中,未来
target watches的采样要均衡覆盖长视频和短视频,避免场景偏置。而
<LFV>、<SFV>这类场景条件tokens只作为输入条件,不计算损失。因为预测它们太简单、信息量低,纳入loss会分散模型容量,最终损害SID user token的预测质量。
Multi-Context Decoding:A unified model 应该在 serving 时为每个 context 生成 scenario-specific user tokens。一种朴素的方法需要为每个 context 单独进行 inference runs,冗余地处理 user history,抵消了 a unified model 的效率提升。
这里的意思是,分别以
<LFV> token和<SFV> token作为condition token,来解码生成结果。用户历史是同一份,但encoder被跑了两次。
我们提出 multi-context decoding 来消除这种冗余。给定 a shared encoder pass over the user history ,我们引入 a context stage,将 decoding 划分为并行的 sub-batches,每个 target context 一个 sub-batch。每个 sub-batch 以其各自的 condition token 来初始化,并通过 beam search 来独立地解码,而所有 sub-batches 完全地共享来自 a single encoder pass 的 encoder hidden states。如 Figure 3 所示,这使得 a single prefill pass 能够进行并发的、context-specific 的解码——从相同的 cached user representation 同时生成 LFV user tokens 和 SFV user tokens。
这里的解决方案是:
先做一次共享的
encoder pass。然后在
encoder输出之后、decoder解码之前,加一个context stage,把解码请求按目标场景分成多个并行的sub-batch:
sub-batch 1:目标场景LFV,用<LFV>初始化。
sub-batch 2:目标场景SFV,用<SFV>初始化。每个
sub-batch内部独立做beam search,生成该场景的多个SID序列。

将 SID-based user tokens 集成到下游模型——特别是需要 dense vectors 的传统 LEMs ——需要将 discrete tokens 投影到 a continuous embedding space 中。我们评估三种 token-to-embedding 的方法(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》):
(1):Prefix Embedding Mapping:将每个 predicted L-prefix SID 映射回 original content embeddings (这个 content embeddings 就是 SID 的 embedding)。由于多个视频可能 collapse 为相同的 coarse-grained prefix,我们对sharing the predicted prefix 的所有视频的 original embeddings 进行均值池化,以获得 SID embedding。
(2):N-gram Embedding:将每个 predicted SID sequence 分割为固定长度的 N-gram sub-words(例如,对于 codeword pairs),每个 N-gram sub-word 被映射到 a learned embedding。SID representation 是 the sum of its sub-word embeddings。
(3):SPM Embedding(《Subword regularization: Improving neural network translation models with multiple subword candidates》):类似于 N-gram,但使用 SentencePiece 从 item distributions 中学习 variable-length sub-words。
我们将 (2) 和 (3) 统称为 Learnable Embeddings: LEs,因为两者都使用随机初始化的 embedding tables,与下游模型一起被端到端地训练,与 (1) 的 static mapping 形成对比。
N-gram Embedding和SPM Embedding的效果要比Prefix Embedding Mapping更好。而N-gram Embedding和SPM Embedding之间效果差不多。
在 serving 时,beam search 为每个用户产生 SID sequences,每个 SID sequence 代表一个 predicted future interest。为了形成 a single user vector,我们通过池化(例如,attention-weighted 池化、均值池化、最大值池化、或top-k concatenation)聚合对应的 SID embeddings。所有策略产生相差无几的下游性能,表明收益主要来自 tokens 本身包含的信息,而非 aggregation choice。
这
个 SID embeddings代表了个兴趣。这里采用均值池化即可用于下游模型。
上述 encoder-derived dense embeddings 和 token-derived SID embeddings 都可以作为下游模型的 direct input features,或作为 cross-attention layers 中的 key-value pairs,其中 candidate items 对 user representations 进行关注(attend)。
为了克服十亿用户规模下的 high serving costs 和 latency,我们将 user representation generation 与 real-time down-stream inference 解耦,如 Figure 4 所示。基于 User Behavior Service: UBS 框架(《Short-form Video Needs Long-term Interests: An Industrial Solution for Serving Large User Sequence Models》)构建,TokenMinds 异步生成 user embeddings 和 user tokens,并将其缓存在快速的 key-value store 中。real-time models 在 scoring 期间检索这些 cached representations,无论 TokenMinds 的底层复杂性如何,都保持恒定的延迟和成本。当 a scoring request 到达时,客户端检查有效的、未过期的 representations:
如果可用,则获取它们并直接馈入下游模型。
否则,立即触发后台刷新:Representation Refresh Service 读取用户最新的 watch history,通过 exported TokenMinds model 运行 inference,并将 updated representations 写回缓存。

我们首先详细说明 TokenMinds 的实验设置,然后从离线指标和 downstream integration 的线上评估考察 recommendation quality。最后,我们分享来自 additional studies 的经验。我们尝试回答以下问题:
RQ1 (Token Adaptation & Viability):如何将 SID-based discrete user tokens 最优地适配用于下游 continuous models;并且当部署在工业规模的 recommendation surfaces 上时,它们是否能产生 measurable improvements ?
RQ2 (Complementary Values):TokenMinds 的 dual-output 性质是否通过同时生成 continuous embeddings 和 discrete tokens 来提供互补的价值?
RQ3 (Cross-Scenario Modeling Impact):统一 cross-scenario modeling 范式(联合训练 LFV 和 SFV)能否在不影响下游 recommendation quality 的情况下实现计算效率提升?
Model:TokenMinds 采用基于 Gemini V1.5 的 encoder-decoder 架构,包含一个 370M-parameter Mixture-of-Experts (MoE) encoder 和一个 370M-parameter dense decoder。两者都从 Continued Pre-Training (CPT) checkpoints(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)初始化。
Training:TokenMinds 在 YouTube LFV and SFV watch histories interleaved with textual search queries 的数据上训练,使用为 Gemini 生态系统开发的内部 JAX 和 Pathways 框架。我们采用 continuous training,在最新的 daily data 上,从而捕获 fresh engagement signals,每天处理数百万样本,这比通常需要 billions of interactions 的 traditional LEMs 更具样本效率。
对于每个 user sample,我们使用 most recent 1,200 watches interleaved with S=10 search queries,最大 input sequence length 为 1,024 tokens(在后续章节中进一步分析)。每个 watch 由一个 condition token、 prefix-L=4 SID tokens、和用于 non-SID features 的 M=1 soft token组成;using a single soft token 以大约 5% 的 offline recall 为代价,显著减少了 sequence length。Target selection 从 24-hour look-ahead window 中采样最多watches。通过对学习率在 grid search,我们发现 linear warmup with cosine decay and cyclic restarts(《Language models are few-shot learners》、《Sgdr: Stochastic gradient descent with warm restarts》)达到了 peak same-day performance,但在 continuous training 下,a constant learning rate 对 day-to-day distribution shifts 更为鲁棒。
Serving:TokenMinds 通过异步的 processing infra 具备 24-hour refresh cadence。对于每个用户:
我们从 encoder 提取 a 1,152-dimensional dense embedding。
并通过 beam search 来解码 B=40 SID sequences(20 sequences 用于 LFV,20 sequences 用于 SFV),所有 sequence 都使用 prefix length L 来形成 user’s discrete token representation。
本节中的所有模型都在相同数据上训练至收敛,以确保公平比较。
Token Accuracy and Training Ablations:我们评估 TokenMinds generated tokens 的预测准确性,并对关键 training design decisions 进行消融。我们在两种协议下评估 token accuracy:
Session Recall:使用近乎完整的 history input 并预测 final watch
Cold-Start Recall:使用 truncated history input 并预测从 future window watch。
所有变体都在两种协议下进行评估以确保可比性。在两种情况下,我们生成 top-10 SID sequences(prefix length L)并计算Recall@10。
Table 1 报告了结果以及三个消融实验,每个消融实验移除一项 training innovation,同时保持 evaluation 固定:
(1):w/o Multiple Targets:将 sampled targets 从 1。
(2):w/o Look-ahead Window:用标准 next-watch prediction( input target look-ahead sampling。
(3):w/o SID Truncation: trains and encodes inputs with full-length SIDs (prefix-L ;在评估时,predictions 仍在 prefix-L granularity 上进行比较以确保指标可比性。
所有三项都降低了准确性,其中 look-ahead removal、以及 full-length SIDs 对冷启动性能尤其有害(分别相对下降-10% 和-17%),确认了 temporal look-ahead sampling 和 coarse-grained prefixes 对于泛化到 immediate session 之外是必不可少的。

Initialization Strategy:为了验证 the choice of initializing from a CPT checkpoint(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》),我们与两个基线进行比较:a general-purpose Pre-Trained Gemini checkpoint (without SID-specific grounding)、以及随机初始化。
Table 2 显示了这些基线在 with and without search queries 的配置下的比较。我们观察到 CPT 始终优于 Pre-Trained Gemini,Pre-Trained Gemini 又优于随机初始化。这证实了在 CPT 期间 the semantic grounding of SIDs 对下游 task-specific fine-tuning 有益,超出了 base LLM 本身的 general sequence modeling 能力。

Impact of Search Queries:为了量化 search queries 的贡献,我们将 our full model 与 a variant trained without search signals进行比较。如 Table 2 所示:
添加 queries 在所有初始化策略下持续改善 recall,确认了 explicit textual intent 提供了超出 watch history 之外的补充信号。
既然
search queries有效,那么为什么只选择个而不是更多个 queries?两个原因:
最大输入长度只有
1,024 tokens,watches已经占了大头。所以 是在有限 tokens预算下,给search queries分配的一个合理配额。如果再增加search queries,每个query本身会tokenize成多个文本tokens,还会加<Search> token,就会挤掉watches。只取 “最近的”
search queries,更早的搜索时效性差。最近的10 searches已经能覆盖短期显式意图,再往前翻,边际信息价值下降。
此外,search 的收益通过更强的 initialization 被放大,表明 an LLM backbone with aligned SID representations 更能够利用 textual search signals 与 behavioral watch histories 相结合。
Token Diversity:在确立了预测准确性并确定了一些关键的驱动因素之后,我们接下来验证这种性能不是通过 redundant beams 来实现的。由于每个 decoded SID 作为下游模型的 a distinct interest signal,beam collapse ——多个 beams 收敛到近乎相同的 outputs ——将严重限制 representational capacity。
我们使用两个指标评估多样性(diversity):
SID Token Collision Rate:测量位置 token 的频率,high collision 表示 semantic beam collapse。
SID Prefix Duplication Rate:统计直到位置 prefixes 的数量,high repetition 表示 isolated exploration of a single branch。
我们在 5K users 的样本上评估这些指标,将 generated tokens 与 ground-truth watches from the same look-ahead 24-hour window 进行比较。为确保尽管规模差异仍能公平比较,我们对较大集合(ground-truth watches or generated SIDs)进行降采样以匹配 smaller set 的大小。为减少 subsampling 的方差,每次比较时,我们对 larger set 重复降采样 10 次,并报告平均 similarity 分数。different token indices 下的结果 Cumulative Distribution Function: CDF 绘制在 Figure 5 中,表明 TokenMinds 达到了与 ground-truth distribution 相当的 prediction diversity,确认了 beam search 产生 diverse interest signals 而不会坍缩为 redundant outputs。
实线是
ground-truth,虚线是ours。

Embedding Quality:我们在 2K random sampled users 上评估 embedding consistency。对于每个用户,我们从其 full history 生成 an embedding watches 生成一个扰动的变体 a random user’s embedding embeddings 在轻微 history 扰动下保持稳定,同时在用户间保持区分性。
我们通过线上 A/B 实验评估 TokenMinds 的现实世界影响,将用 user embeddings 和 user tokens 集成到多个 production ranking models 中。虽然 TokenMinds 已部署在 YouTube 的 ranking system 、retrieval system 和 LLM-based production systems 中,但本文侧重于 ranking integration。所有实验均在 a seven-day period 内进行。我们在 LFV 和SFV 中报告两个关键 quality metrics 上的性能:Engaged Users 和 Satisfied Engagement。在本小节,表格内的粗体数字表示在 95% 置信度下具有统计学显著性。
Token Adaptation Pivot Study:在投入大规模计算到 full-scale dual-output architecture 之前,我们使用一个轻量级的 110M model 在 SFV surface 上进行了线上转折研究,以确定最优的 token adaptation 策略。如 Table 3 所示,我们比较了静态的 Prefix Embedding Mapping: EM 与 Learnable Embeddings: LE ——具体来说是 Unigram (N = 1) 用于 SFV,匹配该 surface 上现有的 item tokenization。LE 始终优于 EM,确认了 allowing downstream models to learn a specialized embedding space 能产生更好的 task adaptation(《Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations》)。在 LFV platform 上使用 SPM-based LE 进行的 auxiliary evaluations 显示出相同的趋势。因此,我们在所有后续 full-scale evaluations 中专用地使用 LE,回答了RQ1 的前半部分。

Downstream Quality:为了评估 TokenMinds 的各自部分的影响和互补的影响,我们比较了单独添加 continuous embeddings、单独添加 discrete SID-based tokens(通过 LE)、以及同时添加两者。Table 4 报告了 primary SFV surfaces 和 primary LFV surfaces 的结果。我们得出两个关键发现。
首先,SID-based user tokens 在 production systems 中提供了 additive value,直接回答了 RQ1 的后半部分。除了 Table 4 中的 primary surfaces,仅在另外两个 LFV surfaces上的 token-only integration 在 Engaged Users 方面实现了 +0.04% / +0.16% 的统计显著提升,在 Satisfied Engagement 方面实现了 +0.07% / +0.11% 的提升,进一步确认了泛化性。
其次,同时部署 embeddings 和 tokens 产生了放大的增益,验证了 RQ2 关于这两种模态的互补性质、以及我们的 dual-output 架构。

Downstream Cost:对于 joint token and embedding generation,TokenMinds 大约需要 339ms per user,完全在后台处理中被吸收。每个用户的 discrete token representation 仅需 1,280 bytes——而 dense embedding 需要 4,608 bytes ——存储减少了 72%。在 serving 时,系统在来自多个 production surfaces 的 1.44M read requests per second 中达到了96.4% 的 cache hit rate。将 pre-computed embeddings 作为 features 所引入带来的计算成本可忽略不计;token integration 开销详见 Table 5。

Cross-Scenario Modeling:我们评估 the unified cross-scenario model 与两个基线:
a LFV-only model(仅在 LFV data 上训练)用于 quality comparison。
separately trained LFV and SFV models 用于 efficiency comparison。
从 Table 6(Panel B)来看,将 LFV 和 SFV 整合到单个框架中产生了可观的成本节约。
训练计算减少了 50%,因为一个模型取代了两个模型。
上游 serving 成本减少了 31%,因为 multi-context decoding 在场景间共享 a single encoder pass(481 chips),而不是 running two full model passes(总共 698 chips)。
下游 integration cost 保持中性,因为 the unified model 产生的 outputs 形状与 LFV-only baseline 相同。
在质量方面,Panel A 显示用 unified representations 替换 LFV-only baseline 持续改善了 freshness 指标(Fresh Engagement 衡量用户与新视频的 interaction),而没有降低 core engagement 指标。值得注意的是,the unified model 在相同的 fixed-length input sequence 中处理 LFV watches 和 SFV watches,几乎将可用的 LFV history 减半。尽管有这种减少,core metrics 仍保持中性这一事实表明,来自 SFV interactions 的 cross-scenario signals 补偿了 reduced LFV context ——这确认了 the unified paradigm 在实现显著效率提升的同时不损害质量。
LFV-only模型中,inputs仅仅包含LFV视频,因此在相同input length下,LFV视频数量相比Cross-Scenario Model翻倍。

为了优化计算效率,这些研究采用 an accelerated offline evaluation protocol(《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》):模型在 7 randomly shuffled days 上训练,并在连续的第 8 天进行评估。
Model Variants & Capacity Allocation:我们探讨 architectural design 和 capacity allocation 对模型质量的影响。受限于可用的 pre-trained Gemini checkpoints,我们研究三种 warm-started 变体:
a Balanced MoE model:370M Enc / 370M MoE Dec。
a Balanced Dense model:370M Enc / 370M Dense Dec。
an Unbalanced Dense model :420M Enc / 110M Dense Dec。
我们围绕 Balanced Dense anchor 对学习率网格搜索 Balanced Dense anchor 的学习率除以 10 到乘以 10)。
Figure 6 显示了 LFV 的 Training 和 8th-Day Recall 与 Iso-FLOPS 的关系( SFV 因趋势相似而省略)。
虽然 the Unbalanced model 显示的 Training Recall 低于 the Balanced Dense baseline,但它达到了相差无几的 8th-Day Recall。这为异步 serving 策略打开了大门——将较重的 encoder 以较低频率刷新以获取稳定的 user embeddings,与较轻的 decoder 以较高频率刷新从而捕获 recent behaviors 相结合。
此外,在相差无几的 FLOPS 下,the Balanced MoE decoder 在 8th-Day Recall 上优于 Dense 对应物,确认了 sparse expert routing 在泛化到 future user behaviors 方面提供了切实的优势。

Input Length Scaling:虽然现代 user modeling 通常受益于 long-term watch histories(《DV365: Extremely Long User History Modeling at Instagram》),但开发 compute-optimal models 需要理解 scaling dynamics。我们分析了用户历史长度(history length: HL)的影响。如 Figure 7 所示:
LFV 和SFV 的 8th-Day Recall@10 在大约 1K watches 处开始饱和。
扩展到 2K watches 在 SFV 上产生相差无几或略微下降的性能。

Batch Size Scaling:我们评估了 4K、8K 和 16K 的 batch sizes,使用 8K 作为 anchor configuration 并进行了广泛调优好的学习率,并通过标准的《Language models are few-shot learners》)扩展到其他 batch sizes。相对于 4K baseline,8th-Day Recall@10 在 8K 时提高了 +2.5% / +5.5%(SFV/LFV),在 16K 时提高了 +7.6% / +13.7%。与 PLUM (《Plum: Adapting pre-trained language models for industrial-scale generative recommendations》)中的发现一致,较大的 batches 加速了收敛,并在相同的 training window 内产生了更强的 representations。这突显了在硬件容量允许的情况下最大化 batch size 对于 compute-optimal training 的重要性。
我们介绍了 TokenMinds ,一个通过从 pre-trained LLMs 改编的 a unified encoder-decoder architecture 来生成 discrete SID-based user tokens 和 dense user embeddings 的框架。Cross-scenario modeling 进一步将 LFV 和 SFV 统一在 a shared SID vocabulary 下,采用 multi-context decoding,大幅减少了上游计算。
作为 YouTube 上已部署的系统,TokenMinds 证明了 SID-based user tokens 是 production ranking systems 的可行的 representation。将这些 tokens 与 dense embeddings 相结合产生了放大的性能增益,确认了 dual-output design 的互补的价值。我们的消融实验进一步验证了 CPT initialization 的重要性、以及 search query signals 对 SID-based user modeling 的协同效益。
这些结果表明,离散的、语义上有据可依的 user representations 为 scaling user modeling 到超越 dense embeddings alone 的 representational constraints 之外提供了一条有希望的方向。