《IAT: Instance-As-Token Compression for Historical User Sequence Modeling in Industrial Recommender Systems》
尽管复杂的 sequence modeling 范式在推荐系统中取得了显著成功,但手工设计的 sequential features 的信息容量限制了性能上限。为了更好地通过 encoding historical interaction patterns 来提升用户体验,本文提出了一种新颖的两阶段 sequence modeling 框架,称为 Instance-As-Token: IAT。
IAT 的第一阶段将每个 historical interaction instance 的所有特征压缩成一个统一的 instance embedding,该 embedding 将 interaction characteristics 编码到一个紧凑但信息丰富的 token 。我们提出了时间顺序(temporal-order)和用户顺序(user-order)两种压缩方案,后者更能满足下游 sequence modeling 的需求。
第二阶段涉及下游任务通过时间戳(timestamps)来获取固定长度的 compressed instance tokens,并采用标准的 sequence modeling 方法来学习长期偏好模式(long-range preferences patterns)。
大量实验表明,IAT 显著优于 SOTA 的方法,并展现出优越的 in-domain、cross-domain 的可迁移性。IAT 已成功部署在真实的工业推荐系统中,包括电子商务广告、商场营销(shopping mall marketing)、以及直播电子商务,在关键业务指标上带来了显著提升。
在现代推荐系统中,modeling short-term and long-term interaction sequences 已被广泛证明可以提高推荐效果和用户满意度。Figure 1 左侧展示了 ranking models 的主流架构。现有研究主要聚焦于设计更复杂的 modeling 范式,为超长序列构建高效的架构,提出 sequential features 与 non-sequential features 之间增强 feature interaction 的机制,以及在推荐系统中验证 scaling laws 和 generative retrieval。

然而,这些努力主要关注于 high-level model design,仍然依赖于 low-level sequential feature engineering 方法。现有的 sequential features 通常源自一组手工设计的特征,如 Figure 1(b) 所示,包括 inherent item features(如 price、category)、user-item interaction features(如 interaction type、frequency)和 context features(如 timestamp)。
一方面,存储、传输和计算方面的 resource constraints 使得大规模构建此类手工特征变得困难,细粒度特征被排除在外,导致 sparse feature representations 和建模性能下降。
另一方面,sequential feature engineering——包括特征设计、开发和验证——通常涉及长周期,导致 feature expansion 的高昂的工业成本(industrial costs)。
这两个方面限制了 historical interaction sequences 中的信息密度,并进一步阻碍了现代架构的 effective scaling。
为解决上述问题并更好地促进 upper interaction modules 的学习过程,我们提出了一种侧重于 enhancing the information capacity of sequences 的新颖的 sequence modeling 框架。其动机是利用一个稠密的且信息丰富的 embedding 来表示 a specific historical interaction,而不是依赖于手工设计的特征,其中手工设计的特征携带 sparse and limited information,如 Figure 1 (c) 所示。现代工业推荐系统中的 training instances 可以全面地描述 historical interaction patterns,通常包含数千个特征。我们的目标是将 historical training instances of a user 压缩成 dense representations,并将这些 representations 用作下游 sequence modeling 的 tokens。该框架被命名为 Instance-As-Token: IAT,由以下两个阶段组成。
IAT Compression:此阶段通过特定的 Compression 机制为每个 training instance 生成紧凑的 instance embedding: InsEmb。作为一个直观的解决方案,我们通过将 compression layer 和 decompression layer 应用于 a base model 的 final feature representation 来训练一个 temporal-order source model,并将 intermediate compressed representations 存储为 InsEmb。
除此之外,我们进一步提出了 user-order source model,该模型按用户顺序来组织 training instances,并引入了一个 Source Instance Transformer: SIT 模块来完成 sequence modeling 过程。SIT 显著提升了 source model 的性能,同时为 generated InsEmb 赋予 sequence modeling capability,因此在应用于下游模型时展现出卓越的 performance transferability。
一些用于描述 historical training instances 的补充特征(例如,multi-task labels )可以选择性地与 generated InsEmb 一起存储。
IAT Sequence Modeling:此阶段在下游模型中检索 stored informative InsEmb 和可选的 key features,这些特征被聚合为 Instance Tokens: InsToken。为了避免未来信息泄露,retrieved InsTokens 根据 request timestamp 被严格截断。然后,可以采用现代架构,如 LONGER 或 Transformer 进行 IAT sequence modeling。得益于 informative InsTokens,下游模型在 offline evaluations 和 online A/B tests 中都取得了显著的性能提升。
总之,我们的主要贡献有四个方面:
(a):新颖的 Sequence Feature Engineering。所提出的 IAT 提出了一种新颖的 sequence feature engineering 机制,以取代低效且效果不佳的手工特征。
(b):独特且合适的 IAT Compression 方案。我们提出了 temporal-order 和 user-order 两种生成 InsEmb 的压缩方案,后者与下游模型更好地对齐。
(c):完整且详细的 IAT 流程。我们从模型架构、streaming training pipeline 、以及 storage deployment 流程的角度详细阐述了 IAT 的流程。
(d):显著的性能提升。大量的离线实验验证了所提出的 IAT 框架的优势,在多个场景下的真实世界工业推荐系统在引入 IAT 后,在线指标均取得了显著提升。
核心思路就是:
先用一个
base model来进行sequence modeling,这一步跟常规方法完全相同。然后,把第一步每个
(user, timestamp, target item)得到的final representation抽取出来,作为downstream model中该用户的行为序列中,对应(user, timestamp, historical item, action_type)的item representation。
user-order source model很强大,但是它无法上线。因此,IAT的作用就是把source model的能力压缩、传递到可在线部署的下游模型中。牺牲一部分离线AUC增益,换取在线可行性和成本可控,这正是两阶段IAT框架的设计目的。但是这种做法需要存储IAT序列,引入了infra的复杂度。
Sequence Modeling 范式:behavior sequence modeling 能够在推荐系统中捕获动态偏好。现有研究主要集中在几个关键方向,以提高 sequence modeling 的性能、效率和可扩展性。
《TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders》 、《Wukong: Towards a scaling law for large-scale recommendation》、《Rankmixer: Scaling up ranking models in industrial recommenders》 的工作增强了 large-scale ranking models 中的 heterogeneous feature interaction。
而 《HyFormer: Revisiting the Roles of Sequence Modeling and Feature Interaction in CTR Prediction》、《OneTrans: Unified Feature Interaction and Sequence Modeling with One Transformer in Industrial Recommender》 的工作则实现了 sequential features 和 non-sequential features 之间的充分的 information integration。
《VQL: An End-to-End Context-Aware Vector Quantization Attention for Ultra-Long User Behavior Modeling》、《Twin v2: Scaling ultra-long user behavior sequence modeling for enhanced ctr prediction at kuaishou》、《Deep Multiple Quantization Network on Long Behavior Sequence for Click-Through Rate Prediction》、《HiSAC: Hierarchical Sparse Activation Compression for Ultra-long Sequence Modeling in Recommenders》 的研究旨在通过高效架构设计(如 clustering 和 sparse attention 技术 《Native sparse attention: Hardware-aligned and natively trainable sparse attention》)处理更长的 user sequences。
在推荐系统中探索 scaling laws(《Scaling laws for neural language models》)、《Scaling recommender transformers to one billion parameters》、《Climber: Toward efficient scaling laws for large recommendation models》)以及推进 generative recommendation(《Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment》、《Mtgr: Industrial- scale generative recommendation framework in meituan》、《Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations》、《OneMall: One Model, More Scenarios–End-to-End Generative Recommender Family at Kuaishou E-Commerce》)的发展也是热门的研究方向。
然而,这些研究未能考虑 traditional hand-crafted sequential features 有限的信息容量。
Sequence Feature Engineering:low-level sequential item features 所覆盖的信息容量决定了 high-level recommendation model 性能的上限。
早期的 sequential recommendation 工作仅利用 ID features(《Self-attentive sequential recommendation》),而后续研究指出其他 sequential features 有助于捕获细粒度的偏好(《Feature-level deeper self-attention network for sequential recommendation》)。
一些研究进一步提出了融合 ID features 之外辅助信息的高级方法(《Sequential modeling with multiple
attributes for watchlist recommendation in e-commerce》、《Aligned side information fusion method for sequential recommendation》、《Decoupled side information fusion for sequential recommendation》)。
利用 items 的多模态信号(multimodal signals)可以有效提高推荐系统的泛化能力,包括 multimodal embedding(《LEMUR: Large scale End-to-end MUltimodal Recommendation》、《MUSE: A Simple Yet Effective Multimodal Search-Based Framework for Lifelong User Interest Modeling》)和 semantic IDs(《Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment》、《Forge: Forming semantic identifiers for generative retrieval in industrial datasets》)。
然而,这些方法要么仍然依赖于手工特征,要么引入过高的开销并遭受信息错位(information misalignment)的问题。相比之下,所提出的 IAT 保留了所有特征中的关键信息,同时实现了与下游任务更好的对齐。
Two-Stage Ranking Frameworks:在 resource usage 和 online latency 的约束下,two-stage ranking framework 可以在模型部署中实现更灵活的 scaling 。
最近的工作(《Exploring Scaling Laws of CTR Model for Online Performance Improvement》、《External Large Foundation Model: How to Efficiently Serve Trillions of Parameters for Online Ads Recommendation》)提出构建 large teacher models 作为 foundation models,并利用它们的 distillation signals 来增强 small-capacity student models 的性能。
HLLM(《Hllm: Enhancing sequential recommendations via hierarchical large language models for item and user modeling》)采用了 a two-tier LLM model,其中第一个 item LLM 提取 item 的丰富 content features ,第二个 user LLM 在这些 item features 上完成 sequence modeling。
LLaTTE(《LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation》)指出 semantic features 是 scaling 的前提,并引入了一个两阶段架构,包括一个生成并缓存 compressed user embeddings 的上游 user model。
我们提出的 IAT 也采用了两阶段训练方法,其中第一阶段生成稠密的且 informative 的 instance embedding,第二阶段利用这些 embeddings 实现更好的 sequence modeling。
我们首先介绍一个流行的 ranking model 的基本架构作为 base model。我们使用小写字母表示单个 instance 的数据,大写字母表示 a batch的数据。features 通常包括 sequential features 和 non-sequential features,分别表示为 sequential features 通常涵盖 historical interaction behaviors,即 sequential features 首先由一个 sequence modeling architecture non-sequential features 组合并进行 deeper feature interaction,即:
其中:feature representation。
click-through rate: CTR 或 conversion rate: CVR )如下:
其中 item
该模型使用二元交叉熵(binary cross-entropy: BCE)loss 进行训练:
其中: training instances 集合,且
现有研究侧重于提高 scaling 能力,而忽略了由手工设计的 information bottleneck)。如 Figure 1 (c) 所示,一个用户的 training instance 通过数千个特征来完整地描述了 historical interactions,这促使我们提出了 Instance-As-Token: IAT ——一种如 Figure 2 所示的新颖的 sequence modeling 框架。
IAT 包含两个阶段:
(1):第一阶段(IAT Compression)。我们将两种类型的 compression modules 应用于 base model,获得 temporal-order source model 或 user-order source model。source models 处理 training instances 并生成紧凑的 instance embeddings: InsEmb,这些 instance embeddings 被存储在一个 centralized repository 中。此阶段将 rich-feature instances 压缩为低维 embeddings ,实现了全面实例信息(comprehensive instance information)的高效存储和传输。
(2):第二阶段(IAT Sequence Modeling)。downstream models 通过将 InsEmb sequence 与补充的关键辅助信息(complementary key side information)相结合来构建 instance tokens: InsToken 从而进行 sequential modeling 。
接下来的部分将详细说明如何基于 1.2.1 节介绍的 ranking model 来构建 source models 和 downstream models 。

此阶段构建 source models 以获得 compressed instance representations。我们专注于 key component 从而用于 InsEmb generation,即两种不同的压缩方案。
为了生成 an instance 的紧凑 representation,一个直观的解决方案是:简单地利用通过 complex feature interaction 所获得的 representation ,即前面公式中定义的 feature representation。直接存储这样的 representation 不可避免地会带来过高的存储开销和传输开销。因此,我们在 base ranking model 中添加了一个 compression layer 和 decompression layer。假设原始维度为 compression dimension 为 compression module 表示如下:
其中:
ReLU、GELU。
feature representation,而 InsEmb。InsEmb 的维度远小于原始维度,即 InsEmb 通过一个额外的轻量级 MLP 进行投影以保留原始维度:
其中:
然后,decompressed feature 参与 final prediction 和 loss optimization,确保 compressed InsEmb(即 high-value information 。该架构如 Figure 2 (a) 所示。
值得注意的是,上述 temporal-order source model 是逐个压缩 training instances 的,得到的 InsEmb 只能包含单个 training instance 的信息。此外,与 base model 相比,compression module 可能会略微降低模型性能,这可能会损害所得到的 InsEmb 的质量。为了生成更高质量的 InsEmb ,我们设计了 user-order source model,该模型引入了一个 Source Instance Transformer: SIT,并使每个 instance 能够在避免未来信息泄露的同时感知同一用户的 historical behaviors。
User-Order Training Instances:与 temporal-order source model 不同,user-order source model 需要重新组织training instances,以加速 source model training 和 InsEmb generation。通常,a ranking model 的 training instances 是根据 all user behaviors 的时间戳来存储和训练的。然而,这导致同一用户的 historical behavior sequence items 产生大量的冗余存储和计算。在 batch training stage ,一些近期工作提出根据 user-order instances 来训练模型(《Make It Long, Keep It Fast: End-to-End 10k-Sequence Modeling at Billion Scale on Douyin》、《Actions speak louder than words: Trillion-parameter sequential transducers for generative recommendations》),这显著提高了效率而没有性能下降。具体来说,我们在特定时间范围内按 user ID 来聚合 training instances,然后根据时间顺序对每个用户的 instances 进行排序。在实际操作中,实际的 batching 策略需要根据 users’ historical training instances 的数量将其分割成固定长度的 batches,并可选地填充 dummy instances。但是,为了公式简洁,我们在下文中假设一个 batch 仅包含单个用户的所有 historical training instances。
Source Instance Transformer:假设给定的用户有 historical training instances,我们将它们组织在同一个 batch 中,并按前面的公式获得 compressed representations,记为 training instances 的 batch size。然后,我们提出 Source Instance Transformer: SIT 来增强 representation 能力,如下所示:
其中:
transformer 架构。
Reshape 操作将整个 batch 的 compressed representations(即 items 的单一序列。这 items 由带有 causal mask function transformer module 来处理,确保每个 compressed InsEmb 只能访问 past instances 的信息。
例如,如 Figure 2 (b) 左侧所示,user instance 1 是最新的 training instance,它可以访问来自所有过去 training instances 的信息流。然后,所得到的 final prediction 和 loss optimization。值得注意的是,我们保存 SIT 之前的 compressed embedding (即 InsEmb,而不是保存
与 temporal-order source model 相比,user-order source model 具有以下优点:
增强 Source Model 的性能。SIT 使得同一用户内的 instances 之间能够进行信息流动,这可以显著提高 base model 的性能,如实验研究所示。然而,temporal-order source model 与 base model 相比可能会经历轻微的性能下降。
增强 InsEmb 的 Sequence Modeling 能力。user-order source model 不仅生成紧凑且信息丰富的 InsEmb,而且在 SIT 的帮助下隐式地提升了 InsEmb 的 sequence modeling 能力。相比之下,temporal-order source model 是为每个 training instance 而单独压缩 representations,从而失去了 user-order source model 的效果。
我们采用了一种训练范式,它同时纳入了 batch training stage 和 stream training stage ,后者在 real-time training instances 上运行。如 Figure 3 所示:
temporal-order source model 逐个压缩 training instances,这与 training instances 的组织方式无关,在 batch training stage 和 stream training stage 保持一致。然而,如前面章节所述,user-order source model 在 batch training 期间依赖于 user-order training instances。在 streaming training 期间,the batch of instances 通常来自不同用户,构建 user-order instances 较为困难。因此,我们需要在 streaming stage 稍微修改 user-order source model 的训练方法。
具体来说,模型架构保持不变,但 batching strategy 不同。我们将 a streaming batch 的 compressed representations 记为 batch size。这 instances 来自不同的用户。为了获得 SIT 所需的 historical InsEmb,我们直接从 store 中为每个用户检索之前的 instances 的 InsEmb,记为 SIT 计算如下:
其中:
stop gradient 操作。
SIT 的 concatenated inputs。
最后,我们利用 decompression layer 和后续 pipelines 保持不变。总之,在 streaming batch 中,只有 instances 的 InsEmb(即 InsEmb 是从 store 中获取的,如 Figure 3 (b) 所示。

在本节中,我们将详细介绍如何利用 stored InsEmb 来促进下游任务。除了 source models 产生的信息丰富的 InsEmb 之外,一些 core side information 也可以被存储并用于 sequence modeling。
Core Side Information:尽管 InsEmb 包含了数千个特征的 compressed information,但它并未显式地涵盖 training instance labels 、时间戳和其他 task-related information 。此模块根据实际场景可选地使用。side information 通过常规 feature engineering 方法来处理,其 embedding representation 记为 batch size,
这是因为,在
IAT Compression阶段,每个target item的label是模型未知的,且需要被预测的。但是在IAT Sequence Modeling阶段,每个historical item的action_type是已知的(对应于item label)。
InsEmb Adaptation Mechanism:stored InsEmb 保持在低维且紧凑的空间中,应投影到高维空间以实现更好的 scaling 和 feature space matching。因此,我们在下游模型中引入了一个 adaptation MLP,从而将 InsEmb 与下游 feature space 进行对齐:
其中:
adaptation MLP 的参数。
store 中获取的历史 InsEmb,其维度为
为什么要从
store中获取历史的个 InsEmb?因为每个InsEmb对应于一个target item,而每个target item又对应于用户的一个historical item。
Instance-As-Token (IAT) Sequence:我们通过拼接 InsEmb 和 core side information 来构建 IAT sequence,如下所示:
其中 sequence modeling 的 IAT sequence embedding 。
Query Token Construction:IAT sequence 是纯 user features,而现代 sequence modeling 方法需要一个 query token,该 token 包含 candidate items 信息以进行 matching prediction。为了更好地与紧凑的 InsEmb 对齐,我们拼接下游模型的 sequential feature tokens 和 non-sequential feature tokens,并将其压缩为 query,即:
其中:所获得的query
其他的 query token construction方法在实验中进行了研究。
就是 。
IAT Sequence Modeling:前面公式中获得的 query token IAT sequence 将首先被投影到相同的维度空间,然后可以根据场景需求,在下游模型中由任何主流的 sequence modeling network 进行建模,包括但不限于 LONGER(《Longer: Scaling up long sequence modeling in industrial recommenders》)或 Transformer 。我们还在实验章节中进行了一些实验研究,以验证 modeling results 在 feature interaction 中的合适位置。
在实践中,我们发现将 modeling results 作为 input token 集成到 upper complex feature interaction(例如,RankMixer 《Rankmixer: Scaling up ranking models in industrial recommenders》、TokenMixer-Large 《TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders》)中可以获得最佳性能,如 Figure 2 (c) 所示。
在 streaming training 阶段,source model 和 down-stream model 都采用实时流式更新。source model 地独立更新,最新的 InsEmb 实时同步到下游模型,确保模型预测的时效性。
Figure 4 详细说明了 IAT 的存储架构,包括以下步骤:
(1):source model 处理 training instances 并生成相应的 InsEmb,然后以 InsID 为 key 将其写入到 embedding parameter server: Emb PS。
(2):同时,feature extraction module 将 (UID, InsID, Side Information) 元组输出到一个 message queue,该队列被异步保存到高性能的 key-value storage system。UID 指的是 hashed user ID。存储系统按时间戳对 InsID sequence 进行排序,并将其截断为长度 historical instance sequence。
(3):当请求到来时,下游模型首先通过 the key of UID 从高性能 key-value storage system 中检索 the sequence of InsID and Side Info。该序列根据 request timestamp 和一些任务相关规则进行截断。
(4):通过 the key of InsID list 从 Emb PS 进一步检索 InsEmb,然后按照公式 IAT sequence,用于下游的 sequence modeling。

InsID 的构建:InsID 是 training instance 的唯一标识符,可以通过 hashing a combination of timestamps and task-related IDs 来构建,(例如,在线广告推荐系统中的 request ID 和 creative ID)。
PS 存储成本估算:PS 中 InsEmb 的存储成本是可量化和可控的。使用 FP32(每个元素 4 bytes)来存储 InsEmb,总存储需求计算如下:
其中:
training instances 的数量。
historical instances 的保留天数。
InsEmb 维度。
此公式同时考虑了每日 training instances 数量和历史数据保留周期,使得存储估算更符合实际工业场景。例如,存储约数亿 samples per day、64 维 InsEmb、为期两年的数据,大约需要几十 TB 的存储空间。
数据集:实验是在从真实世界广告系统收集的工业级 CVR prediction dataset 上进行的。该数据集包含连续 4 个月的训练数据, training instances 总数量达数百亿,代表了广告行业真实的大规模推荐场景。特征空间包括 user features、item features、contextual features 和 traditional behavior sequence features,其中 sequence features 仅保留 sparse core side information(例如,ID、category、action type)。所有训练数据通过移除敏感的 user/item information 和哈希化 feature IDs 进行匿名化处理,确保模型训练和评估期间没有隐私泄露的风险。
baseline 方法:base model 架构符合 Problem Statement 章节的介绍,同时我们为 traditional behavior sequence 选择了三种主流的工业 sequence modeling 方法(即 DIN,LONGER,Transformer),以全面验证 IAT sequence 的有效性和灵活性。
对于所有模型的 feature interaction,我们统一采用 RankMixer 作为 feature interaction module,以消除由不同 feature fusion 策略引起的差异。
对于每个 base model,我们训练相应的 source model 和 downstream model with two kinds of IAT sequence(即 Temporal-Order IAT 和 User-Order IAT),报告它们相对于相应 base model(仅具有 traditional hand-crafted behavior sequence)的 AUC 增益。
具体而言,所有实验均在具有数百个节点的分布式 GPU 集群上使用 batch size
评估指标:为了全面评估模型的有效性和效率,我们报告性能指标和模型效率指标。
我们采用 AUC 和 LogLoss 作为性能指标,这些指标在工业 CVR prediction 任务中被广泛使用。
同时,我们报告 dense parameters 和 training FLOPs 作为主要的效率指标。
我们主要报告 downstream models 的性能,在某些情况下也提供 source models 的指标。
超参数:
在 temporal-order source model 中,feature representation 的原始维度(即 6,000,压缩后的维度为
user-order source model 的压缩后的维度相同,即 SIT 采用 2 layers of transformer,模型维度为 64,intermediate size of FFN 为 128。
读者猜测,
应该也是 6000。
在 user-order source model 的 batch training 阶段,SIT 的 input length 与 batch size 相同,如公式
而在 streaming training 阶段,我们将 SIT 的长度限制为 256 ,即
对于 source models,我们保存 64 维的 compressed representation 作为 InsEmb,如 Figure 2 所示。downstream model 获取 InsEmb 并将其适配到维度
我们对公式
在下面的不同实验研究中,某些设置会略有变化。此外,traditional behavior sequence 的长度固定为 512。
我们首先通过将所提出的 IAT 集成到三种主流的工业 sequential modeling 架构(即 DIN, LONGER 和 Transformer)中,全面验证了其有效性和灵活性。我们将这些架构应用于建模 base model 中的 traditional behavior sequences,而仅在下游模型中使用 Transformer 进行 IAT sequence 建模。在真实场景中,这三种选择是基于可用计算资源和约束条件选择的。
我们分别训练了 temporal-order model 和 user-order model 以生成两种类型的 InsEmb。
temporal-order source model 由于引入了 compression 模块,AUC 略有下降 0.05%。
同时,user-order source model 获得了 0.6% 的 AUC 增益,因为引入的 SIT 增强了 historical user instances 之间的信息流动。
这里没有图表来给出结果?
然而,user-order source model 不能直接部署用于 serving,因为它会产生过高的资源成本,并且 online operation 面临实际限制。
为什么
user-order source model不能直接部署用于serving?因为在线服务路径上会引入用户级长序列计算和大量历史状态访问。user-order source model比temporal-order多了Source Instance Transformer (SIT)。SIT的输入不是一个实例,而是同一用户的历史实例序列,形状类似B × T × D。论文中流式训练时T=256、B=1024、D=64。Transformer注意力复杂度随序列长度平方增长,在线QPS下,每个请求都要做这种用户级序列前向,计算量和显存都很难承受。
利用这两种类型的 InsEmb,我们探索了它们提高下游模型性能的能力。下游模型的性能和效率结果总结在 Table 1 中,其中 "+Temporal/User IAT" 示我们在 base model 中添加了由 temporal-order source model 或 user-order source model 生成的 IAT sequence。
在所有三种序列架构中,集成 IAT sequence 的模型始终优于其对应的 base model。尽管 temporal-order source model 本身经历轻微的性能下降,但对下游模型 AUC 的提升高达 0.15%,这验证了 IAT sequence modeling 的优势。
在真实场景中,0.1% 的 AUC gain 被认为是显著的。根据我们的经验,直观比较,向 base model 中引入长度为 256 的手工序列(hand-crafted sequence)很难获得 0.1% 的 AUC gain。
带有 user-order IAT 的模型始终获得 AUC 的相对提升(高达 0.31%)和 LogLoss 的降低(高达 -0.67%),同时仅伴随模型参数和 FLOPs 的适度增加。这些实验表明,IAT sequence 在各种 sequential modeling 范式中都是有效的。
这里的
base model不是source model,而是下游排序模型,它包含传统的序列特征,但不包含IAT序列。“
user-order source model获得了0.6%的AUC增益“,这里的base model就是Transformer-based Base模型。它的效果反而最好。但是,由于它无法上线,因此,IAT的作用就是把source model的能力压缩、传递到可在线部署的下游模型中。牺牲一部分离线AUC增益,换取在线可行性和成本可控,这正是两阶段IAT框架的设计目的。

我们观察到,将 IAT sequence modeling 架构与 SIT in the user-order source model 进行对齐可以获得更好的结果。具体来说,我们固定 base model 中的 traditional behavior sequence 由 Transformer 建模,并使用 DIN、LONGER 和 Transformer 分别对 IAT sequence 进行建模。
如 Table 2 所示,LONGER 和 Transformer 带来了显著的 AUC 提升和LogLoss 降低,而使用 DIN 对 IAT 进行建模仅带来微弱的增益。这表明 modeling IAT by series of transformer architectures 可能更为合适。
实际上,LONGER 利用 Perceiver(《Perceiver: General perception with iterative attention》)技术和其他一些模块(例如,global token、token merge )来减少繁重的计算负担,这个负担是因为 Transformer 对 long sequence 应用 full attention 而带来的。在随后的实验中,除非另有说明, traditional behavior sequences 和 IAT sequences 均使用 Transformer 架构进行建模。

我们进一步展示了对 user-order IAT 的 AUC gain 的一些分析。
如 Figure 5 (a) 所示,大约 80% 的用户拥有的 historical training instances 少于 256 个。这就是我们将下游模型的 IAT sequence 最大长度设置为 256 的原因。
Figure 5 (b)显示了长度小于 256 的 IAT sequence 的实际长度分布。
如 Figure 5 (c) 所示,如果我们按 IAT sequence 长度对 training instances 进行分组,可以发现较长的 IAT sequence 会带来更显著的 AUC gain,最多约为 0.4 %。

Scaling Up the Base Model:所提出的 IAT 框架可以普遍应用于各种类型的 base model。随着先进的 scaling 方法(《TokenMixer-Large: Scaling Up Large Ranking Models in Industrial Recommenders》、《Scaling recommender transformers to one billion parameters》)的发展,我们持续更新工业场景中的 online base model,引入更多 dense parameters(Dense Scaling),延长 historical behavior sequence 的长度(Seq. Len. Scaling),并利用更复杂的 feature interaction 方法进行 scaling up(Feat. Int. Scaling)。结果列于 Table 3。
总体而言,带有 user-order IAT sequence 的下游模型获得了显著的性能提升,特别是当模型 scale 到 1B 参数时,即使微小的增益也变得更加困难和有价值。

Scaling Up the Source Model:然后我们探索 user-order source model 中 SIT 的 scaling up。我们用 SIT 的 input dimension、SIT 的 transformer layers 层数、以及下游模型中最终使用的 InsEmb 的 compressed dimension 。相对于 base model 的 AUC gain 绘制在 Figure 6中。
增大 InsEmb 的维度、或增加 transformer layers 层数为 source models 带来了显著的增益,即 AUC gain 从 0.52% 提升到 0.80%。下游模型的性能也分别得到了明显提升。

Scaling Up the Downstream Model:对于下游模型中的 user-order IAT sequence modeling,我们针对不同的 maximum IAT sequence length(即 32、64、128、256 和 512)探索了 transformer layers 层数(即 1、2 和 4),并绘制了 AUC gain 相对于 maximum IAT length 或 FLOPs per batch 的 scaling law。下游模型 transformer architecture 的 scaling 符合标准的 scaling law,曲线绘制在 Figure 7中。值得注意的是,随着层数增加,下游模型展现出更好的 scaling 能力。

然后,我们在 Table 4 中对 IAT 的某些设置进行了消融研究,包括 source models 和 downstream models 的设置。
对于 source models ,我们首先将 batch size 1024 改为 2048,这对 temporal-order IAT: T-IAT 没有影响,但略微提升了 user-order IAT: U-IAT 的性能,因为在 user-order batch training stage,一个 batch 可以覆盖每个用户的更多 historical instances。
保存 SIT 模块之后的 InsEmb 而非之前的 InsEmb 会降低下游性能,因为 SIT 之后的 InsEmb 由于 attention 机制丢失了一些基本信息并变得更加相似。
IAT sequence 包含 InsEmb 和一些 core side information: sideinfo,例如表征 behavior actions 的 multitask labels 、以及其他一些 task-related side information。我们在实践中使用了 5 种其他类型的 side information 。
丢弃 label side information 会导致明显的 AUC 下降,因为 InsEmb 包含了关于 labels of training instances的少量信息。
丢弃 other sideinfo 对 temporal-order IAT 有不良影响,但对 user-order IAT 的性能没有影响。
我们也尝试使用由 some ID features of a candidate 构建的更简单的 query,而不是 compressing the sequential and non-sequential tokens(公式 AUC 分别下降了 0.06% 和 0.09%。
最后,我们观察到 IAT sequence modeling 在下游模型中的位置也很重要,将其放置在 feature interaction module(例如,RankMixer)之前,相比较于之后,可以获得更好的结果。
消融研究的更多细节列在附录中。

我们通过在真实世界广告场景上同时训练 temporal-order source model 和 user-order source model 来生成 IAT sequence。然后我们将所生成的 IAT sequence 应用于广告场景本身以及其他几个工业推荐场景。
In-Domain A/B Results:将 IAT sequence 应用于广告场景本身展示了所提出 IAT 的领域内可迁移性(in-domain transferability)。核心收入指标,包括 Advertiser Score: ADSS 和 Advertiser Value: ADVV,报告在 Table 5 中。base model 是一个已经在线上服务了很长时间的强大的广告排序模型。显然,IAT 在这些指标上带来了统计上显著的提升。

Cross-Domain A/B Results:当将所生成的 IAT sequence 应用于其他下游场景时,IAT 也展现出了强大的泛化能力,包括商城广告的 CVR prediction、信息流广告的 CTR prediction 和 CVR prediction、非闭环广告(non-closed-loop advertising)的 CVR prediction以及直播电商的 CT-CVR prediction。广告的核心指标包括 ADSS 和 ADVV,而电商场景则报告 GMV。如 Table 6 所列,IAT 在所有场景中均取得了显著的提升。
这里没有说明是
User-Order IAT还是Temporal-Order IAT。读者猜测应该是User-Order IAT。

我们提出了一种名为 Instance-As-Token: IAT 的新型两阶段 sequence modeling 框架,取代了传统的手工 sequence engineering。
IAT 的第一阶段构建一个 temporal-order source model 或 user-order source model,将每个用户 historical instances 的丰富特征压缩成紧凑且信息丰富的 embeddings (即 InsEmb)。
然后,下游阶段利用这些 InsEmb 进行 SOTA 的 sequence modeling。
大量的离线和在线实验证明了所提出的 IAT 在提高模型质量和业务关键指标方面的优越性。所提出的 IAT 已在多个工业推荐系统中全面部署。未来的工作将进一步提高 IAT 的效率,包括:
(1):探索单阶段框架以降低训练复杂度。
(2):采用更先进的压缩技术以进一步减少存储和传输开销。
我们详细说明正文中介绍的消融研究如下。
在我们的场景中,除了 InsEmb 之外,我们还使用一些 side information,形成用于 IAT sequence modeling 的 InsToken。side information 通常包含两个部分:
Label Side Information:用户的 historical training instance 的 multi-task labels,例如,用户是否购买了 clicked item。这些 labels 与 user behaviors 高度相关,而 InsEmb 仅压缩 bottom feature representations。因此,我们将 label side information 视为重要特征,可以无缝地与 InsEmb 结合,以实现更好的 sequence modeling。
Other Task-Related Side Information:在我们的广告场景中,timestamps 的一些类型对于 sequence modeling 也很重要,这些类型能够描述 a user’s behavior 的归因逻辑(attribution bution logic),我们也考虑了这些 side information。
side information 根据实际场景是可选的,side information 可以由 ID features 来表示,并在模型中通过 embedding tables 进行建模。在我们的模型中,仅存储和使用此类 side information 的 hashed values。
然后,我们展示了 IAT sequence modeling的另外两种变体,它们分别研究了 query construction 和 IAT position 的重要性。这两种选择如 Figure 8 所示。然而,这两种架构都遭受明显的性能下降。
首先,因为 IAT sequence 中的 InsToken 包含信息丰富的信息,因此最好使用也涵盖 informative features of the candidate 的 query token。
其次,将 output token 放置在复杂 feature interaction module 之前可以导致各种 bottom tokens 之间更全面的信息流动,这更有利。

在流式训练阶段,输入到 Source Instance Transformer: SIT 的是一个序列,其中最近的元素是当前实例的 InsEmb(由 compression layer 来生成),其余元素是属于同一用户的 historical instances 的I InsEmb。该序列按顺序排列,当前实例的InsEmb 作为第一个元素,随后是历史实例(按逆时间顺序)。SIT 严格遵循因果约束,以确保每个元素只能访问过去实例的InsEmb。正文仅提供了一种获取 historical InsEmb 的方法,我们将在下面详细介绍另一种方法及其优缺点。
为了获得 SIT 所需的 historical InsEmb,我们提出了两种具有不同 trade-offs 的实用方法:
Method 1: Store Retrieval:直接从 PS 中检索用户之前的 InsEmb (其中
优点:计算成本低,速度快,适用于流式训练场景。
缺点:依赖于过去 training iterations 所生成的 historical InsEmb,这可能与最新的模型参数不一致。
Method 2: Recomputation:检索之前的 source model 的前向传播重新计算其 InsEmb。
优点:使用最新的模型参数重新计算 InsEmb,消除了陈旧历史值的差异,确保了高的建模准确性。具体来说,只允许对历史 InsEmb 进行前向传播而不进行梯度反向传播,而当前 InsEmb 保持正常的梯度流。
缺点:显著增加了计算成本(例如,当 256 倍),仅适用于具有足够计算资源或延迟要求较低的场景。
鉴于 Method 2 由于 historical InsEmb 的重新计算而引入了过高的计算开销,我们在实践中采用 Method 1(Store Retrieval)作为默认方法。
最后,我们考虑一种更复杂的训练方法来增强 user-order source model 的性能。因为在 user-order source model 的 batch training 阶段,training instances 是按用户排序的,earliest trained users 可能产生较弱的 InsEmb,因为 source model 可能尚未收敛,而 finally trained users’ InsEmb 可能受益于收敛且更好的 source model。也就是说,one-pass user-order training 可能导致不同用户之间的不公平。因此,我们建议在两个独立的 jobs 中训练 user-order source model,并在这两个 jobs 中采用完全相反的顺序。Figure 9 提供了说明。
一个模型按正常的用户顺序处理 training instances,但仅存储 the InsEmb of the later user IDs。
相比之下,另一个模型按相反的用户顺序处理 training instances,也仅存储 the InsEmb of the later user IDs。
我们还进行了一些实验研究,以验证这种 two-pass 范式的优势。one-pass user-order source model 本身获得了约 0.6% 的 AUC gain,而所提出的 two-pass user-order source model 进一步获得了高达 0.8% 的 AUC gain。
如果用单模型训练两个
epoches,仍然只能让一半用户在第二个epoch的训练后期生成InsEmb,另一半用户在第二个epoch的前期生成,质量仍然不均匀。
