一、ASIF [2024]

《Aligned Side Information Fusion Method for Sequential Recommendation》

  1. 将 items 的上下文信息(即 side information)与 ID 相结合,已成为提高推荐系统性能的重要途径。现有的 self-attention-based side information fusion方法可分为 early fusion、late fusion 和 hybrid fusion。在实践中:

    • naive early fusion 可能会干扰 the representation of IDs,导致负面影响。

    • 而 late fusion 则缺少 effective interactions between IDs and side information。

    • 人们提出一些 hybrid 方法来解决这些问题,但它们仅在计算 attention scores 时使用 side information,这可能导致 information loss。

    为了在没有噪声干扰(noisy interference)的情况下充分发挥 side information 的潜力,我们提出了一种用于 sequential recommendation 的 Aligned Side Information Fusion: ASIF 方法,由两个部分组成:Fused Attention with Untied Positions 与 Representation Alignment。具体来说:

    • 我们首先解耦 positions,以排除 attention scores 中的噪声干扰。

    • 其次,我们采用 contrastive objective 来保持 ID 和 side information 之间的语义一致性(semantic consistency),然后使用正交分解(orthogonal decomposition)来提取同质部分(homogeneous parts)。

    通过 aligning the representations and fusing them together,ASIF 在不干扰 ID 的情况下充分利用了 side information。在四个数据集上的离线实验结果证明了ASIF 的优越性。此外,我们成功地将该模型部署在 Alipay 的广告系统中,在 clicks 和 Cost Per Mille: CPM 上分别获得了 1.09% 和 1.86% 的提升。

  2. Sequential recommendation 在电子商务、广告和搜索系统等工业场景中发挥着重要作用,其主要目标是对用户的 historical behavior 进行建模,以预测用户可能感兴趣的 next item。在各种解决方案中,attention-based models 由于出色的性能逐渐成为主流。

    早期的 self-attention-based models 如 BERT4Rec(《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》)和 SASRec(《Self-attentive sequential recommendation》)只考虑 item IDs,缺乏捕获 IDs 之外 item attributes 的能力。当 IDs 频繁变化时,这种局限性变得明显。例如,在 Alipay 会员页面的典型推荐场景中,向用户展示可使用积分和金钱兑换的 items。商品池会随着广告活动而频繁更新,导致 item IDs 快速变化。categories 和 brands 等 Attributes 提供了用户 long-term preferences的更稳定的 representation。因此,我们旨在将 side information 纳入推荐模型以提升性能。

    基于 fusion locations 的不同,现有的 self-attention-based side information fusion 方法可分为三类: early fusion、late fusion 和 hybrid fusion。

    • early fusion 在将 ID 与 side information 馈入 attention block 之前将它们融合在一起。

    • 相比之下,late fusion 在 item-level sequences 和 feature-level sequences 上分别应用分离的 self-attention blocks,并在 final stage 才进行融合。

    已有研究指出,early fusion 并不总是能提升性能,反而可能损害 ID 的 representation,导致一种称为信息入侵(information invasion)的现象(《Noninvasive self-attention for side information fusion in sequential recommendation》)。另一方面,late fusion 缺少了 ID 与 side information 之间的交互,并丢失了一些先验信息(prior information)。因此,最近出现了一些 hybrid-fusion 方法。它们通过在 attention score calculation 中纳入 side information 来避免信息入侵,并探索了 attention correlations 的有趣结构。

    尽管取得了显著改进,这些方法仍然存在两个局限:

    • (1):ID 和 attributes 之间的 correlations 可能不同,有些强有些弱,这使得难以消除干扰并有效学习有意义的 correlations。

    • (2):完全将 side information 排除在 final representation 之外以防止信息入侵的方法,可能会无意中丢弃 side information 本身内含的关键信息。

    在这项工作中,我们尝试通过减轻噪声干扰来增强 the utilization of side information。受 《Rethinking Positional Encoding in Language Pre-training》 启发,我们扩展了 early-fusion 方法 SASRecF 的 the fusion form of attention scores。如 Figure 1 所示,ID 与 attributes 有强关联,而 position encoding 与 others 之间的 correlations 相对较弱。这表明,将 position 作为常规 side information 来融合的常见方式,可能会给 attention scores 引入噪声。我们还检查了 Yelp 数据集上 SASRecF 中 ID 和 side information 的 representation spaces,为 information invasion 提供解释。

    • 从宏观角度来看,我们可以观察到两个分布之间存在显著差异(见 Figure 2(a)),表明融合后的 representation space 将显著偏离 original ID space。

    • 从微观角度来看,通过同时将 ID embeddings 和 side information embeddings 投影到坐标系上,我们发现如果两者在某些轴上的方向相反,这些向量段可能会相互抵消,导致信息丢失(见 Figure 2(b))。

    为解决上述问题,我们提出了一种名为 Aligned Side Information Fusion: ASIF的新方法。

    • 首先,我们引入了 Fused Attention with Untied Positions,它在计算 attention score 时将 ID-attributes 与 position encoding 分离,消除 noise interference 并保留 the strong correlation。

    • 其次,我们提出了 Representation Alignment,包括两个步骤:Representation Space Alignment: RSA 和 Homogeneous Information Extraction: HIE。

      • RSA 方法对序列内 interaction 粒度上的 paired ID and attribute 采用 contrastive objective,以确保它们的语义一致性(semantic consistency)。

      • 尽管 RSA 操作使两个分布更加接近,但仍然无法避免异质部分(heterogeneous parts)的存在。因此,HIE 对 ID 和 side information 执行正交分解(orthogonal decomposition)以提取同质部分(homogeneous parts),从而避免信息入侵。

    总之,我们的主要贡献可总结如下:

    • 我们精心设计了 ASIF 框架,基于 Fused Attention With Untied Position 和 Representation Alignment,通过利用 side information 来提升推荐性能。

    • 在 Representation Alignment 方面,我们提出了 RSA 和 HIE。通过使用 contrastive loss 和 orthogonal decomposition,我们在宏观和微观两个方面对齐了 ID 和 side information 的 representation space,有效防止了信息入侵问题。

    • 离线和在线实验证明了我们提出方法的有效性。

    主要技术要点:

    • 修改了 attention 计算,其中计算两个 attention matrix,三次更新:

      HX0=X,HA0=F(X,A),HP0=PCXA(l)=HA(l)Wq,1(l)(HA(l)Wk,1(l))⊤,CP(l)=HP(l)Wq,2(l)(HP0Wk,2(l))⊤HX(l+1)=Softmax(CXA(l)+CP(l)dh)HX(l)Wv,1(l)HA(l+1)=Softmax(CXA(l)dh)HA(l)Wv,2(l)HP(l+1)=Softmax(CP(l)dh)P(l)Wv,3(l)
    • 考虑了item IDs 的 embedding 矩阵和 attributes 的 embedding 矩阵之间的对比损失。

    • 考虑将 attribute representation 投影到 HX 的那一部分 HA∗ 更新到 HX 上面。

      HX←HX+HA∗

1.1 相关工作

  1. Sequential Recommendation:Sequential recommendation 旨在基于用户的 historical behaviors 来预测最有可能被交互的 next item。随着近年来深度学习技术的发展,许多基于神经网络的方法开始涌现,如基于卷积神经网络(CNN)的模型(《Personalized top-n sequential recommendation via convolutional sequence embedding》、《A simple convolutional generative network for next item recommendation》)、基于循环神经网络(RNN)的模型(《Personalizing session-based recommendations with hierarchical recurrent neural networks》)、基于图神经网络(GNN)的模型(《Sequential recommendation with graph neural networks》)、和 attention-based models。其中, self-attention-based 的方法取得了显著进展。

    • SASRec (《Self-attentive sequential recommendation》)将 self-attention 引入 sequential recommendation 模型以捕获 long-range dependencies。

    • BERT4Rec(《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》)采用完形填空目标(Cloze objective),并通过双向自注意力机制提升了性能。

    • 最近的 sequential recommendation 方法也使用 contrastive learning 来增强数据,包括 CL4SRec (《Contrastive learning for sequential recommendation》)和 DuoRec (《Contrastive learning for representation degeneration problem in sequential recommendation》)。

    这些工作仅利用 item ID,忽略了与 item 关联的其他 attributes,而这些 attributes 可能有助于提取 comprehensive sequence patterns。

  2. Side Information Fusion for Sequential Recommendation:不同于仅使用 item IDs 的上述方案,side information(如其他 item attributes 和 ratings)被纳入考虑以捕获有意义的 supervision signals。 S3Rec (《S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization》)注意到 attributes 中包含的重要信息,并设计了四个辅助的 self-supervised tasks 来学习内在关系(intrinsic relationship)。除了在 auxiliary tasks 中利用 side information 外,端到端的 fusion 方法也开始被探索。

    按照 multi-modal fusion 的分类体系(《Multimodal fusion for multimedia analysis: a survey》、《Multi-modal machine learning: A survey and taxonomy》),我们将 self-attention-based side information fusion 方法分为三类: early fusion、late fusion 和 hybrid fusion。

    • 在 early fusion中,ID 和 side information 在模型的 shallow layers 进行组合,然后馈入到网络中并生成 outputs 。例如,SASRecF (《S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization》)将 ID 和 attributes 组合起来并作为 input 馈入到 self-attention block 中(见 Figure 3(a))。

    • 在 late fusion 中,ID 和 side information 的网络是独立的,fusion 发生在 predict layer 之前。FDSA (《Feature-level Deeper Self-Attention Network for Sequential Recommendation》)是 late-fusion 方法,它在 item-level sequences 和 feature-level sequences 分别上应用分离的 self-attention blocks,并在最后阶段才拼接它们的 hidden states(见 Figure 3(b))。

    early fusion 和 late fusion 都有各自的局限。前者无法排除噪声干扰并可能导致信息入侵(information invasion),而后者缺乏 ID 与 attributes 之间的有效交互。Hybrid fusion 介于两者之间,允许 ID 和 side information 在中间层进行交互。

    • NOVA (《Noninvasive self-attention for side information fusion in sequential recommendation》)首次定义了 naive early fusion 导致的 information invasion 问题,并提出仅在 attention scores 的计算中纳入 attributes 来缓解它(见 Figure 3(c) )。然而,它将 position 视为常规的 attribute,给 mixed attention 引入了噪声。

    • 此外,DIF-SR (《Decoupled side information fusion for sequential recommendation》)将 IDs and side information 的 attention scores 解耦,允许 higher-rank attention matrices 和灵活的梯度(见 Figure 3(d))。不幸的是,它放弃了 ID 和 attributes 之间的隐式 cross-relationships。

    这两种方法都仅在 attention scores 中使用 side information,在 value matrices 中完全丢弃它,这可能导致信息损失。我们的工作旨在填补这些空白,在减少噪声干扰的同时增强 side information 的利用。

1.2 方法

  1. ASIF 的整体框架如 Figure 4 所示,接下来将介绍细节。

1.2.1 问题定义

  1. 在 sequential recommendation with side information 中,设 U、V、X 和 Aj 分别表示 user set、item set 、item ID set 和 set of the j-th type of attributes。设 Su=[vu(1),vu(2),⋯,vu(n)] 表示用户 u∈U 按时间顺序构成的 historical sequence of interactions,其中 vu(t)∈V 是 user interaction sequence 中的第 t 个 item,n 是序列的最大长度。假设我们有 m 种 side information,则 vu(t)={xu(t),a1,u(t),a2,u(t),⋯,am,u(t)},其中 xu(t)∈X 是第 t 次交互的 item ID,aj,u(t)∈Aj 表示第 t 次交互的第 j 类属性。

    给 interaction history Su,sequential recommendation 的目标是预测用户 u 可能感兴趣的 next item。它可以形式化为:对于用户 u ,建模该用户在 all candidate items 上的概率:

    P(vu(n+1)=v∣Su)

    .

1.2.2 Fused Attention with Untied Positions

  1. 对于 attention-based models,纳入 side information 的朴素方式是将其与 item ID 融合并馈入到 attention block 中(见 Figure 3(a))。NOVA 遵循此结构但将 side information 从 value matrix 中排除(见 Figure 3(c)),而 DIF-SR 建议对多种 side information representations 和 IDs representations 采用 decoupled attention calculation (见 Figure 3(d)),以确保灵活的梯度。

    然而,根据 Figure 1 ,ID 与 attributes 有强关联性(strong relationship);但 position encoding 与 IDs and attributes 的关联性较弱,这可能使得 position encoding 不适合与其他项进行融合。因此,我们提出了 Fused Attention (FA) with Untied Positions (UP)(见 Figure 3(e))。

  2. 设 X 和 A 分别表示 item IDs 的 embedding matrices 和 item attributes 的 embedding matrices ,我们首先将它们融合并计算 correlation matrix:

    CXA=F(X,A)Wq,1Wk,1⊤F(X,A)⊤∈Rn×n

    其中:

    • Wq,1∈Rd×dh,Wk,1∈Rd×dh。

    • F 表示 fusion 函数,例如:Fsum(X,A)=X+∑j=1mAj∈Rn×d。

    接下来,我们计算 position encoding 的 correlation matrix 为:

    CP=PWq,2Wk,2⊤P⊤∈Rn×n

    其中:

    • P∈Rn×d 表示 absolute position embedding matrix 。

    • Wq,2∈Rd×dh,Wk,2∈Rd×dh。

    然后,我们融合两个 correlation matrices 并得到最终的注意力公式如下:

    HX=FusedAttention(X,A1,⋯,Am,P)=Softmax(CXA+CPdh)XWv,1∈Rn×d

    其中:Wv,1∈Rd×d,HX表示 item IDs 的 hidden state 。

    这里的 correlation matrices 通过加法来融合。

    最后,我们认为 side information 足够重要,值得充分学习,因此我们也在不同的 Transformer layers 之间传递和更新它们,如下所示:

    HA=FusedAttention(X,A1,⋯,Am)=Softmax(CXAdh)F(X,A)Wv,2HP=FusedAttention(P)=Softmax(CPdh)PWv,3

    其中:

    • Wv,2∈Rd×d,Wv,3∈Rd×d 。

    • HA 和 HP 分别表示 attributes 和 hidden states 和 positions 的 hidden states 。

    F(X,A) 为 Item Fused Representation、P 为 position representation,它们作为第一层的 HA0,HP0 。实际上:

    HX0=X,HA0=F(X,A),HP0=PCXA(l)=HA(l)Wq,1(l)(HA(l)Wk,1(l))⊤,CP(l)=HP(l)Wq,2(l)(HP0Wk,2(l))⊤HX(l+1)=Softmax(CXA(l)+CP(l)dh)HX(l)Wv,1(l)HA(l+1)=Softmax(CXA(l)dh)HA(l)Wv,2(l)HP(l+1)=Softmax(CP(l)dh)P(l)Wv,3(l)

    .

1.2.3 Representation Alignment

  1. 从宏观和微观角度来看,入侵现象(invasion phenomenon)的发生可能分别源于过度的分布偏差(distribution deviation)和向量偏移(vector offset)。为解决此问题,我们提出了 Representation Space Alignment: RSA 和 Homogeneous Information Extraction: HIE 来对齐 the representations of IDs and attributes。

    • RSA 的目标是缩小 item IDs 和 attributes 这两者 representation space 的并集,以提高 interaction 粒度上的 semantic consistency(见 Figure 5(a))。

      即,尽可能让这二者的 representation space 重叠起来。

    • HIE 提取 attributes 中与 item IDs 同质的信息,并将其融合到 IDs representation 中(见 Figure 5(b))。

  2. Representation Space Alignment: RSA:受 CLIP 的 alignment operation (Learning transferable visual models from natural language supervision》)的启发,我们利用对比损失(contrastive loss)来对齐 item IDs 和 attributes 的 embedding spaces,旨在使两个分布更加接近(见 Figure 5(a))。然而,与 CLIP 不同的是,我们的 alignment 发生在序列内的 interaction 粒度上,而不是样本粒度上。具体来说,X 和 A 表示 item IDs 的 embedding 矩阵和 attributes 的 embedding 矩阵:

    X=[x→(1)x→(2)⋮x→(n)]∈Rn×d,A=∑j=1mAj=[a→(1)a→(2)⋮a→(n)]∈Rn×d

    其中:x→(t),a→(t)∈R1×d表示序列中第 t 次交互的 item ID embdding 和 attribute embedding。

    接下来,我们计算两个 embeddings 之间的余弦相似度以得到 final matching scores 如下:

    X~=[x→(1)/‖x→(1)‖x→(2)/‖x→(2)‖⋮x→(n)/‖x→(n)‖],A~=[a→(1)/‖a→(1)‖a→(2)/‖a→(2)‖⋮a→(n)/‖a→(n)‖]Y^X=Softmax(X~A~⊤/τ),Y^A=Softmax(A~X~⊤/τ)

    其中:Softmax(⋅)对相似度矩阵的每一行执行,τ 表示可学习的温度系数。

    最后,我们计算 contrastive loss,形式如下:

    Lco=−12N∑i=1N(Yi⊙log⁡Y^Xi+Yi⊙log⁡Y^Ai)

    其中:

    • ⊙ 是逐元素乘积,N 是样本量。

    • Yi 是第 i 个样本的 ground truth,它是一个单位矩阵 Yi=In=diag(1,1,…,1),意味着只有 paired item IDs and attributes 是正例。

    这里的 Lco 是仅仅在 embedding table 上进行,而不是在每一层的 id representation HX 和 attribute representation HA 上进行。

  3. Homogeneous Information Extraction: HIE:space alignment 使两个分布更接近,但仍然无法避免 heterogeneous part 的存在。因此,我们提出对每一层的 hidden states 进行正交分解以提取 homogeneous parts(见 Figure 5(b))。

    直观上,如果 an attribute's representation 与 an ID's representation 的方向相同,则应最大程度地保留它。否则,可能存在冲突,应将其丢弃。因此,我们需要一个 r 维正交坐标系作为 comparison 粒度,该坐标系需要充分容纳 a user's interaction sequence 中所有的 IDs' representations。

    具体来说,我们首先对 ID 的 hidden state 进行 QR 分解:

    HX⊤=QR

    其中:Q∈Rd×n 是正交矩阵,R∈Rn×n 是上三角矩阵。

    然后,我们将两个 hidden states 都映射到 Q 中以获得坐标矩阵(coordinate matrices),如下所示:

    Proj(HX)=HXQ,Proj(HA)=HAQ

    其中:Proj(HX),Proj(HA)∈Rn×n。

    因此我们可以得到 homogeneous part HA∗∈Rn×d如下:

    sign(HX,HA)=ϕ(Proj(HX)⊙Proj(HA))Proj~(HA)=sign(HX,HA)⊙Proj(HA)HA∗=Proj~(HA)Q⊤

    其中:

    • ⊙是逐元素乘积。

    • ϕ(⋅) 是指示函数(indicator function),如果值大于 0 则输出 1,否则输出 0。

    为什么不考虑 negative 部分,即:ϕ(⋅) 当值大于 0 时输出 1、小于 0 时输出 -1?这是因为 HIE 的目标不是 “完整保留属性信息”,而是只抽取属性中与 ID 同质的、同向的那部分信息,用来补充 ID representation。因此 negative 分量按定义就是异质的或冲突的分量,应该被丢弃,而不是翻转为 -1 后继续融合。

    由于 HA∗ 与 HX 是同质的,我们可以直接将其融合到 item representation 中,公式 HX=FusedAttention(X,A1,⋯,Am,P) 可以更新为:

    HX=FusedAttention(X,A1,⋯,Am,P)+HA∗=Softmax(CXA+CPdh)XWv,1+HA∗

    由于用户的平均序列长度通常小于 n,我们可以将 HX⊤ 的维度降为 HX⊤Wr,其中 Wr∈Rn×r ,以减少计算复杂度、以及 QR 分解前的参数冗余。

    QR 分解的原始复杂度 O(dn2),但 n=50 很小;论文还通过降维到 d×r(r≃16)将复杂度降至 O(dr2)。在线部署仅增加 2ms 延迟。

    此外,QR 分解在数学上大多数是可导的,现代深度学习框架也支持其反向传播:只要 HX 的列向量线性无关(即序列中不同位置的 representation 不共线)。

1.3 Model Prediction and Learning

  1. 经过 L 层 Transformer 结构后,我们得到 item ID 的 final hidden state HXL,并计算 prediction score 为:

    Y^=Softmax(HXLV⊤)∈Rn×|V|

    其中:V∈R|V|×d 是 candidate item matrix。

    HXL 给出了每个 position 的 item representation,那么 Softmax(HXLV) 得到的是每个位置的 prediction。在训练期间,这就是自回归训练:训练时对多个历史位置分别做 next-item 监督。但是,在推理时只用最后一个位置的 prediction。

    对于序列推荐任务,我们采用交叉熵损失函数:

    Lce=−1N∑i=1Nyilog⁡y^i

    其中:yi 和 y^i 分别表示第 i 个样本的 ground truth 和预测概率。

    最后,结合 RSA 中的 contrastive loss,我们定义带 balance coefficient λ 的损失函数:

    L=Lce+λ×Lco=−1N∑i=1N(yilog⁡y^i+λ2∑(Yi⊙log⁡Y^Xi+Yi⊙log⁡Y^Ai))

    .

1.4 离线实验

  1. 在本节中,设计了离线实验来评估 ASIF 的性能和有效性。

  2. 数据集:我们在三个公开数据集和一个工业数据集上进行实验:

    • Yelp 数据集是一个知名的商业推荐数据集。Category of business 和 position 被视为 side information。

    • Amazon Beauty 数据集收集自 Amazon 评论数据集。Category of the goods 和 position 信息作为补充的属性。

    • AliEC 是 Alibaba 提供的 Taobao 展示广告数据集。我们使用 category 和 position 作为 side information。

    • Industrial dataset 收集自 Alipay 商业广告系统中的一个场景,经过脱敏和加密处理,不包含任何个人可识别信息(Personal Identifiable Information: PII)。position 以及 item’s entity(如 category 和 brand)被用作 side information。

    遵循相同的数据预处理方式(《Self-attentive sequential recommendation》、《Decoupled side information fusion for sequential recommendation》、《S3-rec: Self-supervised learning for sequential recommendation with mutual information maximization》),我们移除公开数据集中出现次数少于五次的全部 items 和 users。对于 industrial dataset,由于 item IDs 频繁更新,我们保留所有出现过的 users 和 items。所有处理后数据集的统计信息汇总于 table 1。

  3. Baseline 方法:我们将我们的模型与以下 SOTA 的序列推荐方法进行比较。

    • Methods without side information:我们采用如下模型作为基线:

      • GRU-based 的模型 GRU4Rec(《Session-based recommendations with recurrent neural networks》)。

      • self-attention-base 的模型 BERT4Rec(《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》)、SASRec(《Self-attentive sequential recommendation》)、LightSANs(《Lighter and better: low-rank decomposed self-attention networks for next-item recommendation》)。

      • CNN-based 的模型 Caser(《Personalized top-n sequential recommendation via convolutional sequence embedding》)。

      • MLP-based 的模型 FMLP(《Filter-enhanced MLP is all you need for sequential recommendation》)。

    • Naive early-fusion methods:GRU4RecF 、SASRecF、LightSANsF、FMLPF 分别是 GRU4Rec、SASRec、LightSANs 和 FMLP 的 naive early-fusion 变体,它们在馈入网络之前先将 ID 和 side information 融合在一起。

    • Advanced self-attention-based side information fusion methods:我们包括 late-fusion 方法 FDSA (《Feature-level Deeper Self-Attention Network for Sequential Recommendation》)、hybrid-fusion 方法 NOVA (《Noninvasive self-attention for side information fusion in sequential recommendation》)和 DIF-SR(《Decoupled side information fusion for sequential recommendation》),它们与我们的工作高度相关。为了公平比较,我们按照(《Decoupled side information fusion for sequential recommendation》)中的方式基于SASRec 实现 NOVA。

    • 其他相关方法:CL4SRec(《Contrastive learning for sequential recommendation》)和 DuoRec(《Contrastive learning for representation degeneration problem in sequential recommendation》)是使用 contrastive learning objectives 的序列推荐模型。

  4. 评估指标:遵循先前工作(《Self-attentive sequential recommendation》、《Decoupled side information fusion for sequential recommendation》),采用 leave-one-out 策略进行评估。对于每个用户序列,我们使用 the last item 进行测试,the second last item 进行验证,其余 items 用于训练。模型以 full ranking manner 进行评估,如(《Lighter and better: low-rank decomposed self-attention networks for next-item recommendation》、《Noninvasive self-attention for side information fusion in sequential recommendation》、《Decoupled side information fusion for sequential recommendation》)中所做,而非negative sampling ,后者常因 bias 而受批评(《A Case Study on Sampling Strategies for Evaluating Neural Sequential Item Recommendation Models》、《On sampled metrics for item recommendation》)。我们采用两个广泛使用的指标:top-K Hit Rate: HR@K 和 top-K Normalized Discounted Cumulative Gain: NDCG@K ,其中 K = {10, 20}。

  5. 实现细节:我们在开源推荐框架 Recbole (《Recbole: Towards a unified, comprehensive and efficient framework for recommendation algorithms》)上运行所有模型,并使用相同的设置进行评估。

    • 我们将所有数据集的最大序列长度设为 50 ,embedding size 设为 256。

    • 所有网络都是 3 层、4 heads,采用 Adam 优化器,训练 200 epochs,batch size = 2048,学习率为 1e-4。

    • side information fusion methods 的 fusion 函数在 sum、concat 和 gating 中搜索。

    • 对于其他超参数,我们遵循先前论文中提到的最佳设置。

1.4.1 性能比较

  1. 整体性能:Table 2 和 Table 3 报告了三个公开数据集和一个工业数据集的整体性能。我们可以从四个方面观察到:

    • (1):符合直觉的是,一些 fusion 方法比仅使用 ID 的方法表现更好,表明 side information 可以通过捕获更好的 sequence patterns 来提升模型性能。这强调了 side information fusion 工作的重要性。

    • (2):相反,在普通自注意力框架下,SASRecF 考虑了更多种类的 side information,但在所有数据集上与 SASRec 相比都带来了显著的下降,表明 self-attention-based naive early-fusion 方法确实存在信息入侵(information invasion)。

    • (3):NOVA 和 DIF-SR 经过精心设计以缓解 invasion 现象,因此取得了比 SASRec 更好的结果。同时,我们注意到,由于FDSA 将 ID 和 feature 分成两个通道导致缺乏 interaction,其效果并不显著优于 NOVA 和 DIF-SR。

    • (4):可以清楚地看到,ASIF 在所有数据集上都取得了显著优于其他 SOTA 基线方法的结果。这些结果证明了 ASIF 在消除噪声干扰(noisy interference)、以及解决 side information fusion 中的 information invasion 问题方面的效率和有效性。

  2. 消融研究:我们通过消融研究分析了 ASIF 每个组件的有效性。Table 4 显示了 ASIF 及其消融版本在三个公开数据集上的性能。

    • w/o Representation Space Alignment (RSA):我们禁用 contrastive loss 以验证 RSA 的有效性。显著的下降意味着:适当地使两个空间更接近有助于缓解 invasion 现象,从而提升性能。

    • w/o Homogeneous Information Extraction (HIE):没有 HIE 组件,attributes 信息和 position 信息仅仅参与 attention scores 的计算,而不是直接集成到 the hidden state of item representation 中。在这种情况下,所有数据集上的指标均下降。

    • w/o Untied Positions (UP):此版本移除了独立的 position channel,并将 position 视为普通的 attribute。可以观察到,position encoding 与其他项之间的交互增加了噪声,导致性能下降。

    • w/o Fused Attention (FA):我们将 IDs and attributes 的 correlation calculation 解耦,即各自学习其自身的 correlation matrix。结果显示大多数指标下降。这意味着保留 IDs 与 attributes 之间的交叉性(intersectionality)是必要的。

      即移除 HA0=F(X,A) ,而是替代以各自独立更新的 X0=X,A0=A。

    在 Table 4 中,ASIF 的四个消融版本均显著优于 SASRecF。RSA 和 HIE 是 ASIF 中最有效的组件,证明 side information 中确实存在有效信息,应将其仔细地融入 item representation 中。

  3. 超参数研究:

    • loss balance 参数 λ 的影响:我们研究了超参数 λ 的影响,它控制 prediction loss 和 contrastive loss 的平衡。Figure 6 中报告了在三个数据集上 λ∈{0.1,0.5,1,5,10,100} 时的 NDCG@20 。

      • 对于 Beauty 数据集,性能在不同 λ 下表现稳健,方差很小。

      • 而在 Yelp 和 AliEC 数据集上,我们的 ASIF 在 λ=10 时取得最佳性能。

    • 正交基的数量 r 的影响:在三个公开数据集上,ASIF 随基数量 r∈{4,8,12,16,20,24} 变化的性能分别报告于 Figure 6 中。

      较大的基数量通常意味着更精细的正交分解粒度。然而,更精细的粒度并不总是意味着更好。如我们所见,三个公开数据集的最佳正交基数量大约在 16 到 24 之间。

    • fusion function F 的影响:我们比较了三种不同融合函数:Sum、Concat 和 Gate 的性能。Figure 7 展示了结果,显示采用所有三种融合函数的 ASIF 都优于文中所提的 SOTA 基线。这凸显了 ASIF 的稳健性和优越性。

      加法融合的效果更好。这里采用加法。

  4. 案例研究:为了对 ASIF 的可解释性进行更多讨论,我们使用与 Figure 1 中 SASRecF 相同的样本可视化了 ASIF 的 correlations。如 Figure 8 所示,ID-to-ID 、ID-to-Attribute、以及 Attribute-to-Attribute 的 correlations 显示出强模式,表明 ASIF 在排除 position encoding的噪声干扰后具有更好的捕获数据间关联的能力。

    此外,我们在 Figure 9(b) 中可视化了 Yelp 数据集上 ASIF 的 clustered embeddings。与 Figure 9(a) 中 DIF-SR 的和 Figure 2(a) 中 SASRecF 的相比,ID 和 side information 的 representation spaces 在对齐后更加接近;然后经过 Homogeneous Information Extraction 后,the homogeneous part 基本上与 ID representation 对齐了。

1.5 在线部署

  1. 在在线广告系统中,点击率(Click-Through Rate: CTR)prediction 任务是重要组成部分,负责预测用户点击 candidate items 的概率。Xlight 是 Alipay 中的流量平台,为小程序商家等提供广告服务。为了进一步验证所提模型 ASIF 的有效性,我们将其部署到如 Figure 10 所示的广告系统中。在 Alipay 的会员场景中,大多数广告是真实商品,以积分加现金的形式向用户销售。

    • 为了使 ASIF 充分发挥其在 side information fusion 方面的优越性,我们选择商品的类别和品牌作为该场景中被推荐 items 的 side information。

    • 对于离线训练,ASIF 收集过去 7 天内的 recent click samples 作为训练数据集。

    • 对于在线服务,当用户访问会员页面时,系统会从 online feature service platform 发起对 user’s historical behavior 的请求,截断长度为 50。

    ASIF 将对从广告池中检索到的部分广告估算 pCTR。在 Xlight 的实时竞价和排序系统中,每个广告将根据其 Effective Cost Per Mille: eCPM 进行排序,eCPM 基于 pCTR 和 bid 来估算。因此,CTR 的准确预估对 Xlight 平台至关重要。

  2. 由于工业限制,无法在线比较所有基线模型。因此,我们选择 SASRec 作为基线模型进行比较。经过两周的在线 A/B 测试,我们的模型将点击量提高了 1.09%,并在 Cost Per Mille: CPM 上实现了 1.86% 的显著增长。同时,它使多日 online AUC 提升了0.97%,且额外计算成本可忽略不计(p99 latency 2ms)。总之,结合离线评估,ASIF 在真实工业场景中表现出强大的性能。

1.6 结论

  1. 在本文中,我们提出了一种用于序列推荐中 side information fusion 的新方法 ASIF。我们的方法解决了 mixed embedding space 中噪声干扰(noisy interference)和信息入侵(information invasion)的挑战。具体来说:

    • 我们首先引入了 Fused Attention with Untied Positions,它单独计算 position correlations,以避免 mixed attention scores 中的噪声干扰。

    • 其次,我们提出了 Representation Alignment,包括 RSA 和 HIE,以解决信息入侵问题。

      • RSA 使用 contrastive objective 来对齐 IDs embedding space 和 attributes embedding space,以提高它们在 interaction level 上的语义一致性。

      • HIE 采用正交分解提取属性中的 homogeneous part,然后将其集成到 item representation 中,进一步增强 the utilization of side information。

    通过大量实验,我们证明了我们提出的方法在 side information fusion 方面超越了以往的方法,可视化和消融实验证明了其合理性。在 Alipay 广告系统上的在线 A/B 测试显示,ASIF 在点击量上获得了 1.09% 的提升,在 CPM 上获得了 1.86% 的提升。在未来的研究中,我们旨在进一步改进 denoising 技术,并探索自动化的方法来增强 the utilization for side information fusion。