一、NOVA [2021]

《Non-invasive Self-attention for Side Information Fusion in Sequential Recommendation》

  1. 序列推荐系统旨在从用户的 historical behaviors 中建模用户的不断变化的兴趣,从而进行定制化的 time-relevant 推荐。与传统模型相比,CNN 和 RNN 等深度学习方法在推荐任务中取得了显著进展。最近,BERT 框架也作为一种有前景的方法出现,受益于其处理序列数据时的 self-attention 机制。然而,原始 BERT 框架的一个局限是它只考虑自然语言 tokens 这一种输入来源。在 BERT 框架下利用多种类型的信息仍是一个开放问题。尽管如此,利用其他 side information,例如 item category or tag,来进行更全面的描述和更好的推荐,在直觉上是有吸引力的。在我们的试点实验中,我们发现直接将多种 side information 融合到 item embeddings 中的朴素方法通常带来很小甚至负面的效果。因此,在本文中,我们提出非侵入式自注意力机制(NOninVasive self-Attention: NOVA),以在 BERT 框架下有效利用 side information。NOVA 利用 side information 生成更好的注意力分布(attention distribution),而不是直接改变 item embeddings,后者可能导致信息淹没(information overwhelming)。我们在公开数据集和商业数据集上验证了 NOVA-BERT 模型,我们的方法可以稳定地优于 SOTA 的模型,且计算开销可忽略不计。

  2. 推荐系统旨在 model users’ profiles 以进行个性化推荐。它们设计起来具有挑战性,但具有商业价值。序列推荐任务是在给定用户 historical behaviors 的情况下,预测用户下一个会感兴趣的 item(《Deep learning-based sequential recommender systems: Concepts, algorithms, and evaluations》)。与 user-level 或 similarity-based 的静态方法相比,序列推荐系统还对用户不断变化的兴趣进行建模,因此被认为对 real applications 更具吸引力。

    最近,有一种将神经网络方法应用于序列推荐任务的趋势,例如 RNN 和 CNN 框架(《Deep learning-based sequential recommender systems: Concepts, algorithms, and evaluations》)。神经网络现在被广泛应用,并且通常表现强于传统模型,例如 (《The YouTube video recommendation system》)、(《Amazon. com recommendations: Item-to-item collaborative filtering》)、(《BPR: Bayesian personalized ranking from implicit feedback》)、(《Matrix factorization techniques for recommender systems》)、(《Factorizing personalized markov chains for next-basket recommendation》) 和(《Fusing similarity models with markov chains for sparse sequential recommendation》)。

    正如《Deep learning-based sequential recommender systems: Concepts, algorithms, and evaluations》 所述,在人工神经网络家族中,基于 Transformer 的模型(《Attention is all you need》)被认为在处理序列数据方面更先进,因为它们具有 self-attention 机制。像 《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》 、《Behavior sequence transformer for e-commerce recommendation in Alibaba》、《POG: Personalized Outfit Generation for Fashion Recommendation at Alibaba iFashion》 这样将 Transformer 应用于序列推荐任务的研究已经证明了它们优于 CNN 和 RNN 等其他框架(《Session-based recommendations with recurrent neural networks》)。BERT(《Bert: Pre-training of deep bidirectional transformers for language understanding》)是一种领先的双向自注意力 Transformer 模型,《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》 借助它达到了 SOTA 准确率,明显高于单向 Transformer 方法(《Improving language understanding by generative pre-training》、《Self-attentive sequential recommendation》)。

    尽管 BERT 框架已经在许多任务上取得了 SOTA 性能,但它尚未被系统地研究用于利用不同类型的 side information。直观上,除了 item IDs 之外,利用 side information,例如 ratings 和 item descriptions,来进行更全面的描述从而获得更高的预测准确率,是很有吸引力的。然而,BERT 框架最初被设计为只接受一种类型的输入(即 word IDs),这限制了 side information 的使用。通过试点实验,我们发现现有方法通常以侵入式方式利用 side information(Figure 1),但效果甚微。理论上,side information 应该通过提供更多数据而有益。尽管如此,设计能够有效利用额外信息的模型仍然具有挑战性。

    因此,在本研究中,我们研究如何在成功的 BERT 框架下高效利用各种 side information。我们提出一种新颖的非侵入式自注意力机制(Non-inVasive self-Attention: NOVA),它可以持续提高带有 side information 的预测准确率,并在我们所有实验中达到 SOTA 的性能。如 Figure 1 所示,使用 NOVA 时,side information 作为 self-attention 模块的辅助,用于学习更好的注意力分布(attention distribution),而不是被融合到 item representations 中,后者可能导致诸如信息淹没(information overwhelming)之类的副作用。我们在实验室数据集和一个从真实应用商店收集的私有数据集上验证了 NOVA-BERT 设计。结果证明了所提出方法的优越性。我们的三个主要贡献是:

    • 我们提出 NOVA-BERT 框架,它可以为序列推荐任务高效采用各种 side information。

    • 我们还提出非侵入式自注意力(non-invasive self-attention: NOVA)机制,这是一种新颖的设计,使 self-attention 能够用于复合的序列数据。

    • 进行了详细实验和部署以证明 NOVA-BERT 的有效性。我们还加入了可视化分析以获得更好的可解释性。

    核心思想:使用所有信息(包括 side information,item ID)来生成 attention matrix,但是 value projection 仅仅作用在 id representation 上。细节部分主要集中在如何生成 attention matrix。

1.1 相关工作

  1. 推荐系统的技术已经发展了很长时间。最近,研究人员倾向于在推荐系统领域采用神经网络,因为它们具有强大的能力(《Deep learning-based sequential recommender systems: Concepts, algorithms, and evaluations》)。此外,在 DNN 中,趋势也从 CNN 转向 RNN,然后转向 Transformer(《Attention is all you need》)。基于 Transformer 的模型之一 BERT(《Bert: Pre-training of deep bidirectional transformers for language understanding》)被认为是一种先进的序列神经网络,因为它具有双向自注意力机制。《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》 通过达到 SOTA 性能证明了 BERT 的优越性。

    在序列推荐领域,充分利用 side information 以提高准确率也是一个长期讨论的话题。如 Table 1 所示,在 CNN、RNN、注意力模型和 BERT 等不同框架下,若干先前工作尝试利用 side information。尽管如此,大多数工作在使用 side information 时,没有过多研究 side information 应该如何添加。几乎所有工作都使用一种我们称为 “侵入式” 融合('invasive' fusion)的简单实践。

  2. 侵入式方法(Invasive approaches):大多数先前工作,如 《Parallel recurrent neural network architectures for feature-rich session-based recommendations》,直接将 side information 融合到 item representations 中,如 Figure 1(a) 所示。它们通常使用 fusing 操作(例如 summation、concatenation、gated sum)将额外信息与 item ID 信息合并,然后将 mixture 馈入神经网络。我们称这种直接合并实践为 "invasive approaches",因为它们改变了 original representations。

    如 Table 1 所示,先前的 CNN 和 RNN 工作已经尝试通过 concatenation 和 addition 等操作直接将 side information 融合到 item embeddings 中来利用 side information。其他一些工作,如 GRU(《TiSSA: A Time Slice Self-Attention Approach for Modeling Sequential User Behaviors》)和 《Parallel recurrent neural network architectures for feature-rich session-based recommendations》,提出了更复杂的 feature fusion gate 机制和其他训练技巧,试图使 feature selection 成为可学习的过程。然而,根据他们的实验结果,简单方法无法在各种场景下有效利用丰富的 side information。尽管 《Parallel recurrent neural network architectures for feature-rich session-based recommendations》 通过为每种 side information 部署并行的子网络从而提高了预测准确率,但模型变得笨重且不灵活。

  3. 另一项研究 《Incorporating Dwell Time in Session-Based Recommendations with Recurrent Neural Networks》 没有直接改变 item embeddings,而是通过称为 item boosting 的技巧将停留时间(dwell time)纳入 RNN 模型。总体思想是让 loss function 感知停留时间。用户看一个 item 的时间越长,他/她就越感兴趣。然而,这个技巧严重依赖启发式方法,并且仅限于 behavior-related side information。另一方面,一些 item-related side information(例如价格)描述 item 的内在特征,它们不像停留时间,不能轻易被这种方法利用。

1.2 方法

  1. 在本节中,我们详细介绍研究领域和我们的方法。

    • 首先,我们说明研究问题并解释 side information、BERT 和自注意力。

    • 然后,我们提出我们的非侵入式自注意力和不同的融合操作。

    • 最后,我们说明 NOVA-BERT 模型。我们将非侵入式自注意力(NOn-inVasive self-Attention)简称为 NOVA。

1.2.1 问题陈述

  1. 给定用户与系统的 historical interactions,序列推荐任务询问下一个将被交互的 item,或下一个将被做出的 action。令 u 表示一个用户,他/她的 historical interactions 可以表示为一个按时间顺序排列的序列:

    Su=[vu(1),vu(2),⋯,vu(n)]

    其中: vu(j) 表示用户已经做出的第 j 次交互(也称为 behavior)(例如,下载一个 APP)。当只有一种类型的 actions 且没有 side information 时,每次 interaction 可以简单地由一个 item ID 来表示:

    vu(j)=ID(k)

    其中 ID(k)∈I 表示第 k 个 item 的 ID。

    I={ID(1),ID(2),⋯,ID(m)}

    I 是系统中要考虑的所有 items 的词汇表(vocabulary)。m 是词汇表大小,表示问题领域中 items 的总数。

    给定用户的历史 Su,系统预测用户最可能交互的 next item:

    Ipred=ID(k^),k^=arg⁡maxkP(vu(n+1)=ID(k)∣Su)

    .

1.2.2 Side Information

  1. Side information 可以是任何为 recommendations 提供额外有用信息的内容,可以分为两种类型:item-related 或 behavior-related。

    • Item-related side information 是内在的,除了 item ID 之外还描述 item 本身(例如,价格、生产日期、生产者)。

    • Behavior-related side information 与用户发起的 interaction 绑定,例如 action 类型(例如,purchase、rate)、执行时间、或 user feedback scores。每个 interaction 的顺序(即原始 BERT 中的 position ID)也可以被视为一种 behavior-related side information。

  2. 如果包含 side information,一次 interaction 变为:

    vu(j)=(I(k),bu,j(1),⋯,bu,j(q))I(k)=(ID(k),fk(1),⋯,fk(p))

    其中:

    • bu,j(⋅) 表示用户 u 做出的第 j 次 interaction 的一个 behavior-related side information。一共有 q 个 behavior-related side information。

    • I(⋅) 表示一个 item,包含一个 ID 和若干 item-related side information fk(⋅)。一共有 p 个 item-related side information。

    Item-related side information 是静态的,存储每个 particular item 的内在特征。因此词汇表可以重写为:

    I={I(1),I(2),⋯,I(m)}

    目标仍然是预测 next item’s ID:

    Ipred=ID(k^)k^=arg⁡maxkP(vu(n+1)=(I(k),b1,b2,⋯,bq)∣Su)

    其中: b1,b2,⋯,bq 是潜在的 behavior-related side information,如果考虑 behavior-related side information 的话。

    注意:无论 behavior-related side information 是被假设的还是被忽略的,模型仍然应该能够预测 next item。

1.2.3 BERT and Invasive Self-attention

  1. BERT4Rec(《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》)首次将 BERT 框架用于序列推荐任务,达到了 SOTA 性能。如 Figure 2 所示,在 BERT 框架下,items 被表示为向量,称为 embedings 。在训练期间,一些 items 被随机掩码,BERT 模型将尝试使用 multi-head self-attention 机制(《Attention is all you need》)恢复它们的 vector representations,从而恢复 item IDs:

    SA(Q,K,V)=σ(QK⊤dk)V

    其中:

    • σ 是 softmax 函数,dk 是缩放因子。

    • Q、K 和 V 是从 query, key, value 派生而来的组件。

    BERT 遵循 encoder-decoder 设计,为 input sequence 中的每个 item 生成 contextual representation。BERT 采用 embedding layer 存储 m 个向量,每个向量对应词汇表中的一个 item。

  2. 为了利用 side information,像 《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》、《Behavior sequence transformer for e-commerce recommendation in Alibaba》 这样的传统方法使用单独的 embedding layers 将 side information 编码为向量,然后使用 fusion function F 将它们融合到 ID embeddings 中。这种侵入式方法将 side information 注入到原始 embeddings 中,并生成 a mixed representation:

    e→u,j=F(Eid(ID),Ef1(f(1)),⋯,Efp(f(p)),Eb1(bu,j(1)),⋯,Ebq(bu,j(q)))

    其中:

    • e→u,j 是户 u 第 j 次 interaction 的 integrated embedding。

    • E 是将对象编码为向量的 embedding layer。

    The sequence of the integrated embeddings 被馈入模型,作为 the input of user history。

    BERT 框架将通过 self-attention 机制逐层更新 representations:

    Ri+1=BERTLayer(Ri)R1=(e→u,1,e→u,2,⋯,e→u,n)⊤∈Rn×d

    在原始 BERT(《Bert: Pre-training of deep bidirectional transformers for language understanding》)和 Transformer(《Attention is all you need》)中,self-attention 操作是位置不变的函数。因此,a position embedding 被添加到每个 item embedding 中以显式地编码位置信息。Position embeddings 也可以被视为一种 behavior-related side information(即 interaction 的顺序)。从这个角度来看,原始 BERT 也将 positional information 作为唯一的 side information,使用加法作为 fusion function F。

1.2.4 Non-invasive Self-attention (NOVA)

  1. 如果我们端到端地考虑 BERT 框架,它是一个带有堆叠自注意力层(stacked self-attention layers)的一个自编码器(auto-encoder)。相同的 embedding map 同时用于编码 item IDs 和解码 restored vector representations。因此,我们认为侵入式方法具有复合嵌入空间(compound embedding space)的缺点,因为 item IDs 与 other side information 不可逆地融合。混合来自 IDs 和 other side information 的信息可能会使模型不必要地难以解码 item IDs。

    相应地,我们提出一种新颖的方法,称为非侵入式自注意力(non-invasive self-attention: NOVA),以保持 embedding space 的一致性,同时利用 side information 更有效地建模序列。其思想是修改 self-attention 机制,并仔细控制 self-attention 组件的信息来源,即 query Q、key K 和 value V。除了 1.2.3 节中定义的 integrated embedding e→ 之外,NOVA 还保留一个 pure ID embeddings 的分支:

    e→u,j(ID)=Eid(ID)

    因此,对于 NOVA,user history 现在由两组 representations 组成:pure ID embeddings 和 integrated embeddings:

    Ru(ID)=(e→u,1(ID),e→u,2(ID),⋯,e→u,n(ID))⊤∈Rn×dRu=(e→u,1,e→u,2,⋯,e→u,n)⊤∈Rn×d

    NOVA 从 integrated embeddings R 中计算 Q 和 K,从 pure item ID embeddings R(ID) 计算 V。在实践中,我们以张量形式处理整个序列。即,R,R(ID)∈RB×L×h,其中: B 是 batch size,L 是序列长度(即, L=n ),h 是 embedding size (即,h=d )。

  2. NOVA 可以形式化为:

    NOVA(R,R(ID))=σ(QK⊤dk)V

    其中 Q,K,V 通过线性变换计算:

    Q=RWQ∈Rn×d,K=RWK∈Rn×d,V=R(ID)WV∈Rn×d

    其中:WQ,WK,WV∈Rd×d 为待学习的参数。

    NOVA 与侵入式 side information fusing 方式的比较如 Figure 3 所示。逐层来看,沿 NOVA layers 的 representation 保持在一个一致的向量空间 E(ID) (由 e→u,j(ID) 组成)中,其中这个向量空间由 the context of item IDs 纯粹形成。

1.2.5 Fusion Operations

  1. NOVA 以不同于侵入式方法的方式利用 side information,将其视为 an auxiliary,并使用 fusion function F 将 side information 融合到 Keys 和 Queries 中。在本研究中,我们还研究了不同种类的 fusion function 及其性能。

  2. 如上所述,位置信息也是一种 behavior-related side information ,原始 BERT 使用直接的加法操作利用它:

    Fadd(f→1,⋯,f→m)=∑i=1mf→i

    此外,我们定义 'concat' fusor 来拼接所有 side information,随后接一个全连接层以统一维度:

    Fconcat(f→1,⋯,f→m)=FC(f→1⊙⋯⊙f→m)

    其中:FC(⋅) 是一个全连接层。

    受 《TiSSA: A Time Slice Self-Attention Approach for Modeling Sequential User Behaviors》 启发,我们还设计了一种带有 trainable coefficients 的 gating fusor,trainable coefficients 是从 context 中导出的:

    Fgating(f→1,⋯,f→m)=∑i=1mG(i)f→iG=σ(FWF)

    其中:

    • F 是所有特征 (f→1,⋯,f→m)⊤∈Rm×d 的矩阵形式,f→i∈Rd。

    • WF 是 Rd×1 的可训练参数。

1.2.6 NOVA-BERT

  1. 如 Figure 4 所示,我们在 BERT 框架下使用所提出的 NOVA 操作来实现 NOVA-BERT 模型。每个 NOVA layer 接受两个输入:the supplementary side information、the sequences of item representations,然后输出相同形状的 updated representations ,这些 updated representations 将被馈入下一层。

    对于第一层的输入,item representations 是 pure item ID embeddings。由于我们仅使用 side information 作为辅助来学习更好的注意力分布,side information 不会沿 NOVA layer 传播。相同的一组 side information 被显式地提供给每个 NOVA layers。

    NOVA-BERT 遵循 《Bert: Pre-training of deep bidirectional transformers for language understanding》 中原始 BERT 的架构,只是将 self-attention layers 替换为 NOVA layers。因此,额外的参数和计算开销可以忽略不计,主要由轻量级 fusion function 引入。

    我们相信,使用 NOVA-BERT,the hidden representations 被保持在相同的 embedding space 中,这将使解码过程成为 homogeneous vector search 并有利于预测。下一节的结果也从经验上验证了 NOVA-BERT 的有效性。

    注意:

    • 对于 integrated embeddings,相同的 side information 被用于每一层。

      • integrated embeddings 中包含了 id embedding,这个 id embedding 使用逐层 refined ID embedding (即,“这些 updated representations 将被馈入下一层”)。

      • 不同的层使用相同类型的 Fusion Operations ,但是参数不同。例如,如果选择使用 concat 融合,那么它的 FC 函数在每一层是不同的,就跟传统的 BERT Layer 一样。

      • 不同的层使用不同的 WQ,WK,就跟传统的 BERT Layer 一样。

    • 对于 pure ID embeddings,每一层都是用不同的 id representation,后一层的 input 使用前一层的 output。即,“这些 updated representations 将被馈入下一层”)

1.3 实验

  1. 在本节中,我们在公开数据集和私有数据集上评估 NOVA-BERT 框架,旨在从以下方面进一步讨论 NOVA-BERT:

    • 有效性:NOVA-BERT 方法是否优于传统的侵入式 feature fusion 方法?

    • 全面性:NOVA-BERT 的组件如何促进 improvements(例如,不同类型的 side information、fusion function)?

    • 可解释性:side information 如何影响注意力分布从而提高预测准确率?

    • 效率:NOVA-BERT 在未来真实世界部署中的效率如何?

  2. 数据集:我们在公开 MovieLens 数据集和一个称为 APP 的真实世界数据集上评估我们的方法。关于数据集的详细描述请参见 Table 2。

    我们按时间顺序对 rating/downloading records 进行排序,以构建每个用户的 interaction history。遵循 《Self-attentive sequential recommendation》、《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》的做法:

    • 我们使用每个用户记录中除最后两个元素之外的前导子序列(heading subsequence)作为训练集。

    • 序列中倒数第二个 item 被用作验证集,用于调优超参数和寻找最佳 checkpoint。

    • 这些序列中的最后一个元素构成测试集。

    少于五个元素的短序列被丢弃,以消除冷启动问题的干扰。实验旨在聚焦 side information fusion,并与其他方法进行公平比较。

  3. Hyper-parameter Settings:

    • 对于本研究中的所有模型,我们使用 Adam 优化器来训练它们,学习率为 1e-4,训练 200 epochs,batch size = 128。我们固定随机数种子以减轻随机性引起的变化。学习率通过线性衰减调度器来调整,带有 5% 线性预热。

    • 我们还应用网格搜索以最小化实验结果的 bias。搜索空间包含三个超参数:

      hidden size∈{128,256,512},num heads∈{4,8},num layers∈{1,2,3,4}

      最终,我们使用 4 heads 和 3 layers,以及对 MovieLens 数据集采用 512 hidden size、对 APP 数据集采用 256 hidden size。

  4. 评估指标:遵循 《Self-attentive sequential recommendation》,我们采用两个广泛使用的评估指标:Hit Ratio(HR@k) 和 Normalized Discounted Cumulative Gain (NDCG@k),其中 k∈{5,10}。特别地,我们对整个词汇表进行排序,这意味着模型对每个 item 给出预测分数,表示推荐程度。

    一些实践,例如 《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》,使用较小的 candidate set,例如 ground truth item 加上 100 个随机采样的 negative items ,negative items 的概率与流行度成比例。在试点实验中,我们发现这种做法可能导致严重 bias。例如,冷门的 ground-truth item 在 other negative candidates 中很明显,而这些 negative candidates 是因为高流行度而被抽取的。模型倾向于利用宽松的指标(loose metric)。因此,我们使用 ranking-all metric 重新运行基线。

1.3.1 实验结果

  1. 如前所述, 《BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer》 已经证明了 BERT 框架在广泛场景范围内的卓越性能。因此,在本文中,我们选择他们的 SOTA 方法作为基线,并聚焦于在 BERT 框架下利用 side information。我们还检查了加法、拼接和门控的融合函数。

    在 Table 3 中,NOVA-BERT 在三个数据集和所有指标上都优于所有其他方法。与仅利用 position ID 的 Bert4Rec 相比,侵入式方法使用了几种 side information,但改进非常有限甚至不利。相反,NOVA-BERT 可以有效利用 side information,并稳定地优于所有其他方法。

    然而,NOVA 带来的改进在不同数据集上规模不同。在我们的实验中,对于更大、更稠密的数据集,改进的效果减小。

    • 对于 ML-1m 任务,NOVA with gating fuser 在 HR@10 上比基线高出 13.51%,而所有侵入式方法都比基线表现更差。

    • 然而,在 ML-20M 和 APP 任务中,非侵入式方法带来的相对改进分别降至 3.19% 和 1.48%。

    我们假设,在更丰富的语料库中,模型可能仅从 item context 本身就能学习到足够好的 item embeddings,留给 side information 进行补充的空间较小。

    此外,结果还证明了 NOVA-BERT 的鲁棒性。无论采用何种融合函数,NOVA-BERT 都能持续优于基线。最佳融合函数可能取决于数据集。一般来说,门控方法具有很强性能,可能受益于其可训练的门控机制。结果还表明,对于真实世界部署,融合函数的类型可以是一个超参数,应该根据数据集的内在属性进行调优,以达到最佳在线性能。

1.3.2 不同 Side Information 的贡献

  1. 在 1.2.2 节中,我们将 side information 分类为 item-related 和 behavior-related。我们还研究了在 ML-1m 任务下这两类 side information 的贡献。根据定义,出版年份和类别是 item-related 的,而 rating score 是 behavior-related。

    Table 4 显示了不同类型 side information 的结果。

    • 原始 BERT 框架除 position ID 外不考虑其他 side information,在第一行中记为 None。

    • 当给出完整 side information 时,NOVA-BERT 呈现出最显著的改进。

    • 同时,单独来看,item-related side information 和 behavior-related side information 对准确率带来的好处不明显。

      • 由于 item-related side information (年份和类型)是电影的内在特征,它可能从大量 data sequences 中被潜在推导出来,因此导致相似的准确率。

      • 另一方面,如果 behavior-related side information 也被结合,改进明显大于任一类型单独带来的改进之和。

    这一观察表明,不同类型 side information 带来的影响不是独立的。此外,NOVA-BERT 从 synthesis side information 中获益更多,并且具有很强的能力利用丰富数据而不受信息淹没(information overwhelming)所困扰。

    在 ML-1m 的情况下,comprehensive side information 可以在 NOVA-BERT 下最大程度提升准确率。然而,有理由相信,不同类型 side information 的贡献也取决于数据质量。因此,我们也将 side information selection 视为一个数据集相关因素。

1.3.3 注意力分布的可视化

  1. 为了对 NOVA-BERT 的可解释性进行更多讨论,我们可视化了 bottom NOVA layer 的注意力分布。在 Figure 5 中,我们展示了 6 个随机选择样本(列)的注意力分数比较,每一行代表 multi-head self-attention 中的一个 head。每个像素表示从 row item 到 column item 的注意力强度。一行中的所有注意力分数之和为 1。

    • 在 Figure 5(a) 中,可视化了原始 BERT4Rec 模型的注意力分数。

    • 在 Figure 5(b) 中,我们展示了 NOVA-BERT 的注意力分数。

    模型被馈入所有可用类型的 side information。

    如图所示,NOVA-BERT 的注意力分数在局部性方面显示出更强的模式,大致沿对角线集中。另一方面,在基线模型的图中没有观察到这一点。根据我们在整个数据集上的观察,这种对比是普遍的。我们注意到, side information 导致模型在 early layers 形成更明确的注意力。这一观察表明,NOVA-BERT 将 side information 作为 calculating attention matrix 的辅助,可以学习有针对性的注意力分布,从而提高准确率。

    此外,我们还进行了案例研究,以更深入地了解 NOVA 的注意力分布,并进行了更详细的分析。详情请参见附录。

  2. 关于 NOVA 注意力矩阵的更多解释:由于复杂的高维 embedding space 和错综复杂的网络架构,很难完全理解一个 item 如何以及为何关注另一个 item 。然而,从可视化结果中,我们发现,在提供side information 的情况下,NOVA-BERT 倾向于学习更结构化的注意力分数模式。

    由于深度神经网络的有限可解释性,我们不能断言 NOVA-BERT 的注意力模式绝对更好。尽管如此,我们发现观察到的模式类似于 《Revealing the dark secrets of bert》 研究中的模式,该研究呈现了一种相关性:这种注意力分布与自然语言处理任务中更好性能之间的相关性。此外,NOVA-BERT 在序列推荐任务中也显示出这种相关性,表明注意力确实被智能地分布。

1.4 NOVA-BERT 的部署与成本

  1. 根据我们的在线测试,NOVA-BERT 在真实世界场景中持续超越当前部署的方法。在 Table 5 中,列出了 NOVA-BERT 和部分基线模型的效率结果。我们根据 FLOPs 和大小来评估模型,其中 “加法” 为 fusion function。如表所示,NOVA-BERT 几乎没有额外计算开销,并且与侵入式方法具有相同的大小。此外,由于 NOVA-BERT 得到并行计算和 GPU 加速的良好支持,推理时间也大约与原始 BERT 相同。我们还声称,由于参数主要由额外 embedding layers 所带来,对于部署,我们可以使用 look-up table 来替换 embedding layers,从而得到与原始 BERT 相似的模型大小。

1.5 结论与未来工作

  1. 在这项工作中,我们提出了 NOVA-BERT 推荐系统和非侵入式自注意力机制(non-invasive self-attention: NOVA)。所提出的 NOVA 机制没有将 side information 直接融合到 item representations 中,而是将 side information 用作方向指导,并将 item representations 保持在其向量空间中不被掺杂。我们在实验数据集和工业应用上评估了 NOVA-BERT,以可忽略的计算和模型大小开销实现了 SOTA 性能。

    尽管所提出的方法达到了 SOTA 性能,但仍有几个有趣的方向值得未来研究。例如,在每一层融合 side information 可能不是最佳方法,并且也期望更强的 fusing functions。此外,我们将继续研究和改进我们的方法,以获得更高的在线性能,并将其部署到工业产品中。