论文阅读-DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models 背景 在传统 decoder-only 的大模型架构中,在 Self-Attention 层后使用 FFN 层处理输出,所有的 token 都需要经过整个 FFN 层权重矩阵的计算,计算开销大...2025-11-20论文笔记#LLM#DeepSeek
论文阅读-Attention Is All You Need 基本信息 期刊: (发表日期: 2023-08-01) 作者: Ashish Vaswani; Noam Shazeer; Niki Parmar; Jakob Uszkore...2024-08-12论文笔记#Transformer#Attention#Self-Attention
论文阅读-BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 基本信息 期刊: (发表日期: 2019-05-24) 作者: Jacob Devlin; Ming-Wei Chang; Kenton Lee; Kristina Touta...2024-08-04论文笔记#NLP#BERT
论文阅读-Tacotron: Towards End-to-End Speech Synthesis 基本信息 期刊: (发表日期: 2017-04-06) 作者: Yuxuan Wang; R. J. Skerry-Ryan; Daisy Stanton; Yonghui W...2024-07-30论文笔记#Tacotron#TTS
论文阅读-A Survey of Graph Neural Networks for Recommender Systems: Challenges, Methods, and Directions 基本信息 期刊:(发表日期: 2023-01-12) 作者: Chen Gao; Yu Zheng; Nian Li; Yinfeng Li; Yingrong Qin; Ji...2024-07-30论文笔记#Recommender System#GNN