Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Sijun Zhang

dblp:141/4816 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0001-5699-6930ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Deep learning architectures and training · 77% Language models and text generation · 12% Efficient and distributed learning · 12%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › mixture of experts
expert routing
0.912025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Machine learning › Deep learning architectures and training
mixture of experts
0.912025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity
0.312025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency
0.312025
TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025

Methods — techniques the papers use, named apart from their topics

reward loss · 0.9load balance loss · 0.9
YearPublicationVenuePosition
2025 TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice
abstract
The Mixture of Experts (MoE) architecture has emerged as a promising solution to reduce computational overhead by selectively activating subsets of model parameters. The effectiveness of MoE models depends primarily on their routing mechanisms, with the widely adopted Top-K routing scheme used for activating experts. However, the Top-K scheme has notable limitations, including unnecessary activations and underutilization of experts. In this work, rather than modifying the routing mechanism as done in previous studies, we propose the Ternary Choice MoE (TC-MoE), a novel approach that expands the expert space by applying the ternary set {-1, 0, 1} to each expert. This expansion allows more efficient and effective expert activations without incurring significant computational costs. Additionally, given the unique characteristics of the expanded expert space, we introduce a new load balance loss and reward loss to ensure workload balance and achieve a flexible trade-off between effectiveness and efficiency. Extensive experiments demonstrate that TC-MoE achieves an average improvement of over 1.1% compared with traditional approaches, while reducing the average number of activated experts by up to 9%. These results confirm that TC-MoE effectively addresses the inefficiencies of conventional routing schemes, offering a more efficient and scalable solution for MoE-based large language models. Code and models are available at https://github.com/stiger1000/TC-MoE.
Shen Yan 0004, Xingyan Bin, Sijun Zhang, Yisen Wang 0001, Zhouchen Lin
ICLR3
2025 HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization
abstract
Transformers have become the de facto architecture for a wide range of machine learning tasks, particularly in large language models (LLMs). Despite their remarkable performance, many challenges remain in training deep transformer networks, especially regarding the position of the layer normalization. While Pre-Norm structures facilitate more stable training owing to their stronger identity path, they often lead to suboptimal performance compared to Post-Norm. In this paper, we propose **HybridNorm**, a simple yet effective hybrid normalization strategy that integrates the advantages of both Pre-Norm and Post-Norm. Specifically, HybridNorm employs QKV normalization within the attention mechanism and Post-Norm in the feed-forward network (FFN) of each transformer block. We provide both theoretical insights and empirical evidence to demonstrate that HybridNorm improves the gradient flow and the model robustness. Extensive experiments on large-scale transformer models, including both dense and sparse variants, show that HybridNorm consistently outperforms both Pre-Norm and Post-Norm approaches across multiple benchmarks. These findings highlight the potential of HybridNorm as a more stable and effective technique for improving the training and performance of deep transformer models. Code is available at https://github.com/BryceZhuo/HybridNorm.
Zhijian Zhuo 0001, Yutao Zeng, Ya Wang 0002, Sijun Zhang, Jinwen Ma
NeurIPS4
2025 LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders
Xijun Xiao, Huizhi Yang, Bo Han 0014, Sijun Zhang, Wenlin Zhao, Lele Yu, Xionghang Xie, Shiru Ren, Yaocheng Tan, Peng Xu 0017, Yuchao Zheng 0002
RecSys6