VLDB 2026 Research / reviewers in the wild / expert
Sijun Zhang
dblp:141/4816
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0001-5699-6930ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 77% Language models and text generation · 12% Efficient and distributed learning · 12% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › mixture of experts
expert routing |
0.9 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Machine learning › Efficient and distributed learning › model compression › sparsity
activation sparsity |
0.3 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Natural language and speech › Language models and text generation › efficient language model
large language model efficiency |
0.3 | 1 | 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert Choice · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
reward loss · 0.9load balance loss · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TC-MoE: Augmenting Mixture of Experts with Ternary Expert ChoiceabstractThe Mixture of Experts (MoE) architecture has emerged as a promising solution to reduce computational overhead by selectively activating subsets of model parameters.
The effectiveness of MoE models depends primarily on their routing mechanisms, with the widely adopted Top-K routing scheme used for activating experts.
However, the Top-K scheme has notable limitations,
including unnecessary activations and underutilization of experts.
In this work,
rather than modifying the routing mechanism as done in previous studies,
we propose the Ternary Choice MoE (TC-MoE),
a novel approach that expands the expert space by applying the ternary set {-1, 0, 1} to each expert.
This expansion allows more efficient and effective expert activations without incurring significant computational costs.
Additionally,
given the unique characteristics of the expanded expert space,
we introduce a new load balance loss and reward loss to ensure workload balance and achieve a flexible trade-off between effectiveness and efficiency.
Extensive experiments demonstrate that TC-MoE achieves an average improvement of over 1.1% compared with traditional approaches,
while reducing the average number of activated experts by up to 9%.
These results confirm that TC-MoE effectively addresses the inefficiencies of conventional routing schemes,
offering a more efficient and scalable solution for MoE-based large language models.
Code and models are available at https://github.com/stiger1000/TC-MoE. Shen Yan 0004, Xingyan Bin, Sijun Zhang, Yisen Wang 0001, Zhouchen Lin |
ICLR | 3 |
| 2025 | HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid NormalizationabstractTransformers have become the de facto architecture for a wide range of machine learning tasks, particularly in large language models (LLMs). Despite their remarkable performance, many challenges remain in training deep transformer networks, especially regarding the position of the layer normalization. While Pre-Norm structures facilitate more stable training owing to their stronger identity path, they often lead to suboptimal performance compared to Post-Norm. In this paper, we propose **HybridNorm**, a simple yet effective hybrid normalization strategy that integrates the advantages of both Pre-Norm and Post-Norm. Specifically, HybridNorm employs QKV normalization within the attention mechanism and Post-Norm in the feed-forward network (FFN) of each transformer block. We provide both theoretical insights and empirical evidence to demonstrate that HybridNorm improves the gradient flow and the model robustness. Extensive experiments on large-scale transformer models, including both dense and sparse variants, show that HybridNorm consistently outperforms both Pre-Norm and Post-Norm approaches across multiple benchmarks. These findings highlight the potential of HybridNorm as a more stable and effective technique for improving the training and performance of deep transformer models. Code is available at https://github.com/BryceZhuo/HybridNorm. Zhijian Zhuo 0001, Yutao Zeng, Ya Wang 0002, Sijun Zhang, Jinwen Ma |
NeurIPS | 4 |
| 2025 | LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders
Xijun Xiao, Huizhi Yang, Bo Han 0014, Sijun Zhang, Wenlin Zhao, Lele Yu, Xionghang Xie, Shiru Ren, Yaocheng Tan, Peng Xu 0017, Yuchao Zheng 0002 |
RecSys | 6 |