VLDB 2026 Research / reviewers in the wild / expert
Wenlin Zhao
dblp:305/6453
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-0655-4317ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RankMixer: Scaling Up Ranking Models in Industrial RecommendersabstractRecent progress on large language models (LLMs) has spurred interest in scaling up recommendation systems, yet two practical obstacles remain. First, training and serving cost on industrial Recommenders must respect strict latency bounds and high QPS demands. Second, most human-designed feature-crossing modules in ranking models were inherited from the CPU era and fail to exploit modern GPUs, resulting in low Model Flops Utilization (MFU) and poor scalability. We introduce RankMixer, a hardware-aware model design tailored towards a unified and scalable feature-interaction architecture. RankMixer retains the transformer's high parallelism while replacing quadratic self-attention with multi-head token mixing module for higher efficiency. Besides, RankMixer maintains both the modeling for distinct feature subspaces and cross-feature-space interactions with Per-token FFNs. We further extend it to one billion parameters with a Sparse-MoE variant for higher ROI. A dynamic routing strategy is adapted to address the inadequacy and imbalance of experts training. Experiments show RankMixer's superior scaling abilities on a trillion-scale production dataset. By replacing previously diverse handcrafted low-MFU modules with RankMixer, we boost the model MFU from 4.5% to 45%, and scale our online ranking model parameters by two orders of magnitude while maintaining roughly the same inference latency. We verify RankMixer's universality with online A/B tests across two core application scenarios (Recommendation and Advertisement). Finally, we launch 1B Dense-Parameters RankMixer for full traffic serving without increasing the serving cost, which improves user active days by 0.3% and total in-app usage duration by 1.08%. Zhifang Fan, Xiaoxie Zhu, Hangyu Wang, Xintian Han, Xinmin Wang, Wenlin Zhao, Huizhi Yang, Zhe Chen 0015, Yuchao Zheng 0002, Qiwei Chen, Feng Zhang 0047, Peng Xu 0017, Zuotao Liu |
CIKM | 9 |
| 2025 | LONGER: Scaling Up Long Sequence Modeling in Industrial Recommenders
Xijun Xiao, Huizhi Yang, Bo Han 0014, Sijun Zhang, Wenlin Zhao, Lele Yu, Xionghang Xie, Shiru Ren, Yaocheng Tan, Peng Xu 0017, Yuchao Zheng 0002 |
RecSys | 9 |
| 2022 | A Dual-Expert Framework for Event Argument ExtractionabstractEvent argument extraction (EAE) is an important information extraction task, which aims to identify the arguments of an event described in a given text and classify the roles played by them. A key characteristic in realistic EAE data is that the instance numbers of different roles follow an obvious long-tail distribution. However, the training and evaluation paradigms of existing EAE models either prone to neglect the performance on "tail roles'', or change the role instance distribution for model training to an unrealistic uniform distribution. Though some generic methods can alleviate the class imbalance in long-tail datasets, they usually sacrifice the performance of "head classes'' as a trade-off. To address the above issues, we propose to train our model on realistic long-tail EAE datasets, and evaluate the average performance over all roles. Inspired by the Mixture of Experts (MOE), we propose a Routing-Balanced Dual Expert Framework (RBDEF), which divides all roles into "head" and "tail" two scopes and assigns the classifications of head and tail roles to two separate experts. In inference, each encoded instance will be allocated to one of the two experts by a routing mechanism. To reduce routing errors caused by the imbalance of role instances, we design a Balanced Routing Mechanism (BRM), which transfers several head roles to the tail expert to balance the load of routing, and employs a tri-filter routing strategy to reduce the misallocation of the tail expert's instances. To enable an effective learning of tail roles with scarce instances, we devise Target-Specialized Meta Learning (TSML) to train the tail expert. Different from other meta learning algorithms that only search a generic parameter initialization equally applying to infinite tasks, TSML can adaptively adjust its search path to obtain a specialized initialization for the tail expert, thereby expanding the benefits to the learning of tail roles. In experiments, RBDEF significantly outperforms the state-of-the-art EAE models and advanced methods for long-tail data. Rui Li 0044, Wenlin Zhao, Cheng Yang 0002, Sen Su |
SIGIR | 2 |
| 2021 | Treasures Outside Contexts: Improving Event Detection via Global StatisticsabstractEvent detection (ED) aims at identifying event instances of specified types in given texts, which has been formalized as a sequence labeling task.As far as we know, existing neural-based ED models make decisions relying on the contextual semantic features of each word in the input text, which we find is easy to get confused by varied contexts in the test stage.To this end, we come up with the idea of introducing a set of statistical features from word-event co-occurrence frequencies in the entire training set to cooperate with the contextual features.Specifically, we propose a Semantic and Statistic-Joint Discriminative Network (S 2 -JDN) consisting of a semantic feature extractor, a statistical feature extractor, and a joint event discriminator.In experiments, S 2 -JDN effectively exceeds ten recent state-ofthe-art (SOTA) baseline methods on ACE2005 and KBP2015 benchmark datasets.Further, we perform extensive experiments to investigate the effectiveness of S 2 -JDN. Rui Li 0044, Wenlin Zhao, Cheng Yang 0002, Sen Su |
EMNLP (1) | 2 |