VLDB 2026 Research / reviewers in the wild / expert
Wei Wang 0225
dblp:35/7092-225
· DBLP profile ↗
18ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-7028-9845ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment PerspectiveabstractThe low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of problem difficulty and fail to capture the alignment between model competence and problem difficulty in RL training, leading to suboptimal results. To address these challenges, we introduce Competence-Difficulty Alignment Sampling (CDAS). This approach allows for accurate and stable estimation of problem difficulties by aggregating historical performance discrepancies across problems. Subsequently, model competence is quantified to adaptively select problems whose difficulties align with the model's current competence using a fixed-point system. Extensive experiments in mathematical RL training show that CDAS consistently outperforms strong baselines, achieving the highest average accuracy of 45.89%. Furthermore, CDAS reduces the training step time overhead by 57.06% compared to the widely-used Dynamic Sampling strategy, verifying the efficiency of CDAS. Additional experiments on different tasks, model architectures, and model sizes demonstrate the generalization capability of CDAS. Deyang Kong, Xiangyu Xi, Wei Wang 0225, Jingang Wang, Shikun Zhang, Wei Ye 0004 |
AAAI | 4 |
| 2026 | PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal PracticeabstractYuzhen Shi, Huanghai Liu, Yiran HU, Song Gaojie, Xu Xinran, Yubo Ma, Tianyi Tang, Li Zhang, Qingjing Chen, Feng Di, Wenbo Lv, Weiheng Wu, Kexin Yang, Sen Yang, Wei Wang, Rongyao Shi, Qiu Yuanyang, Yuemeng Qi, Zhang Jingwen, Sui Xiaoyu, Yifan Chen, Zhang Yi, An Yang, Bowen Yu, Dayiheng Liu, Junyang Lin, Weixing Shen, Bing Zhao, Charles L. A. Clarke, HU Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuzhen Shi, Huanghai Liu, Yiran Hu, Gaojie Song, Xinran Xu, Yubo Ma, Qingjing Chen, Di Feng, Wenbo Lv, Weiheng Wu, Kexin Yang 0002, Wei Wang 0225, Rongyao Shi, Yuanyang Qiu, Yuemeng Qi, Xiaoyu Sui, Yi Zhang 0101, An Yang, Bowen Yu 0002, Dayiheng Liu, Junyang Lin, Weixing Shen, Charles L. A. Clarke, Hu Wei |
ACL (1) | 15 |
| 2024 | How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data CompositionabstractGuanting Dong, Hongyi Yuan, Keming Lu, Chengpeng Li, Mingfeng Xue, Dayiheng Liu, Wei Wang, Zheng Yuan, Chang Zhou, Jingren Zhou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Guanting Dong 0001, Hongyi Yuan, Keming Lu, Chengpeng Li 0001, Mingfeng Xue, Dayiheng Liu, Wei Wang 0225, Zheng Yuan 0002, Chang Zhou 0005, Jingren Zhou 0001 |
ACL (1) | 7 |
| 2023 | mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoabstractRecent years have witnessed a big convergence of language, vision, and multi-modal pretraining. In this work, we present mPLUG-2, a new unified paradigm with modularized design for multi-modal pretraining, which can benefit from modality collaboration while addressing the problem of modality entanglement. In contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal modules for modality collaboration and disentangling different modality modules to deal with modality entanglement. It is flexible to select different modules for different understanding and generation tasks across all modalities including text, image, and video. Empirical study shows that mPLUG-2 achieves state-of-the-art or competitive results on a broad range of over 30 downstream tasks, spanning multi-modal tasks of image-text and video-text understanding and generation, and uni-modal tasks of text-only, image-only, and video-only understanding. Notably, mPLUG-2 shows new state-of-the-art results of 48.0 top-1 accuracy and 80.3 CIDEr on the challenging MSRVTT video QA and video caption tasks with a far smaller model size and data scale. It also demonstrates strong zero-shot transferability on vision-language and video-language tasks. Code and models will be released in https://github.com/X-PLUG/mPLUG-2. Haiyang Xu 0001, Qinghao Ye, Ming Yan 0008, Yaya Shi, Jiabo Ye, Yuanhong Xu, Chenliang Li 0003, Bin Bi, Qi Qian 0001, Wei Wang 0225, Guohai Xu, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001 |
ICML | 10 |
| 2023 | RRHF: Rank Responses to Align Language Models with Human FeedbackabstractReinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models with human preferences, significantly enhancing the quality of interactions between humans and models.
InstructGPT implements RLHF through several stages, including Supervised Fine-Tuning (SFT), reward model training, and Proximal Policy Optimization (PPO).
However, PPO is sensitive to hyperparameters and requires multiple models in its standard implementation, making it hard to train and scale up to larger parameter counts.
In contrast, we propose a novel learning paradigm called RRHF, which scores sampled responses from different sources via a logarithm of conditional probabilities and learns to align these probabilities with human preferences through ranking loss.
RRHF can leverage sampled responses from various sources including the model responses from itself, other large language model responses, and human expert responses to learn to rank them.
RRHF only needs 1 to 2 models during tuning and can efficiently align language models with human preferences robustly without complex hyperparameter tuning.
Additionally, RRHF can be considered an extension of SFT and reward model training while being simpler than PPO in terms of coding, model counts, and hyperparameters.
We evaluate RRHF on the Helpful and Harmless dataset, demonstrating comparable alignment performance with PPO by reward model score and human labeling.
Extensive experiments show that the performance of RRHF is highly related to sampling quality which suggests RRHF is a best-of-$n$ learner. Hongyi Yuan, Zheng Yuan 0002, Chuanqi Tan, Wei Wang 0225, Songfang Huang |
NeurIPS | 4 |
| 2023 | Achieving Human Parity on Visual Question AnsweringabstractThe Visual Question Answering (VQA) task utilizes both visual image and language analysis to answer a textual question with respect to an image. It has been a popular research topic with an increasing number of real-world applications in the last decade. This paper introduces a novel hierarchical integration of vision and language AliceMind-MMU (ALIbaba’s Collection of Encoder-decoders from Machine IntelligeNce lab of Damo academy - MultiMedia Understanding) , which leads to similar or even slightly better results than a human being does on VQA. A hierarchical framework is designed to tackle the practical problems of VQA in a cascade manner including: (1) diverse visual semantics learning for comprehensive image content understanding; (2) enhanced multi-modal pre-training with modality adaptive attention; and (3) a knowledge-guided model integration with three specialized expert modules for the complex VQA task. Treating different types of visual questions with corresponding expertise needed plays an important role in boosting the performance of our VQA architecture up to the human level. An extensive set of experiments and analysis are conducted to demonstrate the effectiveness of the new research work. Ming Yan 0008, Haiyang Xu 0001, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Luo Si, Rong Jin 0001 |
ACM Trans. Inf. Syst. | 6 |
| 2022 | mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsabstractChenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, He Chen, Guohai Xu, Zheng Cao, Ji Zhang, Songfang Huang, Fei Huang, Jingren Zhou, Luo Si. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Chenliang Li 0003, Haiyang Xu 0001, Wei Wang 0225, Ming Yan 0008, Bin Bi, Jiabo Ye, Guohai Xu, Zheng Cao 0003, Ji Zhang 0011, Songfang Huang, Fei Huang 0002, Jingren Zhou 0001, Luo Si |
EMNLP | 4 |
| 2022 | STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model TrainingabstractDeep neural networks (DNNs) with billion-scale parameters have demonstrated impressive performance in solving many tasks. Unfortunately, training a billion-scale DNN is out of the reach of many data scientists because it requires high-performance GPU servers that are too expensive to purchase and maintain. We present STRONGHOLD, a novel approach for enabling large DNN model training with no change to the user code. STRONGHOLD scales up the largest trainable model size by dynamically offloading data to the CPU RAM and enabling the use of secondary storage. It automatically determines the minimum amount of data to be kept in the GPU memory to minimize GPU memory usage. Compared to state-of-the-art offloading-based solutions, STRONGHOLD improves the trainable model size by 1.9x~6. Sx on a 32GB V100 GPU, with 1.2x~3.7x improvement on the training throughput. It has been deployed into production to successfully support large-scale DNN training. Wei Wang 0225, Shenghao Qiu, Renyu Yang, Songfang Huang, Jie Xu 0007, Zheng Wang 0001 |
SC | 2 |
| 2021 | A Unified Pretraining Framework for Passage Ranking and ExpansionabstractPretrained language models have recently advanced a wide range of natural language processing tasks. Nowadays, the application of pretrained language models to IR tasks has also achieved impressive results. Typical methods either directly apply a pretrained model to improve the re-ranking stage, or use it to conduct passage expansion and term weighting for first-stage retrieval. We observe that the passage ranking and passage expansion tasks share certain inherent relations, and can benefit from each other. Therefore, in this paper, we propose a general pretraining framework to enhance both tasks with Unified Encoder-Decoder networks (UED). The overall ranking framework consists of two parts in a cascade manner: (1) passage expansion with a pretraining-based query generation method; (2) re-ranking of passage candidates from a traditional retrieval method with a pretrained transformer encoder. Both the two parts are based on the same pretrained UED model, where we jointly train the passage ranking and query generation tasks for further improving the full ranking pipeline. An extensive set of experiments have been conducted on two large-scale passage retrieval datasets to demonstrate the state-of-the-art results of the proposed framework in both the first-stage retrieval and the final re-ranking. In addition, we successfully deploy the framework to our online production system, which can stably serve industrial applications with a request volume of up to 100 QPS in less than 300ms. Ming Yan 0008, Chenliang Li 0003, Bin Bi, Wei Wang 0225, Songfang Huang |
AAAI | 4 |
| 2021 | StructuralLM: Structural Pre-training for Form UnderstandingabstractChenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chenliang Li 0003, Bin Bi, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 4 |
| 2021 | VECO: Variable and Flexible Cross-lingual Pre-training for Language Understanding and GenerationabstractFuli Luo, Wei Wang, Jiahao Liu, Yijia Liu, Bin Bi, Songfang Huang, Fei Huang, Luo Si. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fuli Luo, Wei Wang 0225, Bin Bi, Songfang Huang, Fei Huang 0002, Luo Si |
ACL/IJCNLP (1) | 2 |
| 2021 | Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clustersabstractEfficient GPU resource scheduling is essential to maximize resource utilization and save training costs for the increasing amount of deep learning workloads in shared GPU clusters. Existing GPU schedulers largely rely on static policies to leverage the performance characteristics of deep learning jobs. However, they can hardly reach optimal efficiency due to the lack of elasticity. To address the problem, we propose ONES, an ONline Evolutionary Scheduler for elastic batch size orchestration. ONES automatically manages the elasticity of each job based on the training batch size, so as to maximize GPU utilization and improve scheduling efficiency. It determines the batch size for each job through an online evolutionary search that can continuously optimize the scheduling decisions. We evaluate the effectiveness of ONES with 64 GPUs on TACC's Longhorn supercomputers. The results show that ONES can outperform the prior deep learning schedulers with a significantly shorter average job completion time. Zhengda Bian, Shenggui Li, Wei Wang 0225, Yang You 0001 |
SC | 3 |
| 2020 | Generating Well-Formed Answers by Machine Reading with Stochastic Selector Networks
Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
AAAI | 4 |
| 2020 | PALM: Pre-training an Autoencoding&Autoregressive Language Model for Context-conditioned GenerationabstractSelf-supervised pre-training, such as BERT (Devlin et al., 2018), MASS (Song et al., 2019) and BART (Lewis et al., 2019), has emerged as a powerful technique for natural language understanding and generation.Existing pre-training techniques employ autoencoding and/or autoregressive objectives to train Transformer-based models by recovering original word tokens from corrupted text with some masked tokens.The training goals of existing techniques are often inconsistent with the goals of many language generation tasks, such as generative question answering and conversational response generation, for producing new text given context.This work presents PALM with a novel scheme that jointly pre-trains an autoencoding and autoregressive language model on a large unlabeled corpus, specifically designed for generating new text conditioned on context.The new scheme alleviates the mismatch introduced by the existing denoising scheme between pre-training and fine-tuning where generation is more than reconstructing original text.An extensive set of experiments show that PALM achieves new state-of-theart results on a variety of language generation benchmarks covering generative question answering (Rank 1 on the official MARCO leaderboard), abstractive summarization on CNN/DailyMail as well as Gigaword, question generation on SQuAD, and conversational response generation on Cornell Movie Dialogues. Bin Bi, Chenliang Li 0003, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Songfang Huang, Fei Huang 0002, Luo Si |
EMNLP (1) | 5 |
| 2020 | StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
Wei Wang 0225, Bin Bi, Ming Yan 0008, Chen Wu 0006, Jiangnan Xia, Zuyi Bao, Liwei Peng, Luo Si |
ICLR | 1 |
| 2019 | A Deep Cascade Model for Multi-Document Reading ComprehensionabstractA fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i.e., TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms. Ming Yan 0008, Jiangnan Xia, Chen Wu 0006, Bin Bi, Zhongzhou Zhao, Ji Zhang 0011, Luo Si, Rui Wang 0005, Wei Wang 0225, Haiqing Chen |
AAAI | 9 |
| 2019 | Incorporating External Knowledge into Machine Reading for Generative Question AnsweringabstractBin Bi, Chen Wu, Ming Yan, Wei Wang, Jiangnan Xia, Chenliang Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bin Bi, Chen Wu 0006, Ming Yan 0008, Wei Wang 0225, Jiangnan Xia, Chenliang Li 0003 |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Multi-Granularity Hierarchical Attention Fusion Networks for Reading Comprehension and Question AnsweringabstractThis paper describes a novel hierarchical attention network for reading comprehension style question answering, which aims to answer questions for a given narrative paragraph.In the proposed method, attention and fusion are conducted horizontally and vertically across layers at different levels of granularity between question and paragraph.Specifically, it first encode the question and paragraph with fine-grained language embeddings, to better capture the respective representations at semantic level.Then it proposes a multi-granularity fusion approach to fully fuse information from both global and attended representations.Finally, it introduces a hierarchical attention network to focuses on the answer span progressively with multi-level softalignment.Extensive experiments on the large-scale SQuAD and TriviaQA datasets validate the effectiveness of the proposed method.At the time of writing the paper (Jan.12th 2018), our model achieves the first position on the SQuAD leaderboard for both single and ensemble models.We also achieves state-of-the-art results on TriviaQA, AddSent and AddOne-Sent datasets. Wei Wang 0225, Chen Wu 0006, Ming Yan 0008 |
ACL (1) | 1 |