VLDB 2026 Research / reviewers in the wild / expert
Rui Wang 0028
dblp:06/2293-28
· DBLP profile ↗
19ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Computer networks · 7 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal RetrievalabstractVision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowledge beyond the visual content presented in images.The effectiveness of Vision-language RAG systems hinges on multimodal retrieval, which is inherently challenging due to the diverse modalities and knowledge granularities in both queries and knowledge bases.Existing methods have not fully tapped into the potential interplay between these elements.We propose a multimodal RAG system featuring a coarse-to-fine, multi-step retrieval that harmonizes multiple granularities and modalities to enhance efficacy.Our system begins with a broad initial search aligning knowledge granularity for cross-modal retrieval, followed by a multimodal fusion reranking to capture the nuanced multimodal information for top entity selection.A text reranker then filters out the most relevant fine-grained section for augmented generation.Extensive experiments on the InfoSeek and Encyclopedic-VQA benchmarks show our method achieves state-of-the-art retrieval performance and highly competitive answering results, underscoring its effectiveness in advancing KB-VQA systems. Jingjing Fu, Rui Wang 0028, Lei Song 0001, Jiang Bian 0002 |
ACL (1) | 3 |
| 2025 | From Complex to Atomic: Enhancing Augmented Generation via Knowledge-Aware Dual Rewriting and ReasoningabstractRecent advancements in Retrieval-Augmented Generation (RAG) systems have significantly enhanced the capabilities of large language models (LLMs) by incorporating external knowledge retrieval. However, the sole reliance on retrieval is often inadequate for mining deep, domain-specific knowledge and for performing logical reasoning from specialized datasets. To tackle these challenges, we present an approach, which is designed to extract, comprehend, and utilize domain knowledge while constructing a coherent rationale. At the heart of our approach lie four pivotal components: a knowledge atomizer that extracts atomic questions from raw data, a query proposer that generates subsequent questions to facilitate the original inquiry, an atomic retriever that locates knowledge based on atomic knowledge alignments, and an atomic selector that determines which follow-up questions to pose guided by the retrieved information. Through this approach, we implement a knowledge-aware task decomposition strategy that adeptly extracts multifaceted knowledge from segmented data and iteratively builds the rationale in alignment with the initial query and the acquired knowledge. We conduct comprehensive experiments to demonstrate the efficacy of our approach across various benchmarks, particularly those requiring multihop reasoning steps. The results indicate a significant enhancement in performance, up to 12.6% over the second-best method, underscoring the potential of the approach in complex, knowledge-intensive applications. Jingjing Fu, Rui Wang 0028, Lei Song 0001, Jiang Bian 0002 |
ICML | 3 |
| 2025 | Graph Neural Network Enhanced Retrieval for Question Answering of Large Language ModelsabstractZijian Li, Qingyan Guo, Jiawei Shao, Lei Song, Jiang Bian, Jun Zhang, Rui Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zijian Li 0023, Qingyan Guo, Jiawei Shao, Lei Song 0001, Jiang Bian 0002, Jun Zhang 0004, Rui Wang 0028 |
NAACL (Long Papers) | 7 |
| 2024 | Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt OptimizersabstractLarge Language Models (LLMs) excel in various tasks, but they rely on carefully crafted prompts that often demand substantial human effort. To automate this process, in this paper, we propose a novel framework for discrete prompt optimization, called EvoPrompt, which borrows the idea of evolutionary algorithms (EAs) as they exhibit good performance and fast convergence. To enable EAs to work on discrete prompts, which are natural language expressions that need to be coherent and human-readable, we connect LLMs with EAs. This approach allows us to simultaneously leverage the powerful language processing capabilities of LLMs and the efficient optimization performance of EAs. Specifically, abstaining from any gradients or parameters, EvoPrompt starts from a population of prompts and iteratively generates new prompts with LLMs based on the evolutionary operators, improving the population based on the development set. We optimize prompts for both closed- and open-source LLMs including GPT-3.5 and Alpaca, on 31 datasets covering language understanding, generation tasks, as well as BIG-Bench Hard (BBH) tasks. EvoPrompt significantly outperforms human-engineered prompts and existing methods for automatic prompt generation (e.g., up to 25% on BBH). Furthermore, EvoPrompt demonstrates that connecting LLMs with EAs creates synergies, which could inspire further research on the combination of LLMs and conventional algorithms. Qingyan Guo, Rui Wang 0028, Junliang Guo, Kaitao Song, Xu Tan 0003, Jiang Bian 0002, Yujiu Yang 0001 |
ICLR | 2 |
| 2024 | Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient LearningabstractResidual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution.'' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30.95 and 44.27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3.8B DeepNet by an average of 2.9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5.7 accuracy points on the LM Harness Evaluation. Rui Wang 0028, Qingyan Guo, Junliang Guo, Xu Tan 0003, Tong Xiao 0001, Jingang Wang |
NeurIPS | 3 |
| 2023 | SoftCorrect: Error Correction with Soft Detection for Automatic Speech RecognitionabstractError correction in automatic speech recognition (ASR) aims to correct those incorrect words in sentences generated by ASR models. Since recent ASR models usually have low word error rate (WER), to avoid affecting originally correct tokens, error correction models should only modify incorrect words, and therefore detecting incorrect words is important for error correction. Previous works on error correction either implicitly detect error words through target-source attention or CTC (connectionist temporal classification) loss, or explicitly locate specific deletion/substitution/insertion errors. However, implicit error detection does not provide clear signal about which tokens are incorrect and explicit error detection suffers from low detection accuracy. In this paper, we propose SoftCorrect with a soft error detection mechanism to avoid the limitations of both explicit and implicit error detection. Specifically, we first detect whether a token is correct or not through a probability produced by a dedicatedly designed language model, and then design a constrained CTC loss that only duplicates the detected incorrect tokens to let the decoder focus on the correction of error tokens. Compared with implicit error detection with CTC loss, SoftCorrect provides explicit signal about which words are incorrect and thus does not need to duplicate every token but only incorrect tokens; compared with explicit error detection, SoftCorrect does not detect specific deletion/substitution/insertion errors but just leaves it to CTC loss. Experiments on AISHELL-1 and Aidatatang datasets show that SoftCorrect achieves 26.1% and 9.4% CER reduction respectively, outperforming previous works by a large margin, while still enjoying fast speed of parallel generation. Yichong Leng, Xu Tan 0003, Kaitao Song, Rui Wang 0028, Xiang-Yang Li 0001, Tao Qin 0001, Edward Lin, Tie-Yan Liu |
AAAI | 5 |
| 2022 | A Study of Syntactic Multi-Modality in Non-Autoregressive Machine TranslationabstractKexun Zhang, Rui Wang, Xu Tan, Junliang Guo, Yi Ren, Tao Qin, Tie-Yan Liu. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kexun Zhang, Rui Wang 0028, Xu Tan 0003, Junliang Guo, Yi Ren 0006, Tao Qin 0001, Tie-Yan Liu |
NAACL-HLT | 2 |
| 2022 | Museformer: Transformer with Fine- and Coarse-Grained Attention for Music GenerationabstractSymbolic music generation aims to generate music scores automatically. A recent trend is to use Transformer or its variants in music generation, which is, however, suboptimal, because the full attention cannot efficiently model the typically long music sequences (e.g., over 10,000 tokens), and the existing models have shortcomings in generating musical repetition structures. In this paper, we propose Museformer, a Transformer with a novel fine- and coarse-grained attention for music generation. Specifically, with the fine-grained attention, a token of a specific bar directly attends to all the tokens of the bars that are most relevant to music structures (e.g., the previous 1st, 2nd, 4th and 8th bars, selected via similarity statistics); with the coarse-grained attention, a token only attends to the summarization of the other bars rather than each token of them so as to reduce the computational cost. The advantages are two-fold. First, it can capture both music structure-related correlations via the fine-grained attention, and other contextual information via the coarse-grained attention. Second, it is efficient and can model over 3X longer music sequences compared to its full-attention counterpart. Both objective and subjective experimental results demonstrate its ability to generate long music sequences with high quality and better structures. Botao Yu, Peiling Lu, Rui Wang 0028, Xu Tan 0003, Wei Ye 0004, Shikun Zhang, Tao Qin 0001, Tie-Yan Liu |
NeurIPS | 3 |
| 2021 | Lightspeech: Lightweight and Fast Text to Speech with Neural Architecture SearchabstractText to speech (TTS) has been broadly used to synthesize natural and intelligible speech in different scenarios. Deploying TTS in various end devices such as mobile phones or embedded devices requires extremely small memory usage and inference latency. While non-autoregressive TTS models such as FastSpeech have achieved significantly faster inference speed than autoregressive models, their model size and inference latency are still large for the deployment in resource constrained devices. In this paper, we propose LightSpeech, which leverages neural architecture search (NAS) to automatically design more lightweight and efficient models based on FastSpeech. We first profile the components of current Fast-Speech model and carefully design a novel search space containing various lightweight and potentially effective architectures. Then NAS is utilized to automatically discover well performing architectures within the search space. Experiments show that the model discovered by our method achieves 15x model compression ratio and 6.5x inference speedup on CPU with on par voice quality. Audio demos are provided at https://speechresearch.github.io/lightspeech. Renqian Luo, Xu Tan 0003, Rui Wang 0028, Tao Qin 0001, Jinzhu Li, Sheng Zhao 0002, Enhong Chen, Tie-Yan Liu |
ICASSP | 3 |
| 2021 | A Survey on Low-Resource Neural Machine TranslationabstractNeural approaches have achieved state-of-the-art accuracy on machine translation but suffer from the high cost of collecting large scale parallel data. Thus, a lot of research has been conducted for neural machine translation (NMT) with very limited parallel data, i.e., the low-resource setting. In this paper, we provide a survey for low-resource NMT and classify related works into three categories according to the auxiliary data they used: (1) exploiting monolingual data of source and/or target languages, (2) exploiting data from auxiliary languages, and (3) exploiting multi-modal data. We hope that our survey can help researchers to better understand this field and inspire them to design better algorithms, and help industry practitioners to choose appropriate algorithms for their applications. Rui Wang 0028, Xu Tan 0003, Renqian Luo, Tao Qin 0001, Tie-Yan Liu |
IJCAI | 1 |
| 2020 | Semi-Supervised Neural Architecture SearchabstractNeural architecture search (NAS) relies on a good controller to generate better architectures or predict the accuracy of given architectures. However, training the controller requires both abundant and high-quality pairs of architectures and their accuracy, while it is costly to evaluate an architecture and obtain its accuracy. In this paper, we propose SemiNAS, a semi-supervised NAS approach that leverages numerous unlabeled architectures (without evaluation and thus nearly no cost). Specifically, SemiNAS 1) trains an initial accuracy predictor with a small set of architecture-accuracy data pairs; 2) uses the trained accuracy predictor to predict the accuracy of large amount of architectures (without evaluation); and 3) adds the generated data pairs to the original data to further improve the predictor. The trained accuracy predictor can be applied to various NAS algorithms by predicting the accuracy of candidate architectures for them. SemiNAS has two advantages: 1) It reduces the computational cost under the same accuracy guarantee. On NASBench-101 benchmark dataset, it achieves comparable accuracy with gradient-based method while using only 1/7 architecture-accuracy pairs. 2) It achieves higher accuracy under the same computational cost. It achieves 94.02% test accuracy on NASBench-101, outperforming all the baselines when using the same number of architectures. On ImageNet, it achieves 23.5% top-1 error rate (under 600M FLOPS constraint) using 4 GPU-days for search. We further apply it to LJSpeech text to speech task and it achieves 97% intelligibility rate in the low-resource setting and 15% test error rate in the robustness setting, with 9%, 7% improvements over the baseline respectively. Renqian Luo, Xu Tan 0003, Rui Wang 0028, Tao Qin 0001, Enhong Chen, Tie-Yan Liu |
NeurIPS | 3 |
| 2018 | Exploiting Mobility in Cache-Assisted D2D Networks: Performance Analysis and OptimizationabstractCaching popular content at mobile devices, accompanied by device-to-device (D2D) communications, is one promising technology for effective mobile content delivery. User mobility is an important factor when investigating such networks, which unfortunately was largely ignored in most previous works. Preliminary studies have been carried out but the effect of mobility on the caching performance has not been fully understood. In this paper, by explicitly considering users’ contact and inter-contact durations via an alternating renewal process, we first investigate the effect of mobility with a given cache placement. A tractable expression of the data offloading ratio, i.e., the proportion of requested data that can be delivered via D2D links, is derived, which is proved to be increasing with the user moving speed. The analytical results are then used to develop an effective mobility-aware caching strategy to maximize the data offloading ratio. Simulation results are provided to confirm the accuracy of the analytical results and also validate the effect of user mobility. Performance gains of the proposed mobility-aware caching strategy are demonstrated with both stochastic models and real-life data sets. It is observed that the information of the contact durations is critical to design cache placement, especially when they are relatively short or comparable to the inter-contact durations. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 1 |
| 2017 | Mobility increases the data offloading ratio in D2D caching networksabstractCaching at mobile devices, accompanied by device-to-device (D2D) communications, is one promising technique to accommodate the exponentially increasing mobile data traffic. While most previous works ignored user mobility, there are some recent works taking it into account. However, the duration of user contact times has been ignored, making it difficult to explicitly characterize the effect of mobility. In this paper, we adopt the alternating renewal process to model the duration of both the contact and inter-contact times, and investigate how the caching performance is affected by mobility. The data offloading ratio, i.e., the proportion of requested data that can be delivered via D2D links, is taken as the performance metric. We first approximate the distribution of the communication time for a given user by beta distribution through moment matching. With this approximation, an accurate expression of the data offloading ratio is derived. For the homogeneous case where the average contact and intercontact times of different user pairs are identical, we prove that the data offloading ratio increases with the user moving speed, assuming that the transmission rate remains the same. Simulation results are provided to show the accuracy of the approximate result, and also validate the effect of user mobility. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
ICC | 1 |
| 2017 | Real-time Task Scheduling for joint energy efficiency optimization in data centersabstractThe high energy consumption has become one bottleneck in the development of the data centers (DCs), where the main energy consumers are the cooling system and the servers. Therefore, the joint optimization for the energy efficiency of the cooling system and the servers is a crucial problem, while most of previous works on energy saving only studies one of these two components in an isolated manner. In this paper, we propose a real-time strategy, rTCS (real-time Task Classification and Scheduling strategy), to jointly optimize the energy efficiency of these two components in the scenario where the tasks arrive dynamically. Strategy rTCS first labels the tasks to classify them according to their run time and end time with a time complexity of O(1) and a bounded space complexity. Then, rTCS schedules the tasks in real time based on their labels and the energy consumption model of the DC. Simulation results show that rTCS can effectively improve the energy efficiency of DCs. Youshi Wang, Fa Zhang 0001, Rui Wang 0028, Yangguang Shi, Zhiyong Liu 0002 |
ISCC | 3 |
| 2017 | Mobility-Aware Caching in D2D NetworksabstractCaching at mobile devices can facilitate device-to-device (D2D) communications, which may significantly improve spectrum efficiency and alleviate the heavy burden on backhaul links. However, most previous works ignored user mobility, thus having limited practical applications. In this paper, we take advantage of the user mobility pattern by the inter-contact times between different users, and propose a mobility-aware caching placement strategy to maximize thedata offloading ratio, which is defined as the percentage of the requested data that can be delivered via D2D links rather than through base stations. Given the NP-hard caching placement problem, we first propose an optimal dynamic programming algorithm to obtain a performance benchmark with much lower complexity than exhaustive search. We then prove that the problem falls in the category of monotone submodular maximization over a matroid constraint, and propose a time-efficient greedy algorithm, which achieves an approximation ratio as$\frac {1}{2}$. Simulation results with real-life data sets will validate the effectiveness of our proposed mobility-aware caching placement strategy. We observe that users moving at either a very low or very high speed should cache the most popular files, while users moving at a medium speed should cache less popular files to avoid duplication. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 1 |
| 2016 | QoS-aware joint mode selection and channel assignment for D2D communicationsabstractUnderlaying device-to-device (D2D) communications to a cellular network is considered as a key technique to improve spectral efficiency in 5G networks. For such D2D systems, mode selection and resource allocation have been widely utilized for managing interference. However, previous works allowed at most one D2D link to access the same channel, while mode selection and resource allocation are typically separately designed. In this paper, we jointly optimize the mode selection and channel assignment in a cellular network with underlaying D2D communications, where multiple D2D links may share the same channel. Meanwhile, the QoS requirements for both cellular and D2D links are guaranteed, in terms of Signal-to-Interference-plus-Noise Ratio (SINR). We first propose an optimal dynamic programming (DP) algorithm, which provides a much lower computation complexity compared to exhaustive search and serves as the performance bench mark. A bipartite graph based greedy algorithm is then proposed to achieve a polynomial time complexity. Simulation results will demonstrate the advantage of allowing each channel to be accessed by multiple D2D links in dense D2D networks, as well as, the effectiveness of the proposed algorithms. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
ICC | 1 |
| 2016 | Optimal QoS-Aware Channel Assignment in D2D Communications With Partial CSIabstractIn this paper, we propose effective channel assignment algorithms for network utility maximization in a cellular network with underlaying device-to-device (D2D) communications. A major innovation is the consideration of partial channel state information (CSI), i.e., the base station (BS) is assumed to be able to acquire “partial” instantaneous CSI of the cellular and D2D links, as well as, the interference links. In contrast to the existing works, multiple D2D links are allowed to share the same channel, and the quality of service (QoS) requirements for both the cellular and D2D links are enforced. We first develop an optimal channel assignment algorithm based on dynamic programming, which enjoys a much lower complexity compared with exhaustive search and will serve as a performance benchmark. To further reduce complexity, we propose a cluster-based sub-optimal channel assignment algorithm. New closed-form expressions for the expected weighted sum rate and the successful transmission probabilities are also derived. Simulation results verify the effectiveness of the proposed algorithms. Moreover, by comparing different partial CSI scenarios, we observe that the CSI of the D2D communication links and the interference links from the D2D transmitters to the BS significantly affects the network performance, while the CSI of the interference links from the BS to the D2D receivers only has a negligible impact. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
IEEE Trans. Wirel. Commun. | 1 |
| 2015 | QoS-Aware Channel Assignment for Weighted Sum-Rate Maximization in D2D CommunicationsabstractUnderlaying device-to-device (D2D) communication links to a cellular network is a promising way to improve spectrum efficiency, for which the cross- link interference should be carefully controlled. Resource allocation has been widely utilized for managing interference in D2D networks. However, most previous works made simple assumptions by either ignoring the reliability requirement of D2D links or not allowing multiple D2D links to share the same channel. In this paper, we propose effective channel assignment algorithms to maximize the weighted sum-rate in a cellular network with underlaying D2D communications, where multiple D2D links are allowed to share the same channel. Meanwhile, the minimum Signal-to- Interference-plus-Noise Ratio (SINR) requirements for both cellular and D2D links are guaranteed. We first provide an optimal algorithm based on dynamic programming (DP) to serve as the performance benchmark, which enjoys much lower complexity compared to exhaustive search. To further reduce complexity, we then propose a cluster-based near-optimal channel assignment algorithm. Simulation results will demonstrate the advantage of allowing multiple D2D links to share the same channel in dense D2D networks, as well as verifying the effectiveness of the proposed algorithms. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
GLOBECOM | 1 |
| 2014 | Average throughput analysis of downlink cellular networks with multi-antenna base stationsabstractRandom spatial network models have been recently utilized in the performance analysis and system design for multi-cell networks. Such an approach has been mainly adopted to investigate the outage based system performance, such as the outage probability and outage throughput. However, these performance metrics are defined with a fixed-rate transmission, and cannot characterize the performance of data traffic, which normally adopts rate adaptation. In this paper, we will evaluate the average throughput of a space division multiple access (SDMA) based cellular network by considering stochastically distributed base stations (BSs) and mobile terminals (MTs). The major difficulty for the performance analysis is the complicated distribution of the interference links. We shall provide an analytical framework for evaluating the average throughput by using the Moment Generating Function (MGF) based method. Simulations will show that the proposed method is very accurate. In particular, the analytical result can be utilized to determine the optimum number of MTs to be served in SDMA networks that can maximize the network throughput. Rui Wang 0028, Jun Zhang 0004, Shenghui Song 0001, Khaled Ben Letaief |
PIMRC | 1 |