VLDB 2026 Research / reviewers in the wild / expert
Lu Liu 0009
dblp:31/2088-9
· DBLP profile ↗
11ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-1303-9196ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conditional Information Bottleneck for Multimodal Fusion: Overcoming Shortcut Learning in Sarcasm DetectionabstractMultimodal sarcasm detection is a complex task that requires distinguishing subtle complementary signals across modalities while filtering out irrelevant information. Many advanced methods rely on learning shortcuts from datasets rather than extracting intended sarcasm-related features. However, our experiments show that shortcut learning impairs the model's generalization in real-world scenarios. Furthermore, we reveal the weaknesses of current modality fusion strategies for multimodal sarcasm detection through systematic experiments, highlighting the necessity of focusing on effective modality fusion for complex emotion recognition. To address these challenges, we construct MUStARD++R by removing shortcut signals from MUStARD++. Then, a Multimodal Conditional Information Bottleneck (MCIB) model is introduced to enable efficient multimodal fusion for sarcasm detection. Experimental results show that the MCIB achieves the best performance without relying on shortcut learning. Qi Jia 0004, Cong Xu 0001, Feiyu Chen 0005, Yuhan Liu 0014, Haotian Zhang 0017, Lu Liu 0009, Zhichun Wang |
AAAI | 8 |
| 2026 | Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Xiaoli Gong, Jin Zhang 0003 |
Signal Process. | 3 |
| 2026 | Fine-Grained Audio-Visual Event LocalizationabstractAudio-visual event localization (AVEL) aims to recognize events in videos by associating audio-visual information. However, events involved in existing AVEL tasks are usually coarse-grained events. Actually, finer-grained events are sometimes necessary to be distinguished, especially in certain expert-level applications or rich-content-generation studies. However, this is challenging because they are more difficult to detect or distinguish compared with coarse-grained events. To better address this problem, we discuss a new setting of fine-grained AVEL from dataset to method. First, we constructed the first fine-grained audio-visual event dataset, which is called IT-AVE, relying on videos of playing musical instruments, containing 13k video clips and over 52k audio-visual events. All events are labeled from professional music practitioners, and the event categories are all derived from playing techniques, which are fine-grained with little interclass variation. Next, we designed a new fine-grained event localization method, spatial-temporal video event detector (SVED), which focuses on the challenges that fine-grained events are more imperceptible and prone to be disturbed. Finally, we conduct extensive experiments based on the proposed IT-AVE dataset versus fine-grained versions of two existing related datasets, including UnAV-22 derived from UnAV-100 and FineAction-AV derived from FineAction. Experimental results demonstrate the effectiveness of our method. We hope that this work will contribute to the exploration of an integrated understanding of audio-visual videos. Baoyu Fan, Lu Liu 0009, Xiaochuan Li 0001, Jin Zhang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan |
ICCV | 6 |
| 2024 | Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and BaselineabstractExisting video multi-modal sentiment analysis mainly focuses on the sentiment expression of people within the video, yet often neglects the induced sentiment of viewers while watching the videos. Induced sentiment of viewers is essential for inferring the public response to videos and has broad application in analyzing public societal sentiment, effectiveness of advertising and other areas. The micro videos and the related comments provide a rich application scenario for viewers’ induced sentiment analysis. In light of this, we introduces a novel research task, Multimodal Sentiment Analysis for Comment Response of Video Induced(MSA-CRVI), aims to infer opinions and emotions according to comments response to micro video. Meanwhile, we manually annotate a dataset named Comment Sentiment toward to Micro Video (CSMV) to support this research. It is the largest video multi-modal sentiment dataset in terms of scale and video duration to our knowledge, containing 107, 267 comments and 8, 210 micro videos with a video duration of 68.83 hours. To infer the induced sentiment of comment should leverage the video content, we propose the Video Content-aware Comment Sentiment Analysis (VC-CSA) method as a baseline to address the challenges inherent in this new task. Extensive experiments demonstrate that our method is showing significant improvements over other established baselines. We make the dataset and source code publicly available at https://github.com/IEIT-AGI/MSA-CRVI. Qi Jia 0004, Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Guoguang Du 0001, Zhenhua Guo 0003, Yaqian Zhao, Xuanjing Huang 0001, RenGang Li |
NeurIPS | 4 |
| 2022 | Multimodal face aging framework via learning disentangled representation
Lu Liu 0009, Shenghui Wang 0003 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Learning shape and texture progression for young child face aging
Lu Liu 0009, Shenghui Wang 0003 |
Signal Process. Image Commun. | 1 |
| 2020 | Joint beamforming and power allocation using deep learning for D2D communication in heterogeneous networksabstractDevice‐to‐device (D2D) communication plays a significant role in cellular networks as it can increase the capacity, spectrum efficiency and energy efficiency of the system. However, the large computational complexity of D2D resource management optimisation algorithms creates a serious gap between theoretical design and real‐time processing, which leads to the limited use of D2D communication technology. In this study, a novel deep learning‐based optimisation method is proposed to overcome the high computational complexity of joint beamforming design and power allocation optimisation algorithms in D2D communication. Unlike existing approaches, the authors design a convolutional neural network based end‐to‐end network structure to solve complex computing problems for channel state information under a limited feedback scenario. The Max‐SE loss function which indicates quality‐of‐service (QoS) constraint and interference constraint, together with the mean squared error (MSE) function, are designed to maximise the spectral efficiency of the system while minimising the total transmit power. The simulation results show that the proposed approach can achieve performance comparable to the weighted minimum MSE scheme with low computation time. Yuejiao Wang, Shenghui Wang 0003, Lu Liu 0009 |
IET Commun. | 3 |
| 2019 | Generating Responses with a Specific Emotion in DialogabstractIt is desirable for dialog systems to have capability to express specific emotions during a conversation, which has a direct, quantifiable impact on improvement of their usability and user satisfaction.After a careful investigation of real-life conversation data, we found that there are at least two ways to express emotions with language.One is to describe emotional states by explicitly using strong emotional words; another is to increase the intensity of the emotional experiences by implicitly combining neutral words in distinct ways.We propose an emotional dialogue system (EmoDS) that can generate the meaningful responses with a coherent structure for a post, and meanwhile express the desired emotion explicitly or implicitly within a unified framework.Experimental results showed EmoDS performed better than the baselines in BLEU, diversity and the quality of emotional expression. Zhenqiao Song, Xiaoqing Zheng, Lu Liu 0009, Mu Xu, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2018 | Attention-based Belief or Disbelief Feature Extraction for Dependency Parsing
Haoyuan Peng, Lu Liu 0009, Yi Zhou 0018, Junying Zhou, Xiaoqing Zheng |
AAAI | 2 |
| 2018 | RNN-Based Sequence-Preserved Attention for Dependency ParsingabstractRecurrent neural networks (RNN) combined with attention mechanism has proved to be useful for various NLP tasks including machine translation, sequence labeling and syntactic parsing. The attention mechanism is usually applied by estimating the weights (or importance) of inputs and taking the weighted sum of inputs as derived features. Although such features have demonstrated their effectiveness, they may fail to capture the sequence information due to the simple weighted sum being used to produce them. The order of the words does matter to the meaning or the structure of the sentences, especially for syntactic parsing, which aims to recover the structure from a sequence of words. In this study, we propose an RNN-based attention to capture the relevant and sequence-preserved features from a sentence, and use the derived features to perform the dependency parsing. We evaluated the graph-based and transition-based parsing models enhanced with the RNN-based sequence-preserved attention on the both English PTB and Chinese CTB datasets. The experimental results show that the enhanced systems were improved with significant increase in parsing accuracy. Yi Zhou 0018, Junying Zhou, Lu Liu 0009, Jiangtao Feng, Haoyuan Peng, Xiaoqing Zheng |
AAAI | 3 |