VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0185
dblp:10/4661-185
· DBLP profile ↗
16ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-1533-6979ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Missing-Modality-Aware Multimodal Sentiment Analysis via Mutual Information Assessment and Prompted Pre-trained Models
Rongfei Chen, Siwei Cheng, Zhipeng Li 0002, Wei Zhang 0185 |
ICIC (16) | 6 |
| 2026 | A novel two-stage intelligent Kalman filter for maneuvering target tracking
Benqi Zhao, Gaoliang Peng, Wei Zhang 0185, Shiji Zhang |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Understanding domain-specific attribute constraints in multimodal sentiment analysis via ensemble multimodal large language models with structured multistage prompts
Rongfei Chen, Junlong Tong, Xiaoyu Shen 0001, Wei Zhang 0185 |
Expert Syst. Appl. | 6 |
| 2026 | Curriculum-Enhanced Reinforcement Learning for Robust Humanoid LocomotionabstractThe control of humanoid locomotion remains one of the most formidable challenges in robotics. Conventional model-based approaches not only rely extensively on manual design but also exhibit limited generalization across diverse tasks and environments. To overcome these limitations, a curriculum-enhanced reinforcement learning framework is proposed in this work for training robust locomotion. Given the strong temporal dependencies inherent in humanoid locomotion, Mamba-2 is employed as the backbone of the Actor-Critic network, enabling efficient temporal modeling of historical observations and content-based reasoning with linear spatiotemporal complexity. Furthermore, to eliminate the tendency of humanoid robots to favor low-speed motion in order to maintain gait stability and balance, which often leads to inaccurate velocity command tracking, a command-guided curriculum learning (CGCL) method is proposed to improve responsiveness to velocity commands. Experimental results demonstrate that the proposed Mamba-based framework achieves state-of-the-art (SOTA) performance, while CGCL significantly enhances the accuracy of velocity command tracking. Moreover, sim-to-real transfer experiments confirm both the robustness and the seamless deployability of the proposed framework on physical humanoid platforms. Jianbin Qiu, Shixiang Jia, Fenglei Ni, Wei Zhang 0185 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsabstractMultimodal Large Language Models (MLLMs) have achieved strong performance across vision-language tasks, but suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens.Though efforts have been made to prune tokens in MLLMs, they lack a fundamental understanding of how MLLMs process and fuse multimodal information.Through systematic analysis, we uncover a three-stage cross-modal interaction process: (1) Shallow layers recognize task intent, with visual tokens acting as passive attention sinks; (2) Cross-modal fusion occurs abruptly in middle layers, driven by a few critical visual tokens; (3) Deep layers discard vision tokens, focusing solely on linguistic refinement.Based on these findings, we propose VisiPruner, a training-free pruning framework that reduces up to 99% of visionrelated attention computations and 53.9% of FLOPs on LLaVA-v1.5 7B.It significantly outperforms existing token pruning methods and generalizes across diverse MLLMs.Beyond pruning, our insights further provide actionable guidelines for training efficient MLLMs by aligning model architecture with its intrinsic layer-wise processing dynamics. Yingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong, Hui Su, Yijie Pan, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 7 |
| 2025 | PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing TasksabstractWe present PricingLogic, the first benchmark that probes whether Large Language Models (LLMs) can reliably automate tourism-related prices when multiple, overlapping fare rules apply.Travel agencies are eager to offload this error-prone task onto AI systems; however, deploying LLMs without verified reliability could result in significant financial losses and erode customer trust.PricingLogic comprises 300 natural-language questions based on booking requests derived from 42 real-world pricing policies, spanning two levels of difficulty: (i) basic customer-type pricing and (ii) bundled-tour calculations involving interacting discounts.Evaluations of a line of LLMs reveal a steep performance drop on the harder tier, exposing systematic failures in rule interpretation and arithmetic reasoning.These results highlight that, despite their general capabilities, today's LLMs remain unreliable in revenuecritical applications without further safeguards or domain adaptation. Yunuo Liu, Zena Al-Khalili, Dai Cheng, Yanjun Chen 0001, Dietrich Klakow, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 7 |
| 2025 | Context Guided Transformer Entropy Modeling for Video CompressionabstractConditional entropy models effectively leverage spatio-temporal contexts to reduce video redundancy. However, incorporating temporal context often introduces additional model complexity and increases computational cost. In parallel, many existing spatial context models lack explicit modeling the ordering of spatial dependencies, which may limit the availability of relevant context during decoding. To address these issues, we propose the Context Guided Transformer (CGT) entropy model, which estimates probability mass functions of the current frame conditioned on resampled temporal context and dependency-weighted spatial context. A temporal context resampler learns predefined latent queries to extract critical temporal information using transformer encoders, reducing downstream computational overhead. Meanwhile, a teacher-student network is designed as dependency-weighted spatial context assigner to explicitly model the dependency of spatial context order. The teacher generates an attention map to represent token importance and an entropy map to reflect prediction certainty from randomly masked inputs, guiding the student to select the weighted top-k tokens with the highest spatial dependency. During inference, only the student is used to predict undecoded tokens based on high-dependency context. Experimental results demonstrate that our CGT model reduces entropy modeling time by approximately 65% and achieves an 11% BD-Rate reduction compared to the previous state-of-the-art conditional entropy model. Junlong Tong, Wei Zhang 0185, Yaohui Jin, Xiaoyu Shen 0001 |
ICCV | 2 |
| 2025 | Beyond Content Relevance: Evaluating Instruction Following in Retrieval ModelsabstractInstruction-following capabilities in large language models (LLMs) have progressed significantly, enabling more complex user interactions through detailed prompts. However, retrieval systems have not matched these advances, most of them still relies on traditional lexical and semantic matching techniques that fail to fully capture user intent. Recent efforts have introduced instruction-aware retrieval models, but these primarily focus on intrinsic content relevance, which neglects the importance of customized preferences for broader document-level attributes. This study evaluates the instruction-following capabilities of various retrieval models beyond content relevance, including LLM-based dense retrieval and reranking models. We develop InfoSearch, a novel retrieval evaluation benchmark spanning six document-level attributes: Audience, Keyword, Format, Language, Length, and Source, and introduce novel metrics -- Strict Instruction Compliance Ratio (SICR) and Weighted Instruction Sensitivity Evaluation (WISE) to accurately assess the models' responsiveness to instructions. Our findings indicate that although fine-tuning models on instruction-aware retrieval datasets and increasing model size enhance performance, most models still fall short of instruction compliance. We release our dataset and code on https://github.com/EIT-NLP/InfoSearch. Jianqun Zhou, Yuanlei Zheng, Zeyuan Shang, Wei Zhang 0185, Xiaoyu Shen 0001 |
ICLR | 6 |
| 2025 | MAER-Nav: Bidirectional Motion Learning Through Mirror-Augmented Experience Replay for Robot NavigationabstractDeep Reinforcement Learning (DRL) based navigation methods have demonstrated promising results for mobile robots, but suffer from limited action flexibility in confined spaces. Conventional DRL approaches predominantly learn forward-motion policies, causing robots to become trapped in complex environments where backward maneuvers are necessary for recovery. This paper presents MAER-Nav (Mirror-Augmented Experience Replay for Robot Navigation), a novel framework that enables bidirectional motion learning without requiring explicit failure-driven hindsight experience replay or reward function modifications. Our approach integrates a mirror-augmented experience replay mechanism with curriculum learning to generate synthetic backward navigation experiences from successful trajectories. Experimental results in both simulation and real-world environments demonstrate that MAER-Nav significantly outperforms state-of-the-art methods while maintaining strong forward navigation capabilities. The framework effectively bridges the gap between the comprehensive action space utilization of traditional planning methods and the environmental adaptability of learning-based approaches, enabling robust navigation in scenarios where conventional DRL methods consistently fail. Shanze Wang, Mingao Tan, Biao Huang 0016, Xiaoyu Shen 0001, Hailong Huang 0001, Wei Zhang 0185 |
IROS | 7 |
| 2025 | Enhancing Deep Reinforcement Learning-based Robot Navigation Generalization through Scenario AugmentationabstractThis work focuses on enhancing the generalization performance of deep reinforcement learning-based robot navigation in unseen environments. We present a novel data augmentation approach called scenario augmentation, which enables robots to navigate effectively across diverse settings without altering the training scenario. The method operates by mapping the robot’s observation into an imagined space, generating an imagined action based on this transformed observation, and then remapping this action back to the real action executed in simulation. Through scenario augmentation, we conduct extensive comparative experiments to investigate the underlying causes of suboptimal navigation behaviors in unseen environments. Our analysis indicates that limited training scenarios represent the primary factor behind these undesired behaviors. Experimental results confirm that scenario augmentation substantially enhances the generalization capabilities of deep reinforcement learning-based navigation systems. The improved navigation framework demonstrates exceptional performance by producing near-optimal trajectories with significantly reduced navigation time in real-world applications. Shanze Wang, Mingao Tan, Xianghui Wang, Xiaoyu Shen 0001, Hailong Huang 0001, Wei Zhang 0185 |
IROS | 7 |
| 2024 | The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language ModelsabstractReinforcement Learning from Human Feedback significantly enhances Natural Language Processing by aligning language models with human expectations.A critical factor in this alignment is the strength of reward models used during training.This study explores whether stronger reward models invariably lead to better language models.In this paper, through experiments on relevance, factuality, and completeness tasks using the QA-FEEDBACK dataset and reward models based on Longformer, we uncover a surprising paradox: language models trained with moderately accurate reward models outperform those guided by highly accurate ones.This challenges the widely held belief that stronger reward models always lead to better language models, and opens up new avenues for future research into the key factors driving model performance and how to choose the most suitable reward models. Yanjun Chen 0001, Yirong Sun, Xinghao Chen 0009, Wei Zhang 0185, Xiaoyu Shen 0001 |
EMNLP | 5 |
| 2024 | Assessing "Implicit" Retrieval Robustness of Large Language ModelsabstractRetrieval-augmented generation has gained popularity as a framework to enhance large language models with external knowledge.However, its effectiveness hinges on the retrieval robustness of the model.If the model lacks retrieval robustness, its performance is constrained by the accuracy of the retriever, resulting in significant compromises when the retrieved context is irrelevant.In this paper, we evaluate the "implicit" retrieval robustness of various large language models, instructing them to directly output the final answer without explicitly judging the relevance of the retrieved context.Our findings reveal that fine-tuning on a mix of gold and distracting context significantly enhances the model's robustness to retrieval inaccuracies, while still maintaining its ability to extract correct answers when retrieval is accurate.This suggests that large language models can implicitly handle relevant or irrelevant retrieved context by learning solely from the supervision of the final answer in an end-toend manner.Introducing an additional process for explicit relevance judgment can be unnecessary and disrupts the end-to-end approach.1 Xiaoyu Shen 0001, Rexhina Blloshmi, Jiahuan Pei, Wei Zhang 0185 |
EMNLP | 5 |
| 2019 | An End-to-End Learning Approach for Multimodal Emotion Recognition: Extracting Common and Private InformationabstractMultimodal emotion recognition is important for facilitating efficient interaction between humans and machines. To better detect emotional states from multimodal data, we need to effectively extract both the common information that captures dependencies among different modalities, and the private information that characterizes variations in each modality. However, existing works are mostly designed to pursue either one of these objectives but not both. In our work, we propose an end-to-end learning approach to simultaneously extract the common and private information for multimodal emotion recognition. Specifically, we use a correlation loss based on Hirschfeld-Gebelein-Renyi (HGR) maximal correlation and a reconstruction loss based on autoencoders to preserve the common and private information, respectively. Experimental results on eNTERFACE'05 database and RML database demonstrate the effectiveness of our proposed approach. Fei Ma 0006, Wei Zhang 0185, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001 |
ICME | 2 |
| 2018 | Speech Emotion Recognition via Attention-based DNN from Multi-Task LearningabstractSpeech unlocks the huge potentials in emotion recognition. High accurate and real-time understanding of human emotion via speech assists Human-Computer Interaction. Previous works are often limited in either coarse-grained emotion learning tasks or the low precisions on the emotion recognition. To solve these problems, we construct a real-world large-scale corpus composed of 4 common emotions (i.e., anger, happiness, neutral and sadness). We also propose a multi-task attention-based DNN model (i.e., MT-A-DNN) on the emotion learning. MT-A-DNN efficiently learns the high-order dependency and non-linear correlations underlying in the audio data. Extensive experiments show that MT-A-DNN outperforms conventional methods on the emotion recognition. It could take one step further on the real-time acoustic emotion recognition in many smart audio-devices. Fei Ma 0006, Weixi Gu, Wei Zhang 0185, Shiguang Ni, Shao-Lun Huang, Lin Zhang 0001 |
SenSys | 3 |
| 2018 | Multimodal Emotion Recognition by extracting common and modality-specific informationabstractEmotion recognition technologies have been widely used in numerous areas including advertising, healthcare and online education. Previous works usually recognize the emotion from either the acoustic or the visual signal, yielding unsatisfied performances and limited applications. To improve the inference capability, we present a multimodal emotion recognition model, EMOdal. Apart from learning the audio and visual data respectively, EMOdal efficiently learns the common and modality-specific information underlying the two kinds of signals, and therefore improves the inference ability. The model has been evaluated on our large-scale emotional data set. The comprehensive evaluations demonstrate that our model outperforms traditional approaches. Wei Zhang 0185, Weixi Gu, Fei Ma 0006, Shiguang Ni, Lin Zhang 0001, Shao-Lun Huang |
SenSys | 1 |
| 2018 | ACDIN: Bridging the gap between artificial and real bearing damages for bearing fault diagnosis
Yuanhang Chen, Gaoliang Peng, Chaohao Xie, Wei Zhang 0185, Chuanhao Li 0002, Shaohui Liu |
Neurocomputing | 4 |