VLDB 2026 Research / reviewers in the wild / expert
Fei Wang 0073
dblp:52/3194-73
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2026
0009-0004-1142-6434ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XInsight: Integrative Stage-Consistent Psychological Counseling Support Agents for Digital Well-Being
Fei Wang 0073, Jiangnan Yang, Kun Li 0008, Yanyan Wei, Dan Guo 0001, Meng Wang 0001 |
WWW | 1 |
| 2026 | Modeling Long-Term Emotional Support Through Causal World Modeling With Imitation LearningabstractEmotional support conversation systems have emerged as a promising complement to traditional mental health consultations, offering context-aware dialogue to support seekers’ emotional well-being. Despite their potential, two fundamental challenges remain unresolved: 1) modeling long-term emotional trajectory beyond short-term relief; and 2) adapting support strategies to context in a psychologically coherent manner. To address these challenges, we propose CAIWO, a novel framework that integrates world modeling and causality-enhanced imitation learning to systematically support seekers, alleviate psychological stress, and restore emotional balance. Specifically, CAIWO comprises two core components. The first is an emotional world model, which captures long-term emotional trajectories from historical interactions to inform anticipatory guidance. The second is a causality-enhanced imitation learning module, which infers latent causal dependencies to facilitate coherent strategy transitions and mitigate compounding errors typical of conventional imitation learning. By incorporating the final latent variables into the response decoder, CAIWO dynamically adjusts the strategies and generates emotionally resonant responses. Extensive experiments on the ESConv benchmark demonstrate that CAIWO outperforms state-of-the-art baselines by 8.6%, significantly improving the generation of responses that align with seekers’ emotional development and psychological needs. Mingzheng Li, Fei Wang 0073, Kun Li 0008, Yanyan Wei, Yiqi Nie, Yanbin Hao, Xun Yang 0001, Meng Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | TG4MM: Time-Varying Gaussian Splatting for 3D Motion Magnificationabstract3D motion magnification aims to enable us to visualize subtle, imperceptible motions by integrating eulerian video magnification with novel view synthesis. Existing method extracts the variation of feature embeddings using Neural Radiance Fields (NeRF) over time. However, this volume rendering technique suffers from two shortcomings for 3D motion magnification: (1) When reconstructing time-varying scenes through volume rendering, spatial-temporal operations between static and dynamic representations often generate noticeable artifacts, leading to blurred magnified frames. (2) When processing high-resolution dynamic scenes, the intrinsically low rendering efficiency of these techniques causes excessive computational latency, preventing real-time visualization. In this work, instead of NeRF, we propose a novelTime-varying Gaussian Splatting for 3D Motion Magnification(TG4MM) that is capable of achieving real-time rendering while effectively handling blurred magnified frames in dynamic 3D motion magnification scenes. Specifically, we propose a motion-space decoupled triplane modeling approach. The space triplane captures major spatial structures from the first frame, while the motion triplane captures subtle motion information from subsequent frames. Furthermore, we develop a phase-based motion magnification module that enhances subtle motions by applying filters within the embedding space and subtle motion triplane. Experimental results demonstrate the effectiveness of our method, showing that it outperforms existing 3D motion magnification techniques and achieves a speed up to 126 FPS. Jiabao Guo, Fei Wang 0073, Jinyang Huang, Zhi Liu 0002, Dan Guo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Task-Generalized Adaptive Cross-Domain Learning for Multimodal Image FusionabstractMultimodal Image Fusion (MMIF) aims to integrate complementary information from different imaging modalities to overcome the limitations of individual sensors. It enhances image quality and facilitates downstream applications such as remote sensing, medical diagnostics, and robotics. Despite significant advancements, current MMIF methods still face challenges such as modality misalignment, high-frequency detail destruction, and task-specific limitations. To address these challenges, we propose AdaSFFuse, a novel framework for task-generalized MMIF through adaptive cross-domain co-fusion learning. AdaSFFuse introduces two key innovations: the Adaptive Approximate Wavelet Transform (AdaWAT) for frequency decoupling, and the Spatial-Frequency Mamba Blocks for efficient multimodal fusion. AdaWAT adaptively separates the high- and low-frequency components of multimodal images from different scenes, enabling fine-grained extraction and alignment of distinct frequency characteristics for each modality. The Spatial-Frequency Mamba Blocks facilitate cross-domain fusion in both spatial and frequency domains, enhancing this process. These blocks dynamically adjust through learnable mappings to ensure robust fusion across diverse modalities. By combining these components, AdaSFFuse improves the alignment and integration of multimodal features, reduces frequency loss, and preserves critical details. Extensive experiments on four MMIF tasks-Infrared-Visible Image Fusion (IVF), Multi-Focus Image Fusion (MFF), Multi-Exposure Image Fusion (MEF), and Medical Image Fusion (MIF)-demonstrate AdaSFFuse's superior fusion performance, ensuring both low computational cost and a compact network, offering a strong balance between performance and efficiency. The code will be publicly available athttps://github.com/Zhen-yu-Liu/AdaSFFuse. Kun Li 0008, Yu Wang 0224, Yuwei Wang 0002, Yanyan Wei, Fei Wang 0073 |
IEEE Trans. Multim. | 7 |
| 2025 | Hierarchical Matrix-Contrastive Bilateral Fusion for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) seeks to understand human sentiment by leveraging the correlations across multimodal data. Current approaches often employ contrastive learning and text-centric fusion methods to explore the sentiment mapping space, improve the ability to extract and integrate multimodal features, and capture modality correlations. However, these methods typically depend on complex sampling strategies to select predefined positive and negative samples and perform unidirectional fusion of other modalities aligned with the text. This process overlooks the collaborative information that could be shared between modalities, leading to a loss of valuable insights. To address these limitations, we propose Hierarchical MAtrix-Contrastive BiLateral FusiOn (HALO), which integrates two key components: Matrix-Aware Contrastive Learning (MACL) and Hierarchical Bilateral Fusion (HBF). Specifically, MACL uses two supervisory signals to sample positive and negative pairs within the same batch and assigns different weights according to the difficulty of samples, thereby enhancing the cross-modal discrimination ability of the model. In addition, HBF introduces a bilateral fusion method by guiding vision and audio fusion with text, while using vision and audio information to enhance the overall expressive ability of text. Extensive experiments on datasets MOSI and MOSEI demonstrate the effectiveness and superiority of HALO. Chaoxing Tang, Anyang Tong, Fei Wang 0073, Zhangling Duan |
ICMR | 3 |
| 2025 | MAC 2025: The 2nd Micro-Action Analysis Grand ChallengeabstractMicro-Actions (MAs) are a crucial form of non-verbal communication in social interactions, with promising applications in human emotion analysis. Although the topic has attracted considerable research interest, progress has been hindered by the lack of publicly available benchmark datasets. To address this gap, the Micro-Action Analysis Grand Challenge (MAC) is organized annually. This paper presents an overview of the 2nd Micro-Action Analysis Grand Challenge, held in conjunction with ACM Multimedia 2025. We provide a comprehensive summary of the challenge, including its dataset, evaluation protocol, results, and discussion. The top-ranked solutions are highlighted to offer valuable insights for researchers, and potential future directions are outlined to guide ongoing developments in this area. The goal of this grand challenge is to foster innovative research in micro-action analysis and advance research in the human-centric action understanding community. Kun Li 0008, Dan Guo 0001, Haoyu Chen 0001, Pengyu Liu 0005, Fei Wang 0073, Guoying Zhao 0001, Meng Wang 0001 |
ACM Multimedia | 6 |
| 2025 | Motion Matters: Motion-guided Modulation Network for Skeleton-based Micro-Action RecognitionabstractMicro-Actions (MAs) are an important form of non-verbal communication in social interactions, with potential applications in human emotional analysis. However, existing methods in Micro-Action Recognition often overlook the inherent subtle changes in MAs, which limits the accuracy of distinguishing MAs with subtle changes. To address this issue, we present a novel Motion-guided Modulation Network (MMN) that implicitly captures and modulates subtle motion cues to enhance spatial-temporal representation learning. Specifically, we introduce a Motion-guided Skeletal Modulation module (MSM) to inject motion cues at the skeletal level, acting as a control signal to guide spatial representation modeling. In parallel, we design a Motion-guided Temporal Modulation module (MTM) to incorporate motion information at the frame level, facilitating the modeling of holistic motion patterns in micro-actions. Finally, we propose a motion consistency learning strategy to aggregate the motion cues from multi-scale features for micro-action classification. Experimental results on the Micro-Action 52 and iMiGUE datasets demonstrate that MMN achieves state-of-the-art performance in skeleton-based micro-action recognition, underscoring the importance of explicitly modeling subtle motion cues. The code will be available at https://github.com/momiji-bit/MMN Jihao Gu, Kun Li 0008, Fei Wang 0073, Yanyan Wei, Zhiliang Wu, Hehe Fan, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2025 | Temporal Boundary Awareness Network for Repetitive Action CountingabstractRepetitive Action Counting (RAC) is a critical and challenging task in video analysis, aiming to count the number of repeated actions in videos accurately. Existing methods typically generate a Temporal Self-similarity Matrix (TSM) as an intermediate representation to predict the number of repetitive actions. While this simplifies the process, it often overlooks the variable lengths between action cycles and the phenomenon of motion interruptions. The period inconsistency problem caused by the change in the action period and the motion interruption problem resulting from the motion pause are the two main challenges that affect the accuracy of RAC in complex scenes. To address these challenges, we propose a novel framework. First, we construct a boundary-aware encoder equipped with a temporal pyramid structure to build multi-scale video features, capturing the period information of different lengths of repetitive actions to solve the period inconsistency problem. Next, a cycle and boundary attention module is followed by each layer in the pyramid to enhance these multi-scale features with periodic and event boundary information. Finally, we design a gated density estimator to generate the actionness score for each frame that reflects the probability of the corresponding time point being within the motion cycle. These scores are used to weight features to reduce the impact of noise frames without actions present and solve the motion interruption problem for better density prediction. Extensive experiments conducted on public datasets demonstrate the effectiveness of our method. The source code will be available at https://github.com/zqzhang2023/TBANRAC . Zhenqiang Zhang, Kun Li 0008, Shengeng Tang, Yanyan Wei, Fei Wang 0073, Jinxing Zhou, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Repetitive Action Counting with Feature Interaction Enhancement and Adaptive Gate Fusion
Kun Li 0008, Yanyan Wei, Fei Wang 0073, Jinxing Zhou, Dan Guo 0001 |
MMAsia | 4 |