Fei Ma 0006

dblp:22/1199-6 · DBLP profile ↗
← Back
45ranked-venue papers
5as first author
38since 2021 · last 2026
0009-0002-5388-9125ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 2 first-author · 22 since 2021Artificial intelligence and machine learning · 22 · 1 first-author · 20 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
abstract
Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and idealized, failing to reflect the complexity and unpredictability of real-world environments, particularly the presence of anomalies. To bridge this research gap, we propose D-GARA, a dynamic benchmarking framework, to evaluate Android GUI agent robustness in real-world anomalies. D-GARA introduces a diverse set of real-world anomalies that GUI agents commonly face in practice, including interruptions such as permission dialogs, battery warnings, and update prompts. Based on D-GARA framework, we construct and annotate a benchmark featuring commonly used Android applications with embedded anomalies to support broader community research. Comprehensive experiments and results demonstrate substantial performance degradation in state-of-the-art GUI agents when exposed to anomaly-rich environments, highlighting the need for robustness-aware learning. D-GARA is modular and extensible, supporting the seamless integration of new tasks, anomaly types, and interaction scenarios to meet specific evaluation goals.
Yi Bin, Fei Ma 0006, Wenqi Shao, Zheng Wang 0044
AAAI4
2026 A Principle-Driven Adaptive Policy for Group Cognitive Stimulation Dialogue for Elderly with Cognitive Impairment
abstract
Cognitive impairment is becoming a major public health challenge. Cognitive Stimulation Therapy (CST) is an effective intervention for cognitive impairment, but traditional methods are difficult to scale, and existing digital systems struggle with group dialogues and cognitive stimulation principles. While Large Language Models (LLMs) are powerful, their application in this context faces key challenges: cognitive stimulation dialogue paradigms, a lack of therapeutic reasoning, and static-only user modeling. To address these issues, we propose a principle-driven adaptive policy actualized through a Group Cognitive Stimulation Dialogue (GCSD) system. We first construct a dataset with over 500 hours of real-world CST conversations and 10,000+ simulated dialogues generated via our Principle-Guided Scenario Simulation strategy. Our GCSD system then integrates four core modules to overcome LLM limitations: (i) a multi-speaker context controller to resolve role confusion; (ii) dynamic participant cognitive state modeling for personalized interaction; (iii) a cognitive stimulation-focused attention loss to instill cognitive stimulation reasoning; and (iv) a multi-dimensional reward strategy to enhance response value. Experimental results demonstrate that GCSD significantly outperforms baseline models across various evaluation metrics. Future work will focus on long-term clinical validation to bridge the gap between computational performance and clinical efficacy.
Jiyue Jiang, Pengan Chen, Jingqi Zhou, Zheyong Zhu, He Hu 0008, Fei Ma 0006, Qi Tian 0001
AAAI8
2026 TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling
He Hu 0008, Chiyuan Ma, Qianning Wang, Lin Liu 0016, Yucheng Zhou 0001, Laizhong Cui, Fei Ma 0006, Qi Tian 0001
WWW7
2026 Restoring neural radiance fields performance under adverse weather conditions
Ying He 0006, Gan Chen, F. Richard Yu, Ming Li 0073, Fei Ma 0006, Guang Zhou
Eng. Appl. Artif. Intell.5
2026 MindDialog: A large-scale benchmark for counseling dialogue understanding and generation
He Hu 0008, Juzheng Si, Qianning Wang, Tengjin Weng, Yihong Ji, Jiyue Jiang, Fei Ma 0006, Yucheng Zhou 0001, Laizhong Cui, Qi Tian 0001
Pattern Recognit.7
2025 Subgraph Invariant Learning Towards Large-Scale Graph Node Classification
abstract
Graph Neural Networks (GNNs) have shown efficacy in graph node classification, but face computational challenges on large-scale graphs. Although existing graph reduction methods address these issues, they still require high computational resources and fail to prioritize robust performance on out-of-distribution data. To tackle these challenges, we introduce the subgraph invariant learning paradigm, inspired by the small-world phenomenon. This approach enables models trained on specific subgraphs to generalize across diverse subgraphs, reducing computational demands, and enhancing scalability. To promote generalization, we maximize the invariance log-likelihood by deriving a theoretical lower bound of it and formulating the InVar loss. This loss minimizes the discrepancy between node representations and their corresponding invariance representations while maximizing the entropy of the node representation. In response to InVar loss, we propose the Invariance Facilitation Model (IFM), comprising the Invariance Representation Encoder (IRE) and Node Representation Encoder (NRE). IRE, capturing the invariance representations, utilizes Invariance ATTention (InvarATT) to compress long-range dependencies, while NRE learns the node representation, by integrating invariance representations via Telematic ATTention (TeleATT) and exchanging local information within each subgraph through GNNs. Evaluations on four large-scale graph datasets demonstrate the effectiveness, computational efficiency, and interpretability of IFM for large-scale graph node classification.
Leilei Wang, Fei Ma 0006, F. Richard Yu, Pengteng Li, Ying Tiffany He
AAAI3
2025 ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters
abstract
Pose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited notable performance, they tend to encounter issues such as specific region blurring, background sharpening, and decreased identity consistency. In this paper, we introduce ReMask-Animate, which utilizes masks as additional priors to guide the model's local visual attention to specific areas, thereby alleviating feature confusion between different regions of the image. Three distinct mask-guided adapters are designed for cross-condition regional fusion of hand and face pose features, mitigating feature confusion between the foreground and background, and enhancing the visual consistency of character identity. Moreover, these lightweight adapters introduce minimal computational overhead and can be seamlessly integrated into specific layers of the backbone architecture. Extensive experiments show that our method outperforms state-of-the-art methods on five metrics in public datasets. Additionally, qualitative evaluations highlight a significant improvement in the quality of generated videos, demonstrating our approach's superiority.
Xunzhi Xiang, Haiwei Xue, Zonghong Dai, Minglei Li 0001, Ye Yue, Fei Ma 0006, Weijiang Yu, Heng Chang, F. Richard Yu
AAAI7
2025 PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis
abstract
Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of training video. However, due to the limited scale of the training data, these methods often exhibit poor performance in audio-lip synchronization and visual quality. In this paper, we propose a novel 3D Gaussian-based method called PointTalk, which constructs a static 3D Gaussian field of the head and deforms it in sync with the audio. It also incorporates an audio-driven dynamic lip point cloud as a critical component of the conditional information, thereby facilitating the effective synthesis of talking heads. Specifically, the initial step involves generating the corresponding lip point cloud from the audio signal and capturing its topological structure. The design of the dynamic difference encoder aims to capture the subtle nuances inherent in dynamic lip movements more effectively. Furthermore, we integrate the audio-point enhancement module, which not only ensures the synchronization of the audio signal with the corresponding lip point cloud within the feature space, but also facilitates a deeper understanding of the interrelations among cross-modal conditional features. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking head synthesis compared to previous methods.
Xin Zhang 0169, Xiangyang Luo 0002, Weijiang Yu, Heng Chang, Fei Ma 0006, F. Richard Yu
AAAI8
2025 VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
abstract
Visual Language Models (VLMs) have rapidly progressed with the recent success of large language models. However, there have been few attempts to incorporate efficient linear Recurrent Neural Networks (RNNs) architectures into VLMs. In this study, we introduce VisualRWKV, the first application of a linear RNN model to multimodal learning tasks, leveraging the pre-trained RWKV language model. We propose a data-dependent recurrence and sandwich prompts to enhance our modeling capabilities, along with a 2D image scanning mechanism to enrich the processing of visual sequences. Extensive experiments demonstrate that VisualRWKV achieves competitive performance compared to Transformer-based models like LLaVA-1.5 on various benchmarks. Compared to LLaVA-1.5, VisualRWKV has a speed advantage of 3.98 times and can save 54% of GPU memory when reaching an inference length of 24K tokens. To facilitate further research and analysis, we have made the checkpoints and the associated code publicly accessible at the following GitHub repository: https://github.com/howard-hou/VisualRWKV.
Haowen Hou, Peigen Zeng, Fei Ma 0006, F. Richard Yu
COLING3
2025 Training-Free Language-Guided Video Summarization via Multi-Grained Saliency Scoring
Yongwei Nie, Fei Ma 0006, Keke Tang, F. Richard Yu, Hongmin Cai, Ping Li 0016
CVM (3)3
2025 ABM++: Learning Generalizable Manipulation Policies with a Mask-Guided World Model
abstract
Achieving robust generalization across diverse scenarios is crucial for advancing the practical application of robotics. Existing approaches typically rely on interaction object masks as visual inputs to predict subsequent actions, gaining a certain degree of generalization capability. However, these methods primarily map visual inputs and task-relevant object masks to expert actions, overlooking the environmental dynamics that govern physical interactions among objects during manipulation. To overcome these limitations, we introduce ABM++, a novel framework that leverages pre-trained VLMs to build a mask-guided world model (MGWM) within an imitation learning paradigm for generalized robotic manipulation. Specifically, we extend a world model into a coarse-to-fine imitation learning framework to reconstruct future mask-dominated visual features, which enables the model to capture state transitions between the current and next states based on predicted actions, effectively modeling environmental dynamics. Comprehensive experiments demonstrate that ABM++ significantly surpasses established baselines in both simulation and real-world environments, achieving a relative improvement of 12.2% across 8 complex tasks, which underscores the superiority of our method.
Fan Zhuo, Ying He 0006, F. Richard Yu, Pengshuai Yin, Fei Ma 0006
ECAI5
2025 DEP-SLAM: A Dynamic Environment Perception SLAM System with Large Language Models
abstract
Inderscience is a global company, a dynamic leading independent journal publisher disseminates the latest research across the broad fields of science, engineering and technology; management, public and business administration; environment, ecological economics and sustainable development; computing, ICT and internet/web services, and related areas.
Ying He 0006, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
ICASSP3
2025 DictAvatar: Expressive Facial Avatar Reconstruction with Facial Feature Dictionary
Zuyin Wu, Ying Tiffany He, Fei Ma 0006, F. Richard Yu
ICIC (6)5
2025 UniSync: A Unified Framework for Audio-Visual Synchronization
abstract
Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and end-to-end learning techniques. However, these methods often rely on limited audio-visual representations and suboptimal learning strategies, potentially constraining their effectiveness in more complex scenarios. To address these limitations, we present UniSync, a novel approach for evaluating audio-visual synchronization using embedding similarities. UniSync offers broad compatibility with various audio representations (e.g., Mel spectrograms, HuBERT) and visual representations (e.g., RGB images, face parsing maps, facial landmarks, 3DMM), effectively handling their significant dimensional differences. We enhance the contrastive learning framework with a margin-based loss component and cross-speaker unsynchronized pairs, improving discriminative capabilities. UniSync outperforms existing methods on standard datasets and demonstrates versatility across diverse audio-visual representations. Its integration into talking face generation frameworks enhances synchronization quality in both natural and AI-generated content.
Xun Guan, Jiyuan Song, Fei Ma 0006, F. Richard Yu
ICME6
2025 Object Isolated Attention for Consistent Story Visualization
abstract
Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline. One of the main difficulties is maintaining character consistency while creating natural and contextually fitting scenes—an area where many existing methods struggle. In this paper, we propose an enhanced Transformer module that uses separate self attention and cross attention mechanisms, leveraging prior knowledge from pre-trained diffusion models to ensure logical scene creation. The isolated self attention mechanism improves character consistency by refining attention maps to reduce focus on irrelevant areas and highlight key features of the same character. Meanwhile, the isolated cross attention mechanism independently processes each character’s features, avoiding feature fusion and further strengthening consistency. Notably, our method is training-free, allowing the continuous generation of new characters and storylines without re-tuning. Both qualitative and quantitative evaluations show that our approach outperforms current methods, demonstrating its effectiveness.
Xiangyang Luo 0002, Xin Zhang 0169, Fei Ma 0006, F. Richard Yu
ICME7
2025 MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach
abstract
Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in text-driven face editing, they still face significant challenges as none of them simultaneously fulfill the characteristics of diversity, controllability and flexibility. To address this challenge, we propose MuseFace, a text-driven face editing framework, which relies solely on text prompt to enable face editing. Specifically, MuseFace integrates a Text-to-Mask diffusion model and a semantic-aware face editing model, capable of directly generating fine-grained semantic masks from text and performing face editing. The Text-to-Mask diffusion model provides diversity and flexibility to the framework, while the semantic-aware face editing model ensures controllability of the framework. Our framework can create fine-grained semantic masks, making precise face editing possible, and significantly enhancing the controllability and flexibility of face editing models. Extensive experiments demonstrate that MuseFace achieves superior high-fidelity performance.
Xin Zhang 0169, Siting Huang, Xiangyang Luo 0002, Weijiang Yu, Heng Chang, Fei Ma 0006, F. Richard Yu
ICME7
2025 Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction
abstract
Recent advancements in implicit 3D reconstruction methods, e.g., neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https://github.com/Inter3D-ui/Inter3D.
Gan Chen, Ying He 0006, Mulin Yu, F. Richard Yu, Fei Ma 0006, Ming Li 0073, Guang Zhou
IJCAI6
2025 Active Multimodal Distillation for Few-shot Action Recognition
abstract
Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a novel framework that actively identifies reliable modalities for each sample using task-specific contextual cues, thus significantly improving recognition performance. Our framework integrates an Active Sample Inference (ASI) module, which utilizes active inference to predict reliable modalities based on posterior distributions and subsequently organizes them accordingly. Unlike reinforcement learning, active inference replaces rewards with evidence-based preferences, making more stable predictions. Additionally, we introduce an active mutual distillation module that enhances the representation learning of less reliable modalities by transferring knowledge from more reliable ones. Adaptive multimodal inference is employed during the meta-test to assign higher weights to reliable modalities. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing approaches.
Weijia Feng, Ruojia Zhang, Chenyang Wang 0001, Fei Ma 0006, Xiaobao Wang
IJCAI5
2025 VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening
abstract
We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation.
Haiwei Xue, Zhensong Zhang, Minglei Li 0001, Zonghong Dai, F. Richard Yu, Fei Ma 0006, Zhiyong Wu 0001
IJCAI6
2025 GaussianPU: Color Point Cloud Upsampling via 3D Gaussian Splatting
abstract
Dense colored point clouds enhance visual perception and are of significant value in various robotic applications. However, existing learning-based point cloud upsampling methods are constrained by computational resources and batch processing strategies, which often require subdividing point clouds into smaller patches, leading to distortions that degrade perceptual quality. To address this challenge, we propose a novel 2D-3D hybrid colored point cloud upsampling framework (GaussianPU) based on 3D Gaussian Splatting (3DGS) for robotic perception. This approach leverages 3DGS to bridge 3D point clouds with their 2D rendered images in robot vision systems. A dual scale rendered image restoration network transforms sparse point cloud renderings into dense representations, which are then input into 3DGS along with precise robot camera poses and interpolated sparse point clouds to reconstruct dense 3D point clouds. We have made a series of enhancements to the vanilla 3DGS, enabling precise control over the number of points and significantly boosting the quality of the upsampled point cloud for robotic scene understanding. Our framework supports processing entire point clouds on a single consumer-grade GPU, eliminating the need for segmentation and thus producing high-quality, dense colored point clouds with millions of points for robot navigation and manipulation tasks. Extensive experimental results on generating million-level point cloud data validate the effectiveness of our method, substantially improving the quality of colored point clouds and demonstrating significant potential for applications involving large-scale point clouds in autonomous robotics and human-robot interaction scenarios.
Weijing Xie, Chenyang Wang 0001, Fei Ma 0006, F. Richard Yu
IROS6
2025 Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation
abstract
Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing methods often struggle with effectively integrating visual observations and instruction details during navigation, leading to suboptimal path planning and limited success rates. In this paper, we propose OIKG (Observation-graph Interaction and Key-detail Guidance), a novel framework that addresses these limitations through two key components: (1) an observation-graph interaction module that decouples angular and visual information while strengthening edge representations in the navigation space, and (2) a key-detail guidance module that dynamically extracts and utilizes fine-grained location and object information from instructions. By enabling more precise cross-modal alignment and dynamic instruction interpretation, our approach significantly improves the agent’s ability to follow complex navigation instructions. Extensive experiments on the R2R and RxR datasets demonstrate that OIKG achieves state-of-the-art performance across multiple evaluation metrics, validating the effectiveness of our method in enhancing navigation precision through better observation-instruction alignment.
Binkai Ou, Fei Ma 0006
IROS3
2025 Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
abstract
Talking face video generation with arbitrary speech audio is a significant challenge within the realm of digital human technology. The previous studies have emphasized the significance of audio-lip synchronization and visual quality. Currently, limited attention has been given to the learning of visual uncertainty, which creates several issues in existing systems, including inconsistent visual quality and unreliable performance across different input conditions. To address the problem, we propose a Joint Uncertainty Learning Network (JULNet) for high-quality talking face video generation, which incorporates a representation of uncertainty that is directly related to visual error. Specifically, we first design an uncertainty module to individually predict the error map and uncertainty map after obtaining the generated image. The error map represents the difference between the generated image and the ground truth image, while the uncertainty map is used to predict the probability of incorrect estimates. Furthermore, to match the uncertainty distribution with the error distribution through a KL divergence term, we introduce a histogram technique to approximate the distributions. By jointly optimizing error and uncertainty, the performance and robustness of our model can be enhanced. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking face video generation compared to previous methods.
Fei Ma 0006, Yi Bin, Ying He 0006, F. Richard Yu
ICMR2
2025 OnlineHOI: Towards Online Human-Object Interaction Generation and Perception
abstract
The perception and generation of Human-Object Interaction (HOI) are crucial for fields such as robotics, AR/VR, and human behavior understanding. However, current approaches model this task in an offline setting, where information at each time step can be drawn from the entire interaction sequence. In contrast, in real-world scenarios, the information available at each time step comes only from the current moment and historical data, i.e., an online setting. We find that offline methods perform poorly in an online context. Based on this observation, we propose two new tasks: Online HOI Generation and Perception. To address this task, we introduce the OnlineHOI framework, a network architecture based on the Mamba framework that employs a memory mechanism. By leveraging Mamba's powerful modeling capabilities for streaming data and the Memory mechanism's efficient integration of historical information, we achieve state-of-the-art results on the Core4D and OAKINK2 online generation tasks, as well as the online HOI4D perception task.
Yihong Ji, Yiyao Zhuo, Weijiang Yu, Fei Ma 0006, Joshua Zhexue Huang, F. Richard Yu
ACM Multimedia5
2025 HoloTrace: LLM-based Bidirectional Causal Knowledge Graph for Edge-Cloud Video Anomaly Detection
abstract
Video anomaly detection (VAD) is vital for public safety, yet current approaches struggle with limited generalization, low interpretability, and high resource demands. To address these challenges, we propose HoloTrace, an edge-cloud collaborative VAD system that integrates large language models (LLMs) to construct and update a novel bidirectional causal knowledge graph. At the edge, HoloTrace leverages LLM-based cross-modal understanding and employs Hidden Markov Model (HMM) for bidirectional event reasoning, obtaining anomaly boundaries with low computational overhead. On the cloud side, LLMs are leveraged to dynamically update the Bi-CKG graph with key frames sent from the edge, in order to update causal relationships between events. Additionally, we introduce SVAD, a new large-scale VAD dataset comprising 632 real-world surveillance videos across 10 anomaly types and diverse scenes, with manually labeled frame-level annotations. Experimental results demonstrate that HoloTrace not only achieves the highest accuracy but also enhances interpretability and efficiency, paving the way for more generalizable and explainable video anomaly detection systems.
Hanling Wang, Qing Li 0006, Li Chen 0008, Haidong Kang, Fei Ma 0006, Yong Jiang 0001
ACM Multimedia5
2025 Universal Visuo-Tactile Video Understanding for Embodied Interaction
abstract
Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model that enables universal Visuo-Tactile Video (VTV) understanding, bridging the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150,000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, and scenario-based decision-making. Extensive experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile reasoning tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.
Shoujie Li, Xingting Li, Guangyu Chen, Fei Ma 0006, F. Richard Yu, Wenbo Ding 0001
NeurIPS6
2025 RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion Models
abstract
Preserving intellectual property (IP) within a pre-trained diffusion model is critical for protecting the model's copyright and preventing unauthorized model deployment. In this regard, model watermarking is a common practice for IP protection that embeds traceable information within models and allows for further verification. Nevertheless, existing watermarking schemes often face challenges due to their vulnerability to fine-tuning, limiting their practical application in general pre-training and fine-tuning paradigms. Inspired by using mode connectivity to analyze model performance between a pair of connected models, we investigate watermark vulnerability by leveraging Linear Mode Connectivity (LMC) as a proxy to analyze the fine-tuning dynamics of watermark performance. Our results show that existing watermarked models tend to converge to sharp minima in the loss landscape, thus making them vulnerable to fine-tuning. To tackle this challenge, we propose **RoMa**, a **Ro**bust **M**odel w**a**termarking scheme that improves the robustness of watermarks against fine-tuning. Specifically, RoMa decomposes watermarking into two components, including *Embedding Functionality*, which preserves reliable watermark detection capability, and *Path-specific Smoothness*, which enhances the smoothness along the watermark-connected path to improve robustness. Extensive experiments on benchmark datasets MS-COCO-2017 and CUB-200-2011 demonstrate that RoMa significantly improves watermark robustness against fine-tuning while maintaining generation quality, outperforming baselines. The code is available at [https://github.com/xiekks/RoMa](https://github.com/xiekks/RoMa).
Yingsha Xie, Zeyu Qin, Fei Ma 0006, Li Shen 0008, F. Richard Yu, Xiaochun Cao
NeurIPS4
2025 Rethinking Efficient and Effective Point-Based Networks for Event Camera Classification and Regression
abstract
Event cameras draw inspiration from biological systems, boasting low latency and high dynamic range while consuming minimal power. The most current approach to processing Event Cloud often involves converting it into frame-based representations, which neglects the sparsity of events, loses fine-grained temporal information, and increases the computational burden. In contrast, Point Cloud is a popular representation for processing 3-dimensional data and serves as an alternative method to exploit local and global spatial features. Nevertheless, previous point-based methods show an unsatisfactory performance compared to the frame-based method in dealing with spatio-temporal event streams. In order to bridge the gap, we propose EventMamba, an efficient and effective framework based on Point Cloud representation by rethinking the distinction between Event Cloud and Point Cloud, emphasizing vital temporal information. The Event Cloud is subsequently fed into a hierarchical structure with staged modules to process both implicit and explicit temporal features. Specifically, we redesign the global extractor to enhance explicit temporal extraction among a long sequence of events with temporal aggregation and State Space Model (SSM) based Mamba. Our model consumes minimal computational resources in the experiments and still exhibits SOTA point-based performance on six different scales of action recognition datasets. It even outperformed all frame-based methods on both Camera Pose Relocalization (CPR) and eye-tracking regression tasks.
Yue Zhou 0010, Jiadong Zhu, Xiaopeng Lin, Haotian Fu, Yulong Huang 0001, Yuetong Fang, Fei Ma 0006, Hao Yu 0001, Bojun Cheng
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 Human Motion Video Generation: A Survey
abstract
Human motion video generation has garnered significant research interest due to its broad applications, enabling innovations such as photorealistic singing heads or dynamic avatars that seamlessly dance to music. However, existing surveys in this field focus on individual methods, lacking a comprehensive overview of the entire generative process. This paper addresses this gap by providing an in-depth survey of human motion video generation, encompassing over ten sub-tasks, and detailing the five key phases of the generation process: input, motion planning, motion video generation, refinement, and output. Notably, this is the first survey that discusses the potential of large language models in enhancing human motion video generation. Our survey reviews the latest developments and technological trends in human motion video generation across three primary modalities: vision, text, and audio. By covering over two hundred papers, we offer a thorough overview of the field and highlight milestone works that have driven significant technological breakthroughs. Our goal for this survey is to unveil the prospects of human motion video generation and serve as a valuable resource for advancing the comprehensive applications of digital humans.
Haiwei Xue, Xiangyang Luo 0002, Zhanghao Hu, Xin Zhang 0169, Xunzhi Xiang, Yuqin Dai, Jianzhuang Liu, Zhensong Zhang, Minglei Li 0001, Jian Yang 0003, Fei Ma 0006, Zhiyong Wu 0001, Changpeng Yang, Zonghong Dai, F. Richard Yu
IEEE Trans. Pattern Anal. Mach. Intell.11
2025 MLLM-TA: Leveraging Multimodal Large Language Models for Precise Temporal Video Grounding
abstract
In untrimmed video tasks, identifying temporal boundaries in videos is crucial for temporal video grounding. With the emergence of multimodal large language models (MLLMs), recent studies have focused on endowing these models with the capability of temporal perception in untrimmed videos. To address the challenge, in this paper, we introduce a multimodal large language model named MLLM-TA with precise temporal perception to obtain temporal attention. Unlike the traditional MLLMs, answering temporal questions through one or two words related to temporal information, we leverage the text description proficiency of MLLMs to acquire video temporal attention with description. Specifically, we design a dual temporal-aware generative branches aimed at the visual space of the entire video and the textual space of global descriptions, simultaneously generating mutually supervised consistent temporal attention, thereby enhancing the video temporal perception capabilities of MLLMs. Finally, we evaluate our approach on both video grounding task and highlight detection task on three popular benchmarks, including Charades-STA, ActivityNet Captions and QVHighlights. The extensive results show that our MLLM-TA significantly outperforms previous approaches both on zero-shot and supervised setting, achieving state-of-the-art performance.
Yi Liu 0081, Haowen Hou, Fei Ma 0006, Shiguang Ni, F. Richard Yu
IEEE Signal Process. Lett.3
2025 A Review of Human Emotion Synthesis Based on Generative Technology
abstract
Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer interactions. Recent advancements in generative models, such as Autoencoders, Generative Adversarial Networks, Diffusion Models, Large Language Models, and Sequence-to-Sequence Models, have significantly contributed to the development of this field. However, there is a notable lack of comprehensive reviews in this field. To address this problem, this paper aims to address this gap by providing a thorough and systematic overview of recent advancements in human emotion synthesis based on generative models. Specifically, this review will first present the review methodology, the emotion models involved, the mathematical principles of generative models, and the datasets used. Then, the review covers the application of different generative models to emotion synthesis based on a variety of modalities, including facial images, speech, and text. It also examines mainstream evaluation metrics. Additionally, the review presents some major findings and suggests future research directions, providing a comprehensive understanding of the role of generative technology in the nuanced domain of emotion synthesis.
Fei Ma 0006, Yukan Li, Ying He 0006, Fuji Ren, F. Richard Yu, Shiguang Ni
IEEE Trans. Affect. Comput.1
2024 A Language-Driven Navigation Strategy Integrating Semantic Maps and Large Language Models
abstract
Accurate perception of semantic and spatial information is crucial for robots performing language-driven navigation tasks. Existing approaches utilize visual-language models to extract semantic information from the environment and construct maps. However, constrained by the generalization and accuracy of these models themselves, the constructed maps may not be accurate and comprehensive, thereby affecting the accuracy of navigation tasks. Inspired by foundational models’ outstanding classification and segmentation capabilities, this study introduces a semantic map constructed using foundational models. We leverage a foundational model to semantically segment objects in the robot’s video stream and fuse semantics onto the map. Furthermore, this map is used in conjunction with large language models (LLMs) that receive natural language instructions to complete the navigation task. A substantial number of experiments in a simulated environment demonstrate that our method outperforms existing ones in language-driven navigation tasks.
Zhengjun Zhong, Ying He 0006, Pengteng Li, F. Richard Yu, Fei Ma 0006
IROS5
2024 CodeSwap: Symmetrically Face Swapping Based on Prior Codebook
abstract
Face swapping, the technique of transferring the identity from one face to another, merges as a field with significant practical applications. However, previous swapping methods often result in visible artifacts. To address this issue, in our paper, we propose CodeSwap, a symmetrical framework to achieve face swapping with high-fidelity and realism. Specifically, our method firstly utilizes a codebook that captures the knowledge of high quality facial features. Building on this foundation, the face swapping is then converted into the code manipulation task in a code space. To achieve this, we design a Transformer-based architecture to update each code independently, which enable more precise manipulations. Furthermore, we incorporate a mask generator to achieve seamless blending of the generated face with the background of target image. A distinctive characteristic of our method is its symmetrical approach to processing both target and source images, simultaneously extracting information from each to improve the quality of face swapping. This symmetry also simplifies the bidirectional exchange of faces in a singular operation. Through extensive experiments on ClelebA-HQ and FF++, our method is proven to not only achieve efficient identity transfer but also substantially reduce the visible artifacts.
Xiangyang Luo 0002, Xin Zhang 0169, Xinyi Tong 0002, Weijiang Yu, Heng Chang, Fei Ma 0006, F. Richard Yu
ACM Multimedia7
2024 SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
abstract
Audio-driven talking face generation aims to synthesize video with lip movements synchronized to input audio. However, current generative techniques face challenges in preserving intricate regional textures (skin, teeth). To address the aforementioned challenges, we propose a novel framework called SegTalker to decouple lip movements and image textures by introducing segmentation as intermediate representation. Specifically, given the mask of image employed by a parsing network, we first leverage the speech to drive the mask and generate talking segmentation. Then we disentangle semantic regions of image into style codes using a mask-guided encoder. Ultimately, we inject the previously generated talking segmentation and style codes into a mask-guided StyleGAN to synthesize video frame. In this way, most of textures are fully preserved. Moreover, our approach can inherently achieve background separation and facilitate mask-guided facial local editing. In particular, by editing the mask and swapping the region textures from a given reference image (e.g. hair, lip, eyebrows), our approach enables facial editing seamlessly when generating talking face video. Experiments demonstrate that our proposed approach can effectively preserve texture details and generate temporally consistent video while remaining competitive in lip synchronization. Quantitative and qualitative results on the HDTF and MEAD datasets illustrate the superior performance of our method over existing methods.
Lingyu Xiong, Xize Cheng, Jintao Tan, Xianjia Wu, Xiandong Li, Fei Ma 0006, Minglei Li 0001, Huang Xu 0003
ACM Multimedia7
2024 Dependency-Aware Microservice Deployment for Edge Computing: A Deep Reinforcement Learning Approach With Network Representation
abstract
The popularity of microservices in industry has sparked much attention in the research community. Despite significant progress in microservice deployment for resource-intensive services and applications at the network edge, the intricate dependencies among microservices are often overlooked, and some studies underestimate the importance of system context extraction in deployment strategies. This paper addresses these issues by formulating the microservice deployment problem as a max-min problem, considering system cost and quality of service (QoS) jointly. We first study the attention-based microservice representation (AMR) method to achieve effective system context extraction. In this way, the contributions of different computing power providers (users, edge servers, or cloud servers) in the networks can be effectively paid attention to. Subsequently, we propose the attention-modified soft actor-critic (ASAC) algorithm to tackle the microservice deployment problem. ASAC leverages attention mechanisms to enhance decision-making and adapt to changing system dynamics. Our simulation results demonstrate ASAC's effectiveness, prioritizing average system cost and reward compared to the other state-of-the-art algorithms.
Chenyang Wang 0001, Hao Yu 0013, Xiuhua Li 0001, Fei Ma 0006, Xiaofei Wang 0001, Tarik Taleb, Victor C. M. Leung
IEEE Trans. Mob. Comput.4
2022 Poster Abstract: Representation Learning from Multimodal Sensor Data with Maximally Correlated Autoencoders
abstract
With the development of sensing technology, multiple sensors are widely used in Internet of Things (IoT) devices. A key challenge is to learn feature representations from multimodal sensor data to combine the information of different sensors. Although progress has been made by previous works, the correlation between different sensors is still not well exploited, which may limit the performance of representation learning. To address this problem, we propose a deep learning approach to learn representations from multimodal sensor data with maximally correlated autoencoders (MCA). It can efficiently capture high dependence between different modalities at different feature levels. The learned representations are further used for the recognition task. Experimental results on the real-world RGB- D dataset demonstrate the high effectiveness of MCA.
Fei Ma 0006, Weixi Gu, Shiguang Ni, Lin Zhang 0001
IPSN1
2021 Semi-Supervised Multimodal Image Translation for Missing Modality Imputation
abstract
Missing data is a common problem in multimodal and multi-view learning. It raises a critical challenge for most multimodal algorithms, which are unable to deal with incomplete datasets. Rather than discarding entries with missing modalities, this paper aims to reconstruct the complete image-based multimodal data by imputing missing modalities. We solve the imputation problem as an image translation task, which transforms images in one domain to other domains. Existing image translation techniques either can not fully utilize the information contained in partially complete entries or are limited to the bimodal situation. We propose a semi-supervised algorithm for multimodal learning with missing data, namely Cyclic Autoencoder (CycAE). Specifically, a novel cyclical structure, as well as the correlation among modalities, is integrated to leverage infoπnation from complete entries to incomplete ones. Experiments on two multimodal datasets show that our model outperforms state-of-the-art models. Downstream tasks can also benefit from the completed datasets.
Wangbin Sun, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Shiguang Ni, Lin Zhang 0001
ICASSP2
2021 An Efficient Approach for Audio-Visual Emotion Recognition With Missing Labels And Missing Modalities
abstract
Audio-visual emotion recognition is important for human-machine interaction systems by combining the information of audio and visual modalities. Although great progress has been made by previous works using multimodal learning compared with unimodal learning, they still cannot effectively deal with two key challenges. Firstly, it is difficult or expensive to acquire labeled emotional data, which results in a large amount of data with missing labels. Secondly, emotional data often has missing modalities. To address these problems, we propose a unified deep learning framework to efficiently handle missing labels and missing modalities for audio-visual emotion recognition through correlation analysis. Specifically, we consider four types of emotional data during the training stage: complete, label missing, visual missing, and audio missing. We propose a correlation loss based on Hirschfeld-Gebelein-Ŕenyi (HGR) maximal correlation to effectively capture the common information in different types of training data for emotion prediction. Experiments on the eNTERFACE’05 and RAVDESS datasets show that our deep learning approach has high effectiveness for audio-visual emotion recognition.
Fei Ma 0006, Shao-Lun Huang, Lin Zhang 0001
ICME1
2021 A Semi-supervised Learning Approach for Visual Question Answering based on Maximal Correlation
abstract
In this paper, we propose a semi-supervised learning approach for the Visual Question Answering (VQA) task based on maximal correlation. Instead of training the VQA model with just classification loss like cross-entropy, we propose a semi-supervised loss function to incorporate Soft-HGR, a training approach based on Hirschfeld-Gebelein-Rényi (HGR) maximal correlation, to realize semi-supervised model training. With Soft-HGR, the high-order correlation from cross-modal common information of VQA image-question pairs is utilized to improve VQA model performance even without discriminative supervision from answer labels. We conduct experiments on the VQA v2 dataset by training the VQA model with different percentages of unlabeled samples. Experimental results show that our approach is efficient and model-agnostic for this semi-supervised learning task.
Sikai Yin, Fei Ma 0006, Shao-Lun Huang
SMC2
2020 Person Recognition with HGR Maximal Correlation on Multimodal Data
abstract
Multimodal person recognition is a common task in video analysis and public surveillance, where information from multiple modalities, such as images and audio extracted from videos, are used to jointly determine the identity of a person. Previous person recognition techniques either use only uni-modal data or only consider shared representations between different input modalities, while leaving the extraction of their relationship with identity information to downstream tasks. Furthermore, real-world data often contain noise, which makes recognition more challenging practical situations. In our work, we propose a novel correlation-based multimodal person recognition framework that is relatively simple but can efficaciously learn supervised information in multimodal data fusion and resist noise. Specifically, our framework learns a discriminative embeddings of persons by joint learning visual features and audio features while maximizing HGR maximal correlation among multimodal input and persons' identities. Experiments are done on a subset of Voxceleb2. Compared with state-of-the-art methods, the proposed method demonstrates an improvement of accuracy and robustness to noise.
Yihua Liang, Fei Ma 0006, Yang Li 0104, Shao-Lun Huang
ICPR2
2019 An End-to-End Learning Approach for Multimodal Emotion Recognition: Extracting Common and Private Information
abstract
Multimodal emotion recognition is important for facilitating efficient interaction between humans and machines. To better detect emotional states from multimodal data, we need to effectively extract both the common information that captures dependencies among different modalities, and the private information that characterizes variations in each modality. However, existing works are mostly designed to pursue either one of these objectives but not both. In our work, we propose an end-to-end learning approach to simultaneously extract the common and private information for multimodal emotion recognition. Specifically, we use a correlation loss based on Hirschfeld-Gebelein-Renyi (HGR) maximal correlation and a reconstruction loss based on autoencoders to preserve the common and private information, respectively. Experimental results on eNTERFACE'05 database and RML database demonstrate the effectiveness of our proposed approach.
Fei Ma 0006, Wei Zhang 0185, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001
ICME1
2019 Info-Detection: An Information-Theoretic Approach to Detect Outlier
Fei Ma 0006, Yang Li 0104, Shao-Lun Huang, Lin Zhang 0001
ICONIP (5)2
2019 Unsupervised anomaly detection via generative adversarial networks: poster abstract
abstract
Unsupervised anomaly detection is a fundamental problem in various research areas and application domains, namely the discrimination of abnormal samples from normal samples where training data are only composed of one class (normal) while testing data contains both among which the majority are normal samples. However, previous works can not effectively fit the distribution of high dimensional data and suffers from low AUC scores which measures the classification performance of imbalanced data. To solve these problems, we propose an unsupervised anomaly detection model based on GAN, i.e., UAD-GAN. Specifically, we adopt transfer learning to extract visual features with pre-trained Inception-v3 model and use the discriminator to detect anomalies. UAD-GAN can fit the data distribution and detect anomalies efficiently. Extensive experiments show that UAD-GAN achieves state-of-the-art performance compared to other approaches.
Hanling Wang, Fei Ma 0006, Shao-Lun Huang, Lin Zhang 0001
IPSN3
2018 Real-Time Emotion Detection via E-See
abstract
Real-time emotion detection has being attracted to human attention recently. Recognizing the inner emotion not only assists people to communicate and understand with each other, but also prevents the occurrence of the serious diseases (e.g., autism) and the emergency (i.e., child abuse, sexual invasion). Existing works usually adopt the professional and cumbersome devices to learn the emotions, and therefore limited in the daily usage. In this work, we design a pervasive and wearable device E-See that enables to recognize the emotion in real time. The prototype of the device is deployed in a microcomputer currently, and it can be resized as a small button worn on the collar or extend as a platform to detect the real-time emotion.
Weixi Gu, Yue Zhang 0044, Fei Ma 0006, Khalid M. Mosalam, Lin Zhang 0001, Shiguang Ni
SenSys3
2018 Speech Emotion Recognition via Attention-based DNN from Multi-Task Learning
abstract
Speech unlocks the huge potentials in emotion recognition. High accurate and real-time understanding of human emotion via speech assists Human-Computer Interaction. Previous works are often limited in either coarse-grained emotion learning tasks or the low precisions on the emotion recognition. To solve these problems, we construct a real-world large-scale corpus composed of 4 common emotions (i.e., anger, happiness, neutral and sadness). We also propose a multi-task attention-based DNN model (i.e., MT-A-DNN) on the emotion learning. MT-A-DNN efficiently learns the high-order dependency and non-linear correlations underlying in the audio data. Extensive experiments show that MT-A-DNN outperforms conventional methods on the emotion recognition. It could take one step further on the real-time acoustic emotion recognition in many smart audio-devices.
Fei Ma 0006, Weixi Gu, Wei Zhang 0185, Shiguang Ni, Shao-Lun Huang, Lin Zhang 0001
SenSys1
2018 Multimodal Emotion Recognition by extracting common and modality-specific information
abstract
Emotion recognition technologies have been widely used in numerous areas including advertising, healthcare and online education. Previous works usually recognize the emotion from either the acoustic or the visual signal, yielding unsatisfied performances and limited applications. To improve the inference capability, we present a multimodal emotion recognition model, EMOdal. Apart from learning the audio and visual data respectively, EMOdal efficiently learns the common and modality-specific information underlying the two kinds of signals, and therefore improves the inference ability. The model has been evaluated on our large-scale emotional data set. The comprehensive evaluations demonstrate that our model outperforms traditional approaches.
Wei Zhang 0185, Weixi Gu, Fei Ma 0006, Shiguang Ni, Lin Zhang 0001, Shao-Lun Huang
SenSys3