Huimin Ma 0001

dblp:69/7694-1 · DBLP profile ↗
← Back
146ranked-venue papers
2as first author
105since 2021 · last 2026
0000-0001-5383-5667ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 66 since 2021Artificial intelligence and machine learning · 65 · 2 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021
YearPublicationVenuePosition
2026 Object Detection Data Synthesis via Box-to-Image Generation Based on Diffusion Models
abstract
Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain limited. This paper presents ODGEN, a novel method to generate high-quality images conditioned on bounding boxes, thereby facilitating data synthesis for object detection. Given a domain-specific object detection dataset, we first fine-tune a pre-trained diffusion model on both cropped foreground objects and entire images to fit target distributions. Then we propose to control the diffusion model using synthesized visual prompts with spatial constraints and object-wise textual descriptions. ODGEN exhibits robustness in handling complex scenes and specific domains. Further, we design a dataset synthesis pipeline to evaluate ODGEN on 7 domain-specific benchmarks to demonstrate its effectiveness. Adding training data generated by ODGEN improves up to 25.3% [email protected]:.95 with object detectors like YOLOv5 and YOLOv7, outperforming prior controllable generative methods. We also design an evaluation protocol based on COCO-2014 to validate the synthetic data of ODGEN in general domains and observe an advantage up to 5.6% in [email protected]:.95 against existing methods. In addition, we employ a series of large-scale object detection datasets to train a general model named Stable Box Diffusion, which covers thousands of object categories in most common scenes.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Implicit alignment and query refinement for RGB-T semantic segmentation
Chang Liu 0136, Haizhuang Liu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Qianchuan Zhao, Huimin Ma 0001
Pattern Recognit.7
2026 ME-TST+: Micro-Expression Analysis via Temporal State Transition With ROI Relationship Awareness
abstract
Micro-expressions (MEs) are regarded as important indicators of an individual’s intrinsic emotions, preferences, and tendencies. ME analysis requires spotting of ME intervals within long video sequences and recognition of their corresponding emotional categories. Previous deep learning approaches commonly employ sliding-window classification networks. However, the use of fixed window lengths and hard classification presents notable limitations in practice. Furthermore, these methods typically treat ME spotting and recognition as two separate tasks, overlooking the essential relationship between them. To address these challenges, this paper proposes two state space model-based architectures, namely ME-TST and ME-TST+, which utilize temporal state transition mechanisms to replace conventional window-level classification with video-level regression. This enables a more precise characterization of the temporal dynamics of MEs and supports the modeling of MEs with varying durations. In ME-TST+, we further introduce multi-granularity ROI modeling and the SlowFast Mamba framework to alleviate information loss associated with treating ME analysis as a time series task. Additionally, we propose a synergy strategy for spotting and recognition at both the feature and result levels, leveraging their intrinsic relationship to enhance overall analysis performance. Extensive experiments demonstrate that the proposed methods achieve state-of-the-art performance. The code is available at https://github.com/zizheng-guo/ME-TST.
Zizheng Guo 0002, Bochao Zou, Junbao Zhuo, Huimin Ma 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 MLSTP: A Mamba-Based Long-Short Term Fusion Network for Improved Trajectory Prediction in Autonomous Driving
abstract
Traffic trajectory prediction plays a crucial role in ensuring the safety and efficiency of autonomous driving systems. With increasing traffic complexity, accurate and real-time prediction of surrounding vehicle trajectories has become a major challenge. Existing methods often define trajectory prediction as a static task, predicting future trajectories at once based solely on historical states. However, this approach leads to significant long-term prediction errors and struggles to capture potential hazards in complex, dynamic traffic conditions. In this paper, we propose a lightweight end-to-end trajectory prediction model integrating a long-term guidance module and a short-term feedback mechanism. The long-term guidance module predicts the final driving goal that remains unchanged over time, guiding the vehicle’s eventual position. The short-term feedback mechanism captures the dynamic driving environment, continuously updating the model with immediate changes in vehicle interactions and road conditions to avoid unexpected dangers. At each time step, our approach leverages accumulated historical information and incorporates both short-term and long-term future information to iteratively predict multi-modal trajectories, effectively reducing long-term prediction errors. Additionally, our method introduces the Mamba module, which enhances real-time prediction performance and reduces historical information loss during long-term prediction. Experiments on the Argoverse 1 and Argoverse 2 datasets show that our method, with fewer parameters and higher inference speed, achieves results comparable to or surpassing previous state-of-the-art methods.
Jiaxing Ren, Bochao Zou, Juntao Lyu, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.4
2026 Multi-Scale Spatial Channel Joint Representation for General Multi-Modality Image Fusion With Self-Supervision
abstract
The rapid advancement of multi-modality image fusion technology enables researchers to simultaneously acquire information from different modalities within a single fused image. In existing methods, some general approaches can implement both infrared and visible image fusion (IVIF) and medical image fusion (MIF) in the same framework. Nevertheless, these methods often ignore the learning of specific features in different modalities, resulting in unsatisfactory performance in fused results. To overcome this issue, we propose a multi-scale joint framework with self-supervision for general multi-modality image fusion, abbreviated as SCSFusion. It enables more targeted and robust implementation of IVIF and MIF. Specifically, in the fusion network, a joint attention module is employed to parallelly capture self-attention features in spatial and channel domains, which can keep fused results accurate in visual representation. Meanwhile, we utilize source images of different modalities to generate visual-focused maps as pseudo labels for self-supervised training of the fusion results. It effectively preserves the salient details in each fused image from being disrupted by other extracted information. Moreover, a medical dataset with segmentation labels, termed M2DF, is reorganized for fusion and down-stream tasks in MIF. With the help of M2DF, a pre-trained segmentation model can be cascaded with the fusion network, aiming to obtain high-level semantic features from inputs and enhance the data generalization in our general framework. We have conducted extensive experiments and analyses on SCSFusion in M$\rm ^{3}$FD, FMB, and M2DF datasets, respectively. The results indicate that the fused images generated by SCSFusion can not only achieve visually appealing results and superior performance metrics in MIF and IVIF, but also exhibit satisfactory performance in down-stream tasks.
Jiawei Li 0016, Jiansheng Chen 0001, Jinyuan Liu 0001, Xinlong Ding, Huimin Ma 0001
IEEE Trans. Multim.6
2025 RRT-MVS: Recurrent Regularization Transformer for Multi-View Stereo
abstract
Learning-based multi-view stereo methods aim to predict depth maps for reconstructing dense point clouds. These methods rely on regularization to reduce redundancy in the cost volume. However, existing methods have limitations: CNN-based regularization is restricted to local receptive fields, while Transformer-based regularization struggles with handling depth discontinuities. These limitations often result in inaccurate depth maps with significant noise, particularly noticeable in the boundary and background regions. In this paper, we propose a Recurrent Regularization Transformer for Multi-View Stereo (RRT-MVS), which addresses these limitations by regularizing the cost volume separately for depth and spatial dimensions. Specifically, we introduce Recurrent Self-Attention (R-SA) to aggregate global matching costs within and across the cost maps and filter out noisy feature correlations. Additionally, we present Depth Residual Attention (DRA) to aggregate depth correlations within the cost volume and a Positional Adapter (PA) to enhance 3D positional awareness in each 2D cost map, further augmenting the effectiveness of R-SA. Experimental results demonstrate that RRT-MVS achieves state-of-the-art performance on the DTU and Tanks-and-Temples datasets. Notably, RRT-MVS ranks first on both the Tanks-and-Temples intermediate and advanced benchmarks among all published methods.
Jianfei Jiang 0006, Liyong Wang, Haochen Yu, Huimin Ma 0001
AAAI6
2025 A²RNet: Adversarial Attack Resilient Network for Robust Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion (IVIF) is a crucial technique for enhancing visual performance by integrating unique information from different modalities into one fused image. Exiting methods pay more attention to conducting fusion with undisturbed data, while overlooking the impact of deliberate interference on the effectiveness of fusion results. To investigate the robustness of fusion models, in this paper, we propose a novel adversarial attack resilient network, called A2RNet. Specifically, we develop an adversarial paradigm with an anti-attack loss function to implement adversarial attacks and training. It is constructed based on the intrinsic nature of IVIF and provide a robust foundation for future research advancements. We adopt a Unet as the pipeline with a transformer-based defensive refinement module (DRM) under this paradigm, which guarantees fused image quality in a robust coarse-to-fine manner. Compared to previous works, our method mitigates the adverse effects of adversarial perturbations, consistently maintaining high-fidelity fusion results. Furthermore, the performance of downstream tasks can also be well maintained under adversarial attacks.
Jiawei Li 0016, Jiansheng Chen 0002, Xinlong Ding, Jinyuan Liu 0001, Bochao Zou, Huimin Ma 0001
AAAI8
2025 ProtoCar: Learning 3D Vehicle Prototypes from Single-View and Unconstrained Driving Scene Images
abstract
Reconstructing 3D models from sensor data is a valuable and promising direction for developing testing and validation environments in applications like autonomous driving. However, existing methods for 3D modeling often rely on extensive multi-view data or controlled conditions, making them difficult and expensive to scale. Furthermore, these methods, particularly those based on neural radiance fields, typically produce implicit models that can be challenging to manipulate and suffer from slow rendering speeds. In this paper, we introduce ProtoCar, a novel approach that overcomes these limitations by learning 3D vehicle prototypes from single-view images with diverse and unconstrained visual conditions. ProtoCar uses real-world driving data from LiDAR and image sensors, and employs 3D Gaussian splatting techniques to represent explicit geometric and texture. Extensive experiments demonstrate that ProtoCar generates high-quality 3D models and adapts well to various vehicle types and challenging visual scenarios, offering a scalable and effective solution for 3D modeling in environments with limited and variable visual information.
Hongyuan Liu 0007, Haochen Yu, Bochao Zou, Juntao Lyu, Qi Mei, Huimin Ma 0001
AAAI7
2025 Image-to-video Adaptation with Outlier Modeling and Robust Self-learning
abstract
The image-to-video adaptation task seeks to effectively harness both labeled images and unlabeled videos for achieving effective video recognition. The modality gap of the image and video modalities and the domain discrepancy across the two domains are the two essential challenges in this task. Existing methods reduce the domain discrepancy via close-set domain adaptation techniques, resulting in inaccurate domain alignment as there exist outlier target frames. To tackle this issue, we extend the vanilla classifier with outlier classes, where each outlier class responsible for capturing outlier frames for a specific class via batch nuclear norm maximization loss. We further propose a new loss by treating the source images apart from class c as instances from outlier class specific for c. As for the modality gap, existing methods usually utilize the pseudo labels obtained from an image-level adapted model to learn a video-level model. Rare efforts are dedicated to handling the noise in pseudo labels. We proposed a new metric based on label propagation consistency to select samples for training a better video-level model. Experiments on 3 benchmarks validating the effectiveness of our method.
Junbao Zhuo, Shuhui Wang, Zhenghan Chen, Li Shen 0005, Qingming Huang, Huimin Ma 0001
AAAI6
2025 RhythmMamba: Fast, Lightweight, and Accurate Remote Physiological Measurement
abstract
Remote photoplethysmography (rPPG) is a method for non-contact measurement of physiological signals from facial videos, holding great potential in various applications such as healthcare, affective computing, and anti-spoofing. Existing deep learning methods struggle to address two core issues of rPPG simultaneously: understanding the periodic pattern of rPPG among long contexts and addressing large spatiotemporal redundancy in video segments. These represent a trade-off between computational complexity and the ability to capture long-range dependencies. In this paper, we introduce RhythmMamba, a state space model-based method that captures long-range dependencies while maintaining linear complexity. By viewing rPPG as a time series task through the proposed frame stem, the periodic variations in pulse waves are modeled as state transitions. Additionally, we design multi-temporal constraint and frequency domain feed-forward, both aligned with the characteristics of rPPG time series, to improve the learning capacity of Mamba for rPPG signals. Extensive experiments show that RhythmMamba achieves state-of-the-art performance with 319% throughput and 23% peak GPU memory.
Bochao Zou, Zizheng Guo 0002, Xiaocheng Hu, Huimin Ma 0001
AAAI4
2025 SAM2Object: Consolidating View Consistency via SAM2 for Zero-Shot 3D Instance Segmentation
abstract
In the field of zero-shot 3D instance segmentation, existing 2D-to-3D lifting methods typically obtain 2D segmentation across multiple RGB frames using vision foundation models, which are then projected and merged into 3D space. However, since the inference of vision foundation models on a single frame is not integrated with adjacent frames, the masks of the same object may vary across different frames, leading to a lack of view consistency in the 2D segmentation. Furthermore, current lifting methods average the 2D segmentation from multiple views during the projection into 3D space, causing low-quality masks and high-quality masks to share the same weight. These factors can lead to fragmented 3D segmentation. In this paper, we present SAM2Object, a novel zero-shot 3D instance segmentation method that effectively utilizes the Segment Anything Model 2 to segment and track objects, consolidating view consistency across frames. Our approach combines these consistent 2D masks with 3D geometric priors, improving the robustness of 3D segmentation. Additionally, we introduce mask consolidation module to filter out low-quality masks across frames, which enables more precise 2D-to-3D matching. Comprehensive evaluations on Scan-NetV2, ScanNet++ and ScanNet200 demonstrate the robustness and effectiveness of SAM2Object, showcasing its ability to outperform previous methods. Our project page is at https://jihuaizhaohd.github.io/SAM2Object.
Jihuai Zhao, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
CVPR4
2025 A novel multimodal personality prediction method based on pretrained models and graph relational transformer network
abstract
Multimodal personality analysis aims to identify and express human personality traits in videos. However, RNN and its variants have a limited ability to learn long-term temporal dependencies and existing methods neglect bimodal association features. Based on the fact that visual modalities play a dominant role in this task. Therefore, we propose a personality prediction method catering to learning and fusing intra-modal and intermodal feature dynamics. We first utilize pretrained models’ encoders to extract unimodal spatial scene features from videos. Then, we use xLSTM to capture sequence dependencies between different scene frames used as scene features. Meanwhile, we design a graph relational transformer network to learn longer intra-modal temporal interaction in three unimodal spatial features. Then, we calculate the similarity scores between visual and audio or text features as bimodal association features. Second, we design a multimodal attention feature fusion module to determine the contribution of each feature and aggregate these features. Finally, the MLP model is trained and used to predict scores for personality traits. Experiments on two benchmark datasets demonstrate that our method outperforms the existing methods and achieves state-of-the-art performance. Our code is available at https://github.com/RongquanWang/MP-PMGRT.
Rongquan Wang, Xianyu Xu, Huimin Ma 0001
ICASSP5
2025 Enhancing Autonomous Driving through Dual-Process Learning with Behavior and Reflection Integration
abstract
Contemporary autonomous driving (AD) methodologies, which predominantly convert visual features into control directives, face long-tail challenges due to constraints imposed by limited data distribution. Conversely, human drivers exhibit proficiency in such conditions, underscoring the significance of emulating human cognition in AD systems. Therefore, we introduce Dual-Process Learning (D-PL) approach for cognitive-enhanced decision-making. Inspired by dual-process theory, the D-PL method combines Behavior Pattern Learning (BPL) and Self-Reflective Learning (SRL) to integrate quick, intuitive decisions with deliberate, analytical reasoning, constructing a hierarchical decision model for sophisticated trajectory planning. Our approach improves decision-making, enhances adaptability, and tackles the crucial open-world generalization challenge encountered by current AD methods. Comprehensive evaluations on the nuScenes dataset validate the robustness of our method, demonstrating its superior performance in navigating the intricacies of real-world contrasting with conventional models.
Kangsheng Wang, Huimin Ma 0001
ICASSP4
2025 Synergistic Spotting and Recognition of Micro-Expression via Temporal State Transition
abstract
Micro-expressions are involuntary facial movements that cannot be consciously controlled, conveying subtle cues with substantial real-world applications. The analysis of micro-expressions generally involves two main tasks: spotting micro-expression intervals in long videos and recognizing the emotions associated with these intervals. Previous deep-learning methods have primarily relied on classification networks utilizing sliding windows. However, fixed window sizes and window-level hard classification introduce numerous constraints. Additionally, these methods have not fully exploited the potential of complementary pathways for spotting and recognition. In this paper, we present a novel temporal state transition architecture grounded in the state space model, which replaces conventional window-level classification with video-level regression. Furthermore, by leveraging the inherent connections between spotting and recognition tasks, we propose a synergistic strategy that enhances overall analysis performance. Extensive experiments demonstrate that our method achieves state-of-the-art performance. The codes are available at https://github.com/zizheng-guo/ME-TST.
Bochao Zou, Zizheng Guo 0002, Wenfeng Qin, Xin Li 0034, Kangsheng Wang, Huimin Ma 0001
ICASSP6
2025 Kaleidoscopic Background Attack: Disrupting Pose Estimation With Multi-Fold Radial Symmetry Textures
abstract
Camera pose estimation is a fundamental computer vision task that is essential for applications like visual localization and multi-view stereo reconstruction. In the object-centric scenarios with sparse inputs, the accuracy of pose estimation can be significantly influenced by background textures that occupy major portions of the images across different viewpoints. In light of this, we introduce the Kaleidoscopic Background Attack (KBA), which uses identical segments to form discs with multi-fold radial symmetry. These discs maintain high similarity across different viewpoints, enabling effective attacks on pose estimation models even with natural texture segments. Additionally, a projected orientation consistency loss is proposed to optimize the kaleidoscopic segments, leading to significant enhancement in the attack effectiveness. Experimental results show that optimized adversarial kaleidoscopic backgrounds can effectively attack various camera pose estimation models.
Xinlong Ding, Jiawei Li 0016, Bochao Zou, Huimin Ma 0001
ICCV7
2025 MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network
abstract
Learning-based Multi-View Stereo (MVS) methods aim to predict depth maps for a sequence of calibrated images to recover dense point clouds. However, existing MVS methods often struggle with challenging regions, such as textureless regions and reflective surfaces, where feature matching fails. In contrast, monocular depth estimation inherently does not require feature matching, allowing it to achieve robust relative depth estimation in these regions. To bridge this gap, we propose MonoMVSNet, a novel monocular feature and depth guided MVS network that integrates powerful priors from a monocular foundation model into multi-view geometry. Firstly, the monocular feature of the reference view is integrated into source view features by the attention mechanism with a newly designed cross-view position encoding. Then, the monocular depth of the reference view is aligned to dynamically update the depth candidates for edge regions during the sampling procedure. Finally, a relative consistency loss is further designed based on the monocular depth to supervise the depth prediction. Extensive experiments demonstrate that MonoMVSNet achieves state-of-the-art performance on the DTU and Tanks-and-Temples datasets, ranking first on the Tanks-and-Temples Intermediate and Advanced benchmarks. The source code is available at https://github.com/JianfeiJ/MonoMVSNet.
Jianfei Jiang 0006, Qiankun Liu 0001, Haochen Yu, Hongyuan Liu 0007, Liyong Wang, Huimin Ma 0001
ICCV7
2025 DADet: Safeguarding Image Conditional Diffusion Models Against Adversarial and Backdoor Attacks via Diffusion Anomaly Detection
Xinlong Ding, Jiawei Li 0016, Yudong Zhang 0008, Rongquan Wang, Huimin Ma 0001, Jiansheng Chen 0001
ICCV7
2025 Adversarial Iterative Pre-enactment Framework for Air Combat Based on Mental Simulation Theory
Songde Han, Huimin Ma 0001
ICIG (1)4
2025 Cross-Domain Adaptation for Few-Shot 3D Shape Generation
Huimin Ma 0001
ICIG (2)2
2025 Puzzle-MAE: A Puzzle-Inspired Mask Autoencoder for Multi-Modal Fusion
abstract
Most unsupervised methods in the video domain rely on simple encoder-decoder structures, often resulting in discrepancies between the features extracted from unmasked patches and those from the original patches. To address this issue, we propose a novel self-supervised learning framework, PuzzleMAE, which extracts features from both masked and unmasked patches and aligns them with original image representations to improve feature consistency. Inspired by the human ability to solve puzzles through holistic image recognition and the exploitation of spatial adjacency, we propose the Global-Local Attention Module, which effectively integrates global contextual information with local feature representations. Furthermore, we introduce 3D Relative Position Embedding and Structural Position Embedding to emulate human-like spatial and structural awareness of positional relationships during the puzzle-solving process. The effectiveness of our method is validated on two downstream tasks: the First Impression V2 and DFEW datasets.
Xin Li 0034, Bochao Zou, Rongquan Wang, Huimin Ma 0001
ICME4
2025 CSCE: Boosting LLM Reasoning by Simultaneous Enhancing of Causal Significance and Consistency
abstract
Chain-based reasoning methods like chain of thought (CoT) play a rising role in solving reasoning tasks for large language models (LLMs). However, the causal hallucinations between a step of reasoning and corresponding state transitions are becoming a significant obstacle to advancing LLMs’ reasoning capabilities, especially in long-range reasoning tasks. This paper proposes a non-chain-based reasoning framework for simultaneous consideration of causal significance and consistency, i.e., the Causal Significance and Consistency Enhancer (CSCE). We customize LLM’s loss function utilizing treatment effect assessments to enhance its reasoning ability from two aspects: causal significance and consistency. This ensures that the model captures essential causal relationships and maintains robust and consistent performance across various scenarios. Additionally, we transform the reasoning process from the cascading multiple one-step reasoning commonly used in Chain-Based methods, like CoT, to a causal-enhanced method that outputs the entire reasoning process in one go, further improving the model’s reasoning efficiency. Extensive experiments show that our method improves both the reasoning success rate and speed. These improvements further demonstrate that non-chain-based methods can also aid LLMs in completing reasoning tasks.
Kangsheng Wang, Juntao Lyu, Huimin Ma 0001
ICME5
2025 Efficient Knowledge Transfer in Multi-Task Learning through Task-Adaptive Low-Rank Representation
abstract
Pre-trained language models (PLMs) demonstrate remarkable intelligence but struggle with emerging tasks unseen during training in real-world applications. Training separate models for each new task is usually impractical. Multi-task learning (MTL) addresses this challenge by transferring shared knowledge from source tasks to target tasks. As an dominant parameter-efficient fine-tuning method, prompt tuning (PT) enhances MTL by introducing an adaptable vector that captures task-specific knowledge, which acts as a prefix to the original prompt that preserves shared knowledge, while keeping PLM parameters frozen. However, PT struggles to effectively capture the heterogeneity of task-specific knowledge due to its limited representational capacity. To address this challenge, we propose Task-Adaptive Low-Rank Representation (TA-LoRA), an MTL method built on PT, employing the low-rank representation to model task heterogeneity and a fast-slow weights mechanism where the slow weight encodes shared knowledge, while the fast weight captures task-specific nuances, avoiding the mixing of shared and task-specific knowledge, caused by training low-rank representations from scratch. Moreover, a zero-initialized attention mechanism is introduced to minimize the disruption of immature low-rank components on original prompts during warm-up epochs. Experiments on 16 tasks demonstrate that TA-LoRA achieves state-of-the-art performance in full-data and few-shot settings while maintaining superior parameter efficiency.1
Kangsheng Wang, Huimin Ma 0001
ICME4
2025 From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
abstract
As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interpretative exploration of their internal mechanisms. In response, this study adopts an approach based on internal mechanisms to provide an interpretability-driven assessment of ToM in multimodal large language models (MLLMs). Specifically, we first construct a multimodal ToM test dataset, GridToM, which incorporates diverse belief testing tasks and perceptual information from multiple perspectives. Next, our analysis shows that attention heads in multimodal large models can distinguish cognitive information across perspectives, providing evidence of ToM capabilities. Furthermore, we present a lightweight, training-free approach that significantly enhances the model’s exhibited ToM by adjusting in the direction of the attention head.
Siqi Liu 0010, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
ICML5
2025 Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression
abstract
Micro-expressions (MEs) are involuntary, low-intensity, and short-duration facial expressions that often reveal an individual's genuine thoughts and emotions. Most existing ME analysis methods rely on window-level classification with fixed window sizes and hard decisions, which limits their ability to capture the complex temporal dynamics of MEs. Although recent approaches have adopted video-level regression frameworks to address some of these challenges, interval decoding still depends on manually predefined, window-based methods, leaving the issue only partially mitigated. In this paper, we propose a prior-guided video-level regression method for ME analysis. We introduce a scalable interval selection strategy that comprehensively considers the temporal evolution, duration, and class distribution characteristics of MEs, enabling precise spotting of the onset, apex, and offset phases. In addition, we introduce a synergistic optimization framework, in which the spotting and recognition tasks share parameters except for the classification heads. This fully exploits complementary information, makes more efficient use of limited data, and enhances the model's capability. Extensive experiments on multiple benchmark datasets demonstrate the state-of-the-art performance of our method, with an STRS of 0.0562 on CAS(ME)3 and 0.2000 on SAMMLV. The code is available at https://github.com/zizheng-guo/BoostingVRME.
Zizheng Guo 0002, Bochao Zou, Yinuo Jia, Huimin Ma 0001
ACM Multimedia5
2025 AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios
abstract
By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous work focus on vehicle-to-vehicle and vehicle-to-infrastructure collaboration, with limited attention to aerial perspectives provided by UAVs, which uniquely offer dynamic, top-down views to alleviate occlusions and monitor large-scale interactive environments. A major reason for this is the lack of high-quality datasets for aerial-ground collaborative scenarios. To bridge this gap, we present AGC-Drive, the first large-scale real-world dataset for Aerial-Ground Cooperative 3D perception. The data collection platform consists of two vehicles, each equipped with five cameras and one LiDAR sensor, and one UAV carrying a forward-facing camera and a LiDAR sensor, enabling comprehensive multi-view and multi-agent perception. Consisting of approximately 80K LiDAR frames and 360K images, the dataset covers 14 diverse real-world driving scenarios, including urban roundabouts, highway tunnels, and on/off ramps. Notably, 17\% of the data comprises dynamic interaction events, including vehicle cut-ins, cut-outs, and frequent lane changes. AGC-Drive contains 350 scenes, each with approximately 100 frames and fully annotated 3D bounding boxes covering 13 object categories. We provide benchmarks for two 3D perception tasks: vehicle-to-vehicle collaborative perception and vehicle-to-UAV collaborative perception. Additionally, we release an open-source toolkit, including spatiotemporal alignment verification tools, multi-agent visualization systems, and collaborative annotation utilities. The dataset and code are available at https://github.com/PercepX/AGC-Drive.
Yunhao Hou, Bochao Zou, Shangdong Yang, Junbao Zhuo, Siheng Chen, Jiansheng Chen 0001, Huimin Ma 0001
NeurIPS10
2025 MVSMamba: Multi-View Stereo with State Space Model
abstract
Robust feature representations are essential for learning-based Multi-View Stereo (MVS), which relies on accurate feature matching. Recent MVS methods leverage Transformers to capture long-range dependencies based on local features extracted by conventional feature pyramid networks. However, the quadratic complexity of Transformer-based MVS methods poses challenges to balance performance and efficiency. Motivated by the global modeling capability and linear complexity of the Mamba architecture, we propose MVSMamba, the first Mamba-based MVS network. MVSMamba enables efficient global feature aggregation with minimal computational overhead. To fully exploit Mamba's potential in MVS, we propose a Dynamic Mamba module (DM-module) based on a novel reference-centered dynamic scanning strategy, which enables: (1) Efficient intra- and inter-view feature interaction from the reference to source views, (2) Omnidirectional multi-view feature representations, and (3) Multi-scale global feature aggregation. Extensive experimental results demonstrate MVSMamba outperforms state-of-the-art MVS methods on the DTU dataset and the Tanks-and-Temples benchmark with both superior performance and efficiency. The source code is available at https://github.com/JianfeiJ/MVSMamba.
Jianfei Jiang 0006, Qiankun Liu 0001, Hongyuan Liu 0007, Haochen Yu, Liyong Wang, Huimin Ma 0001
NeurIPS7
2025 Few Annotated Pixels and Point Cloud Based Weakly Supervised Semantic Segmentation of Driving Scenes
Huimin Ma 0001, Jiansheng Chen 0001, Yu Wang 0002
Int. J. Comput. Vis.1
2025 Correction: Few Annotated Pixels and Point Cloud Based Weakly Supervised Semantic Segmentation of Driving Scenes
Huimin Ma 0001, Jiansheng Chen 0001, Yu Wang 0002
Int. J. Comput. Vis.1
2025 DomainStudio: Fine-Tuning Diffusion Models for Domain-Driven Image Generation Using Limited Data
Huimin Ma 0001
Int. J. Comput. Vis.2
2025 AwareTrack: Object awareness for visual tracking via templates interaction
Hong Zhang 0018, Jianbo Song, Yifan Yang 0003, Huimin Ma 0001
Image Vis. Comput.6
2025 Occlusion-guided multi-modal fusion for vehicle-infrastructure cooperative 3D object detection
Huazhen Chu, Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.5
2025 SparseComm: An Efficient Sparse Communication Framework for Vehicle-Infrastructure Cooperative 3D Detection
Haizhuang Liu, Huazhen Chu, Junbao Zhuo, Bochao Zou, Jiansheng Chen 0001, Huimin Ma 0001
Pattern Recognit.6
2025 RhythmFormer: Extracting patterned rPPG signals based on periodic sparse attention
Bochao Zou, Zizheng Guo 0002, Jiansheng Chen 0001, Junbao Zhuo, Weiran Huang 0001, Huimin Ma 0001
Pattern Recognit.6
2025 GET3DGS: Generate 3D Gaussians Based on Points Deformation Fields
abstract
The 3D Gaussian Splatting method has recently shown significant advancements in rendering speed and scene composition quality, enhancing its industrial applications and boosting the demand for 3D Gaussian asset generation. However, existing mature 3D generation technologies predominantly rely on implicit representations, which often struggle to balance geometric quality with editability. The production of 3D Gaussian assets generally involves diffusion models that require a dual-stage process of reconstruction and generation, resulting in substantial training and inference costs. To overcome these challenges, we introduce GET3DGS, an innovative approach that combines 3D-aware GANs with 3D Gaussian Splatting representations. This method facilitates the manipulation of the physical attributes of 3D Gaussians, such as geometry and texture, via point deformation fields. Offering faster inference speeds and end-to-end training capabilities, our model outperforms existing diffusion model-based methods. By deriving high-quality Gaussian point cloud geometric representations from 2D images, our approach reduces material accumulation costs and produces data compatible with 3D Gaussian rendering engines. We have evaluated the generative performance of our model on ShapeNet and OmniObject3D and demonstrate competitive results in terms of image and geometric quality relative to previous methods.
Haochen Yu, Weixi Gong, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 HomuGAN: A 3D-Aware GAN With the Method of Cylindrical Spatial-Constrained Sampling
abstract
Controllable 3D-aware scene synthesis seeks to disentangle the various latent codes in the implicit space enabling the generation network to create highly realistic images with 3D consistency. Recent approaches often integrate Neural Radiance Fields with the upsampling method of StyleGAN2, employing Convolutions with style modulation to transform spatial coordinates into frequency domain representations. Our analysis indicates that this approach can give rise to a bubble phenomenon in StyleNeRF. We argue that the style modulation introduces extraneous information into the implicit space, disrupting 3D implicit modeling and degrading image quality. We introduce HomuGAN, incorporating two key improvements. First, we disentangle the style modulation applied to implicit modeling from that utilized for super-resolution, thus alleviating the bubble phenomenon. Second, we introduce Cylindrical Spatial-Constrained Sampling and Parabolic Sampling. The latter sampling method, as an alternative method to the former, specifically contributes to the performance of foreground modeling of vehicles. We evaluate HomuGAN on publicly available datasets, comparing its performance to existing methods. Empirical results demonstrate that our model achieves the best performance, exhibiting relatively outstanding disentanglement capability. Moreover, HomuGAN addresses the training instability problem observed in StyleNeRF and reduces the bubble phenomenon.
Haochen Yu, Weixi Gong, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Image Process.4
2025 Individualized Driving Intention Prediction With Inverse Reinforcement Learning
abstract
Advanced Driver Assistance Systems (ADAS) are designed to prevent collisions, identify the condition of drivers while operating vehicles, and provide additional information to enhance drivers’ awareness of potential hazards on the road. Today, ADAS are capable of predicting drivers’ actions several seconds in advance, preparing for potential future hazards to prevent accidents or reduce injuries to occupants. Most previous works have achieved prediction results by analyzing and processing a vast amount of driving data from multiple drivers, based on the macro intention preferences of multiple drivers and external environmental features, collectively referred to as the generalized intention prediction network. This network utilizes extensive driving data to predict the common driving intentions of the overall driving population, without considering individualized driving styles. However, according to our research, different drivers exhibit distinct latent preferences in real-world driving scenarios. The generalized intention prediction network is influenced by these latent preferences, resulting in poor generalization capabilities and inaccurate predictions across different drivers. In this study, we propose a individualized driver intention prediction network. Based on Inverse Reinforcement Learning (IRL), it extracts individualized driving intention feature preferences that influence driving intentions from the driver’s historical behavior to improve generalized prediction results and achieve individualized driving intention prediction. We demonstrate that preferences vary among different drivers in the driving domain, leading to biases in model predictions. Upon experimental validation, the method we have proposed demonstrates remarkable efficacy on both the Brain4Cars and IESDD datasets, thereby showcasing its enhanced applicability in real-world scenarios.
Siqi Liu 0010, Jiansheng Chen 0001, Chenghao Guo, Jiehui Wu, Qifeng Luo, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.7
2025 Improving Adversarial Robustness Against Universal Patch Attacks Through Feature Norm Suppressing
abstract
Universal adversarial patch attacks, which are readily implemented, have been validated to be able to fool real-world deep convolutional neural networks (CNNs), posing a serious threat to practical computer vision systems based on CNNs. Unfortunately, current defending approaches are severely understudied facing the following problems. Patch detection-based methods suffer from dramatic performance drops against white-box or adaptive attacks since they rely heavily on empirical clues. Methods based on adversarial training or certified defense are difficult to be scaled up to large-scale datasets or complex practical networks due to prohibitively high computational overhead or over strong assumptions on the network structure. In this article, we focus on two cases of widely adopted universal adversarial patch attacks, namely the universal targeted attack on image classifiers and the universal vanishing attack on object detectors. We find that, for popular CNNs, the attacking success of the adversarial patch relies on feature vectors centered at the patch location with large norm in classifiers and large channel-aware norm (CA-Norm) in detectors, and further present a mathematical explanation for this phenomenon. Based on this, we propose a simple but effective defending method using the feature norm suppressing (FNS) layer, which can renormalize the feature norm by nonincreasing functions. As a differentiable module, FNS can be adaptively inserted in various CNN architectures to achieve multistage suppression of the generation of large norm feature vectors. Moreover, FNS is efficient with no trainable parameters and very low computational overhead. We evaluate our proposed defending method across multiple CNN architectures and datasets against the strong adaptive white-box attacks in both visual classification and detection tasks. In both tasks, FNS significantly outperforms previous defending methods on adversarial robustness with a relatively low influence on the performance of benign images. Code is available at https://github.com/jschenthu/FNS.
Jiansheng Chen 0001, Yu Wang 0002, Youze Xue, Huimin Ma 0001
IEEE Trans. Neural Networks Learn. Syst.5
2025 Isolated Diffusion: Optimizing Multi-Concept Text-to-Image Generation Training-Freely With Isolated Diffusion Guidance
abstract
Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still struggle to deal with multi-concept generation accurately in many cases. This phenomenon is known as "concept bleeding" and displays as the unexpected overlapping or merging of various concepts. This paper presents a general approach for text-to-image diffusion models to address the mutual interference between different subjects and their attachments in complex scenes, pursuing better text-image consistency. The core idea is to isolate the synthesizing processes of different concepts. We propose to bind each attachment to corresponding subjects separately with split text prompts. Besides, we introduce a revision method to fix the concept bleeding problem in multi-subject synthesis. We first depend on pre-trained object detection and segmentation models to obtain the layouts of subjects. Then we isolate and resynthesize each subject individually with corresponding text prompts to avoid mutual interference. Overall, we achieve a training-free strategy, named Isolated Diffusion, to optimize multi-concept text-to-image synthesis. It is compatible with the latest Stable Diffusion XL (SDXL) and prior Stable Diffusion (SD) models. We compare our approach with alternative methods using a variety of multi-concept text prompts and demonstrate its effectiveness with clear advantages in text-image consistency and user study.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Vis. Comput. Graph.2
2024 Transferable Adversarial Attacks for Object Detection Using Object-Aware Significant Feature Distortion
abstract
Transferable black-box adversarial attacks against classifiers by disturbing the intermediate-layer features have been extensively studied in recent years. However, these methods have not yet achieved satisfactory performances when directly applied to object detectors. This is largely because the features of detectors are fundamentally different from that of the classifiers. In this study, we propose a simple but effective method to improve the transferability of adversarial examples for object detectors by leveraging the properties of spatial consistency and limited equivariance of object detectors’ features. Specifically, we combine a novel loss function and deliberately designed data augmentation to distort the backbone features of object detectors by suppressing significant features corresponding to objects and amplifying the surrounding vicinal features corresponding to object boundaries. As such the target object and background area on the generated adversarial samples are more likely to be confused by other detectors. Extensive experimental results show that our proposed method achieves state-of-the-art black-box transferability for untargeted attacks on various models, including one/two-stage, CNN/Transformer-based, and anchor-free/anchor-based detectors.
Xinlong Ding, Yining Qin, Huimin Ma 0001
AAAI6
2024 Step Vulnerability Guided Mean Fluctuation Adversarial Attack against Conditional Diffusion Models
abstract
The high-quality generation results of conditional diffusion models have brought about concerns regarding privacy and copyright issues. As a possible technique for preventing the abuse of diffusion models, the adversarial attack against diffusion models has attracted academic attention recently. In this work, utilizing the phenomenon that diffusion models are highly sensitive to the mean value of the input noise, we propose the Mean Fluctuation Attack (MFA) to introduce mean fluctuations by shifting the mean values of the estimated noises during the reverse process. In addition, we reveal that the vulnerability of different reverse steps against adversarial attacks actually varies significantly. By modeling the step vulnerability and using it as guidance to sample the target steps for generating adversarial examples, the effectiveness of adversarial attacks can be substantially enhanced. Extensive experiments show that our algorithm can steadily cause the mean shift of the predicted noises so as to disrupt the entire reverse generation process and degrade the generation results significantly. We also demonstrate that the step vulnerability is intrinsic to the reverse process by verifying its effectiveness in an attack method other than MFA. Code and Supplementary is available at https://github.com/yuhongwei22/MFA
Jiansheng Chen 0001, Xinlong Ding, Yudong Zhang 0008, Ting Tang, Huimin Ma 0001
AAAI6
2024 Upper-Body Hierarchical Graph for Skeleton Based Emotion Recognition in Assistive Driving
Jiehui Wu, Jiansheng Chen 0001, Qifeng Luo, Siqi Liu 0010, Youze Xue, Huimin Ma 0001
ECCV (26)6
2024 Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy
Hong Zhang 0018, Yixuan Lyu, Qian Yu 0008, Huimin Ma 0001, Ding Yuan 0001, Yifan Yang 0003
ECCV (55)5
2024 Center of Pressure Estimation by Analyzing Walking Videos
abstract
Center of pressure (COP) serves as a widely utilized indicator for evaluating balance-related issues, e.g., gait quality of neurological disorders, fall risk of the elderly, and recovery of the injured. Existing methods for acquiring COP mostly rely on expensive force platforms or wearable force-sensing sensors that are relatively difficult to deploy. We propose a novel low-cost strategy to estimate COP with only visual information. Specifically, we first collect a database with walking videos. The plantar pressure of the subjects during walking is synchronously collected for calculating COP ground truth. Then, we propose an encoder-decoder model for estimating COPs using human pose and body shape as inputs. A perceptual codebook is used to tolerate the error in human pose estimation to improve the accuracy of the COP estimation. Experiments demonstrate the effectiveness of our proposed method. Our proposal achieves a correlation coefficient of 0.962, which is 0.8% better than the baseline model using Multi-layer Perceptron (MLP). The normalized RMSEs are improved by 13.9% and 19.0% in the anterior-posterior and medial-lateral directions, respectively. Compared to existing methods using wearable sensors or plantar pressure plates, our method is cheaper and easier to deploy.
Jiansheng Chen 0002, Yining Qin, Poyu Lin, Jiawei Li 0016, Youze Xue, Huimin Ma 0001
ICASSP6
2024 PLS: Unsupervised Domain Adaptation for 3d Object Detection Via Pseudo-Label Sizes
abstract
3D object detection has gained increasing attention in modern autonomous driving systems. However, the performance of the detector significantly degrades during cross-domain deployment due to domain shift. The detector is inevitably biased towards its training dataset when employed on a target dataset, particularly towards object sizes. State-of-the-art unsupervised domain adaptation approaches explicitly address the variation in object sizes by appropriately scaling the source data. However, such methods require additional target domain statistics information, which contradicts the original unsupervised assumption. In this work, we present PLS, a novel unsupervised domain adaptation method for 3D object detection to overcome the object sizes bias via Pseudo-Label Sizes, which utilizes only source domain annotations. PLS alternates between generating high-quality pseudo-label sizes through the detector and model training with the pseudo-label sizes to scale and augment the source data. This iterative process enables the detector to be trained with augmented data that resembles the target domain sizes, thereby improving the performance of detector in cross-domain scenarios. Our experimental results show the outstanding performance of our PLS in various scenarios. In addition, PLS is a plug-and-play module that can be used to directly replace existing weakly-supervised scaling methods. Experimental results show that existing excellent architectures with PLS are able to achieve better performance, and making them completely unsupervised.
Rongquan Wang, Xin Li 0034, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
ICASSP7
2024 Enhancing Adversarial Transferability in Object Detection with Bidirectional Feature Distortion
abstract
Previous works have shown that perturbing internal-layer features can significantly enhance the transferability of black-box attacks in classifiers. However, these methods have not achieved satisfactory performance when applied to detectors due to the inherent differences in features between detectors and classifiers. In this paper, we introduce a concise and practical untargeted adversarial attack in a label-free manner, which leverages only the feature extracted from the backbone model. By implicitly suppressing the critical feature elements for detection while enhancing the candidate object-relevant elements corresponding to possible detection boxes, we conduct a Bidirectional Feature Distortion Attack (BFDA). Experimental results show that BFDA achieves state-of-the-art black-box transferability on various detector architectures.
Xinlong Ding, Huimin Ma 0001
ICASSP5
2024 Refining 3D Human Mesh via Model-Free Offsets Estimation
abstract
3D human mesh reconstruction from a single RGB image is a challenging task. Existing methods either utilize parametric mesh models to restrain the 3D human structures or directly regress the 3D coordinates of the mesh vertices. The former ones, called model-based methods, usually fail to recover the high variance of human mesh due to limited capacity of parametric models, whereas the latter ones called model-free methods suffer from unrealistic human structures because of lack of 3D priors. To mitigate the drawbacks of them, we propose that the model-based reconstruction can serve as a good starting point for the model-free refinement, so that the 3D structure priors of the parametric model and the high representation capability of model-free methods can be both inherited. By building a model-free refinement head upon a pretrained model-based regressor, our method reduces the reconstruction errors of 3D human mesh on public datasets H36M and 3DPW, demonstrating the advantage of combining model-based and model-free methods together.
Youze Xue, Hongbing Ma, Huimin Ma 0001
ICASSP4
2024 Invisible Pedestrians: Synthesizing Adversarial Clothing Textures To Evade Industrial Camera-Based 3D Detection
abstract
Recent studies have explored attacking perception models of autonomous driving systems and creating adversarial textures to evade 2D vision or infrared pedestrian detectors. However, purely camera-based 3D detectors remain to be explored, especially in complex real-world traffic scenarios and road conditions where pedestrians are prevalent. In this paper, we propose a pipeline that accurately renders 3D human models onto 2D scenes while maintaining physical realism and spatial coordinate constraints. By leveraging this pipeline, we design a soft-weighted confidence loss to optimize adversarial textures on clothing, effectively suppressing predicted boxes near the human model. A comprehensive evaluation of adversarial textures is conducted on a test set comprising 100 scenes with complex environments. Our attack can consistently evade detection by the industrial camera-based 3D detector SMOKE across multiple viewpoints and scenes.
Xinlong Ding, Jintai Du, Huimin Ma 0001
ICME6
2024 Analyzing Behavior and Intention in Multi-Agent Systems Using Graph Neural Networks
abstract
Multi-agent behavior and intention analysis has been applied in many aspects of our daily lives such as driving maneuver anticipation and assistive driving perception. However, the research in this field suffers from the lack of publicly available datasets, which is mainly caused by the low availability and high complexity of the multi-agent behavior and intention data. In this paper, we propose MBI, a dataset that contains five categories of agents in eight scenarios based on an open-source simulated platform called HarFang3D DogFight SandBox. In MBI, five types of behaviors and five types of intentions are collected. Additionally, we benchmark different Graph Neural Networks (GNNs) on the proposed dataset to test their performance on the tasks of analyzing different behaviors and intentions in multi-agent systems. We also propose a new method called D2RGAT and find it can achieve the best results on MBI.
Jintai Du, Xinlong Ding, Jiehui Wu, Huimin Ma 0001
ICME7
2024 EMo Transformer: Transformer-Based Depression Detection via Eye Movements
abstract
Depressive disorder has become a prevalent psychological illness that significantly impacts individuals’ daily lives. Traditional questionnaire assessment and clinical interviews suffer from issues such as subjectivity and a high consumption of medical resources. With the advancement of artificial intelligence, there is a growing number of depression detection methods based on statistical features. However, these methods have problems of insufficient stimulus extraction and neglecting temporal information. In order to solve these problems, we propose a transformer-based model named EMo Transformer, designed for detecting depression by effectively extracting features from stimuli and combining them with eye movements. Additionally, due to challenge in collecting data from depression patients, we design a simple and effective data augmentation method to solve this challenge. Subsequently, we design an ensemble model using the models with and without data augmentation. The experimental results of accuracy 91.95% demonstrate that our method is effective.
Xin Li 0034, Haizhuang Liu, Rongquan Wang, Bochao Zou, Huimin Ma 0001
ICME6
2024 Unveiling the Dynamics of Information Interplay in Supervised Learning
abstract
In this paper, we use matrix information theory as an analytical tool to analyze the dynamics of the information interplay between data representations and classification head vectors in the supervised learning process. Specifically, inspired by the theory of Neural Collapse, we introduce matrix mutual information ratio (MIR) and matrix entropy difference ratio (HDR) to assess the interactions of data representation and class classification heads in supervised learning, and we determine the theoretical optimal values for MIR and HDR when Neural Collapse happens. Our experiments show that MIR and HDR can effectively explain many phenomena occurring in neural networks, for example, the standard supervised training dynamics, linear mode connectivity, and the performance of label smoothing and pruning. Additionally, we use MIR and HDR to gain insights into the dynamics of grokking, which is an intriguing phenomenon observed in supervised training, where the model demonstrates generalization capabilities long after it has learned to fit the training data. Furthermore, we introduce MIR and HDR as loss terms in supervised and semi-supervised learning to optimize the information interactions among samples and classification heads. The empirical results provide evidence of the method’s effectiveness, demonstrating that the utilization of MIR and HDR not only aids in comprehending the dynamics throughout the training process but can also enhances the training procedure itself.
Kun Song 0004, Zhiquan Tan, Bochao Zou, Huimin Ma 0001, Weiran Huang 0001
ICML4
2024 CMT: Co-training Mean-Teacher for Unsupervised Domain Adaptation on 3D Object Detection
Junbao Zhuo, Xin Li 0034, Haizhuang Liu, Rongquan Wang, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia7
2024 Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video Generation
Kenan Huang, Junbao Zhuo, Shuhui Wang, Chi Su, Qingming Huang, Huimin Ma 0001
ACM Multimedia6
2024 Affinity3D: Propagating Instance-Level Semantic Affinity for Zero-Shot Point Cloud Semantic Segmentation
abstract
Zero-shot point cloud semantic segmentation aims to recognize novel classes at the point level. Previous methods mainly transfer excellent zero-shot generalization capabilities from images to point clouds. However, directly transferring knowledge from images to point clouds faces two ambiguous problems. On the one hand, 2D models will generate wrong predictions when the image changes. On the other hand, directly mapping 3D points to 2D pixels by perspective projection fails to consider the visibility of 3D points in camera view. The wrong geometric alignment of 3D points and 2D pixels causes semantic ambiguity. To tackle these two problems, we propose a framework named Affinity3D that intends to empower 3D semantic segmentation models to perceive novel samples. Our framework aggregates instances in 3D and recognizes them in 2D, leveraging the excellent geometric separation in 3D and the zero-shot capabilities of 2D models. Affinity3D involves an affinity module that rectifies the wrong predictions by comparing them with similar instances and a visibility module preventing knowledge transfer from visible 2D pixels to invisible 3D points. Extensive experiments have been conducted on the SemanticKITTI and nuScenes datasets. Our framework achieves state-of-the-art performance on both two datasets. Code is available at https://github.com/opjang5/Affinity3D.
Haizhuang Liu, Junbao Zhuo, Jiansheng Chen 0001, Huimin Ma 0001
ACM Multimedia5
2024 ODGEN: Domain-specific Object Detection Data Generation with Diffusion Models
abstract
Modern diffusion-based image generative models have made significant progress and become promising to enrich training data for the object detection task. However, the generation quality and the controllability for complex scenes containing multi-class objects and dense objects with occlusions remain limited. This paper presents ODGEN, a novel method to generate high-quality images conditioned on bounding boxes, thereby facilitating data synthesis for object detection. Given a domain-specific object detection dataset, we first fine-tune a pre-trained diffusion model on both cropped foreground objects and entire images to fit target distributions. Then we propose to control the diffusion model using synthesized visual prompts with spatial constraints and object-wise textual descriptions. ODGEN exhibits robustness in handling complex scenes and specific domains. Further, we design a dataset synthesis pipeline to evaluate ODGEN on 7 domain-specific benchmarks to demonstrate its effectiveness. Adding training data generated by ODGEN improves up to 25.3% [email protected]:.95 with object detectors like YOLOv5 and YOLOv7, outperforming prior controllable generative methods. In addition, we design an evaluation protocol based on COCO-2014 to validate ODGEN in general domains and observe an advantage up to 5.6% in [email protected]:.95 against existing methods.
Yuxuan Liu 0014, Jiulong Shan, Huimin Ma 0001
NeurIPS7
2024 Enhancing pseudo label quality for pedestrian and cyclist in weakly supervised 3D object detection
Haizhuang Liu, Huazhen Chu, Bochao Zou, Huimin Ma 0001
Neurocomputing6
2024 Image paragraph captioning with topic clustering and topic shift prediction
Ting Tang, Jiansheng Chen 0001, Huimin Ma 0001, Yudong Zhang 0008
Knowl. Based Syst.4
2024 High-Quality and Diverse Few-Shot Image Generation via Masked Discrimination
abstract
Few-shot image generation aims to generate images of high quality and great diversity with limited data. However, it is difficult for modern GANs to avoid overfitting when trained on only a few images. The discriminator can easily remember all the training samples and guide the generator to replicate them, leading to severe diversity degradation. Several methods have been proposed to relieve overfitting by adapting GANs pre-trained on large source domains to target domains using limited real samples. This work presents masked discrimination to realize few-shot GAN adaptation, which is the first feature-level augmentation method for generative tasks. Random masks are applied to features extracted by the discriminator from input images. We aim to encourage the discriminator to judge various images that share partially common features with training samples as realistic. Correspondingly, the generator is guided to generate diverse images instead of replicating training samples. In addition, we employ a cross-domain consistency loss for the discriminator to keep relative distances between generated samples in its feature space. It strengthens global image discrimination and guides adapted GANs to preserve more information learned from source domains for higher image quality, resulting in better cross-domain correspondence. The effectiveness of our approach is demonstrated both qualitatively and quantitatively with higher quality and greater diversity on a series of few-shot image generation tasks than prior methods.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Image Process.2
2023 Defending Against Universal Patch Attacks by Restricting Token Attention in Vision Transformers
abstract
Previous works reveal that similar to CNNs, vision transformers (ViT) are also vulnerable to universal adversarial patch attacks. In this paper, we empirically reveal and mathematically explain that the shallow tokens in the transformer and the attention of the network can largely influence the classification result. Adversarial patches usually produce large feature norm for the corresponding shallow token vectors which can attract the attention anomalously. Inspired by this, we propose a restriction operation on the attention matrix, which effectively reduces the influence of the patch region. Experiments on ImageNet validate that our proposal can effectively improve ViT’s robustness towards white-box universal patch attacks while maintaining satisfactory classification accuracy for clean samples.
Huimin Ma 0001, Xinlong Ding
ICASSP3
2023 Learning a Graph Neural Network with Cross Modality Interaction for Image Fusion
abstract
Infrared and visible image fusion has gradually proved to be a vital fork in the field of multi-modality imaging technologies. In recent developments, researchers not only focus on the quality of fused images but also evaluate their performance in downstream tasks. Nevertheless, the majority of methods seldom put their eyes on mutual learning from different modalities, resulting in fused images lacking significant details and textures. To overcome this issue, we propose an interactive graph neural network (GNN)-based architecture between cross modality for fusion, called IGNet. Specifically, we first apply a multi-scale extractor to achieve shallow features, which are employed as the necessary input to build graph structures. Then, the graph interaction module can construct the extracted intermediate features of the infrared/visible branch into graph structures. Meanwhile, the graph structures of two branches interact for cross-modality and semantic learning, so that fused images can maintain the important feature expressions and enhance the performance of downstream tasks. Besides, the proposed leader nodes can improve information propagation in the same modality. Finally, we merge all graph features to get the fusion result. Extensive experiments on different datasets (i.e. TNO, MFNet, and M3FD) demonstrate that our IGNet can generate visually appealing fused images while scoring averagely 2.59% [email protected] and 7.77% mIoU higher in detection and segmentation than the compared state-of-the-art methods. The source code of the proposed IGNet can be available at https://github.com/lok-18/IGNet.
Jiawei Li 0016, Jiansheng Chen 0002, Jinyuan Liu 0001, Huimin Ma 0001
ACM Multimedia4
2023 Micro-Expression Spotting with Face Alignment and Optical Flow
abstract
Facial expression spotting holds significant importance as it can signify emotional changes. Particularly, micro-expressions possess the potential to reveal genuine emotions, making them even more valuable in practical domains such as public safety and finance. However, spotting micro-expressions proves challenging due to their subtle movements and brief duration. This paper proposes an expression spotting method based on face alignment and optical flow. We first use a finer crop-align technique to preprocess the facial videos by aligning the face and the nose tip. Then, regions of interest (ROIs) are defined by analyzing the statistics of action units. The optical flow features are then extracted and subjected to low-pass filtering to eliminate high-frequency noise. Furthermore, candidate expression segments are identified based on the magnitude of the processed optical flows. Finally, non-maximum suppression is utilized to remove overlapping segments. The effectiveness of the proposed method is evaluated on the challenge test set, resulting in an overall F1-score of 0.19. Additional results obtained from CAS(ME)2 and SAMM Long videos provide further verification of the method's efficacy. The code is available online.
Wenfeng Qin, Bochao Zou, Xin Li 0034, Weiping Wang 0007, Huimin Ma 0001
ACM Multimedia5
2023 Enhancing Adversarial Robustness of Multi-modal Recommendation via Modality Balancing
abstract
Recently multi-modal recommender systems have been widely applied in real scenarios such as e-commerce businesses. Existing multi-modal recommendation methods exploit the multi-modal content of items as auxiliary information and fuse them to boost performance. Despite the superior performance achieved by multi-modal recommendation models, there's currently no understanding of their robustness to adversarial attacks. In this work, we first identify the vulnerability of existing multi-modal recommendation models. Next, we show the key reason for such vulnerability is modality imbalance, i.e., the prediction score margin between positive and negative samples in the sensitive modality will drop dramatically facing adversarial attacks and fail to be compensated by other modalities. Finally, based on this finding we propose a novel defense method to enhance the robustness of multi-modal recommendation models through modality balancing. Specifically, we first adopt an embedding distillation to obtain a pair of content-similar but prediction-different item embeddings in the sensitive modality and calculate the score margin reflecting the modality vulnerability. Then we optimize the model to utilize the score margin between positive and negative samples in other modalities to compensate for the vulnerability. The proposed method can serve as a plug-and-play module and is flexible to be applied to a wide range of multi-modal recommendation models. Extensive experiments on two real-world datasets demonstrate that our method significantly improves the robustness of multi-modal recommendation models with nearly no performance degradation on clean data.
Chen Gao 0001, Jiansheng Chen 0001, Depeng Jin, Huimin Ma 0001, Yong Li 0008
ACM Multimedia5
2023 Synthesizing Videos from Images for Image-to-Video Adaptation
abstract
We address the image-to-video adaptation task that aims to leverage labeled images and unlabeled videos for video recognition. There are two major challenges in this task, including the domain discrepancy between the two domains, and the modality gap between the image and video modalities. Existing methods mainly employ a two-stage paradigm by first adopting frame-level adaptation to reduce the domain discrepancy and then learning a spatio-temporal model to bridge the modality gap. In this paper, we provide a new perspective and propose a single-stage method that synthesizes video from the source static image and converts the image-to-video adaptation problem into a video-to-video adaptation problem. With the synthesized video, we present a simple baseline that a spatio-temporal model is trained with cross entropy loss with source labels and the Batch Nuclear norm Maximization loss to encourage the classification responses of target videos maintain the discriminability and diversity. We further propose a new pseudo label generation method that inherits the robustness of class prototype and the effectiveness of the small loss criterion. Based on the constructed baseline and the proposed pseudo label generation method, we train a model that achieves state-of-the-art performances or gets comparable performances on three standard benchmarks. Our codes are publicly available at https://github.com/junbaoZHUO/ST-I2V.
Junbao Zhuo, Xingyu Zhao 0005, Shuhui Wang, Huimin Ma 0001, Qingming Huang
ACM Multimedia4
2023 FD-Align: Feature Discrimination Alignment for Fine-tuning Pre-Trained Models in Few-Shot Learning
abstract
Due to the limited availability of data, existing few-shot learning methods trained from scratch fail to achieve satisfactory performance. In contrast, large-scale pre-trained models such as CLIP demonstrate remarkable few-shot and zero-shot capabilities. To enhance the performance of pre-trained models for downstream tasks, fine-tuning the model on downstream data is frequently necessary. However, fine-tuning the pre-trained model leads to a decrease in its generalizability in the presence of distribution shift, while the limited number of samples in few-shot learning makes the model highly susceptible to overfitting. Consequently, existing methods for fine-tuning few-shot learning primarily focus on fine-tuning the model's classification head or introducing additional structure. In this paper, we introduce a fine-tuning approach termed Feature Discrimination Alignment (FD-Align). Our method aims to bolster the model's generalizability by preserving the consistency of spurious features across the fine-tuning process. Extensive experimental results validate the efficacy of our approach for both ID and OOD tasks. Once fine-tuned, the model can seamlessly integrate with existing methods, leading to performance improvements. Our code can be found in https://github.com/skingorz/FD-Align.
Kun Song 0004, Huimin Ma 0001, Bochao Zou, Huishuai Zhang, Weiran Huang 0001
NeurIPS2
2023 Multi-source Information Fusion for Depression Detection
Rongquan Wang, Huiwei Wang, Huimin Ma 0001
PRCV (5)5
2023 Semi-Structural Interview-Based Chinese Multimodal Depression Corpus Towards Automatic Preliminary Screening of Depressive Disorders
abstract
Depression is a common psychiatric disorder worldwide. However, in China, a considerable number of patients with depression are not diagnosed, and most of them are not aware of their depression. Despite increasing efforts, the goal of automatic depression screening from behavioral indicators has not been achieved. A major limitation is the lack of available multimodal depression corpus in Chinese since linguistic knowledge is crucial in clinical practice. Therefore, we first carried out a comprehensive survey with psychiatrists from a renowned psychiatric hospital to identify key interview topics which are highly related to the diagnosis of depression. Then, a semi-structural interview study was conducted over a year with subjects who have undergone clinical diagnosis and professional assessment. After that, Visual, acoustic, and textual features were extracted and analyzed between the two groups, statistically significant differences were observed in all three modalities. Benchmark evaluations of both single modal and multimodal fusion methods of depression assessment were also performed. A multimodal transformer-based fusion approach achieved the best performance. Finally, the proposed Chinese Multimodal Depression Corpus (CMDC) was made publicly available after de-identification and annotation. Hopefully, the release of this corpus would promote the research progress and practical applications of automatic depression screening.
Bochao Zou, Jiali Han, Xiangwen Lyu, Huimin Ma 0001
IEEE Trans. Affect. Comput.8
2023 A Hybrid Deep Transfer Learning Model With Kernel Metric for COVID-19 Pneumonia Classification Using Chest CT Images
abstract
Coronavirus disease-2019 (COVID-19) as a new pneumonia which is extremely infectious, the classification of this coronavirus is essential to effectively control the development of the epidemic. Pathological changes in the chest computed tomography (CT) scans are often used as one of the diagnostic criteria of COVID-19. Meanwhile, deep learning-based transfer learning is currently an effective strategy for computer-aided diagnosis (CAD). To further improve the performance of deep transfer learning model used for COVID-19 classification with CT images, in this article, we propose a hybrid model combined with a semi-supervised domain adaption model and extreme learning machine (ELM) classifier, and the application of a novel multikernel correntropy induced loss function in transfer learning is also presented. The proposed model is evaluated on open-source datasets. The experimental results are compared to some baseline models to verify the effectiveness, while adopting accuracy, precision, recall,$F_{1}$score and area under curve (AUC) as the evaluation metrics. Experimental results show that the proposed method improves the performance of original model and is more suitable for CT images analysis.
Jianyuan Li, Xiong Luo, Huimin Ma 0001, Wenbing Zhao 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Temporal Information Fusion Network for Driving Behavior Prediction
abstract
Since enormous hazards are caused by traffic crashes every year, ensuring safe driving is a hot topic in transportation. Technologies related to the Advanced Driver Assistance System (ADAS) are evolving rapidly. But without an adequate understanding of driving intention, ADAS usually can’t help the driver prepare for the danger in advance. This paper focuses on the fusion strategy of driver and environment information and proposes a lightweight end-to-end model, temporal information fusion network (TIFN). Driving behavior is the interactive result of the driver and the external world. To better understand the driver’s intention, the state update cell (STU) is proposed to introduce the influence of environment information into the driver’s state modeling, inspired by the selective attention of the human cognition process. Meanwhile, semantic segmentation features are extracted to offer clear clues affecting driver attention in place of motion optical flow images and binary value vectors. Finally, the driver’s intention and environment state are combined to make a joint prediction. The experiments evaluated on Brain4cars and IESDD show that the proposed approach has superior performance than other approaches that only use camera data.
Chenghao Guo, Haizhuang Liu, Jiansheng Chen 0001, Huimin Ma 0001
IEEE Trans. Intell. Transp. Syst.4
2023 Virtual Try-On With Garment Self-Occlusion Conditions
abstract
Image-based virtual try-on focuses on changing the model's garment item to the target ones and preserving other visual features. To preserve the texture detail of the given in-shop garment, former methods use geometry-based methods (e.g., Thin-plate-spline interpolation) to realize garment warping. However, due to limited degree of freedom, geometry-based methods perform poorly when garment self-occlusion occurs, which is common in daily life. To address this challenge, we propose a novel occlusion-focused virtual try-on system. Compared to previous ones, our system contains three critical submodules, namely, Garment Part Modeling (GPM), a group of Garment Part Generators (GPGs), and Overlap Relation Estimator (ORE). GPM takes the pose landmarks as input, and progressively models the mask of body parts and garments. Based on these masks, GPGs are introduced to generate each garment part. Finally, ORE is proposed to model the overlap relationships between each garment part, and we bind the generated garments under the guidance of overlap relationships predicted by ORE. To make the most of extracted overlap relationships, we proposed an IoU-based hard example mining method for loss terms to handle the sparsity of the self-occlusion samples in the dataset. Furthermore, we introduce part affinity field as pose representation instead of landmark used widely by previous methods and achieve accuracy improvement on try-on layout estimation stage. We evaluate our model on the VITON dataset and found it can outperform previous approaches, especially on samples with garment self-occlusion.
Zhening Xing, Si Liu 0001, Shangzhe Di, Huimin Ma 0001
IEEE Trans. Multim.5
2023 MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs
abstract
Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents but varying in motions. However, most previous video synthesis frameworks based on pre-trained image generators treat content and motion generation separately, leading to unrealistic generated videos. Therefore, we design a novel framework to build the motion space, aiming to achieve content consistency and fast convergence for video generation. We present MotionVideoGAN, a novel video generator synthesizing videos based on the motion space learned by pre-trained image pair generators. Firstly, we propose an image pair generator named MotionStyleGAN to generate image pairs sharing the same contents and producing various motions. Then we manage to acquire motion codes to edit one image in the generated image pairs and keep the other unchanged. The motion codes help us edit images within the motion space since the edited image shares the same contents with the other unchanged one in image pairs. Finally, we introduce a latent code generator to produce latent code sequences using motion codes for video generation. Our approach achieves state-of-the-art performance on the most complex video dataset ever used for unconditional video generation evaluation, UCF101.
Huimin Ma 0001, Jiansheng Chen 0001
IEEE Trans. Multim.2
2022 Gestalt-Guided Image Understanding for Few-Shot Learning
Kun Song 0004, Jiansheng Chen 0001, Huimin Ma 0001
ACCV (2)5
2022 Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial Labels
abstract
Previous weakly-supervised methods of 3D object detection in driving scenes mainly rely on spatial labels, which provide the location, dimension, or orientation information. The annotation of 3D spatial labels is time-consuming. There also exist methods that do not require spatial labels, but their detections may fall on object parts rather than entire objects or backgrounds. In this paper, a novel cross-modal weakly-supervised 3D progressive refinement framework (WS3DPR) for 3D object detection that only needs image-level class annotations is introduced. The proposed framework consists of two stages: 1) classification refinement for potential objects localization and 2) regression refinement for spatial pseudo labels reasoning. In the first stage, a region proposal network is trained by cross-modal class knowledge transferred from 2D image to 3D point cloud and class information propagation. In the second stage, the locations, dimensions, and orientations of 3D bounding boxes are further refined with geometric reasoning based on 2D frustum and 3D region. When only image-level class labels are available, proposals with different 3D locations become overlapped in 2D, leading to the misclassification of foreground objects. Therefore, a 2D-3D semantic consistency block is proposed to disentangle different 3D proposals after projection. The overall framework progressively learns features in a coarse to fine manner. Comprehensive experiments on the KITTI3D dataset demonstrate that our method achieves competitive performance compared with previous methods with a lightweight labeling process.
Haizhuang Liu, Huimin Ma 0001, Bochao Zou, Rongquan Wang, Jiansheng Chen 0001
ACM Multimedia2
2022 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
abstract
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
Youze Xue, Jiansheng Chen 0001, Yudong Zhang 0008, Huimin Ma 0001, Hongbing Ma
ACM Multimedia5
2022 Detecting protein complexes with multiple properties by an adaptive harmony search algorithm
abstract
BACKGROUND: Accurate identification of protein complexes in protein-protein interaction (PPI) networks is crucial for understanding the principles of cellular organization. Most computational methods ignore the fact that proteins in a protein complex have a functional similarity and are co-localized and co-expressed at the same place and time, respectively. Meanwhile, the parameters of the current methods are specified by users, so these methods cannot effectively deal with different input PPI networks. RESULT: To address these issues, this study proposes a new method called MP-AHSA to detect protein complexes with Multiple Properties (MP), and an Adaptation Harmony Search Algorithm is developed to optimize the parameters of the MP algorithm. First, a weighted PPI network is constructed using functional annotations, and multiple biological properties and the Markov cluster algorithm (MCL) are used to mine protein complex cores. Then, a fitness function is defined, and a protein complex forming strategy is designed to detect attachment proteins and form protein complexes. Next, a protein complex filtering strategy is formulated to filter out the protein complexes. Finally, an adaptation harmony search algorithm is developed to determine the MP algorithm's parameters automatically. CONCLUSIONS: Experimental results show that the proposed MP-AHSA method outperforms 14 state-of-the-art methods for identifying protein complexes. Also, the functional enrichment analyses reveal that the protein complexes identified by the MP-AHSA algorithm have significant biological relevance.
Rongquan Wang, Huimin Ma 0001
BMC Bioinform.3
2022 Correction: Detecting protein complexes with multiple properties by an adaptive harmony search algorithm
Rongquan Wang, Huimin Ma 0001
BMC Bioinform.3
2022 LIDAR: learning from imperfect demonstrations with advantage rectification
Huimin Ma 0001, Xiong Luo
Frontiers Comput. Sci.2
2022 Visibility of points: Mining occlusion cues for monocular 3D object detection
Huazhen Chu, Lisha Mo, Rongquan Wang, Huimin Ma 0001
Neurocomputing5
2022 Attribute assisted teacher-critical training strategies for image captioning
Jiansheng Chen 0001, Huimin Ma 0001, Hongbing Ma, Wanli Ouyang
Neurocomputing3
2022 Co-attention dictionary network for weakly-supervised semantic segmentation
Weitao Wan, Jiansheng Chen 0001, Ming-Hsuan Yang 0001, Huimin Ma 0001
Neurocomputing4
2022 Weakly-supervised semantic segmentation with superpixel guided local and global consistency
Huimin Ma 0001, Xiang Wang 0003, Xi Li 0010, Yu Wang 0002
Pattern Recognit.2
2022 Concordance between facial micro-expressions and physiological signals under emotion elicitation
Bochao Zou, Xiangwen Lyu, Huimin Ma 0001
Pattern Recognit. Lett.5
2022 Cross-Modal Cross-Domain Dual Alignment Network for RGB-Infrared Person Re-Identification
abstract
RGB-Infrared cross-modal person re-identification (Re-ID) has drawn increasing attention due to its application value in practice. Most of the current works rely on a supervised training manner. However, in real-world applications, manual collection of pair-wise RGB-Infrared (IR) person data is labor-intensive and time-consuming. Moreover, when a trained model is directly used in another domain, there is usually a significant performance drop. To overcome the above problems, we make the first attempt to transfer the learned model to a new RGB-IR domain which is unlabeled. The practical problem covers two kinds of challenges, i.e., cross-modal (RGB-Infrared) and cross-domain (different dataset) person Re-ID. Previous works have often considered only one of them either cross-modal or cross-domain. In this work, we propose a dual alignment network (DAN) to solve the RGB-Infrared cross-modal cross-domain person Re-ID problem. This network consists of three parts: Domain Adversarial Alignment component (DAA), Pseudo Label Generation module for target domain (PLG), and Cross-Modal Alignment component (CMA). These three modules complement and promote the model to learn domain-invariant and modality-invariant person representations. Further, we propose a protocol of cross-modal cross-domain person Re-ID by synthesizing target domains by adding random noise, adjusting the lighting intensity, and changing the background color, respectively. Experiments on real and synthetic datasets under the same cross-modalities across domains demonstrate the effectiveness of our method.
Xiaowei Fu, Fuxiang Huang, Huimin Ma 0001, Xin Xu 0001, Lei Zhang 0038
IEEE Trans. Circuits Syst. Video Technol.4
2022 Self-Attention-Based Machine Theory of Mind for Electric Vehicle Charging Demand Forecast
abstract
The popularization of electric vehicles (EVs) and charging stations has been threatening the distribution network’s reliability and efficiency. The prediction of EV charging demand can benefit the optimization of the operation of energy-transportation nexus and improve social welfare toward a low carbon future. In this article, a short-term probabilistic charging demand forecast model is proposed to estimate the quantiles of future charging demand of a charging station 15 min ahead, i.e., the self-attention-based machine theory of mind (SAMToM). The SAMToM has considered both the users’ historical charging habits (schedules) and the current trend of charging demand variation using the framework of machine theory of mind (MToM), and real-world-data-based case studies have verified its superiority in EV charging demand forecast over state-of-the-arts. Moreover, analyses show that the advantage of SAMToM lies in the following aspects. 1) The self-attention layers have mitigated the long-range forgetting in SAMToM. 2) The MToM architecture enables SAMToM to balance historical charging habits and current charging demand variation trends well. 3) Using a quantile forecast evaluation metric as the loss function, i.e., the continuous ranked probability score (CRPS), enables SAMToM to aim directly at the highest quality of forecasted quantiles.
Huimin Ma 0001, Hongbin Sun 0002, Kailong Liu
IEEE Trans. Ind. Informatics2
2022 A Transferred Recurrent Neural Network for Battery Calendar Health Prognostics of Energy-Transportation Systems
abstract
Battery-based energy storage system is a key component to achieve low carbon industrial and social economy, where battery health status plays a vital role in determining the safety and reliability of energy-transportation nexus. This article proposes a transferred recurrent neural network (RNN)-based framework to achieve efficient calendar capacity prognostics under both witnessed and unwitnessed storage conditions. Specifically, this transferred RNN framework contains a base model part and a transfer model part. The base model is first trained by using the easily collected and time-saving accelerated ageing dataset from high temperature and state-of-charge (SOC) cases. Then the transfer part is tuned by using only a small portion of starting capacity data from unwitnessed condition of interest. The developed framework is evaluated under a well-rounded ageing dataset with three different storage SOCs (20%, 50%, and 90%) and temperatures (10 °C, 25 °C, and 45 °C). Experimental results demonstrate that the derived transferred RNN framework is capable of providing satisfactory calendar capacity health prognostics under different storage cases. A model structure with the impact factor terms of SOC and temperature outperforms other counterparts especially for the unwitnessed conditions. The proposed framework could assist engineers to significantly reduce battery ageing experiment burden and is also promising to capture future capacity information for battery health and life-cycle cost analysis of energy-transportation applications.
Kailong Liu, Hongbin Sun 0002, Minrui Fei, Huimin Ma 0001
IEEE Trans. Ind. Informatics5
2022 Boosting Monocular 3D Human Pose Estimation With Part Aware Attention
abstract
Monocular 3D human pose estimation is challenging due to depth ambiguity. Convolution-based and Graph-Convolution-based methods have been developed to extract 3D information from temporal cues in motion videos. Typically, in the lifting-based methods, most recent works adopt the transformer to model the temporal relationship of 2D keypoint sequences. These previous works usually consider all the joints of a skeleton as a whole and then calculate the temporal attention based on the overall characteristics of the skeleton. Nevertheless, the human skeleton exhibits obvious part-wise inconsistency of motion patterns. It is therefore more appropriate to consider each part's temporal behaviors separately. To deal with such part-wise motion inconsistency, we propose the Part Aware Temporal Attention module to extract the temporal dependency of each part separately. Moreover, the conventional attention mechanism in 3D pose estimation usually calculates attention within a short time interval. This indicates that only the correlation within the temporal context is considered. Whereas, we find that the part-wise structure of the human skeleton is repeating across different periods, actions, and even subjects. Therefore, the part-wise correlation at a distance can be utilized to further boost 3D pose estimation. We thus propose the Part Aware Dictionary Attention module to calculate the attention for the part-wise features of input in a dictionary, which contains multiple 3D skeletons sampled from the training set. Extensive experimental results show that our proposed part aware attention mechanism helps a transformer-based model to achieve state-of-the-art 3D pose estimation performance on two widely used public datasets. The codes and the trained models are released at https://github.com/thuxyz19/3D-HPE-PAA.
Youze Xue, Jiansheng Chen 0001, Xiangming Gu, Huimin Ma 0001, Hongbing Ma
IEEE Trans. Image Process.4
2022 Learning representative viewpoints in 3D shape recognition
Huazhen Chu, Chao Le, Rongquan Wang, Xi Li 0010, Huimin Ma 0001
Vis. Comput.5
2021 Effect of Depression Severity on Emotion Context Insensitivity Revealed by Facial Activities Analysis
abstract
Major Depression Disorder (MDD) is defined as a mood condition. The Emotion Context Insensitivity (ECI) hypothesis of depression argues that the reaction sensitivity of depressed people to the changes of emotional context is attenuated. This paper designs an experiment under this hypothesis to more effectively distinguish facial cues for different levels of depression severity. By adopting publicly validated video stimuli for emotion elicitation, statistical analysis was performed to reveal significant facial features for depression severity assessment. The results found that the high severity group showed a decrease in specific action units and an increase in others, and made more intense facial movements for negative stimuli than that for positive and neutral stimuli. Hopefully, this study could benefit the interpretation of emotional functioning in depression and provide a reference for the evaluation of depression severity based on facial characteristics.
Bochao Zou, Xiangwen Lyu, Huimin Ma 0001
BIBM6
2021 Depression Detection by Analysing Eye Movements on Emotional Images
abstract
To achieve an objective and efficient depression detection system, we propose a cognitive psychology experimental paradigm based on the attentional bias theory and eye movements in this paper. We select images of three different emotions (positive, neutral, and negative) as experimental stimulus. Comparing with the traditional free viewing paradigm, the paradigm we proposed adds a stage of frame tracking to analyse the process of attention disengagement. Based on extracted psychological features from eye movement data, we train a mental state classifier of Support Vector Machine to classify people with depression and normal controls, and the model achieves 77.0% of accuracy, which achieve state-of-the-art under the same data condition. Our model is interpretable and our results demonstrate the theory of attention bias.
Ruizhe Shen, Qi Zhan, Yu Wang 0002, Huimin Ma 0001
ICASSP4
2021 Defending against Universal Adversarial Patches by Clipping Feature Norms
abstract
Physical-world adversarial attacks based on universal adversarial patches have been proved to be able to mislead deep convolutional neural networks (CNNs), exposing the vulnerability of real-world visual classification systems based on CNNs. In this paper, we empirically reveal and mathematically explain that the universal adversarial patches usually lead to deep feature vectors with very large norms in popular CNNs. Inspired by this, we propose a simple yet effective defending approach using a new feature norm clipping (FNC) layer which is a differentiable module that can be flexibly inserted in different CNNs to adaptively suppress the generation of large norm deep feature vectors. FNC introduces no trainable parameter and only very low computational overhead. However, experiments on multiple datasets validate that it can effectively improve the robustness of different CNNs towards white-box universal patch attacks while maintaining a satisfactory recognition accuracy for clean samples.
Youze Xue, Weitao Wan, Jiayu Bao, Huimin Ma 0001
ICCV7
2021 Device-Adaptive 2D Gaze Estimation: A Multi-Point Differential Framework
Runtong Li, Huimin Ma 0001, Rongquan Wang
ICIG (2)2
2021 Frequency Transfer Model: Generating High Frequency Components for Fluid Simulation Details Reconstruction
Huimin Ma 0001
ICIG (3)2
2021 Depression Detection by Combining Eye Movement with Image Semantics
abstract
Depression is a common mental disorder that affects patients’ daily life. Most existing depression detection methods consume a lot of medical resources and exist at risk of subjective judgment. Therefore, we propose an objective and convenient experimental paradigm. Firstly, it selects emotional images as stimuli and records the subjects’ eye movement data. Secondly, we establish a connection between image processing and subjects’ psychological conditions analysis. Rather than some AI-based methods focus on feature engineering of recorded data, we design the saliency difference detection network and semantic segmentation network to explore the images’ deep semantic features and combine them with the subjects’ gaze pattern. Finally, we train a mental state classifier of Support Vector Machine to detect depression. The experimental results demonstrate that it achieves accuracy up to 90.06%, which outperforms previous methods.
Huimin Ma 0001, Zeyu Pan, Rongquan Wang
ICIP2
2021 PLNL-3DSSD: Part-Aware 3D Single Stage Detector Using Local And Non-Local Attention
abstract
3D object detection in the real crowded scene is still a challenging task due to occlusion and density change. We propose a part-aware 3D single-stage detector with local and non-local attention (PLNL-3DSSD) to fully use part information and inter-object relation. A primary part feature fusion is proposed for encoding the entire box feature vector by introducing semantic parts dividing. We develop a parallel part branch for robust and accurate object detection. We also develop 10-cal and non-local attention in set abstraction for enhancing data flow transfer between objects. Our method ranks second in single-stage 3D object detector on the KITTI 3D car detection benchmark while ensuring satisfactory efficiency.
Haizhuang Liu, Huimin Ma 0001, Yanxian Chen, Xi Li 0010
ICIP2
2021 Enhancing Adversarial Robustness For Image Classification By Regularizing Class Level Feature Distribution
abstract
Recent researches have shown that deep neural networks (DNNs) are vulnerable to adversarial examples. Adversarial training is practically the most effective approach to improve the robustness of DNNs against adversarial examples. However, conventional adversarial training methods only focus on the classification results or the instance level relationship on feature representations for adversarial examples. Inspired by the fact that adversarial examples break the distinguishability of the feature representations of DNNs for different classes, we propose Intra and Inter Class Feature Regularization $(\mathrm{I}^{2}$ FR) to make the feature distribution of adversarial examples maintain the same classification property as clean examples. On the one hand, the intra-class regularization restricts the distance of features between adversarial examples and both the corresponding clean data and samples for the same class. On the other hand, the inter-class regularization prevents the feature of adversarial examples from getting close to other classes. By adding $\mathrm{I}^{2}$ FR in both adversarial example generation and model training steps in adversarial training, we can get stronger and more diverse adversarial examples, and the neural network learns a more distinguishable and reasonable feature distribution. Experiments on various adversarial training frameworks demonstrate that $\mathrm{I}^{2}$ FR is adaptive for multiple training frameworks and outperforms the state-of-the-art methods for classification of both clean data and adversarial examples.
Youze Xue, Jiansheng Chen 0001, Yu Wang 0002, Huimin Ma 0001
ICIP5
2021 Semantic Tag Augmented XlanV Model for Video Captioning
abstract
The key of video captioning is to leverage the cross-modal information from both vision and language perspectives. We propose to leverage the semantic tags to bridge the gap between these modalities rather than directly concatenating or attending to the visual and linguistic features as the previous works. The semantic tags are the object tags and the action tags detected in videos, which can be viewed as partial captions for the input video. To effectively exploit the semantic tags, we design a Semantic Tag augmented XlanV (ST-XlanV) model which encodes 4 kinds of visual and semantic features with X-Linear Attention based cross-attention modules. Moreover, tag related tasks are also designed in the pre-training stage to aid the model more fruitfully exploits the cross-modal information. The proposed model reaches the 5th place in the pre-training for video captioning challenge with the help of the semantic tags. Our codes will be available at: https://github.com/RubickH/ST-XlanV.
Hongwei Xue, Huimin Ma 0001, Hongbing Ma
ACM Multimedia4
2021 Variational Automatic Curriculum Learning for Sparse-Reward Cooperative Multi-Agent Problems
abstract
We introduce an automatic curriculum algorithm, Variational Automatic Curriculum Learning (VACL), for solving challenging goal-conditioned cooperative multi-agent reinforcement learning problems. We motivate our curriculum learning paradigm through a variational perspective, where the learning objective can be decomposed into two terms: task learning on the current curriculum, and curriculum update to a new task distribution. Local optimization over the second term suggests that the curriculum should gradually expand the training tasks from easy to hard. Our VACL algorithm implements this variational paradigm with two practical components, task expansion and entity curriculum, which produces a series of training tasks over both the task configurations as well as the number of entities in the task. Experiment results show that VACL solves a collection of sparse-reward problems with a large number of agents. Particularly, using a single desktop machine, VACL achieves 98% coverage rate with 100 agents in the simple-spread benchmark and reproduces the ramp-use behavior originally shown in OpenAI’s hide-and-seek project.
Jiayu Chen 0005, Yuanxin Zhang, Yuanfan Xu, Huimin Ma 0001, Huazhong Yang, Jiaming Song, Yu Wang 0002, Yi Wu 0013
NeurIPS4
2021 Improving Adversarial Robustness of Detector via Objectness Regularization
Jiayu Bao, Hongbing Ma, Huimin Ma 0001
PRCV (4)4
2021 LiDAR-Based Symmetrical Guidance for 3D Object Detection
Huazhen Chu, Huimin Ma 0001, Haizhuang Liu, Rongquan Wang
PRCV (4)2
2021 Face Anti-spoofing Based on Cooperative Pose Analysis
Poyu Lin, Huimin Ma 0001, Hongbing Ma
PRCV (3)4
2021 Distance-Based Class Activation Map for Metric Learning
Yeqing Shen, Huimin Ma 0001, Yuhan Dong
PRCV (4)2
2021 MVAD-Net: Learning View-Aware and Domain-Invariant Representation for Baggage Re-identification
Huimin Ma 0001, Ruiqi Lu, Yanxian Chen
PRCV (1)2
2021 STA-GCN: Spatio-Temporal AU Graph Convolution Network for Facial Micro-expression Recognition
Xinhui Zhao, Huimin Ma 0001, Rongquan Wang
PRCV (1)2
2021 Label Disentangled Analysis for unsupervised visual domain adaptation
Ni Xiao, Lei Zhang 0038, Xin Xu 0001, Tan Guo, Huimin Ma 0001
Knowl. Based Syst.5
2021 Single annotated pixel based weakly supervised semantic segmentation under driving scenes
Xi Li 0010, Huimin Ma 0001, Yanxian Chen, Hongbing Ma
Pattern Recognit.2
2021 Pedestrian instance segmentation with prior structure of semantic parts
Huazhen Chu, Huimin Ma 0001, Xi Li 0010
Pattern Recognit. Lett.2
2021 Driving Behavior Prediction Considering Cognitive Prior and Driving Context
abstract
Driving behavior plays a key role in the interaction between vehicle and driver in transportation systems. Some applications about driving behavior in Advanced Driver Assistance Systems (ADAS) improve driving safety significantly. This paper introduces the driving context and models driving behavior in a combination of cognitive perspective and data-driven perspective. First, we use a cognitive fusion method by adding a delay time module to fuse the environmental information and inside information. To better capture the driving context relationship between outside and inside features, we transfer the behavior prediction task to the sequence labeling task by introducing the visual inertia hypothesis. We propose the Predictive-Bi-LSTM-CRF algorithm which used the Bidirectional Long-Short Term Memory Networks (Bi-LSTM) and Conditional Random Field (CRF) as the loss layer to model the driving behavior. Besides, we define a new comprehensive evaluation metric for the prediction task considering F1-score and the prediction time before maneuver together. Our experiment results achieve the state of art performance on the Brain4Cars dataset and demonstrate the applicability of our theory.
Huimin Ma 0001, Xiang Wang 0003, Yuhan Dong
IEEE Trans. Intell. Transp. Syst.3
2020 S-VoteNet: Deep Hough Voting with Spherical Proposal for 3D Object Detection
abstract
Current 3D object detection methods adopt an analogous box prediction structure with the 2D methods, which predict center and size of the object simultaneously in a box regression procedure, leading to the poor performance of 3D detector to a great extent. In this work, we propose S-VoteNet, which converts the prediction of 3D bounding box into two parts: center prediction and size prediction. By introducing a novel spherical proposal, S-VoteNet uses vote groups to predict the center and radius of object rather than all parameters of 3D bounding box. The prediction of radius is used to constrain the object size, and the radius-based spherical center loss is applied to measure the geometric distance between the proposal and ground-truth. To make better use of the geometric information provided by point cloud, S-VoteNet aggregates seeds by the votes indices to generate seed groups. The seed groups are then used for box size regression and orientation estimation. By decoupling the localization and size estimation, our method effectively reduces the regression burden of the 3D detector. Experimental results on SUN RGB-D 3D detection benchmark demonstrate that our S-VoteNet achieves state-of-the-art performance by using only point cloud as input.
Yanxian Chen, Huimin Ma 0001, Xi Li 0010, Xiong Luo
ICPR2
2020 AVD-Net: Attention Value Decomposition Network For Deep Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) is of importance for variable real-world applications but remains more challenges like stationarity and scalability. While recently value function factorization methods have obtained empirical good results in cooperative multi-agent environment, these works mostly focus on the decomposable learning structures. Inspired by the application of attention mechanism in machine translation and other related domains, we propose an attention based approach called attention value decomposition network (AVD-Net), which capitalizes on the coordination relations between agents. AVD-Net employs centralized training with decentralized execution (CTDE) paradigm, which factorizes the joint action-value functions with only local observations and actions of agents. Our method is evaluated on multi-agent particle environment (MPE) and StarCraft micromanagement environment (SMAC). The experiment results show the strength of our approach compared to existing methods with state-of-the-art performance in cooperative scenarios.
Yuanxin Zhang, Huimin Ma 0001, Yu Wang 0002
ICPR2
2020 Cross-Domain Disentangle Network for Image Manipulation
Zhening Xing, Jinghuan Wen, Huimin Ma 0001
PRCV (3)3
2020 Weakly-Supervised Semantic Segmentation by Iterative Affinity Learning
Xiang Wang 0003, Sifei Liu, Huimin Ma 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2020 Semantic head enhanced pedestrian detection in a crowd
Ruiqi Lu, Huimin Ma 0001, Yu Wang 0002
Neurocomputing2
2020 Deep clustering for weakly-supervised semantic segmentation in autonomous driving scenes
Xiang Wang 0003, Huimin Ma 0001, Shaodi You
Neurocomputing2
2020 WSODPB: Weakly supervised object detection with PCSNet and box regression module
Huimin Ma 0001, Xi Li 0010, Yu Wang 0002
Neurocomputing2
2020 OccGAN: Semantic image augmentation for driving scenes
Lisha Mo, Huimin Ma 0001
Pattern Recognit. Lett.3
2020 Weaklier Supervised Semantic Segmentation With Only One Image Level Annotation per Category
abstract
Image semantic segmentation tasks and methods based on weakly supervised conditions have been proposed and achieve better and better performance in recent years. However, the purpose of these tasks is mainly to simplify the labeling work. In this paper, we establish a new and more challenging task condition: weaklier supervision with one image level annotation per category, which only provides prior knowledge that humans need to recognize new objects, and aims to achieve pixel-level object semantic understanding. In order to solve this problem, a three-stage semantic segmentation framework is put forward, which realizes image level, pixel level, and object common features learning from coarse to fine grade, and finally obtains semantic segmentation results with accurate and complete object regions. Researches on PASCAL VOC 2012 dataset demonstrates the effectiveness of the proposed method, which makes an obvious improvement compared to baselines. Based on fewer supervised information, the method also provides satisfactory performance compared to weakly supervised learning-based methods with complete image-level annotations.
Xi Li 0010, Huimin Ma 0001, Xiong Luo
IEEE Trans. Image Process.2
2020 Deep generative smoke simulator: connecting simulated and real data
Jinghuan Wen, Huimin Ma 0001, Xiong Luo
Vis. Comput.2
2019 Semantic Inference Network for Human-Object Interaction Detection
Lisha Mo, Huimin Ma 0001
ICIG (1)3
2019 Challenges Driven Network for Visual Tracking
Jiaming Wei, Huimin Ma 0001, Ruiqi Lu
ICIG (1)2
2019 Modified Capsule Network for Object Classification
Huimin Ma 0001, Xi Li 0010
ICIG (1)2
2019 Occluded Pedestrian Detection with Visible IoU and Box Sign Predictor
abstract
Training a robust classifier and an accurate box regressor are difficult for occluded pedestrian detection. Traditionally adopted Intersection over Union (IoU) measurement does not consider the occluded region of the object and leads to improper training samples. To address such issue, a modification called visible IoU is proposed in this paper to explicitly incorporate the visible ratio in selecting samples. Then a newly designed box sign predictor is placed in parallel with box regressor to separately predict the moving direction of training samples. It leads to higher localization accuracy by introducing sign prediction loss during training and sign refining in testing. Following these novelties, we obtain state-of-the-art performance on CityPersons benchmark for occluded pedestrian detection.
Ruiqi Lu, Huimin Ma 0001
ICIP2
2019 Essential element-region driven model in image recognition
Lisha Mo, Xiong Luo, Huimin Ma 0001
Neurocomputing4
2019 Real-time smoke simulation based on vorticity preserving lattice Boltzmann method
Jinghuan Wen, Huimin Ma 0001
Vis. Comput.2
2018 Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features
abstract
Weakly-supervised semantic segmentation under image tags supervision is a challenging task as it directly associates high-level semantic to low-level appearance. To bridge this gap, in this paper, we propose an iterative bottom-up and top-down framework which alternatively expands object regions and optimizes segmentation network. We start from initial localization produced by classification networks. While classification networks are only responsive to small and coarse discriminative object regions, we argue that, these regions contain significant common features about objects. So in the bottom-up step, we mine common object features from the initial localization and expand object regions with the mined features. To supplement non-discriminative regions, saliency maps are then considered under Bayesian framework to refine the object regions. Then in the top-down step, the refined object regions are used as supervision to train the segmentation network and to predict object masks. These object masks provide more accurate localization and contain more regions of object. Further, we take these object masks as initial localization and mine common object features from them. These processes are conducted iteratively to progressively produce fine object masks and optimize segmentation networks. Experimental results on Pascal VOC 2012 dataset demonstrate that the proposed method outperforms previous state-of-the-art methods by a large margin.
Xiang Wang 0003, Shaodi You, Xi Li 0010, Huimin Ma 0001
CVPR4
2018 Key Parts Context and Scene Geometry in Human Head Detection
abstract
Due to the relatively fixed shape and color, head detection is widely used in many computer vision tasks, such as finding people and crowd counting. In this paper, we propose a method to detect human heads including: (1) A pairwise CNN model based on key parts context of human head and shoulder. (2) A fusion algorithm by using the priority of scene geometry structure. (3) To further test our approach, we collect a dataset about the passengers inside the bus with head annotations. This dataset composed of 2316 representative images extracted from a total of twenty hours of video. We evaluate our method on two indoor human head datasets and achieve state-of-the-art performance.
Chao Le, Huimin Ma 0001, Xiang Wang 0003, Xi Li 0010
ICIP2
2018 Region Proposal Ranking via Fusion Feature for Object Detection
abstract
Region proposals followed by a classification network is a famous structure in object detection task. Recently, there are a lot of proposal methods that work well via enough samplings. However, how to use as few samples as possible to achieve high recall performance is still a difficult problem. In this paper, we propose a novel region proposal method via fusion features for general object detection. By establishing an encoding of object attributes, an objectness candidate inference method based on energy function is proposed for region proposal ranking, followed by a multi-level non maximum suppression strategy for further adjusting. Experiment on PASCAL VOC shows that our method achieves better performance on region proposal and object detection tasks.
Xi Li 0010, Huimin Ma 0001, Xiang Wang 0003
ICIP2
2018 Driving Maneuvers Prediction Based on Cognition-driven and Data-driven Method
abstract
Advanced Driver Assistance Systems (ADAS) improve driving safety significantly. They alert drivers from unsafe traffic conditions when a dangerous maneuver appears. Traditional methods to predict driving maneuvers are mostly based on data-driven models alone. However, existing methods to understand the driver's intention remain an ongoing challenge due to a lack of intersection of human cognition and data analysis. To overcome this challenge, we propose a novel method that combines both the cognition-driven model and the data-driven model. We introduce a model named Cognitive Fusion-RNN (CF-RNN) which fuses the data inside the vehicle and the data outside the vehicle in a cognitive way. The CF-RNN model consists of two Long Short-Term Memory (LSTM) branches regulated by human reaction time. Experiments on the Brain4Cars benchmark dataset demonstrate that the proposed method outperforms previous methods and achieves state-of-the-art performance.
Huimin Ma 0001, Yuhan Dong
VCIP2
2018 3D Object Proposals Using Stereo Imagery for Accurate Object Class Detection
abstract
The goal of this paper is to perform 3D object detection in the context of autonomous driving. Our method aims at generating a set of high-quality 3D object proposals by exploiting stereo imagery. We formulate the problem as minimizing an energy function that encodes object size priors, placement of objects on the ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. We then exploit a CNN on top of these proposals to perform object detection. In particular, we employ a convolutional neural net (CNN) that exploits context and depth information to jointly regress to 3D bounding box coordinates and object pose. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. When combined with the CNN, our approach outperforms all existing results in object detection and orientation estimation tasks for all three KITTI object classes. Furthermore, we experiment also with the setting where LIDAR information is available, and show that using both LIDAR and stereo leads to the best result.
Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Feature proposal model on multidimensional data clustering and its application
Xi Li 0010, Huimin Ma 0001, Xiang Wang 0003
Pattern Recognit. Lett.2
2018 Saliency detection via alternative optimization adaptive influence matrix model
Xi Li 0010, Huimin Ma 0001, Xiang Wang 0003
Pattern Recognit. Lett.2
2018 Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object Detection
abstract
In this paper, we propose a novel edge preserving and multi-scale contextual neural network for salient object detection. The proposed framework is aiming to address two limits of the existing CNN based methods. First, region-based CNN methods lack sufficient context to accurately locate salient object since they deal with each region independently. Second, pixel-based CNN methods suffer from blurry boundaries due to the presence of convolutional and pooling layers. Motivated by these, we first propose an end-to-end edge-preserved neural network based on Fast R-CNN framework (named RegionNet) to efficiently generate saliency map with sharp object boundaries. Later, to further improve it, multi-scale spatial context is attached to RegionNet to consider the relationship between regions and the global scenes. Furthermore, our method can be generally applied to RGB-D saliency detection by depth refinement. The proposed framework achieves both clear detection boundary and multi-scale contextual robustness simultaneously for the first time, and thus achieves an optimized performance. Experiments on six RGB and two RGB-D benchmark datasets demonstrate that the proposed method achieves state-of-the-art performance.
Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen, Shaodi You
IEEE Trans. Image Process.2
2018 Traffic Light Recognition for Complex Scene With Fusion Detections
abstract
Traffic light recognition is one of the important tasks in the studies of intelligent transport system. In this paper, a robust traffic light recognition model based on vision information is introduced for on-vehicle camera applications. Our contribution mainly includes three aspects. First, in order to reduce computational redundancy, the aspect ratio, area, location, and context of traffic lights are utilized as prior information, which establishes a task model for traffic light recognition. Second, in order to improve the accuracy, we propose a series of improved methods based on an aggregate channel feature method, including modifying the channel feature for each types of traffic light and establishing a structure of fusion detectors. Third, we introduce a method of inter-frame information analysis, utilizing detection information of previous frame to modify original proposal regions, which makes the accuracy further improved. In the comparison of other traffic light detection algorithms, our model achieves competitive results on the complex scene VIVA data set. Furthermore, an analysis of small target luminous object detection tasks is given.
Xi Li 0010, Huimin Ma 0001, Xiang Wang 0003
IEEE Trans. Intell. Transp. Syst.2
2017 Multi-view 3D Object Detection Network for Autonomous Driving
abstract
This paper aims at high-accuracy 3D object detection in autonomous driving scenario. We propose Multi-View 3D networks (MV3D), a sensory-fusion framework that takes both LIDAR point cloud and RGB images as input and predicts oriented 3D bounding boxes. We encode the sparse 3D point cloud with a compact multi-view representation. The network is composed of two subnetworks: one for 3D object proposal generation and another for multi-view feature fusion. The proposal network generates 3D candidate boxes efficiently from the birds eye view representation of 3D point cloud. We design a deep fusion scheme to combine region-wise features from multiple views and enable interactions between intermediate layers of different paths. Experiments on the challenging KITTI benchmark show that our approach outperforms the state-of-the-art by around 25% and 30% AP on the tasks of 3D localization and 3D detection. In addition, for 2D detection, our approach obtains 14.9% higher AP than the state-of-the-art on the hard data among the LIDAR-based methods.
Xiaozhi Chen, Huimin Ma 0001, Ji Wan, Bo Li 0018
CVPR2
2017 Single Image Action Recognition Using Semantic Body Part Actions
abstract
In this paper, we propose a novel single image action recognition algorithm based on the idea of semantic part actions. Unlike existing part-based methods, we argue that there exists a mid-level semantic, the semantic part action; and human action is a combination of semantic part actions and context cues. In detail, we divide human body into seven parts: head, torso, arms, hands and lower body. For each of them, we define a few semantic part actions (e.g. head: laughing). Finally, we exploit these part actions to infer the entire body action (e.g. applauding). To make the proposed idea practical, we propose a deep network-based framework which consists of two subnetworks, one for part localization and the other for action prediction. The action prediction network jointly learns part-level and body-level action semantics and combines them for the final decision. Extensive experiments demonstrate our proposal on semantic part actions as elements for entire body action. Our method reaches mAP of 93.9% and 91.2% on PASCAL VOC 2012 and Stanford-40, which outperforms the state-of-the-art by 2.3% and 8.6%.
Zhichen Zhao, Huimin Ma 0001, Shaodi You
ICCV2
2017 Survival-Oriented Reinforcement Learning Model: An Effcient and Robust Deep Reinforcement Learning Algorithm for Autonomous Driving Problem
Changkun Ye, Huimin Ma 0001, Shaodi You
ICIG (2)2
2017 Boundary-aware box refinement for object proposal generation
Xiaozhi Chen, Huimin Ma 0001, Chenzhuo Zhu, Xiang Wang 0003, Zhichen Zhao
Neurocomputing2
2017 Generalized symmetric pair model for action classification in still images
Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen
Pattern Recognit.2
2016 Monocular 3D Object Detection for Autonomous Driving
abstract
The goal of this paper is to perform 3D object detection from a single monocular image in the domain of autonomous driving. Our method first aims to generate a set of candidate class-specific object proposals, which are then run through a standard CNN pipeline to obtain high-quality object detections. The focus of this paper is on proposal generation. In particular, we propose an energy minimization approach that places object candidates in 3D using the fact that objects should be on the ground-plane. We then score each candidate box projected to the image plane via several intuitive potentials encoding semantic segmentation, contextual information, size and location priors and typical object shape. Our experimental evaluation demonstrates that our object proposal generation approach significantly outperforms all monocular approaches, and achieves the best detection performance on the challenging KITTI benchmark, among published monocular competitors.
Xiaozhi Chen, Kaustav Kundu, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun
CVPR4
2016 Salient object detection via fast R-CNN and low-level cues
abstract
Recent advances in salient object detection have exploited the deep Convolutional Neural Network (CNN) to represent high-level semantic, however, due to the presence of convolutional and pooling layers, it is difficult for CNN to generate saliency map with sharp boundaries. In this paper, we propose multi-scale mask-based Fast R-CNN framework which generate saliency score of each region. Since the regions are segmented using edge-preserved methods, the results are naturally with sharp boundaries. To consider context information, we also propose low-level contrast and backgroundness prior which are complementary with high-level semantic. Finally, an edge-based propagation method which takes advantages of edge information is proposed to refine the saliency map. Experiments on three benchmark datasets demonstrate that the proposed method outperforms previous methods and achieves state-of-the-art performance.
Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen
ICIP2
2016 Multi-scale region candidate combination for action recognition
abstract
In still images, multi-scale regions contain rich information of different granularity. However, only semantically meaningful regions provide auxiliary cues for action recognition. Moreover, regions at different scales contribute differently. Motivated by the two observations, we propose an approach that is composed of three components: 1) detecting semantic region candidates at multiple scales, 2) training networks at each scale, 3) extracting features and learning to fuse them. The proposed approach captures multi-scale cues and highlights the optimal scale for each action, Experimental results show that our approach reaches the state-of-the-art performance on two challenging benchmarks: 1) PASCAL VOC 2012 and 2) Stanford-40.
Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen
ICIP2
2016 Geodesic weighted Bayesian model for saliency optimization
Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen
Pattern Recognit. Lett.2
2016 Semantic parts based top-down pyramid for action recognition
Zhichen Zhao, Huimin Ma 0001, Xiaozhi Chen
Pattern Recognit. Lett.2
2015 Improving object proposals with multi-thresholding straddling expansion
abstract
Recent advances in object detection have exploited object proposals to speed up object searching. However, many of existing object proposal generators have strong localization bias or require computationally expensive diversification strategies. In this paper, we present an effective approach to address these issues. We first propose a simple and useful localization bias measure, called superpixel tightness. Based on the characteristics of superpixel tightness distribution, we propose an effective method, namely multi-thresholding straddling expansion (MTSE) to reduce localization bias via fast diversification. Our method is essentially a box refinement process, which is intuitive and beneficial, but seldom exploited before. The greatest benefit of our method is that it can be integrated into any existing model to achieve consistently high recall across various intersection over union thresholds. Experiments on PASCAL VOC dataset demonstrates that our approach improves numerous existing models significantly with little computational overhead.
Xiaozhi Chen, Huimin Ma 0001, Xiang Wang 0003, Zhichen Zhao
CVPR2
2015 Geodesic weighted Bayesian model for salient object detection
abstract
In recent years, a variety of salient object detection methods under Bayesian framework have been proposed and many achieved state of the art. However, those ignore spatial relationships and thus background regions similar to the objects are also highlighted. In this paper, we propose a novel geodesic weighted Bayesian model to address this issue. We consider spatial relationships by attaching more importance to regions which are more likely to be parts of a salient object, thus suppressing background regions. First, we learn a combined similarity via multiple features to measure similarity of adjacent regions. Then, we apply the combined similarity as edge weight to construct an undirected weighted graph and compute geodesic distance. Last, we utilize the geodesic distance to weight the observation likelihood to infer a more precise saliency map. Experiments on several benchmark datasets demonstrate the effectiveness of our model.
Xiang Wang 0003, Huimin Ma 0001, Xiaozhi Chen
ICIP2
2015 3D Object Proposals for Accurate Object Class Detection
abstract
The goal of this paper is to generate high-quality 3D object proposals in the context of autonomous driving. Our method exploits stereo imagery to place proposals in the form of 3D bounding boxes. We formulate the problem as minimizing an energy function encoding object size priors, ground plane as well as several depth informed features that reason about free space, point cloud densities and distance to the ground. Our experiments show significant performance gains over existing RGB and RGB-D object proposal methods on the challenging KITTI benchmark. Combined with convolutional neural net (CNN) scoring, our approach outperforms all existing results on all three KITTI object classes.
Xiaozhi Chen, Kaustav Kundu, Yukun Zhu, Andrew G. Berneshawi, Huimin Ma 0001, Sanja Fidler, Raquel Urtasun
NIPS5
2014 Learning a compact latent representation of the Bag-of-Parts model
abstract
The Bag-of-Parts (BoP) model, which employs distinctive parts to represent images, has shown superior performance in vision recognition tasks. Our work is motivated by the need of reducing redundancy in tens of thousands parts. We propose a novel method to learn a compact latent representation from redundant part responses. We address this problem by employing spectral clustering and a multi-column coding scheme. The BoP model is viewed as a multi-scale convolutional model and additional sparse autoencoders are used to infer the latent patterns embedded in high-dimensional part-based representations. Spatial and semantic information is preserved by sparse learning on multiple spatial regions individually. Experiments demonstrate that the learnt representation achieves competitive performance with state-of-the-art methods on PASCAL VOC 2007 dataset.
Xiaozhi Chen, Huimin Ma 0001
ICIP2
2011 Manifold topological multi-resolution analysis method
Shaodi You, Huimin Ma 0001
Pattern Recognit.2
2009 A Solution to Efficient Viewpoint Space Partition in 3D Object Recognition
abstract
Viewpoint Space Partition based on Aspect Graph is one of the core techniques of 3D object recognition. Projection images obtained from critical viewpoint following this approach can efficiently provide topological information of an object. Computational complexity has been a huge challenge for obtaining the representation viewpoints used in 3D recognition. In this paper, we discuss inefficiency of calculation due to redundant nonexistent visual events; propose a systematic criterion for edge selection involved in EEE events. Pruning algorithm based on concave-convex property is demonstrated. We further introduce intersect relation into our pruning algorithm. These two methods not only enable the calculation of EEE events, but also can be implemented before viewpoint calculation, hence realizes view-independent pruning algorithm. Finally, analysis on simple representative models supports the effectiveness of our methods. Further investigations on Princeton Models, including airplane, automobile, etc, show a two orders of magnitude reduction in the number of EEE events on average.
Huimin Ma 0001, Shaodi You, Ze Yuan
ICIG2