Lihua Zhang 0002

dblp:46/1089-2 · DBLP profile ↗
← Back
107ranked-venue papers
1as first author
105since 2021 · last 2026
0000-0003-0467-4347ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 53 · 53 since 2021Artificial intelligence and machine learning · 52 · 1 first-author · 51 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 11 since 2021Systems, architecture and hardware · 11 · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
abstract
Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks.
Haiwei Xue, Minghao Han, Mingcheng Li, Xiaolu Hou, Dingkang Yang, Lihua Zhang 0002, Xu Zheng 0002
AAAI7
2026 UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation
abstract
Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately, making it difficult to accurately handle occlusions and transparency. Moreover, the deformed 3DGS still suffers from visual artifacts due to the sensitivity to the topology quality of the proxy mesh. These issues pose serious obstacles to the joint use of 3DGS and meshes, making it difficult to adapt 3DGS to conventional mesh-oriented graphics pipelines. We propose UniMGS, the first unified framework for rasterizing mesh and 3DGS in a single-pass anti-aliased manner, with a novel binding strategy for 3DGS deformation based on proxy mesh. Our key insight is to blend the colors of both triangle and Gaussian fragments by anti-aliased α-blending in a single pass, achieving visually coherent results with precise handling of occlusion and transparency. To improve the visual appearance of the deformed 3DGS, our Gaussian-centric binding strategy employs a proxy mesh and spatially associates Gaussians with the mesh faces, significantly reducing rendering artifacts. With these two components, UniMGS enables the visualization and manipulation of 3D objects represented by mesh or 3DGS within a unified framework, opening up new possibilities in embodied AI, virtual reality, and gaming. We will release our source code to facilitate future research.
Zeyu Xiao 0001, Yimin Cong, Dongliang Kou, Zhenyi Wu, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
AAAI10
2026 GeoPro-Depth: Geometrically Consistent Prompting for Robust Metric Depth Completion
abstract
The emergence of vision foundation models has significantly advanced monocular depth estimation; however, these models inherently suffer from scale ambiguity, limiting their utility in downstream metric applications. To address this, we propose GeoPro-Depth, a novel framework that reformulates depth completion as a geometric prompting task. Instead of training reconstruction networks from scratch, our method leverages a lightweight PromptNet to encode sparse depth measurements, which are then integrated into a pre-trained foundation model via a zero-initialization strategy to safeguard semantic features. To mitigate the impact of sensor noise, an Instance-Aware Geometric Filtering (IAGF) module is introduced, utilizing instance constraints to robustly filter outliers. Furthermore, a geometric consistency training objective is formulated to align metric predictions with relative depth priors in unobserved regions, ensuring structural fidelity. Extensive experiments on the ScanNet++, TUM-RGBD, and Replica datasets demonstrate state-of-the-art performance. Downstream evaluation on a 3D Gaussian Splatting SLAM system further confirms the superior geometric consistency and reconstruction quality of our model.
Xun Fang, Zixuan Hua, Lihua Zhang 0002
ICMR3
2026 Integrating channel priors and shared pattern learning for enhanced motor imagery EEG classification
Yangjie Luo, Zihua Chen, Junkongshuai Wang, Zhongxue Gan 0001, Lihua Zhang 0002, Xiaoyang Kang 0001
Neurocomputing6
2026 Nested resolution mesh-graph CNN for automated extraction of liver surface anatomical landmarks
abstract
The anatomical landmarks on the liver (mesh) surface, including the falciform ligament and liver ridge, are composed of triangular meshes of varying shapes, sizes, and positions, making them highly complex. Extracting and segmenting these landmarks is critical for augmented reality-based intraoperative navigation and monitoring. The key to this task lies in comprehensively understanding the overall geometric shape and local topological information of the liver mesh. However, due to the liver's variations in shape and appearance, coupled with limited data, deep learning methods often struggle with automatic liver landmark segmentation. To address this, we propose a two-stage automatic framework combining mesh-CNN and graph-CNN. In the first stage, dynamic graph convolution (DGCNN) is employed on low-resolution meshes to achieve rapid global understanding, generating initial landmark proposals at two levels, "dilation" and "erosion", and mapping them onto the original high-resolution surface. Subsequently, a refinement network based on mesh convolution fuses these landmark proposals from edge features along the local topology of the high-resolution mesh surface, producing refined segmentation results. Additionally, we incorporate an anatomy-aware Dice loss to address resolution imbalance and better handle sparse anatomical regions. Extensive experiments on two liver datasets, both in-distribution and out-of-distribution, demonstrate that our method accurately processes liver meshes of different resolutions, outperforming state-of-the-art methods. The reconstructed liver mesh dataset and the source code are available at https://github.com/xukun-zhang/MeshGraphCNN.
Xukun Zhang, Jinghui Feng, Peng Liu 0074, Minghao Han, Yanlan Kang, Sharib Ali, Lihua Zhang 0002
Medical Image Anal.10
2026 Towards unified molecule-enhanced pathology image representation learning via integrating spatial transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Zizhi Chen, Haopeng Kuang, Lihua Zhang 0002
Pattern Recognit.7
2025 BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
abstract
With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods rely on pre-trained text-to-image diffusion models and monocular depth estimators. However, the generated scenes occupy large amounts of storage space and often lack effective regularisation methods, leading to geometric distortions. To this end, we propose BloomScene, a lightweight structured 3D Gaussian splatting for crossmodal scene generation, which creates diverse and high-quality 3D scenes from text or image inputs. Specifically, a crossmodal progressive scene generation framework is proposed to generate coherent scenes utilizing incremental point cloud reconstruction and 3D Gaussian splatting. Additionally, we propose a hierarchical depth prior-based regularization mechanism that utilizes multi-level constraints on depth accuracy and smoothness to enhance the realism and continuity of the generated scenes. Ultimately, we propose a structured context-guided compression mechanism that exploits structured hash grids to model the context of unorganized anchor attributes, which significantly eliminates structural redundancy and reduces storage overhead. Comprehensive experiments across multiple scenes demonstrate the significant potential and advantages of our framework compared with several baselines.
Xiaolu Hou, Mingcheng Li, Dingkang Yang, Jiawei Chen 0012, Ziyun Qian, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
AAAI10
2025 Debiased Multimodal Understanding for Human Language Sequences
abstract
Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module.
Zhi Xu 0010, Dingkang Yang, Mingcheng Li, Zhaoyu Chen 0001, Jiawei Chen 0012, Jinjie Wei, Lihua Zhang 0002
AAAI8
2025 Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators
abstract
Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the internal factual knowledge of target LLMs. In addition, hallucinations typically exhibit multifaceted patterns in downstream tasks, limiting the model's holistic performance across tasks. In this paper, we propose a Comparator-driven Decoding-Time (CDT) framework to alleviate the response hallucination. Firstly, we construct hallucinatory and truthful comparators with multi-task fine-tuning samples. In this case, we present an instruction prototype-guided mixture of experts strategy to enhance the ability of the corresponding comparators to capture different hallucination or truthfulness patterns in distinct task instructions. CDT constrains next-token predictions to factuality-robust distributions by contrasting the logit differences between the target LLMs and these comparators. Systematic experiments on multiple downstream tasks show that our framework can significantly improve the model performance and response factuality.
Dingkang Yang, Dongling Xiao, Jinjie Wei, Mingcheng Li, Zhaoyu Chen 0001, Ke Li 0015, Lihua Zhang 0002
AAAI7
2025 MMPF: Multi-Modal Perception Framework for Abnormal Medical Condition Detection
abstract
As the global population ages and the incidence of chronic diseases increases, the demand for early detection of abnormal medical conditions is increasing. Traditional health monitoring methods often require significant resources and specialized personnel, limiting their widespread use. Leveraging advancements in AI technologies, this study proposes a non-invasive method for detecting abnormal medical conditions from image data. A multimodal perception framework is introduced, integrating features from various modalities, including facial expressions and body postures, to enhance detection accuracy. The framework employs a Cascaded Squeeze-Excitation (CSE) module, consisting of Adaptive and Multi-modal Squeeze-Excitation components, to capture complex feature dependencies and improve cross-modal performance. Extensive experiments demonstrate the effectiveness of this approach, showing improved performance over existing methods. In addition, a new dataset that encompasses a wide range of medical conditions has been released, providing a valuable resource for future research in this domain.
Chuyi Zhong, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
AAAI4
2025 MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
abstract
Diffusion models have shown excellent performance in text-to-image generation. Nevertheless, existing methods often suffer from performance bottlenecks when handling complex prompts that involve multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Compositional Diffusion (MCCD) for text-to-image generation for complex scenes. Specifically, we design a multi-agent collaboration-based scene parsing module that generates an agent system comprising multiple agents with distinct tasks, utilizing MLLMs to extract various scene elements effectively. In addition, Hierarchical Compositional diffusion utilizes a Gaussian mask and filtering to refine bounding box regions and enhance objects through region enhancement, resulting in the accurate and high-fidelity generation of complex scenes. Comprehensive experiments demonstrate that our MCCD significantly improves the performance of the baseline models in a training-free manner, providing a substantial advantage in complex scene generation.
Mingcheng Li, Xiaolu Hou, Dingkang Yang, Ziyun Qian, Jiawei Chen 0012, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
CVPR10
2025 CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
abstract
Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise clinical comprehension and disease diagnosis inference. Furthermore, owing to the limited accessibility of medical data and the imbalanced distribution of diseases, the underrepresentation of rare diseases in training data makes large-scale medical visual language models prone to hallucinations, such as omissions or fabrications, severely undermining diagnostic performance and further intensifying the challenges for MRG in practice. In this study, to effectively mitigate hallucinations in medical report generation, we propose a chain-of-medical-thought approach (CoMT), which intends to imitate the cognitive process of human doctors by decomposing diagnostic procedures. The radiological features with different importance are structured into fine-grained medical thought chains to enhance the inferential ability during diagnosis, thereby alleviating hallucination problems and enhancing the diagnostic accuracy of MRG.
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Ke Li 0015, Lihua Zhang 0002
ICASSP8
2025 V-Fusion: 2D Detection-enhanced Multimodal 3D BEV Object Detection
abstract
Integrating information from multiple sensors enhances the performance of autonomous vehicle perception systems. However, current multimodal 3D object detection methods focus on unifying modalities into a bird’s-eye view (BEV) representation, which overlooks the inherent characteristics of camera perspective view (PV), where 2D detection performance significantly surpasses that of state-of-the-art 3D detectors. In this paper, we propose V-Fusion, a high-quality 2D detection-enhanced multimodal BEV object detection method. By leveraging the 2D priors of PV, we construct 3D query proposals that complement BEV 3D queries. To address the modal discrepancy in generating 3D queries from 2D priors, we propose a depth-robust 2D-to-3D query generation strategy. Additionally, we introduce a novel geometry-constrained self-attention mechanism to enhance the interaction of BEV 3D queries and employ an additional set of learnable 3D queries to account for potentially missed objects. Notably, V-Fusion achieves 74.1 NDS performance on the challenging nuScenes dataset, outperforming SparseFusion in 1.0 NDS and offering comparable inference speed.
Jingwei Bian, Wei Li 0055, Lihua Zhang 0002
ICASSP6
2025 MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion
abstract
Motion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine details in complex motions, leading to models that capture only coarse-grained style characteristics. To overcome this limitation, we introduce the Motion Adaptive Fusion Diffusion (MAFD) framework, which leverages adaptive signal fusion to highlight essential style-defining features while minimizing redundant information. Moreover, current diffusion-based denoisers often fail to effectively capture the temporal relationships in motion sequences, producing rigid and fragmented stylized motions. Drawing inspiration from the Mamba model, we propose the Style Mamba Denoiser (SMD), which adopts a selection mechanism to preserve long-range dependencies and maintain temporal coherence. Extensive experiments show that our approach outperforms state-of-the-art methods in both qualitative and quantitative evaluations, achieving more refined and coherent stylized motions.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Dongliang Kou, Lihua Zhang 0002
ICASSP5
2025 SAF: Local Shape-aware Face-based Garment Collision Handling via Neural SDFs
abstract
Learning-based garment prediction presents an appealing alternative to physics-based methods owing to its high efficiency. However, the predicted garments can exhibit noticeable penetrations into the body. Many collision-handling methods operate in the point domain, which has an inherent limitation in addressing the penetration issues with the mesh faces. In light of this, we propose a local Shape-Aware, Face-based collision handling approach (SAF) that can be applied to garment prediction networks to achieve real-time collision handling. Considering the nature of garments, we design a compact formulation to model the continuity of the garment surface, and utilize neural Signed Distance Fields (SDFs) to accomplish penetration resolution. Recognizing the significant impact of the body’s local shape on collision handling, we further propose using the angles between SDF gradients to characterize the sharpness of the body. Our approach can compensate for the inaccuracy of neural SDFs and preserve local smoothness and details. Extensive experiments demonstrate the outstanding performance and generalizability of our method.
Minzhe Tang, Ruisheng Yuan, Dongliang Kou, Lihua Zhang 0002
ICASSP5
2025 Think4CPP: Reinforcement Learning by Thinking with Latent World Model for Safe Coverage Path Planning
Zhentang Liao, Zhongxue Gan 0001, Lihua Zhang 0002, Zhiyan Dong
ICIC (20)4
2025 MAS-VINS: Motion-Aware Semantic Visual-Inertial SLAM for Dynamic Environments of Varying Motion Intensity
Qining Zhang, Zhiyan Dong, Lihua Zhang 0002
ICIC (14)4
2025 VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction
abstract
Multimodal learning combining pathology images and genomic sequences enhances cancer survival analysis but faces clinical implementation barriers due to limited access to genomic sequencing in under-resourced regions. To enable survival prediction using only whole-slide images (WSI), we propose the Visual-Genomic Answering-Guided Transformer (VGAT), a framework integrating Visual Question Answering (VQA) techniques for genomic modality reconstruction. By adapting VQA’s text feature extraction approach, we derive stable genomic representations that circumvent dimensionality challenges in raw genomic data. Simultaneously, a cluster-based visual prompt module selectively enhances discriminative WSI patches, addressing noise from unfiltered image regions. Evaluated across five TCGA datasets, VGAT outperforms existing WSI-only methods, demonstrating the viability of genomic-informed inference without sequencing. This approach bridges multimodal research and clinical feasibility in resource-constrained settings. The code link is https://github.com/CZZZZZZZZZZZZZZZZZ/VGAT.
Zizhi Chen, Minghao Han, Xukun Zhang, Shuwei Ma, Tao Liu 0050, Lihua Zhang 0002
ICME7
2025 QCG-SLAM: Quadtree-based Condensed Gaussian Splatting for Visual SLAM
abstract
Recent studies highlight the potential of 3D Gaussian Splatting in visual SLAM systems. However, current methods often initialize and densify Gaussians on a per-pixel basis, leading to redundancy and high storage demands. To address these limitations, we propose QCG-SLAM, an innovative Gaussian-based SLAM system introducing three key contributions: (1) a redesigned Gaussian initialization and densification process leveraging a quadtree segmentation scheme, effectively reducing the number of Gaussians required for scene representation; (2) a coarse-to-fine scene reconstruction paradigm that accelerates convergence and enhances rendering quality; and (3) a Confidence Sampling-based Keyframe Selection (CSKS) strategy that mitigates catastrophic forgetting and tracking cumulative error issues, demonstrating clear advantages in both accuracy and efficiency. Extensive experiments demonstrate that QCG-SLAM requires only 32.18% of SplaTAM’s memory on the Replica dataset and 25.70% on the TUM-RGBD dataset on average, while maintaining state-of-the-art (SOTA) performance in camera tracking and scene reconstruction.
Xun Fang, Zixuan Hua, Lihua Zhang 0002
ICME4
2025 A Time-Frequency Feature Fusion Approach to Silent Speech Signal Recognition
Yangjie Luo, Zihua Chen, Lihua Zhang 0002, Xiaoyang Kang 0001
ICONIP (1)4
2025 Multi-Objective Partial Computation Offloading for Edge Intelligence with Heterogeneous Components
abstract
Edge intelligence, the fusion of edge computing and artificial intelligence (AI), drives the advancement of intelligent Internet of Things (IoT). Since AI applications are often data-and computation-intensive, resource-scarce edge devices need to migrate data to resource-rich edge servers through computation offloading to meet requirements such as energy efficiency and low latency. Existing studies often focus on CPU-based edge systems and neglect the impacts of other components, such as memory, on offloading. From a parallel processing perspective, this article establishes a system model and a multi-objective optimization model for edge intelligence systems with heterogeneous components, including diverse processors, memory, network, and applications, to minimize system energy consumption, total execution time, and the workload ratio of edge servers. A multi-objective optimization algorithm integrating archive initialization, hybrid perturbation, clustering, and modified simulated annealing is proposed and validated through experiments using real-world software and hardware. Results demonstrate that the proposed algorithm significantly outperforms comparative algorithms in terms of inverted generational distance, pure diversity, and run-time while revealing the influence of application characteristics on offloading performance.
Baoyu Xu, Yancheng Ruan, Tianyu Qi, Guobing Zou, Xiaoyang Kang 0001, Lihua Zhang 0002
ICPADS6
2025 Music-Driven Legged Robots: Synchronized Walking to Rhythmic Beats
abstract
We address the challenge of effectively controlling the locomotion of legged robots by incorporating precise frequency and phase characteristics, which is often ignored in locomotion policies that do not account for the periodic nature of walking. We propose a hierarchical architecture that integrates a low-level phase tracker, oscillators, and a high-level phase modulator. This controller allows quadruped robots to walk in a natural manner that is synchronized with external musical rhythms. Our method generates diverse gaits across different frequencies and achieves real-time synchronization with music in the physical world. This research establishes a foundational framework for enabling real-time execution of accurate rhythmic motions in legged robots. The video and code are available at https://music-walker.github.io/.
Taixian Hou, Xiaoyi Wei, Zhiyan Dong, Jiafu Yi, Peng Zhai, Lihua Zhang 0002
ICRA7
2025 Continuous Control of Diverse Skills in Quadruped Robots Without Complete Expert Datasets
abstract
Learning diverse skills for quadruped robots presents significant challenges, such as mastering complex transitions between different skills and handling tasks of varying difficulty. Existing imitation learning methods, while successful, rely on expensive datasets to reproduce expert behaviors. Inspired by introspective learning, we propose Progressive Adversarial Self-Imitation Skill Transition (PASIST), a novel method that eliminates the need for complete expert datasets. PASIST autonomously explores and selects high-quality trajectories based on predefined target poses instead of demonstrations, leveraging the Generative Adversarial Self-Imitation Learning (GASIL) framework. To further enhance learning, We develop a skill selection module to mitigate mode collapse by balancing the weights of skills with varying levels of difficulty. Through these methods, PASIST is able to reproduce skills corresponding to the target pose while achieving smooth and natural transitions between them. Evaluations on both simulation platforms and the Solo 8 robot confirm the effectiveness of PASIST, offering an efficient alternative to expert-driven learning.
Jiaxin Tu, Xiaoyi Wei, Taixian Hou, Xiaofei Gao, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002
ICRA8
2025 $\mathrm{A}^{3} \mathrm{E}^{2}$ Net: Integrating Waypoints, Lanes, and Traces for Multi-Modal Trajectory Prediction
abstract
Due to increasing complexity of traffic conditions, accurately predicting future trajectories of traffic agents is essential for driving efficiency and collision avoidance, making it vital to the decision-making of autonomous vehicles. However, it presents a significant challenge in trajectory prediction owing to the complex road topology and uncertainty of drivers' intention. In this paper, we introduce$\mathrm{A}^{3} \mathrm{E}^{2} \text{Net}$, a multi-modal trajectory prediction network that integrates waypoints, lanes, and motion traces, which collectively constitute the driving environment. It employs self-attention mechanism for waypoint representation, a Graph Neural Network (GNN) for road topology, and a spatiotemporal Transformer to encode motion dynamics. Subsequently, a cross-attention module fuses context information across granularities and captures agent-environment interactions. Finally, inspired by the outstanding performance of the autoregressive model in large language models (LLMs), we propose a goal-driven autoregression-based trajectory generation module to introduce both temporal logic and latent intentions, and thus enable precise and diverse multi-modal trajectory prediction (MTP). Evaluations on the Argoverse motion forecasting benchmark highlight our method's state-of-the-art performance and computational efficiency.
Xiaoyi Wei, Peng Zhai, Lihua Zhang 0002
ICTAI5
2025 Robust Signed Distance Fields for Articulated Human Body Reconstruction via Multiresolution Hash Encoding
abstract
The Signed Distance Fields (SDF) of the human body has broad applications in shape representation, collision handling, and medical image analysis, etc. However, due to the inherently high complexity of human motion, computing the SDFs of dynamic human bodies both accurately and efficiently has long been a challenging problem in computer graphics. In this paper, we demonstrate that by decomposing a widely used explicit human body model (SMPL) and modeling each component in a targeted manner, we can simultaneously realize both efficiency and accuracy. From a high level, the pipeline of the explicit model can be divided into Linear Blend Skinning (LBS) and Pose Space Deformation (PSD). By partitioning the human body into multiple parts and using the transformation matrix of each part, we apply inverse transformations to map spatial points from the posed space back to the canonical space. This eliminates the need for learning transformations and significantly reduces the difficulty of learning PSD. We observe that PSD is essentially a weighted sum of a series of fixed corrective shapes, where the only variable is the coefficient. We propose using Multiresolution Hash Encoding (MHE) to accurately capture the influence of each corrective shape on the SDF and aggregate the features in a manner similar to the explicit model. Our experiments show that our method is robust, effective, and highly efficient.
Minzhe Tang, Dongliang Kou, Mingcheng Li, Lihua Zhang 0002
IJCNN4
2025 LSTPointGMN: Lightweight Spatial-Temporal Graph Mamba Network for Gesture Recognition Using Millimeter-Wave Radar
abstract
Gesture recognition enhances human-computer interaction by providing intuitive and natural ways of interaction, with widespread applications in virtual reality, smart homes, and other domains. However, traditional camera-based methods are sensitive to lighting variations and carry the risk of privacy breaches. To address these limitations, millimeter-wave radar has become a viable alternative. The current advanced methods employ attention mechanisms, which have excellent global modeling capabilities, but have quadratic complexity. In this paper, we introduce LSTPointGMN, a lightweight spatial-temporal graph Mamba network for millimeter-wave radar gesture recognition. We design a spatial-temporal Mamba block based on GNN, which has strong global modeling capabilities and linear complexity, in order to more efficiently extract the spatial and temporal dimension features of the point cloud produced by the millimeter-wave radar. Comprehensive evaluations demonstrate that LSTPointGMN achieves superior performance on two public datasets, while significantly reducing GPU memory usage and computational cost, and improving inference speed.
Lihua Zhang 0002
IJCNN4
2025 Robust Reinforcement Learning based on Momentum Adversarial Training
abstract
Reinforcement learning (RL) is a fundamental and pivotal algorithm in the advancement of autonomous intelligence, including Embodied Intelligence and Physical Intelligence. The performance of RL directly influences the quality and efficiency of a robot’s decision-making and execution during interactions with its environment. Moreover, the robustness of RL remains a critical challenge that needs to be addressed. A promising approach to enhancing robustness is adversarial reinforcement learning. However, the existing methods primarily focus on perturbations in the state space, while perturbations in the action space have been relatively underexplored. The action space in RL is as crucial as the state space in autonomous intelligence. Furthermore, action-space perturbations provide a more comprehensive evaluation of RL robustness. Therefore, it is necessary and valuable to investigate RL robustness under action-space perturbations for the development of autonomous intelligence. To this end, we propose an adversarial learning framework that employs momentum-based gradient descent to model perturbations in the action space, such as actuator disturbances. Furthermore, we introduce an improved optimization method that integrates historical gradient information into conventional Stochastic Gradient Descent (SGD). This approach enhances training stability and improves perturbation efficiency. The proposed method is evaluated through simulations in the MuJoCo environment and UAV control experiments in GymFC, demonstrating significant improvements in robustness and adaptability under action-space perturbations. Additionally, real-world UAV flight tests are conducted to further validate the effectiveness of the proposed framework. The results confirm that the Sim-to-Real transfer is successful, providing empirical evidence for the applicability of our method in real-world scenarios. This study shows that enhancing RL robustness through action-space perturbations is feasible and effective. More importantly, our findings contribute to the future development of autonomous intelligence, particularly in improving its resilience to uncertainties and dynamic environments.
Hanchen Liu, Junru Sheng, Lihua Zhang 0002, Zhiyan Dong
IROS4
2025 PIG: Physically-based Multi-Material Interaction with 3D Gaussians
abstract
3D Gaussian Splatting has achieved remarkable success in reconstructing both static and dynamic 3D scenes. However, in a scene represented by 3D Gaussian primitives, interactions between objects suffer from inaccurate 3D segmentation, imprecise deformation among different materials, and severe rendering artifacts. To address these challenges, we introduce PIG: Physically-Based Multi-Material Interaction with 3D Gaussians, a novel approach that combines 3D object segmentation with the simulation of interacting objects in high precision. Firstly, our method facilitates fast and accurate mapping from 2D pixels to 3D Gaussians, enabling precise 3D object-level segmentation. Secondly, we assign unique physical properties to correspondingly segmented objects within the scene for multi-material coupled interactions. Finally, we have successfully embedded constraint scales into deformation gradients, specifically clamping the scaling and rotation properties of the Gaussian primitives to eliminate artifacts and achieve geometric fidelity and visual consistency. Experimental results demonstrate that our method not only outperforms the state-of-the-art (SOTA) in terms of visual quality, but also opens up new directions and pipelines for the field of physically realistic scene generation.
Zeyu Xiao 0001, Zhenyi Wu, Qipeng Yan, Zhuoer Liang, Lihua Zhang 0002
ICMR7
2025 VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
abstract
In histopathology, tissue sections are typically stained using common H&E staining or special stains (MAS, PAS, PASM, etc. ) to clearly visualize specific tissue structures. The rapid advancement of deep learning offers an effective solution for generating virtually stained images, significantly reducing the time and labor costs associated with traditional histochemical staining. However, a new challenge arises in separating the fundamental visual characteristics of tissue sections from the visual differences induced by staining agents. Additionally, virtual staining often overlooks essential pathological knowledge and the physical properties of staining, resulting in only style-level transfer. To address these issues, we introduce, for the first time in virtual staining tasks, a pathological vision-language large model (VLM) as an auxiliary tool. We integrate contrastive learnable prompts, foundational concept anchors for tissue sections, and staining-specific concept anchors to leverage the extensive knowledge of the pathological VLM. This approach is designed to describe, frame, and enhance the direction of virtual staining. Furthermore, we have developed a data augmentation method based on the constraints of the VLM. This method utilizes the VLM's powerful image interpretation capabilities to further integrate image style and structural information, proving beneficial in high-precision pathological diagnostics. Extensive evaluations on publicly available multi-domain unpaired staining datasets demonstrate that our method can generate highly realistic images and enhance the accuracy of downstream tasks, such as glomerular detection and segmentation. Our code. https://github.com/CZZZZZZZZZZZZZZZZZ/VPGAN-HARBOR is available.
Zizhi Chen, Minghao Han, Yizhou Liu 0002, Ziyun Qian, Xukun Zhang, Jingwei Wei, Lihua Zhang 0002
ACM Multimedia9
2025 UMSD: High Realism Motion Style Transfer via Unified Mamba-based Diffusion
abstract
Motion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer.
Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002
ACM Multimedia9
2025 Diverse object placement with dual interaction
Xianhe Cheng, Peng Zhai, Dingkang Yang, Lihua Zhang 0002
Neurocomputing6
2025 Multilevel and Energy-Efficient Partial Computation Offloading in Heterogeneous Edge Intelligence
abstract
Due to the diversity of edge devices (EDs) and applications, edge systems are heterogeneous and have been applied in artificial intelligence fields, such as smart factories and intelligent transportation, which is called heterogeneous edge intelligence. Many studies employ computation offloading to transfer processing data from resource-scarce EDs to resource-rich edge servers. These studies primarily focus on the overall resource consumption of homogeneous edge systems, neglecting the system heterogeneity and the details of resource consumption. In this article, we construct a system model from a parallel perspective for the heterogeneous edge system with different processors, memory, and applications, which perceives the cost of energy and delay from three levels: system, application, and component. A hybrid metaheuristic algorithm combined with a greedy rule, hybrid mutation, and whale optimization algorithm (GHMWOA) is proposed to realize partial computation offloading. A partial offloading architecture of heterogeneous edge intelligence is proposed to validate our model and algorithm with real-world hardware and software. Experiment results not only show GHMWOA outperforms multiple classical optimization algorithms in minimizing energy consumption, but also discover on which system component energy consumption depends, and how properties of application and system influence the cost of energy.
Baoyu Xu, Yancheng Ruan, Chenghu Qiu, Shuibing He, Xiaoyang Kang 0001, Lihua Zhang 0002
IEEE Internet Things J.7
2025 Role play: Learning adaptive role-specific strategies in multi-agent interactions
Weifan Long, Wen Wen 0005, Peng Zhai, Lihua Zhang 0002
Knowl. Based Syst.4
2025 An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challenge
abstract
Augmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain.
Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli
Medical Image Anal.7
2025 Enhanced multi-modal abdominal image registration via structural awareness and region-specific optimization
Xukun Zhang, Lihua Zhang 0002
Pattern Recognit. Lett.4
2025 MSTDF: Motion Style Transfer Towards High Visual Fidelity Based on Dynamic Fusion
abstract
Emotion-guided motion style transfer is a novel research direction, enabling the efficient generation of motion in various emotional styles for use in films, games, and other domains. However, existing methods primarily rely on global feature statistics for motion style transfer, neglecting local semantic structure and resulting in the degradation of motion content structure. This letter proposes a novel Motion Style Transfer based on Dynamic Fusion (MSTDF) framework, which treats content and style motion as distinct signals and employs dynamic fusion for high-fidelity motion style transfer. Additionally, to address the challenge of traditional discriminators capturing subtle motion style features, we propose the Motion Dynamic Fusion (MDF) discriminator to capture the details and fine-grained style characteristics of motion sequences, assisting the generator in producing higher-fidelity stylized motion. Finally, extensive experiments on the Xia dataset demonstrate that our method surpasses state-of-the-art methods in qualitative and quantitative comparisons.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Zeyu Xiao 0001, Lihua Zhang 0002
IEEE Signal Process. Lett.5
2025 Multidimensional Cockpit Perception via Mutual Information-Guided Signal Graph Fusion
abstract
Multidimensional Cockpit Perception (MCP) is an indispensable technology in assisted driving systems for modern vehicles, which requires joint recognition of the driver's emotional behaviors, traffic context, and vehicle condition. Despite promising progress in 2D to 3D neural networks, semantic discrepancies and intricate temporal dependencies among different signal streams still cause serious performance bottlenecks. To address these challenges, this paper proposes a Mutual Information-guided signal Graph fusion (MIG) framework for robust cockpit perception. MIG introduces an adaptive mutual information constraint to maximize the mutual information among different perceived signal streams, thus reinforcing the task-related feature semantic consistency. In addition, a signal graph fusion module is designed to capture intra- and inter-signal elemental correlations at a fine-grained level to learn joint multi-signal representations used for different downstream tasks. Systematic experiments on the MCP benchmark verify the effectiveness of the proposed framework and the necessity of our components.
Zhi Xu 0010, Xiaoyang Kang 0001, Lihua Zhang 0002
IEEE Signal Process. Lett.3
2025 Robust Multi-Agent Collaborative Perception via Spatio-Temporal Awareness
abstract
As an emerging application in autonomous driving, multi-agent collaborative perception has recently received significant attention. Despite promising advances from previous efforts, several unavoidable challenges that cause performance bottlenecks remain, including the single-frame detection dilemma, communication redundancy, and defective collaboration process. To this end, we proposeSCOPE++, a versatile collaborative perception framework aggregating spatio-temporal information across on-road agents to tackle these issues. We introduce four components inSCOPE++for robust collaboration by seeking a reasonable trade-off between perception performance and communication bandwidth. First, we devise a context-aware information aggregation to capture valuable semantic cues in the temporal context and enhance the current local representation of the ego agent. Second, an exclusivity-aware sparse communication is introduced to filter perceptually unnecessary information from collaborators and transmit complementary features relative to the ego agent. Third, we present an importance-aware cross-agent collaboration to incorporate semantic representations of spatially critical locations across agents flexibly. Finally, a contribution-aware adaptive fusion is designed to integrate multi-source representations based on dynamic contributions. Our framework is evaluated on multiple LiDAR-based collaborative detection datasets in real-world and simulated scenarios, and comprehensive experiments show thatSCOPE++outperforms state-of-the-art methods on all datasets.
Kun Yang 0010, Zhi Xu 0010, Dingkang Yang, Lihua Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.7
2025 Denoising Transformer for BEV 3D Object Detection via Multiview Multiscale Cross-Attention
Xukun Zhang, Xiaoyi Wei, Zhi Xu 0010, Youxing Wang, Peng Zhai, Lihua Zhang 0002
IEEE Trans. Intell. Transp. Syst.10
2025 Safety measures in automotive operating system: a comprehensive review of trends and defense frameworks
abstract
With the rapid development of intelligent electric vehicles (IEVs), an increasing number of algorithms are being deployed on software platforms, progressively leading to the formation of automotive operating systems (OS). This paper introduces firstly the status of automotive OS and analyzes their development trends. As an emerging technology, automotive OS is a safety-critical system that plays a vital role in driving safety. To address the challenges associated with ensuring the safety of automotive OS, this paper proposes systematically safety measures across multiple perspectives, including software architecture, time protection, memory protection, software monitoring, and communication. Furthermore, a system safety mode is proposed from an engineering implementation perspective. Finally, experimental results validate the feasibility of deploying multiple communication protocols within automotive OS, effectively addressing the challenges posed by big data and high-concurrency demands in IEVs.
Jingwei Bian, Wei Li 0055, Lihua Zhang 0002
J. Supercomput.4
2025 MSCPT: Few-Shot Whole Slide Image Classification With Multi-Scale and Context-Focused Prompt Tuning
abstract
Multiple instance learning (MIL) has become a standard paradigm for the weakly supervised classification of whole slide images (WSIs). However, this paradigm relies on using a large number of labeled WSIs for training. The lack of training data and the presence of rare diseases pose significant challenges for these methods. Prompt tuning combined with pre-trained Vision-Language models (VLMs) is an effective solution to the Few-shot Weakly Supervised WSI Classification (FSWC) task. Nevertheless, applying prompt tuning methods designed for natural images to WSIs presents three significant challenges: 1) These methods fail to fully leverage the prior knowledge from the VLM's text modality; 2) They overlook the essential multi-scale and contextual information in WSIs, leading to suboptimal results; and 3) They lack exploration of instance aggregation methods. To address these problems, we propose a Multi-Scale and Context-focused Prompt Tuning (MSCPT) method for FSWC task. Specifically, MSCPT employs the frozen large language model to generate pathological visual language prior knowledge at multiple scales, guiding hierarchical prompt tuning. Additionally, we design a graph prompt tuning module to learn essential contextual information within WSI, and finally, a non-parametric cross-guided instance aggregation module has been introduced to derive the WSI-level features. Extensive experiments, visualizations, and interpretability analyses were conducted on five datasets and three downstream tasks using three VLMs, demonstrating the strong performance of our MSCPT. All codes have been made publicly accessible at https://github.com/Hanminghao/MSCPT.
Minghao Han, Linhao Qu, Dingkang Yang, Xukun Zhang, Lihua Zhang 0002
IEEE Trans. Medical Imaging6
2024 A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities
abstract
Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion approaches. To this end, we propose a Unified multimodal Missing modality self-Distillation Framework (UMDF) to handle the problem of uncertain missing modalities in MSA. Specifically, a unified self-distillation mechanism in UMDF drives a single network to automatically learn robust inherent representations from the consistent distribution of multimodal data. Moreover, we present a multi-grained crossmodal interaction module to deeply mine the complementary semantics among modalities through coarse- and fine-grained crossmodal attention. Eventually, a dynamic feature integration module is introduced to enhance the beneficial semantics in incomplete modalities while filtering the redundant information therein to obtain a refined and robust multimodal representation. Comprehensive experiments on three datasets demonstrate that our framework significantly improves MSA performance under both uncertain missing-modality and complete-modality testing conditions.
Mingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang 0001, Shuaibing Wang, Liuzhen Su, Kun Yang 0010, Lihua Zhang 0002
AAAI10
2024 SceneWeaver: Text-Driven Scene Generation with Geometry-aware Gaussian Splatting
Xiaolu Hou, Mingcheng Li, Jiawei Chen 0012, Dingkang Yang, Ziyun Qian, Lihua Zhang 0002
ACML6
2024 Large Vision-Language Models as Emotion Recognizers in Context Awareness
Yuxuan Lei, Dingkang Yang, Zhaoyu Chen 0001, Jiawei Chen 0012, Peng Zhai, Lihua Zhang 0002
ACML6
2024 CPR-Coach: Recognizing Composite Error Actions Based on Single-Class Training
abstract
Fine- grained medical action analysis plays a vital role in improving medical skill training efficiency, but it faces the problems of data and algorithm shortage. Cardiopul-monary Resuscitation (CPR) is an essential skill in emer-gency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this pa-per constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compres-sion and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper in-vestigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable “Single-class Training & Multi-class Testing” problem, we propose a human-cognition-inspired framework named ImagineNet to improve the model's multi-error recognition performance under restricted supervision. Extensive comparison and actual deployment experiments verify the effectiveness of the framework. We hope this work could bring new inspiration to the computer vision and medical skills training communities simultaneously. The dataset and the code are publicly available on https://github.com/Shunli-Wang/CPR-Coach.
Shunli Wang 0001, Shuaibing Wang, Dingkang Yang, Mingcheng Li, Haopeng Kuang, Liuzhen Su, Peng Zhai, Lihua Zhang 0002
CVPR9
2024 Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
abstract
Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines.
Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002
CVPR10
2024 De-Confounded Data-Free Knowledge Distillation for Handling Distribution Shifts
abstract
Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset.
Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi
CVPR7
2024 Robust Emotion Recognition in Context Debiasing
abstract
Context-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the target person's emotional state. Despite advancements, the biggest challenge remains due to context bias interference. The harmful bias forces the models to rely on spurious correlations between background contexts and emotion labels in likelihood estimation, causing severe performance bottlenecks and confounding valuable context priors. In this paper, we propose a counterfactual emotion inference (CLEF) framework to address the above issue. Specifically, we first formulate a generalized causal graph to decouple the causal relationships among the variables in CAER. Following the causal graph, CLEF introduces a non-invasive context branch to capture the adverse direct effect caused by the context bias. During the inference, we eliminate the direct context effect from the total causal effect by comparing factual and counterfactual outcomes, resulting in bias mitigation and robust prediction. As a model-agnostic framework, CLEF can be readily integrated into existing methods, bringing consistent performance gains.
Dingkang Yang, Kun Yang 0010, Mingcheng Li, Shunli Wang 0001, Shuaibing Wang, Lihua Zhang 0002
CVPR6
2024 Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002
ECCV (58)10
2024 MISS: A Generative Pre-training and Fine-Tuning Approach for Med-VQA
Jiawei Chen 0012, Dingkang Yang, Yuxuan Lei, Lihua Zhang 0002
ICANN (8)5
2024 PLOD-YOLO: Premium Lightweight Object Detection for Autonomous Following Robot
Wenqing Deng, WeiJie Zhang, Zhiyan Dong, Lihua Zhang 0002
ICIC (4)7
2024 Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide Images
abstract
Survival prediction is a complex ordinal regression task that aims to predict the survival coefficient ranking among a cohort of patients, typically achieved by analyzing patients’ whole slide images. Existing deep learning approaches mainly adopt multiple instance learning or graph neural networks under weak supervision. Most of them are unable to uncover the diverse interactions between different types of biological entities(e.g., cell cluster and tissue block) across multiple scales, while such interactions are crucial for patient survival prediction. In light of this, we propose a novel multi-scale heterogeneity-aware hypergraph representation framework. Specifically, our framework first constructs a multi-scale heterogeneity-aware hypergraph and assigns each node with its biological entity type. It then mines diverse interactions between nodes on the graph structure to obtain a global representation. Experimental results demonstrate that our method outperforms state-of-the-art approaches on three benchmark datasets. Code is publicly available at https://github.com/Hanminghao/H2GT.
Minghao Han, Xukun Zhang, Dingkang Yang, Tao Liu 0050, Haopeng Kuang, Jinghui Feng, Lihua Zhang 0002
ICME7
2024 IIPC: Intra-Inter Patch Correlations for Garment Collision Handling
abstract
Realistic garment simulation is critical for digital humans. However, noticeable penetrations still exist in current learning-based garment simulation techniques. To reduce penetrations in predicted garments, we resort to the garment geometry and neural Signed Distance Fields (SDFs) for effective collision handling. The key idea of our method is that we divide the garment into patches and model the local and global garment geometry through Intra- and Inter-Patch Correlations (IIPC), which can be easily learned through the powerful context-understanding ability of Transformers. The geometry information is then utilized to predict a per-vertex moving offset, according to which we move the penetrating vertices along the SDF’s gradient directions to solve collisions. Our module can be coupled with learning-based backbones to effectively solve penetrations while retaining real-time performance. Extensive experiments show that the proposed method excels the prior works significantly.
Ruisheng Yuan, Minzhe Tang, Dongliang Kou, Dingkang Yang, Lihua Zhang 0002
ICME7
2024 STGCN-DHD: Spatio-Temporal Graph Convolutional Network for EEG-Based Driving Hazard Detection
Jialong Liang, Weifan Long, Peng Zhai, Lihua Zhang 0002
ICONIP (4)5
2024 T-GET3D: A Generative Model of High-Quality 3D Textured Shapes Guided by Texts
Xinxin Shi, Xianhe Cheng, Peixuan Zhang, Dingkang Yang, Lihua Zhang 0002
ICONIP (2)6
2024 Multi-Task Learning of Active Fault-Tolerant Controller for Leg Failures in Quadruped robots
abstract
Electric quadruped robots used in outdoor exploration are susceptible to leg-related electrical or mechanical failures. Unexpected joint power loss and joint locking can immediately pose a falling threat. Typically, controllers lack the capability to actively sense the condition of their own joints and take proactive actions. Maintaining the original motion patterns could lead to disastrous consequences, as the controller may produce irrational output within a short period of time, further creating the risk of serious physical injuries. This paper presents a hierarchical fault-tolerant control scheme employing a multi-task training architecture capable of actively perceiving and overcoming two types of leg joint faults. The architecture simultaneously trains three joint task policies for health, power loss, and locking scenarios in parallel, introducing a symmetric reflection initialization technique to ensure rapid and stable gait skill transformations. Experiments demonstrate that the control scheme is robust in unexpected scenarios where a single leg experiences concurrent joint faults in two joints. Furthermore, the policy retains the robot’s planar mobility, enabling rough velocity tracking. Finally, zero-shot Sim2Real transfer is achieved on the real-world SOLO8 robot, countering both electrical and mechanical failures.
Taixian Hou, Jiaxin Tu, Xiaofei Gao, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002
ICRA6
2024 MDRPC: Music-Driven Robot Primitives Choreography
abstract
Dance has been an important art form and means of communication since the dawn of human civilization. Equipping humanoid robots with the ability to perform smooth dance movements to music is a key research priority in artificial intelligence, robotics and human-computer interaction. However, existing kinematics-based dance generation methods often violate real-world physical laws as they do not consider physical constraints, leading to unrealistic movements. Additionally, due to the diversity and dynamic variability of input music, most existing physics-based methods, which rely on task-specific reward functions, face significant challenges in effectively handling music-driven dance generation tasks. To address these issues, we introduce MDRPC, the first physics-based, music-driven dance generation method for humanoid robots. Inspired by human choreographic principles, MDRPC is defined as a two-phase framework. The initial phase utilizes adversarial imitation learning to acquire a rich set of reusable dance primitives from a music-dance dataset. In the subsequent phase, these dance primitives are orchestrated under the guidance of musical theory and choreographic rules to generate complex humanoid dance sequences. Specifically, we propose beat alignment and dance diversity reward functions to synchronize motion rhythms with music beats and enhance the diversity of dance movements. We implement MDRPC on a simulated humanoid robot, and the results confirm that our method effectively controls the humanoid, enabling it to perform dance movements harmoniously synchronized with the music.
Haiyang Guan, Xiaoyi Wei, Weifan Long, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
ICTAI6
2024 Bio-Inspired Feature Selection via an Improved Binary Golden Jackal Optimization Algorithm
Jinghui Feng, Xukun Zhang, Lihua Zhang 0002
KSEM (2)3
2024 Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, Lihua Zhang 0002
MICCAI (5)7
2024 Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Models
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Xiaolu Hou, Lihua Zhang 0002
ACM Multimedia7
2024 IF-Garments: Reconstructing Your Intersection-Free Multi-Layered Garments from Monocular Videos
abstract
Reconstructing garments from monocular videos has attracted considerable attention as it provides a convenient and low-cost solution for clothing digitization. In reality, people wear clothing with countless variations and multiple layers. Existing studies attempt to extract garments from a single video. They either behave poorly in generalization due to reliance on limited clothing templates or struggle to handle the intersections of multi-layered clothing leading to the lack of physical plausibility. Besides, there are inevitable and undetectable overlaps for a single video that hinder researchers from modeling complete and intersection-free multi-layered clothing. To address the above limitations, in this paper, we propose a novel method to reconstruct multi-layered clothing from multiple monocular videos sequentially, which surpasses existing work in generalization and robustness against penetration. For each video, neural fields are employed to implicitly represent the clothed body, from which the meshes with frame-consistent structures are explicitly extracted. Next, we implement a template-free method for extracting a single garment by back-projecting the image segmentation labels of different frames onto these meshes. In this way, multiple garments can be obtained from these monocular videos and then aligned to form the whole outfit. However, intersection always occurs due to overlapping deformation in the real world and perceptual errors in monocular videos. To this end, we innovatively introduce a physics-aware module that combines neural fields with a position-based simulation framework to fine-tune the penetrating vertices of garments, ensuring robustly intersection-free. Additionally, we collect a mini dataset with fashionable garments to evaluate the quality of clothing reconstruction comprehensively. We release our code and data at https://github.com/SMY19999/IF-Garments.
Qipeng Yan, Zhuoer Liang, Dongliang Kou, Dingkang Yang, Ruisheng Yuan, Mingcheng Li, Lihua Zhang 0002
ACM Multimedia9
2024 MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
abstract
Accurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms designed for specific perception tasks, leading to a lack of complementary learning among tasks and decreased performance in multi-task learning (MTL) due to joint training. In this paper, we propose MaskBEV, a masked attention-based MTL paradigm that unifies 3D object detection and bird's eye view (BEV) map segmentation. MaskBEV introduces a task-agnostic Transformer decoder to process these diverse tasks, enabling MTL to be completed in a unified decoder without requiring additional design of specific task heads. To fully exploit the complementary information between BEV map segmentation and 3D object detection tasks in BEV space, we propose spatial modulation and scene-level context aggregation strategies. These strategies consider the inherent dependencies between BEV segmentation and 3D detection, naturally boosting MTL performance. Extensive experiments on nuScenes dataset show that compared with previous state-of-the-art MTL methods, MaskBEV achieves 1.3 NDS improvement in 3D object detection and 2.7 mIoU improvement in BEV map segmentation, while also demonstrating slightly leading inference speed.
Xukun Zhang, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Lihua Zhang 0002
ACM Multimedia7
2024 Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
abstract
Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases.
Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002
NeurIPS14
2024 PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
abstract
Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatric applications due to inadequate instruction data and vulnerable training procedures. To address the above issues, this paper builds PedCorpus, a high-quality dataset of over 300,000 multi-task instructions from pediatric textbooks, guidelines, and knowledge graph resources to fulfil diverse diagnostic demands. Upon well-designed PedCorpus, we propose PediatricsGPT, the first Chinese pediatric LLM assistant built on a systematic and robust training pipeline. In the continuous pre-training phase, we introduce a hybrid instruction pre-training mechanism to mitigate the internal-injected knowledge inconsistency of LLMs for medical domain adaptation. Immediately, the full-parameter Supervised Fine-Tuning (SFT) is utilized to incorporate the general medical knowledge schema into the models. After that, we devise a direct following preference optimization to enhance the generation of pediatrician-like humanistic responses. In the parameter-efficient secondary SFT phase, a mixture of universal-specific experts strategy is presented to resolve the competency conflict between medical generalist and pediatric expertise mastery. Extensive results based on the metrics, GPT-4, and doctor evaluations on distinct downstream tasks show that PediatricsGPT consistently outperforms previous Chinese medical LLMs. The project and data will be released at https://github.com/ydk122024/PediatricsGPT.
Dingkang Yang, Jinjie Wei, Dongling Xiao, Shunli Wang 0001, Mingcheng Li, Shuaibing Wang, Jiawei Chen 0012, Qingyao Xu, Ke Li 0015, Peng Zhai, Lihua Zhang 0002
NeurIPS14
2024 3DLaneFormer: End-to-End 3D Lane Detection with Voxel Descriptors
Qiangbin Xie, Xukun Zhang, Shunli Wang 0001, Lihua Zhang 0002
PRCV (4)6
2024 CASSTIMP: Cascaded Architecture for Symptom Status Tracking with Inquiry-Aware Attention and Multi-Perception Pooling
abstract
Symptom status tracking poses a significant challenge due to the intricate nature of symptom identification and inference from medical doctor-patient dialogues. Numerous prior studies in this domain have relied on approaches involving multi-label classification and multi-task learning. Multi-label classification methods typically consider symptoms and statuses within a unified label space. Nevertheless, this approach frequently results in sparse predictions, eroding semantic relationships among labels and causing instability in prediction outcomes. In contrast, multi-task learning segregates symptom prediction and status prediction into separate tasks, thereby improving performance relative to conventional multi-label classification methods. Nonetheless, despite these advancements, the imbalance in training task weights persists, leading to suboptimal performance. To tackle these challenges, we employ a cascaded model structure rooted in the Question-Answering (QA) paradigm in this study. Our approach utilizes dialogue content to create context-inquiry pairs and introduces two novel modules: inquiry-aware attention and multi-perception pooling. Inquiry-aware attention enhances the contextual relationship between inquiries and dialogues, while multi-perception pooling extracts diverse semantics from the dialogue. The experimental results unequivocally demonstrate our method's efficiency, surpassing state-of-the-art techniques in symptom status tracking and indicating its superior effectiveness.
Haowen Yu, Mingcheng Li, Lihua Zhang 0002
SMC3
2024 Towards heart infarction detection via image-based dataset and three-stream fusion framework
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
Comput. Commun.4
2024 Expression guided medical condition detection via the Multi-Medical Condition Image Dataset
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002
Eng. Appl. Artif. Intell.5
2024 Dual knowledge-guided two-stage model for precise small organ segmentation in abdominal CT images
abstract
Abstract Multi‐organ segmentation from abdominal CT scans is crucial for various medical examinations and diagnoses. Despite the remarkable achievements of existing deep‐learning‐based methods, accurately segmenting small organs remains challenging due to their small size and low contrast. This article introduces a novel knowledge‐guided cascaded framework that utilizes two types of knowledge—image intrinsic (anatomy) and clinical expertise (radiology)—to improve the segmentation accuracy of small abdominal organs. Specifically, based on the anatomical similarities in abdominal CT scans, the approach employs entropy‐based registration techniques to map high‐quality segmentation results onto inaccurate results from the first stage, thereby guiding precise localization of small organs. Additionally, inspired by the practice of annotating images from multiple perspectives by radiologists, novel Multi‐View Fusion Convolution (MVFC) operator is developed, which can extract and adaptively fuse features from various directions of CT images to refine segmentation of small organs effectively. Simultaneously, the MVFC operator offers a seamless alternative to conventional convolutions within diverse model architectures. Extensive experiments on the Abdominal Multi‐Organ Segmentation (AMOS) dataset demonstrate the superiority of the method, setting a new benchmark in the segmentation of small organs.
Tao Liu 0050, Xukun Zhang, Zhongwei Yang, Minghao Han, Haopeng Kuang, Shuwei Ma, Lihua Zhang 0002
IET Image Process.9
2024 Preference detection of the humanoid robot face based on EEG and eye movement
Pengchao Wang, Gege Zhan, Aiping Wang, Zuoting Song, Xueze Zhang, Junkongshuai Wang, Lan Niu, Jianxiong Bin, Lihua Zhang 0002, Jie Jia 0002, Xiaoyang Kang 0001
Neural Comput. Appl.11
2024 Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded Training
abstract
Understanding emotions from diverse contexts has received widespread attention in computer vision communities. The core philosophy of Context-Aware Emotion Recognition (CAER) is to provide valuable semantic cues for recognizing the emotions of target persons by leveraging rich contextual information. Current approaches invariably focus on designing sophisticated structures to extract perceptually critical representations from contexts. Nevertheless, a long-neglected dilemma is that a severe context bias in existing datasets results in an unbalanced distribution of emotional states among different contexts, causing biased visual representation learning. From a causal demystification perspective, the harmful bias is identified as a confounder that misleads existing models to learn spurious correlations based on likelihood estimation, limiting the models' performance. To address the issue, we embrace causal inference to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a customized causal graph. Subsequently, we present a Contextual Causal Intervention Module (CCIM) to de-confound the confounder, which is built upon backdoor adjustment theory to facilitate seeking approximate causal effects during model training. As a plug-and-play component, CCIM can easily integrate with existing approaches and bring significant improvements. Systematic experiments on three datasets demonstrate the effectiveness of our CCIM.
Dingkang Yang, Kun Yang 0010, Haopeng Kuang, Zhaoyu Chen 0001, Lihua Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 3D human pose estimation with single image and inertial measurement unit (IMU) sequence
Liujun Liu, Jiewen Yang, Peixuan Zhang, Lihua Zhang 0002
Pattern Recognit.5
2024 Dual-stream framework for image-based heart infarction detection using convolutional neural networks
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
Soft Comput.5
2024 Adaptive Multiphase Liver Tumor Segmentation With Multiscale Supervision
abstract
The segmentation of liver tumors using multi-phase computed tomography (CT) images has garnered considerable attention in medical signal processing. However, existing multi-phase liver tumor segmentation methods primarily concentrate on feature integration across various phases, neglecting a comprehensive exploration of synergistic relationships among these phases and constraints on features across different scales. This limitation has led to performance bottlenecks in existing approaches. This article proposes a robust multi-phase liver tumor segmentation framework designed to address the aforementioned challenges. Specifically, we introduce a novel multi-phase and channel-stacked dual attention module, seamlessly integrated within a multi-scale architecture. This module adaptively captures essential semantic information among different phases, enhancing the segmentation network's feature extraction capabilities. A scale-weighted loss function for multi-scale supervision is also designed to mitigate false positives in the segmentation results. To facilitate a systematic evaluation of our model's performance on multi-phase data, we curate a new dataset comprising samples from four distinct phases. Our proposed framework is rigorously assessed through comprehensive quantitative and qualitative experiments, highlighting its compelling performance.
Haopeng Kuang, Xue Yang 0013, Jingwei Wei, Lihua Zhang 0002
IEEE Signal Process. Lett.5
2024 CPR-CLIP: Multimodal Pre-Training for Composite Error Recognition in CPR Training
abstract
The expensive cost of the medical skill training paradigm hinders the development of medical education, which has attracted widespread attention in the intelligent signal processing community. To address the issue of composite error action recognition in Cardiopulmonary Resuscitation (CPR) training, this letter proposes a multimodal pre-training framework named CPR-CLIP based on prompt engineering. Specifically, we design three prompts to fuse multiple errors naturally on the semantic level and then align linguistic and visual features via the contrastive pre-training loss. Extensive experiments verify the effectiveness of the CPR-CLIP. Ultimately, the CPR-CLIP is encapsulated to an electronic assistant, and four doctors are recruited for evaluation. Nearly four times efficiency improvement is observed in comparative experiments, which demonstrates the practicality of the system. We hope this work brings new insights to the intelligent medical skill training and signal processing communities simultaneously. Code is available onhttps://github.com/Shunli-Wang/CPR-CLIP.
Shunli Wang 0001, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
IEEE Signal Process. Lett.4
2024 Towards Asynchronous Multimodal Signal Interaction and Fusion via Tailored Transformers
abstract
The signals from human expressions are usually multimodal, including natural language, facial gestures, and acoustic behaviors. A key challenge is how to fuse multimodal time-series signals with temporal asynchrony. To this end, we present a Transformer-driven Signal Interaction and Fusion (TSIF) approach to effectively model asynchronous multimodal signal sequences. TSIF consists of linear and cross-modal transformer modules with different duties. The linear transformer module efficiently performs the global interaction for multimodal signals, and the vital philosophy is to replace the dot product similarity with the Exponential Kernel while achieving linear complexity by a low-rank matrix decomposition. By targeting the language modality, the cross-modal transformer module aims to capture reliable element correlations among distinct signals and mitigate noise interference in audio and visual modalities. Numerous experiments on two multimodal benchmarks show that our TSIF comparably outperforms previous state-of-the-art models with lower space-time complexities. The systematic analysis also proves the effectiveness of the proposed modules.
Dingkang Yang, Haopeng Kuang, Kun Yang 0010, Mingcheng Li, Lihua Zhang 0002
IEEE Signal Process. Lett.5
2024 Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
abstract
Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach.
Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.7
2023 STAPointGNN: Spatial-Temporal Attention Graph Neural Network for Gesture Recognition Using Millimeter-Wave Radar
Shunli Wang 0001, Lihua Zhang 0002
CollaborateCom (2)4
2023 Context De-Confounded Emotion Recognition
abstract
Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight.
Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002
CVPR11
2023 Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism
abstract
Accurate visualization of liver tumors and their surrounding blood vessels is essential for noninvasive diagnosis and prognosis prediction of tumors. In medical image segmentation, there is still a lack of in-depth research on the simultaneous segmentation of liver tumors and peritumoral blood vessels. To this end, we collect the first liver tumor, and vessel segmentation benchmark datasets containing 52 portal vein phase computed tomography images with liver, liver tumor, and vessel annotations. In this case, we propose a 3D U-shaped Cross-Attention Network (UCA-Net) that utilizes a tailored cross-attention mechanism instead of the traditional skip connection to effectively model the encoder and decoder feature. Specifically, the UCA-Net uses a channel-wise cross-attention module to reduce the semantic gap between encoder and decoder and a slice-wise cross-attention module to enhance the contextual semantic learning ability among distinct slices. Experimental results show that the proposed UCA-Net can accurately segment 3D medical images and achieve state-of-the-art performance on the liver tumor and intrahepatic vessel segmentation task.
Haopeng Kuang, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
ICASSP5
2023 D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection
abstract
Although CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backbone. In this paper, we propose to fuse convolution and transformer, and simultaneously considering the different contributions of non-empty voxels at different positions in 3D space to object detection, it is not consistent with applying standard convolution and transformer directly on voxels. Specifically, we design a novel deformable sparse transformer to perform long-range information interaction on fine-grained local detail semantics aggregated by focal sparse convolution, termed D-Conformer. D-Conformer learns valuable voxels with position-wise in sparse space and can be applied to most voxel-based detectors as a backbone. Extensive experiments demonstrate that our method achieves satisfactory detection results and outperforms state-of-the-art 3D detection methods by a large margin.
Liuzhen Su, Xukun Zhang, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002
ICASSP8
2023 AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception
abstract
Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE.
Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002
ICCV15
2023 HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
abstract
We propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose information, which contains hand prior knowledge that is strongly correlated with occluded regions. Thus, in this work, we propose a novel 3D hand mesh reconstruction network HandGCAT, that can fully exploit hand prior as compensation information to enhance occluded region features. Specifically, we designed the Knowledge-Guided Graph Convolution (KGC) module and the Cross-Attention Transformer (CAT) module. KGC extracts hand prior information from 2D hand pose by graph convolution. CAT fuses hand prior into occluded regions by considering their high correlation. Extensive experiments on popular datasets with challenging hand-object occlusions, such as HO3D v2, HO3D v3, and DexYCB demonstrate that our HandGCAT reaches state-of-the-art performance. The code is available at https://github.com/heartStrive/HandGCAT.
Shuaibing Wang, Shunli Wang 0001, Dingkang Yang, Mingcheng Li, Ziyun Qian, Liuzhen Su, Lihua Zhang 0002
ICME7
2023 MTSAN-MI: Multiscale Temporal-Spatial Convolutional Self-attention Network for Motor Imagery Classification
Junkongshuai Wang, Yangjie Luo, Lihua Zhang 0002, Xiaoyang Kang 0001
ICONIP (9)4
2023 FE-YOLOv5: Improved YOLOv5 Network for Multi-scale Drone-Captured Scene Detection
Zhiyan Dong, Dingkang Yang, Lihua Zhang 0002
ICONIP (2)5
2023 A Scalable Die-to-Die Interconnect with Replay and Repair Schemes for 2.5D/3D Integration
abstract
Chiplet is a critical technology in the post-Moore era, and the die-to-die (D2D) interconnect is essential for communication between chiplets. Meanwhile, several edge-computing devices based on 2.5D/3D chiplet have recently emerged. However, a lightweight D2D interconnect for 2.5D/3D edge-computing systems is lacking. Given the differences between 2.5D/3D integration, a scalable D2D interconnect with replay and repair schemes is presented in this paper. A credit-based flow control scheme and a custom replay scheme are presented for high efficiency. An effective detection and repair scheme is proposed to enhance fault tolerance for the D2D interconnect. Compared with a previous D2D interconnect design, the proposed D2D interconnect delivers 1.07/1.09Gbps throughput ($\sim 2.4\times/\sim 3.9\times \text{for}\ \text{write}/\text{read}$) and significantly reduced energy/bit with only ∼1.7× increased hardware cost. Additionally, compared with a previous chip-to-chip interconnect design, the proposed D2D interconnect can be configured down to power consumption as low as 0.55pJ/bit and 38.40Gbps throughput, achieving ∼2.5× throughput and significantly reduced latency with a negligible increase in hardware cost.
Bo Jiao 0003, Jinshan Zhang 0006, Shiwei Liu 0002, Hao Jiang 0024, Jun Tao 0001, Wenning Jiang, Qi Liu 0010, Lihua Zhang 0002, Haozhe Zhu, Chixiao Chen
ISCAS9
2023 Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002
MICCAI (3)12
2023 How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception
abstract
Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https://github.com/ydk122024/How2comm.
Dingkang Yang, Kun Yang 0010, Jing Liu 0050, Zhi Xu 0010, Rongbin Yin, Peng Zhai, Lihua Zhang 0002
NeurIPS8
2023 FlyTransformer: A Cross-Modal Fusion Policy for UAV End-to-End Trajectory Planning
abstract
The ability to perform efficient trajectory planning is crucial for UAV to carry out tasks autonomously. However, existing research on UAV trajectory planning often employs the cascade process method that involves high-precision maps, real-time positioning and path planning. These methods have limitations such as high computational complexity and time delay, which hinder the efficiency of trajectory planning. End-to-end trajectory planning methods offer a promising solution to this problem. As the core of these end-to-end methods, perception-end plays a decisive role in trajectory planning. But current multimodal fusion of perception is only post-fusion, lacks intermediate feature-level fusion and lacks attention to global visuospatial information. To solve these problems, we propose a new network architecture called FlyTransformer, which fuses the proprioceptive state and visual perception in feature-level for end-to-end trajectory planning. And the key visuospatial information can be attentioned in this architecture. We evaluate our method in forest and cuboid scenarios and their corresponding outdoor scenarios. The results show that FlyTransformer outperforms other baseline algorithms in terms of efficiency and performance.
Wenxiang Shi, Kailei Tang, Junru Sheng, Zhiyan Dong, Lihua Zhang 0002, Xiaoyang Kang 0001
SMC6
2023 Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002
Knowl. Based Syst.10
2023 Towards Robust Multimodal Sentiment Analysis Under Uncertain Signal Missing
abstract
Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of signal completeness. However, many inevitable factors in real applications lead to uncertain signal missing, causing significant degradation of model performance. To this end, we propose a Robust multimodal Missing Signal Framework (RMSF) to handle the problem of uncertain signal missing for MSA tasks and can be generalized to other multimodal patterns. Specifically, a hierarchical cross modal interaction module in RMSF exploits potential complementary semantics among modalities via coarse- and fine-grained cross modal attention. Furthermore, we design an adaptive feature refinement module to enhance the beneficial semantics of modalities and filter redundant features. Finally, we propose a knowledge integrated self-distillation module that enables dynamic knowledge integration and bidirectional knowledge transfer within a single network to precisely reconstruct missing semantics. Comprehensive experiments are conducted on two datasets, indicating that RMSF significantly improves MSA performance under both uncertain missing-signal and complete-signal cases.
Mingcheng Li, Dingkang Yang, Lihua Zhang 0002
IEEE Signal Process. Lett.3
2022 Robust Adversarial Reinforcement Learning with Dissipation Inequation Constraint
abstract
Robust adversarial reinforcement learning is an effective method to train agents to manage uncertain disturbance and modeling errors in real environments. However, for systems that are sensitive to disturbances or those that are difficult to stabilize, it is easier to learn a powerful adversary than establish a stable control policy. An improper strong adversary can destabilize the system, introduce biases in the sampling process, make the learning process unstable, and even reduce the robustness of the policy. In this study, we consider the problem of ensuring system stability during training in the adversarial reinforcement learning architecture. The dissipative principle of robust H-infinity control is extended to the Markov Decision Process, and robust stability constraints are obtained based on L2 gain performance in the reinforcement learning system. Thus, we propose a dissipation-inequation-constraint-based adversarial reinforcement learning architecture. This architecture ensures the stability of the system during training by imposing constraints on the normal and adversarial agents. Theoretically, this architecture can be applied to a large family of deep reinforcement learning algorithms. Results of experiments in MuJoCo and GymFc environments show that our architecture effectively improves the robustness of the controller against environmental changes and adapts to more powerful adversaries. Results of the flight experiments on a real quadcopter indicate that our method can directly deploy the policy trained in the simulation environment to the real environment, and our controller outperforms the PID controller based on hardware-in-the-loop. Both our theoretical and empirical results provide new and critical outlooks on the adversarial reinforcement learning architecture from a rigorous robust control perspective.
Peng Zhai, Zhiyan Dong, Lihua Zhang 0002, Shunli Wang 0001, Dingkang Yang
AAAI4
2022 Emotion Recognition for Multiple Context Awareness
Dingkang Yang, Shunli Wang 0001, Yang Liu 0246, Peng Zhai, Liuzhen Su, Mingcheng Li, Lihua Zhang 0002
ECCV (37)8
2022 Stpointgcn: Spatial Temporal Graph Convolutional Network for Multiple People Recognition Using Millimeter-Wave Radar
abstract
Gait recognition is a new biometric technology, which aims to identify people by their walking posture. Compared with fingerprint recognition, face recognition and other technologies, gait recognition usually has the characteristics of long-distance non-contact and difficulty in camouflage. And compared with the camera-based method, using millimeter-wave radar for gait recognition is immune to light and weather conditions. Moreover, due to the non-invasive feature of millimeter-wave radar, we can design products without privacy risk. In this paper, we propose an end-to-end STPoint-GCN structure, which can extract and aggregate the features of sparse point clouds collected by millimeter-wave radar from the dimensions of space and time. In order to verify our method, we collect and disclose our own gait recognition dataset based on millimeter-wave radar. After comparing with the existing mainstream algorithms, we find that our method is superior to the existing mainstream methods for single-person scenarios and multi-person co-existing scenarios.
Peixian Gong, Lihua Zhang 0002
ICASSP3
2022 Curriculum Adversarial Training for Robust Reinforcement Learning
abstract
Reinforcement learning with adversarial training is currently a key method for improving the robustness of DRL. However, in adversarial training, especially for unstable or disturbance-sensitive systems, the adversary always learns the policy significantly faster than the DRL agent and thus easily generates powerful perturbations. The agent cannot effectively adapt to the overly powerful adversary, which leads to unstable training and even failure to learn the robust policy. In this work, we propose a novel adversarial training method, called Curriculum Adversarial Training, inspired by the idea of curriculum learning. The method dynamically adjusts the strength of the adversary through natural curriculum learning for progressive adversarial training. Thus, the DRL system considers to reasonable learning rules, and the agent faces a suitable learning process. Furthermore, we adopt an advanced action space perturbation method with a most attractive ability as the adversary during training. The proposed method is compared with popular baseline methods through MuJoCo tasks. Experimental results show that our method can improve the robustness of the policy significantly and adapt to uncertain environment effectively.
Junru Sheng, Peng Zhai, Zhiyan Dong, Xiaoyang Kang 0001, Chixiao Chen, Lihua Zhang 0002
IJCNN6
2022 CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in Space
abstract
Reliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to complete robust 6D pose estimation of the space-borne targets under complicated background. Specifically, conventional methods are adopted to extract the features of the whole image in the factual case. In the counterfactual case, a non-existent image without the target but only the background is imagined. Side effect caused by background interference is reduced by counterfactual analysis, which leads to unbiased prediction in final results. In addition, we also carry out low-bit-width quantization for CA-SpaceNet and deploy part of the framework to a Processing-In-Memory (PIM) accelerator on FPGA. Qualitative and quantitative results demonstrate the effectiveness and efficiency of our proposed method. To our best knowledge, this paper applies causal inference and network quantization to the 6D pose estimation of space-borne targets for the first time. The code is available at https://github.com/Shunli-Wang/CA-SpaceNet.
Shunli Wang 0001, Shuaibing Wang, Bo Jiao 0003, Dingkang Yang, Liuzhen Su, Peng Zhai, Chixiao Chen, Lihua Zhang 0002
IROS8
2022 Disentangled Representation Learning for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition aims to identify human emotions from text, audio, and visual modalities. Previous methods either explore correlations between different modalities or design sophisticated fusion strategies. However, the serious problem is that the distribution gap and information redundancy often exist across heterogeneous modalities, resulting in learned multimodal representations that may be unrefined. Motivated by these observations, we propose a Feature-Disentangled Multimodal Emotion Recognition (FDMER) method, which learns the common and private feature representations for each modality. Specifically, we design the common and private encoders to project each modality into modality-invariant and modality-specific subspaces, respectively. The modality-invariant subspace aims to explore the commonality among different modalities and reduce the distribution gap sufficiently. The modality-specific subspaces attempt to enhance the diversity and capture the unique characteristics of each modality. After that, a modality discriminator is introduced to guide the parameter learning of the common and private encoders in an adversarial manner. We achieve the modality consistency and disparity constraints by designing tailored losses for the above subspaces. Furthermore, we present a cross-modal attention fusion module to learn adaptive weights for obtaining effective multimodal representations. The final representation is used for different downstream tasks. Experimental results show that the FDMER outperforms the state-of-the-art methods on two multimodal emotion recognition benchmarks. Moreover, we further verify the effectiveness of our model via experiments on the multimodal humor detection task.
Dingkang Yang, Haopeng Kuang, Yangtao Du, Lihua Zhang 0002
ACM Multimedia5
2022 Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language Sequences
abstract
Understanding human behaviors and intents from videos is a challenging task. Video flows usually involve time-series data from different modalities, such as natural language, facial gestures, and acoustic information. Due to the variable receiving frequency for sequences from each modality, the collected multimodal streams are usually unaligned. For multimodal fusion of asynchronous sequences, the existing methods focus on projecting multiple modalities into a common latent space and learning the hybrid representations, which neglects the diversity of each modality and the commonality across different modalities. Motivated by this observation, we propose a Multimodal Fusion approach for learning modality-Specific and modality-Agnostic representations (MFSA) to refine multimodal representations and leverage the complementarity across different modalities. Specifically, a predictive self-attention module is used to capture reliable contextual dependencies and enhance the unique features over the modality-specific spaces. Meanwhile, we propose a hierarchical cross-modal attention module to explore the correlations between cross-modal elements over the modality-agnostic space. In this case, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, the modality-specific and -agnostic multimodal representations are used together for downstream tasks. Comprehensive experiments on three multimodal datasets clearly demonstrate the superiority of our approach.
Dingkang Yang, Haopeng Kuang, Lihua Zhang 0002
ACM Multimedia4
2022 Contextual and Cross-Modal Interaction for Multi-Modal Speech Emotion Recognition
abstract
Speech emotion recognition combining linguistic content and audio signals in the dialog is a challenging task. Nevertheless, previous approaches have failed to explore emotion cues in contextual interactions and ignored the long-range dependencies between elements from different modalities. To tackle the above issues, this letter proposes a multimodal speech emotion recognition method using audio and text data. We first present a contextual transformer module to introduce contextual information via embedding the previous utterances between interlocutors, which enhances the emotion representation of the current utterance. Then, the proposed cross-modal transformer module focuses on the interactions between text and audio modalities, adaptively promoting the fusion from one modality to another. Furthermore, we construct associative topological relation over mini-batch and learn the association between deep fused features with graph convolutional network. Experimental results on the IEMOCAP and MELD datasets show that our method outperforms current state-of-the-art methods.
Dingkang Yang, Yang Liu 0246, Lihua Zhang 0002
IEEE Signal Process. Lett.4
2021 A 0.57-GOPS/DSP Object Detection PIM Accelerator on FPGA
abstract
The paper presents an object detection accelerator featuring a processing-in-memory (PIM) architecture on FPGAs. PIM architectures are well known for their energy efficiency and avoidance of the memory wall. In the accelerator, a PIM unit is developed using BRAM and LUT based counters, which also helps to improve the DSP performance density. The overall architecture consists of 64 PIM units and three memory buffers to store inter-layer results. A shrunk and quantized Tiny-YOLO network is mapped to the PIM accelerator, where DRAM access is fully eliminated during inference. The design achieves a throughput of 201.6 GOPs at 100MHz clock rate and correspondingly, a performance density of 0.57 GOPS/DSP.
Bo Jiao 0003, Jinshan Zhang 0006, Yuanyuan Xie, Shunli Wang 0001, Haozhe Zhu, Xiaoyang Kang 0001, Zhiyan Dong, Lihua Zhang 0002, Chixiao Chen
ASP-DAC8
2021 Computing Utilization Enhancement for Chiplet-based Homogeneous Processing-in-Memory Deep Learning Processors
abstract
This paper presents a design strategy of chiplet-based processing-in-memory systems for deep neural network applications. Monolithic silicon chips are area and power limited, failing to catch the recent rapid growth of deep learning algorithms. The paper first demonstrates a straightforward layer-wise method that partitions the workload of a monolithic accelerator to a multi-chiplet pipeline. A quantitative analysis shows that the straightforward separation degrades the overall utilization of computing resources due to the reduced on-chiplet memory size, thus introducing a higher memory wall. A tile interleaving strategy is proposed to overcome such degradation. This strategy can segment one layer to different chiplets which maximizes the computing utilization. To facilitate the strategy, the modification of the chiplet system hardware is also discussed. To validate the proposed strategy, a nine-chiplet processing-in-memory system is evaluated with a custom-designed object detection network. Each chiplet can achieve a peak performance of 204.8GOPS at a 100-MHz rate. The peak performance of the overall system is 1.711TOPS, where no off-chip memory access is needed. By the tile interleaving strategy, the utilization is improved from 53.9 to 92.8
Bo Jiao 0003, Haozhe Zhu, Jinshan Zhang 0006, Shunli Wang 0001, Xiaoyang Kang 0001, Lihua Zhang 0002, Mingyu Wang 0001, Chixiao Chen
ACM Great Lakes Symposium on VLSI6
2021 ALPINE: An Agile Processing-in-Memory Macro Compilation Framework
abstract
Processing-in-Memory architectures and circuit designs are playing significant roles in the recent energy-efficient machine learning chips. This paper proposes a PIM macro compilation framework called ALPINE to speed up previously tedious and error-prone PIM design flow, paving the way towards open-source and process-portable PIM chips. Relying on an extensible PIM standard cell library, ALPINE can generate the corresponding topology according to the specification, and process placement and routing. The proposed PIM macro is compatible with different storage devices such as SRAM and RRAM, and can support various quantization bit-widths and dataflows. To verify the effectiveness, a 128×128 SRAM-based PIM macro instance is implemented, and the simulation results show that it can achieve an energy efficiency of 19.05TOPS/W under 65nm CMOS technology. The macro performance is not inferior to the state-of-the-art custom PIM designs.
Jinshan Zhang 0006, Bo Jiao 0003, Yunzhengmao Wang, Haozhe Zhu, Lihua Zhang 0002, Chixiao Chen
ACM Great Lakes Symposium on VLSI5
2021 Learning Associative Representation for Facial Expression Recognition
abstract
The main inherent challenges with the Facial Expression Recognition (FER) are high intra-class variations and high inter-class similarities, while existing methods pay little attention to the association within inter- and intra-class expressions. This paper introduces a novel Expression Associative Network (EAN) to learn association of facial expression, specifically, from two aspects: 1) associative topological relation over mini-batch is constructed by similarity matrix with an adjacent regularization, and 2) learning association of expressions with Graph Convolutional Network (GCN). Besides, an auxiliary module as invariant feature generator based on Generative Adversarial Networks (GAN) is designed to suppress pose variations, illumination changes, and occlusions. Results on public benchmarks achieve comparable or better performance compared with current state-of-the-art methods, with 90.07% on FERPlus, 86.36% on RAF-DB, and improve by 3.92% over SOTA on synthetic wrong labeling datasets.
Yangtao Du, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
ICIP5
2021 MMPoint-GNN: Graph Neural Network with Dynamic Edges for Human Activity Recognition through a Millimeter-Wave Radar
abstract
Human activity recognition has a wide range of application prospects and research significance in intelligent monitoring, assisted driving and human-computer interaction, such as intelligent monitoring of the elderly living alone, warning of dangerous behaviors of drivers and development of somatosensory games. Traditionally, human activity recognition is realized by cameras or wearable devices. However, in privacy-sensitive areas such as wards and cars, users may not be willing to share too many private videos. In this paper, we use millimeterwave radar to collect point clouds of human activities, design a novel graph neural network MMPoint-GNN with dynamic edges for the first time to process sparse point clouds, and combine it with Bidirectional LSTM to build a human activity recognition framework. We transform the logic operation into a differentiable function by edge selection network, and achieve the dynamic edge selection in MMPoint-GNN. Finally, we evaluate our method by comparing it with other methods on MMActivity dataset and MMGesture dataset. The results show that MMPoint-GNN outperforms all other baselines. The code is available at https://github.com/gongpx20069/mmRadar_for_HAR_VS
Peixian Gong, Lihua Zhang 0002
IJCNN3
2021 TSA-Net: Tube Self-Attention Network for Action Quality Assessment
abstract
In recent years, assessing action quality from videos has attracted growing attention in computer vision community and human-computer interaction. Most existing approaches usually tackle this problem by directly migrating the model from action recognition tasks, which ignores the intrinsic differences within the feature map such as foreground and background information. To address this issue, we propose a Tube Self-Attention Network (TSA-Net) for action quality assessment (AQA). Specifically, we introduce a single object tracker into AQA and propose the Tube Self-Attention Module (TSA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions. The TSA module is embedded in existing video networks to form TSA-Net. Overall, our TSA-Net is with the following merits: 1) High computational efficiency, 2) High flexibility, and 3) The state-of-the-art performance. Extensive experiments are conducted on popular action quality assessment datasets including AQA-7 and MTL-AQA. Besides, a dataset named Fall Recognition in Figure Skating (FR-FS) is proposed to explore the basic action assessment in the figure skating scene. Our TSA-Net achieves the Spearman's Rank Correlation of 0.8476 and 0.9393 on AQA-7 and MTL-AQA, respectively, which are the new state-of-the-art results. The results on FR-FS also verify the effectiveness of the TSA-Net. The code and FR-FS dataset are publicly available at https://github.com/Shunli-Wang/TSA-Net.
Shunli Wang 0001, Dingkang Yang, Peng Zhai, Chixiao Chen, Lihua Zhang 0002
ACM Multimedia5
2003 Genetic algorithm for affine point pattern matching
Lihua Zhang 0002, Wenli Xu
Pattern Recognit. Lett.1
2001 A geometric reasoning based algorithm for point pattern matching
Wenli Xu, Lihua Zhang 0002
Sci. China Ser. F Inf. Sci.2