Dingkang Yang

dblp:304/1099 · DBLP profile ↗
← Back
85ranked-venue papers
15as first author
85since 2021 · last 2026
0000-0003-1829-5671ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 11 first-author · 55 since 2021Artificial intelligence and machine learning · 50 · 10 first-author · 50 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
abstract
Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.
Weitao Jia, Jinghui Lu, Haiyang Yu 0004, Guozhi Tang, An-Lan Wang, Weijie Yin, Dingkang Yang, Yuxiang Nie, Bin Shan, Hao Feng 0009, Irene Li, Kun Yang 0010, Jingqun Tang, Teng Fu 0001, Changhong Jin, Xiaohui Lv, Can Huang 0002
AAAI8
2026 SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
abstract
Satire, a form of artistic expression combining humor with implicit critique, holds significant social value by illuminating societal issues. Despite its cultural and societal significance, satire comprehension, particularly in purely visual forms, remains a challenging task for current vision-language models. This task requires not only detecting satire but also deciphering its nuanced meaning and identifying the implicated entities. Existing models often fail to effectively integrate local entity relationships with global context, leading to misinterpretation, comprehension biases, and hallucinations. To address these limitations, we propose SatireDecoder, a training-free framework designed to enhance satirical image comprehension. Our approach proposes a multi-agent system performing visual cascaded decoupling to decompose images into fine-grained local and global semantic representations. In addition, we introduce a chain-of-thought reasoning strategy guided by uncertainty analysis, which breaks down the complex satire comprehension process into sequential subtasks with minimized uncertainty. Our method significantly improves interpretive accuracy while reducing hallucinations. Experimental results validate that SatireDecoder outperforms existing baselines in comprehending visual satire, offering a promising direction for vision-language reasoning in nuanced, high-level semantic tasks.
Haiwei Xue, Minghao Han, Mingcheng Li, Xiaolu Hou, Dingkang Yang, Lihua Zhang 0002, Xu Zheng 0002
AAAI6
2026 UniMGS: Unifying Mesh and 3D Gaussian Splatting with Single-Pass Rasterization and Proxy-Based Deformation
abstract
Joint rendering and deformation of mesh and 3D Gaussian Splatting (3DGS) have significant value as both representations offer complementary advantages for graphics applications. However, due to differences in representation and rendering pipelines, existing studies render meshes and 3DGS separately, making it difficult to accurately handle occlusions and transparency. Moreover, the deformed 3DGS still suffers from visual artifacts due to the sensitivity to the topology quality of the proxy mesh. These issues pose serious obstacles to the joint use of 3DGS and meshes, making it difficult to adapt 3DGS to conventional mesh-oriented graphics pipelines. We propose UniMGS, the first unified framework for rasterizing mesh and 3DGS in a single-pass anti-aliased manner, with a novel binding strategy for 3DGS deformation based on proxy mesh. Our key insight is to blend the colors of both triangle and Gaussian fragments by anti-aliased α-blending in a single pass, achieving visually coherent results with precise handling of occlusion and transparency. To improve the visual appearance of the deformed 3DGS, our Gaussian-centric binding strategy employs a proxy mesh and spatially associates Gaussians with the mesh faces, significantly reducing rendering artifacts. With these two components, UniMGS enables the visualization and manipulation of 3D objects represented by mesh or 3DGS within a unified framework, opening up new possibilities in embodied AI, virtual reality, and gaming. We will release our source code to facilitate future research.
Zeyu Xiao 0001, Yimin Cong, Dongliang Kou, Zhenyi Wu, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
AAAI7
2026 Delving into the adversarial robustness of semantic segmentation with decision-based black-box attacks
Zhaoyu Chen 0001, Zhengyang Shan, Jingwen Chang, Kaixun Jiang, Dingkang Yang, Yiting Cheng 0001
Neural Networks5
2026 Towards unified molecule-enhanced pathology image representation learning via integrating spatial transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Zizhi Chen, Haopeng Kuang, Lihua Zhang 0002
Pattern Recognit.2
2026 Privacy-Preserving Video Anomaly Detection: A Survey
abstract
The video anomaly detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However, vision-based surveillance systems such as closed-circuit television (CCTV) often capture personally identifiable information. The lack of transparency and interpretability in video transmission and usage raises public concerns about privacy and ethics, limiting the real-world application of VAD. Recently, researchers have focused on privacy concerns in VAD by conducting systematic studies from various perspectives, including data, features, and systems, making privacy-preserving VAD (P2VAD) a hotspot in the AI community. However, the current research in P2VAD is fragmented, and prior reviews have mostly focused on methods using RGB sequences, overlooking privacy leakage and appearance bias considerations. To address this gap, this article is the first to systematically review the progress of P2VAD, defining its scope and providing an intuitive taxonomy. We outline the basic assumptions, learning frameworks, and optimization objectives of various approaches, analyzing their strengths, weaknesses, and potential correlations. In addition, we provide open access to research resources such as benchmark datasets and available code. Finally, we discuss key challenges and future opportunities from the perspectives of AI development and P2VAD deployment, aiming to the guide future work in the field.
Yang Liu 0246, Siao Liu, Xiaoguang Zhu, Hao Yang 0055, Juncen Guo, Liangyu Teng, Dingkang Yang, Yan Wang 0068, Jing Liu 0050
IEEE Trans. Neural Networks Learn. Syst.8
2025 BloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene Generation
abstract
With the widespread use of virtual reality applications, 3D scene generation has become a new challenging research frontier. 3D scenes have highly complex structures and need to ensure that the output is dense, coherent, and contains all necessary structures. Many current 3D scene generation methods rely on pre-trained text-to-image diffusion models and monocular depth estimators. However, the generated scenes occupy large amounts of storage space and often lack effective regularisation methods, leading to geometric distortions. To this end, we propose BloomScene, a lightweight structured 3D Gaussian splatting for crossmodal scene generation, which creates diverse and high-quality 3D scenes from text or image inputs. Specifically, a crossmodal progressive scene generation framework is proposed to generate coherent scenes utilizing incremental point cloud reconstruction and 3D Gaussian splatting. Additionally, we propose a hierarchical depth prior-based regularization mechanism that utilizes multi-level constraints on depth accuracy and smoothness to enhance the realism and continuity of the generated scenes. Ultimately, we propose a structured context-guided compression mechanism that exploits structured hash grids to model the context of unorganized anchor attributes, which significantly eliminates structural redundancy and reduces storage overhead. Comprehensive experiments across multiple scenes demonstrate the significant potential and advantages of our framework compared with several baselines.
Xiaolu Hou, Mingcheng Li, Dingkang Yang, Jiawei Chen 0012, Ziyun Qian, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
AAAI3
2025 Debiased Multimodal Understanding for Human Language Sequences
abstract
Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures or fusion strategies to achieve impressive improvements. Unfortunately, they all suffer from the subject variation problem due to data distribution discrepancies among subjects. Concretely, MLU models are easily misled by distinct subjects with different expression customs and characteristics in the training data to learn subject-specific spurious correlations, limiting performance and generalizability across new subjects. Motivated by this observation, we introduce a recapitulative causal graph to formulate the MLU procedure and analyze the confounding effect of subjects. Then, we propose SuCI, a simple yet effective causal intervention module to disentangle the impact of subjects acting as unobserved confounders and achieve model training via true causal effects. As a plug-and-play component, SuCI can be widely applied to most methods that seek unbiased predictions. Comprehensive experiments on several MLU benchmarks clearly show the effectiveness of the proposed module.
Zhi Xu 0010, Dingkang Yang, Mingcheng Li, Zhaoyu Chen 0001, Jiawei Chen 0012, Jinjie Wei, Lihua Zhang 0002
AAAI2
2025 Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators
abstract
Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the internal factual knowledge of target LLMs. In addition, hallucinations typically exhibit multifaceted patterns in downstream tasks, limiting the model's holistic performance across tasks. In this paper, we propose a Comparator-driven Decoding-Time (CDT) framework to alleviate the response hallucination. Firstly, we construct hallucinatory and truthful comparators with multi-task fine-tuning samples. In this case, we present an instruction prototype-guided mixture of experts strategy to enhance the ability of the corresponding comparators to capture different hallucination or truthfulness patterns in distinct task instructions. CDT constrains next-token predictions to factuality-robust distributions by contrasting the logit differences between the target LLMs and these comparators. Systematic experiments on multiple downstream tasks show that our framework can significantly improve the model performance and response factuality.
Dingkang Yang, Dongling Xiao, Jinjie Wei, Mingcheng Li, Zhaoyu Chen 0001, Ke Li 0015, Lihua Zhang 0002
AAAI1
2025 MMPF: Multi-Modal Perception Framework for Abnormal Medical Condition Detection
abstract
As the global population ages and the incidence of chronic diseases increases, the demand for early detection of abnormal medical conditions is increasing. Traditional health monitoring methods often require significant resources and specialized personnel, limiting their widespread use. Leveraging advancements in AI technologies, this study proposes a non-invasive method for detecting abnormal medical conditions from image data. A multimodal perception framework is introduced, integrating features from various modalities, including facial expressions and body postures, to enhance detection accuracy. The framework employs a Cascaded Squeeze-Excitation (CSE) module, consisting of Adaptive and Multi-modal Squeeze-Excitation components, to capture complex feature dependencies and improve cross-modal performance. Extensive experiments demonstrate the effectiveness of this approach, showing improved performance over existing methods. In addition, a new dataset that encompasses a wide range of medical conditions has been released, providing a valuable resource for future research in this domain.
Chuyi Zhong, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
AAAI2
2025 MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
abstract
Diffusion models have shown excellent performance in text-to-image generation. Nevertheless, existing methods often suffer from performance bottlenecks when handling complex prompts that involve multiple objects, characteristics, and relations. Therefore, we propose a Multi-agent Collaboration-based Compositional Diffusion (MCCD) for text-to-image generation for complex scenes. Specifically, we design a multi-agent collaboration-based scene parsing module that generates an agent system comprising multiple agents with distinct tasks, utilizing MLLMs to extract various scene elements effectively. In addition, Hierarchical Compositional diffusion utilizes a Gaussian mask and filtering to refine bounding box regions and enhance objects through region enhancement, resulting in the accurate and high-fidelity generation of complex scenes. Comprehensive experiments demonstrate that our MCCD significantly improves the performance of the baseline models in a training-free manner, providing a substantial advantage in complex scene generation.
Mingcheng Li, Xiaolu Hou, Dingkang Yang, Ziyun Qian, Jiawei Chen 0012, Jinjie Wei, Qingyao Xu, Lihua Zhang 0002
CVPR4
2025 Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models
abstract
Multi-modal keyphrase prediction (MMKP) aims to advance beyond text-only methods by incorporating multiple modalities of input information to produce a set of conclusive phrases.Traditional multi-modal approaches have been proven to have significant limitations in handling the challenging absence and unseen scenarios.Additionally, we identify shortcomings in existing benchmarks that overestimate model capability due to significant overlap in training tests.In this work, we propose leveraging vision-language models (VLMs) for the MMKP task.Firstly, we use two widely-used strategies, e.g., zero-shot and supervised fine-tuning (SFT) to assess the lower bound performance of VLMs.Next, to improve the complex reasoning capabilities of VLMs, we adopt Fine-tune-CoT, which leverages high-quality CoT reasoning data generated by a teacher model to finetune smaller models.Finally, to address the "overthinking" phenomenon, we propose a dynamic CoT strategy which adaptively injects CoT data during training, allowing the model to flexibly leverage its reasoning capabilities during the inference stage.We evaluate the proposed strategies on various datasets and the experimental results demonstrate the effectiveness of the proposed approaches.The code is available at https://github.com/bytedance/DynamicCoT. 6.
Qihang Ma, Dingkang Yang, Chenshaodong, Ran Jiao
EMNLP4
2025 CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
abstract
Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, generating accurate reports remains arduous due to the requirement for precise clinical comprehension and disease diagnosis inference. Furthermore, owing to the limited accessibility of medical data and the imbalanced distribution of diseases, the underrepresentation of rare diseases in training data makes large-scale medical visual language models prone to hallucinations, such as omissions or fabrications, severely undermining diagnostic performance and further intensifying the challenges for MRG in practice. In this study, to effectively mitigate hallucinations in medical report generation, we propose a chain-of-medical-thought approach (CoMT), which intends to imitate the cognitive process of human doctors by decomposing diagnostic procedures. The radiological features with different importance are structured into fine-grained medical thought chains to enhance the inferential ability during diagnosis, thereby alleviating hallucination problems and enhancing the diagnostic accuracy of MRG.
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Ke Li 0015, Lihua Zhang 0002
ICASSP3
2025 Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection
abstract
Although multi-scale alignment has improved domain adaptive object detection by addressing data distribution differences and annotation challenges, little attention has been given to class distribution differences between domains. Additionally, the utilization of feature information across different alignment levels is limited. To alleviate these issues, this paper proposes a novel domain adaptive method that leverages salient features for inter-domain class balancing. Our method consists of three core modules. Specifically, 1) the pixel feature salience-guided module enhances target focus and guides alignment at other scales, improving overall alignment capability; 2) the spatial domain feature purification module filters the noise, extracts salient features, and provides high-quality samples for alignment; 3) the instance relationship adaptive adjustment module adjusts instance weights for different classes to alleviate class distribution differences between different domains. Extensive experiments on multiple datasets demonstrate the effectiveness of our method.
Haiming Peng, Dingkang Yang, Weilong Lin, Xinhua Zeng
ICASSP2
2025 MAFD: Fine-Grained Motion Style Transfer with Adaptive Signal Fusion
abstract
Motion style transfer allows for the swift switching of different styles within the same motion for virtual avatars, offering significant efficiency gains and enhanced motion diversity compared to traditional motion capture methods. However, many existing methods struggle with controlling fine details in complex motions, leading to models that capture only coarse-grained style characteristics. To overcome this limitation, we introduce the Motion Adaptive Fusion Diffusion (MAFD) framework, which leverages adaptive signal fusion to highlight essential style-defining features while minimizing redundant information. Moreover, current diffusion-based denoisers often fail to effectively capture the temporal relationships in motion sequences, producing rigid and fragmented stylized motions. Drawing inspiration from the Mamba model, we propose the Style Mamba Denoiser (SMD), which adopts a selection mechanism to preserve long-range dependencies and maintain temporal coherence. Extensive experiments show that our approach outperforms state-of-the-art methods in both qualitative and quantitative evaluations, achieving more refined and coherent stylized motions.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Dongliang Kou, Lihua Zhang 0002
ICASSP2
2025 UMSD: High Realism Motion Style Transfer via Unified Mamba-based Diffusion
abstract
Motion style transfer is a significant research area in computer vision, enabling the rapid switching of stylistic variations for the same motion in virtual digital humans. This dramatically enhances the richness and realism of motions, making it widely applicable in multimedia contexts such as film, gaming, and the Metaverse. However, most existing methods employ a two-stream structure, which often overlooks the intrinsic relationships between content and style motions, resulting in information loss and misalignment. Additionally, these methods struggle to capture temporal dependencies in long-range motion sequences, resulting in less natural outputs. To address these limitations, we propose a Unified Motion Style Diffusion (UMSD) Framework that simultaneously extracts features from content and style motions, achieving comprehensive information interaction. We also introduce the Motion Style Mamba (MSM) denoiser, which, for the first time in motion style transfer, leverages Mamba's powerful sequence modelling capability to produce more temporally coherent stylized motion sequences. Furthermore, we design a diffusion-based content consistency loss and a style consistency loss to ensure that the framework preserves content motion while effectively learning style motion features. Extensive experiments demonstrate that our approach outperforms State-Of-The-Art (SOTA) methods qualitatively and quantitatively, achieving more realistic and coherent motion style transfer.
Ziyun Qian, Zeyu Xiao 0001, Xingliang Jin, Dingkang Yang, Mingcheng Li, Zhenyi Wu, Dongliang Kou, Peng Zhai, Lihua Zhang 0002
ACM Multimedia4
2025 Boosting Adversarial Transferability with Spatial Adversarial Alignment
abstract
Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods still show limited transferability, partiovovocularly in cross-architecture scenarios, such as from CNN to ViT. To achieve high transferability, we propose a technique termed Spatial Adversarial Alignment (SAA), which employs an alignment loss and leverages a witness model to fine-tune the surrogate model. Specifically, SAA consists of two key parts: spatial-aware alignment and adversarial-aware alignment. First, we minimize the divergences of features between the two models in both global and local regions, facilitating spatial alignment. Second, we introduce a self-adversarial strategy that leverages adversarial examples to impose further constraints, aligning features from an adversarial perspective. Through this alignment, the surrogate model is trained to concentrate on the common features extracted by the witness model. This facilitates adversarial attacks on these shared features, thereby yielding perturbations that exhibit enhanced transferability. Extensive experiments on various architectures on ImageNet show that aligned surrogate models based on SAA can provide higher transferable adversarial examples, especially in cross-architecture attacks.
Zhaoyu Chen 0001, Haijing Guo, Kaixun Jiang, Jiyuan Fu, Xinyu Zhou 0006, Dingkang Yang, Hao Tang 0005, Bo Li 0115
NeurIPS6
2025 AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
abstract
Learning rate is widely regarded as crucial for effective foundation model pretraining. Recent research explores and demonstrates the transferability of learning rate configurations across varying model and dataset sizes, etc. Nevertheless, these approaches are constrained to specific training scenarios and typically necessitate extensive hyperparameter tuning on proxy models. In this work, we propose \textbf{AdaLRS}, a plug-in-and-play adaptive learning rate search algorithm that conducts online optimal learning rate search via optimizing loss descent velocities. We provide theoretical and experimental analyzes to show that foundation model pretraining loss and its descent velocity are both convex and share the same optimal learning rate. Relying solely on training loss dynamics, AdaLRS involves few extra computations to guide the search process, and its convergence is guaranteed via theoretical analysis. Experiments on both LLM and VLM pretraining show that AdaLRS adjusts suboptimal learning rates to the neighborhood of optimum with marked efficiency and effectiveness, with model performance improved accordingly. We also show the robust generalizability of AdaLRS across varying training scenarios, such as different model sizes, training paradigms, base learning rate scheduler choices, and hyperparameter settings.
Hongyuan Dong, Dingkang Yang, Ran Jiao
NeurIPS2
2025 DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
abstract
We introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hindering progress in models that require joint reasoning over temporal, textual, and visual information. To address this gap, DanmakuTPPBench comprises two complementary components:(1) DanmakuTPP-Events, a novel dataset derived from the Bilibili video platform, where user-generated bullet comments (Danmaku) naturally form multi-modal events annotated with precise timestamps, rich textual content, and corresponding video frames;(2) DanmakuTPP-QA, a challenging question-answering dataset constructed via a novel multi-agent pipeline powered by state-of-the-art LLMs and multi-modal LLMs (MLLMs), targeting complex temporal-textual-visual reasoning. We conduct extensive evaluations using both classical TPP models and recent MLLMs, revealing significant performance gaps and limitations in current methods’ ability to model multi-modal event dynamics. Our benchmark establishes strong baselines and calls for further integration of TPP modeling into the multi-modal language modeling landscape. Project page: https://github.com/FRENKIE-CHIANG/DanmakuTPPBench.
Jichu Li, Yang Liu 0246, Dingkang Yang, Quyu Kong
NeurIPS4
2025 Diverse object placement with dual interaction
Xianhe Cheng, Peng Zhai, Dingkang Yang, Lihua Zhang 0002
Neurocomputing3
2025 MSTDF: Motion Style Transfer Towards High Visual Fidelity Based on Dynamic Fusion
abstract
Emotion-guided motion style transfer is a novel research direction, enabling the efficient generation of motion in various emotional styles for use in films, games, and other domains. However, existing methods primarily rely on global feature statistics for motion style transfer, neglecting local semantic structure and resulting in the degradation of motion content structure. This letter proposes a novel Motion Style Transfer based on Dynamic Fusion (MSTDF) framework, which treats content and style motion as distinct signals and employs dynamic fusion for high-fidelity motion style transfer. Additionally, to address the challenge of traditional discriminators capturing subtle motion style features, we propose the Motion Dynamic Fusion (MDF) discriminator to capture the details and fine-grained style characteristics of motion sequences, assisting the generator in producing higher-fidelity stylized motion. Finally, extensive experiments on the Xia dataset demonstrate that our method surpasses state-of-the-art methods in qualitative and quantitative comparisons.
Ziyun Qian, Dingkang Yang, Mingcheng Li, Zeyu Xiao 0001, Lihua Zhang 0002
IEEE Signal Process. Lett.2
2025 Robust Multi-Agent Collaborative Perception via Spatio-Temporal Awareness
abstract
As an emerging application in autonomous driving, multi-agent collaborative perception has recently received significant attention. Despite promising advances from previous efforts, several unavoidable challenges that cause performance bottlenecks remain, including the single-frame detection dilemma, communication redundancy, and defective collaboration process. To this end, we proposeSCOPE++, a versatile collaborative perception framework aggregating spatio-temporal information across on-road agents to tackle these issues. We introduce four components inSCOPE++for robust collaboration by seeking a reasonable trade-off between perception performance and communication bandwidth. First, we devise a context-aware information aggregation to capture valuable semantic cues in the temporal context and enhance the current local representation of the ego agent. Second, an exclusivity-aware sparse communication is introduced to filter perceptually unnecessary information from collaborators and transmit complementary features relative to the ego agent. Third, we present an importance-aware cross-agent collaboration to incorporate semantic representations of spatially critical locations across agents flexibly. Finally, a contribution-aware adaptive fusion is designed to integrate multi-source representations based on dynamic contributions. Our framework is evaluated on multiple LiDAR-based collaborative detection datasets in real-world and simulated scenarios, and comprehensive experiments show thatSCOPE++outperforms state-of-the-art methods on all datasets.
Kun Yang 0010, Zhi Xu 0010, Dingkang Yang, Lihua Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2025 MSCPT: Few-Shot Whole Slide Image Classification With Multi-Scale and Context-Focused Prompt Tuning
abstract
Multiple instance learning (MIL) has become a standard paradigm for the weakly supervised classification of whole slide images (WSIs). However, this paradigm relies on using a large number of labeled WSIs for training. The lack of training data and the presence of rare diseases pose significant challenges for these methods. Prompt tuning combined with pre-trained Vision-Language models (VLMs) is an effective solution to the Few-shot Weakly Supervised WSI Classification (FSWC) task. Nevertheless, applying prompt tuning methods designed for natural images to WSIs presents three significant challenges: 1) These methods fail to fully leverage the prior knowledge from the VLM's text modality; 2) They overlook the essential multi-scale and contextual information in WSIs, leading to suboptimal results; and 3) They lack exploration of instance aggregation methods. To address these problems, we propose a Multi-Scale and Context-focused Prompt Tuning (MSCPT) method for FSWC task. Specifically, MSCPT employs the frozen large language model to generate pathological visual language prior knowledge at multiple scales, guiding hierarchical prompt tuning. Additionally, we design a graph prompt tuning module to learn essential contextual information within WSI, and finally, a non-parametric cross-guided instance aggregation module has been introduced to derive the WSI-level features. Extensive experiments, visualizations, and interpretability analyses were conducted on five datasets and three downstream tasks using three VLMs, demonstrating the strong performance of our MSCPT. All codes have been made publicly accessible at https://github.com/Hanminghao/MSCPT.
Minghao Han, Linhao Qu, Dingkang Yang, Xukun Zhang, Lihua Zhang 0002
IEEE Trans. Medical Imaging3
2024 A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities
abstract
Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion approaches. To this end, we propose a Unified multimodal Missing modality self-Distillation Framework (UMDF) to handle the problem of uncertain missing modalities in MSA. Specifically, a unified self-distillation mechanism in UMDF drives a single network to automatically learn robust inherent representations from the consistent distribution of multimodal data. Moreover, we present a multi-grained crossmodal interaction module to deeply mine the complementary semantics among modalities through coarse- and fine-grained crossmodal attention. Eventually, a dynamic feature integration module is introduced to enhance the beneficial semantics in incomplete modalities while filtering the redundant information therein to obtain a refined and robust multimodal representation. Comprehensive experiments on three datasets demonstrate that our framework significantly improves MSA performance under both uncertain missing-modality and complete-modality testing conditions.
Mingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang 0001, Shuaibing Wang, Liuzhen Su, Kun Yang 0010, Lihua Zhang 0002
AAAI2
2024 Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation
abstract
Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impractical to deploy in natural high-security scenes due to these methods rely entirely on original or publicly available data with a similar distribution. In fact, these data are almost always private, specific, and distinctive for scenes that require high robustness. To tackle these issues, we propose a challenging but significant task called Data-Free Adversarial Robustness Distillation (DFARD), which aims to train small, easily deployable, robust models without relying on data. We demonstrate that the challenge lies in the lower upper bound of knowledge transfer information, making it crucial to mining and transferring knowledge more efficiently. Inspired by human education, we design a plug-and-play Interactive Temperature Adjustment (ITA) strategy to improve the efficiency of knowledge transfer and propose an Adaptive Generator Balance (AGB) module to retain more data information. Our method uses adaptive hyperparameters to avoid a large number of parameter tuning, which significantly outperforms the combination of existing techniques. Meanwhile, our method achieves stable and reliable performance on multiple benchmarks.
Zhaoyu Chen 0001, Dingkang Yang, Pinxue Guo, Kaixun Jiang, Lizhe Qi
AAAI3
2024 SceneWeaver: Text-Driven Scene Generation with Geometry-aware Gaussian Splatting
Xiaolu Hou, Mingcheng Li, Jiawei Chen 0012, Dingkang Yang, Ziyun Qian, Lihua Zhang 0002
ACML4
2024 Large Vision-Language Models as Emotion Recognizers in Context Awareness
Yuxuan Lei, Dingkang Yang, Zhaoyu Chen 0001, Jiawei Chen 0012, Peng Zhai, Lihua Zhang 0002
ACML2
2024 CPR-Coach: Recognizing Composite Error Actions Based on Single-Class Training
abstract
Fine- grained medical action analysis plays a vital role in improving medical skill training efficiency, but it faces the problems of data and algorithm shortage. Cardiopul-monary Resuscitation (CPR) is an essential skill in emer-gency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this pa-per constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compres-sion and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper in-vestigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable “Single-class Training & Multi-class Testing” problem, we propose a human-cognition-inspired framework named ImagineNet to improve the model's multi-error recognition performance under restricted supervision. Extensive comparison and actual deployment experiments verify the effectiveness of the framework. We hope this work could bring new inspiration to the computer vision and medical skills training communities simultaneously. The dataset and the code are publicly available on https://github.com/Shunli-Wang/CPR-Coach.
Shunli Wang 0001, Shuaibing Wang, Dingkang Yang, Mingcheng Li, Haopeng Kuang, Liuzhen Su, Peng Zhai, Lihua Zhang 0002
CVPR3
2024 Correlation-Decoupled Knowledge Distillation for Multimodal Sentiment Analysis with Incomplete Modalities
abstract
Multimodal sentiment analysis (MSA) aims to understand human sentiment through multimodal data. Most MSA efforts are based on the assumption of modality completeness. However, in real-world applications, some practical factors cause uncertain modality missingness, which drastically degrades the model's performance. To this end, we propose a Correlation-decoupled Knowledge Distillation (CorrKD) framework for the MSA task under uncertain missing modalities. Specifically, we present a sample-level contrastive distillation mechanism that transfers comprehensive knowledge containing cross-sample correlations to reconstruct missing semantics. Moreover, a category-guided prototype distillation mechanism is introduced to capture cross-category correlations using category prototypes to align feature distributions and generate favorable joint representations. Eventually, we design a response-disentangled consistency distillation strategy to optimize the sentiment decision boundaries of the student network through response disentanglement and mutual information maximization. Comprehensive experiments on three datasets indicate that our framework can achieve favorable improvements compared with several baselines.
Mingcheng Li, Dingkang Yang, Shuaibing Wang, Yan Wang 0068, Kun Yang 0010, Dongliang Kou, Ziyun Qian, Lihua Zhang 0002
CVPR2
2024 De-Confounded Data-Free Knowledge Distillation for Handling Distribution Shifts
abstract
Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset.
Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi
CVPR2
2024 Robust Emotion Recognition in Context Debiasing
abstract
Context-aware emotion recognition (CAER) has recently boosted the practical applications of affective computing techniques in unconstrained environments. Mainstream CAER methods invariably extract ensemble representations from diverse contexts and subject-centred characteristics to perceive the target person's emotional state. Despite advancements, the biggest challenge remains due to context bias interference. The harmful bias forces the models to rely on spurious correlations between background contexts and emotion labels in likelihood estimation, causing severe performance bottlenecks and confounding valuable context priors. In this paper, we propose a counterfactual emotion inference (CLEF) framework to address the above issue. Specifically, we first formulate a generalized causal graph to decouple the causal relationships among the variables in CAER. Following the causal graph, CLEF introduces a non-invasive context branch to capture the adverse direct effect caused by the context bias. During the inference, we eliminate the direct context effect from the total causal effect by comparing factual and counterfactual outcomes, resulting in bias mitigation and robust prediction. As a model-agnostic framework, CLEF can be readily integrated into existing methods, bringing consistent performance gains.
Dingkang Yang, Kun Yang 0010, Mingcheng Li, Shunli Wang 0001, Shuaibing Wang, Lihua Zhang 0002
CVPR1
2024 Pathology-Knowledge Enhanced Multi-instance Prompt Learning for Few-Shot Whole Slide Image Classification
Linhao Qu, Dingkang Yang, Qinhao Guo, Rongkui Luo, Shaoting Zhang 0001, Xiaosong Wang 0001
ECCV (11)2
2024 Self-cooperation Knowledge Distillation for Novel Class Discovery
Zhaoyu Chen 0001, Dingkang Yang, Yunquan Sun, Lizhe Qi
ECCV (70)3
2024 Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002
ECCV (58)1
2024 Align Before Collaborate: Mitigating Feature Misalignment for Robust Multi-agent Perception
Kun Yang 0010, Dingkang Yang, Ke Li 0015, Dongling Xiao, Zedian Shao, Peng Sun 0007
ECCV (4)2
2024 MISS: A Generative Pre-training and Fine-Tuning Approach for Med-VQA
Jiawei Chen 0012, Dingkang Yang, Yuxuan Lei, Lihua Zhang 0002
ICANN (8)2
2024 Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide Images
abstract
Survival prediction is a complex ordinal regression task that aims to predict the survival coefficient ranking among a cohort of patients, typically achieved by analyzing patients’ whole slide images. Existing deep learning approaches mainly adopt multiple instance learning or graph neural networks under weak supervision. Most of them are unable to uncover the diverse interactions between different types of biological entities(e.g., cell cluster and tissue block) across multiple scales, while such interactions are crucial for patient survival prediction. In light of this, we propose a novel multi-scale heterogeneity-aware hypergraph representation framework. Specifically, our framework first constructs a multi-scale heterogeneity-aware hypergraph and assigns each node with its biological entity type. It then mines diverse interactions between nodes on the graph structure to obtain a global representation. Experimental results demonstrate that our method outperforms state-of-the-art approaches on three benchmark datasets. Code is publicly available at https://github.com/Hanminghao/H2GT.
Minghao Han, Xukun Zhang, Dingkang Yang, Tao Liu 0050, Haopeng Kuang, Jinghui Feng, Lihua Zhang 0002
ICME3
2024 IIPC: Intra-Inter Patch Correlations for Garment Collision Handling
abstract
Realistic garment simulation is critical for digital humans. However, noticeable penetrations still exist in current learning-based garment simulation techniques. To reduce penetrations in predicted garments, we resort to the garment geometry and neural Signed Distance Fields (SDFs) for effective collision handling. The key idea of our method is that we divide the garment into patches and model the local and global garment geometry through Intra- and Inter-Patch Correlations (IIPC), which can be easily learned through the powerful context-understanding ability of Transformers. The geometry information is then utilized to predict a per-vertex moving offset, according to which we move the penetrating vertices along the SDF’s gradient directions to solve collisions. Our module can be coupled with learning-based backbones to effectively solve penetrations while retaining real-time performance. Extensive experiments show that the proposed method excels the prior works significantly.
Ruisheng Yuan, Minzhe Tang, Dongliang Kou, Dingkang Yang, Lihua Zhang 0002
ICME5
2024 T-GET3D: A Generative Model of High-Quality 3D Textured Shapes Guided by Texts
Xinxin Shi, Xianhe Cheng, Peixuan Zhang, Dingkang Yang, Lihua Zhang 0002
ICONIP (2)5
2024 MDRPC: Music-Driven Robot Primitives Choreography
abstract
Dance has been an important art form and means of communication since the dawn of human civilization. Equipping humanoid robots with the ability to perform smooth dance movements to music is a key research priority in artificial intelligence, robotics and human-computer interaction. However, existing kinematics-based dance generation methods often violate real-world physical laws as they do not consider physical constraints, leading to unrealistic movements. Additionally, due to the diversity and dynamic variability of input music, most existing physics-based methods, which rely on task-specific reward functions, face significant challenges in effectively handling music-driven dance generation tasks. To address these issues, we introduce MDRPC, the first physics-based, music-driven dance generation method for humanoid robots. Inspired by human choreographic principles, MDRPC is defined as a two-phase framework. The initial phase utilizes adversarial imitation learning to acquire a rich set of reusable dance primitives from a music-dance dataset. In the subsequent phase, these dance primitives are orchestrated under the guidance of musical theory and choreographic rules to generate complex humanoid dance sequences. Specifically, we propose beat alignment and dance diversity reward functions to synchronize motion rhythms with music beats and enhance the diversity of dance movements. We implement MDRPC on a simulated humanoid robot, and the results confirm that our method effectively controls the humanoid, enabling it to perform dance movements harmoniously synchronized with the music.
Haiyang Guan, Xiaoyi Wei, Weifan Long, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
ICTAI4
2024 Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Ziyun Qian, Lihua Zhang 0002
MICCAI (5)3
2024 Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Models
Jiawei Chen 0012, Dingkang Yang, Mingcheng Li, Jinjie Wei, Xiaolu Hou, Lihua Zhang 0002
ACM Multimedia2
2024 IF-Garments: Reconstructing Your Intersection-Free Multi-Layered Garments from Monocular Videos
abstract
Reconstructing garments from monocular videos has attracted considerable attention as it provides a convenient and low-cost solution for clothing digitization. In reality, people wear clothing with countless variations and multiple layers. Existing studies attempt to extract garments from a single video. They either behave poorly in generalization due to reliance on limited clothing templates or struggle to handle the intersections of multi-layered clothing leading to the lack of physical plausibility. Besides, there are inevitable and undetectable overlaps for a single video that hinder researchers from modeling complete and intersection-free multi-layered clothing. To address the above limitations, in this paper, we propose a novel method to reconstruct multi-layered clothing from multiple monocular videos sequentially, which surpasses existing work in generalization and robustness against penetration. For each video, neural fields are employed to implicitly represent the clothed body, from which the meshes with frame-consistent structures are explicitly extracted. Next, we implement a template-free method for extracting a single garment by back-projecting the image segmentation labels of different frames onto these meshes. In this way, multiple garments can be obtained from these monocular videos and then aligned to form the whole outfit. However, intersection always occurs due to overlapping deformation in the real world and perceptual errors in monocular videos. To this end, we innovatively introduce a physics-aware module that combines neural fields with a position-based simulation framework to fine-tune the penetrating vertices of garments, ensuring robustly intersection-free. Additionally, we collect a mini dataset with fashionable garments to evaluate the quality of clothing reconstruction comprehensively. We release our code and data at https://github.com/SMY19999/IF-Garments.
Qipeng Yan, Zhuoer Liang, Dongliang Kou, Dingkang Yang, Ruisheng Yuan, Mingcheng Li, Lihua Zhang 0002
ACM Multimedia5
2024 Sampling to Distill: Knowledge Transfer from Open-World Data
abstract
Data-Free Knowledge Distillation (DFKD) is a novel task that aims to train high-performance student models using only the pre-trained teacher network without original training data. Most of the existing DFKD methods rely heavily on additional generation modules to synthesize the substitution data resulting in high computational costs and ignoring the massive amounts of easily accessible, low-cost, unlabeled open-world data. Meanwhile, existing methods ignore the domain shift issue between the substitution data and the original data, resulting in knowledge from teachers not always trustworthy and structured knowledge from data becoming a crucial supplement. To tackle the issue, we propose a novel Open-world Data Sampling Distillation (ODSD) method for the DFKD task without the redundant generation process. First, we try to sample open-world data close to the original data's distribution by an adaptive sampling module and introduce a low-noise representation to alleviate the domain shift issue. Then, we build structured relationships of multiple data examples to exploit data knowledge through the student model itself and the teacher's structured representation. Extensive experiments on CIFAR-10, CIFAR-100, NYUv2, and ImageNet show that our ODSD method achieves state-of-the-art performance with lower FLOPs and parameters. Especially, we improve 1.50%-9.59% accuracy on the ImageNet dataset and avoid training the separate generator for each class.
Zhaoyu Chen 0001, Jie Zhang 0107, Dingkang Yang, Zuhao Ge, Yang Liu 0246, Siao Liu, Yunquan Sun, Lizhe Qi
ACM Multimedia4
2024 MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
abstract
Accurate and robust multimodal multi-task perception is crucial for modern autonomous driving systems. However, current multimodal perception research follows independent paradigms designed for specific perception tasks, leading to a lack of complementary learning among tasks and decreased performance in multi-task learning (MTL) due to joint training. In this paper, we propose MaskBEV, a masked attention-based MTL paradigm that unifies 3D object detection and bird's eye view (BEV) map segmentation. MaskBEV introduces a task-agnostic Transformer decoder to process these diverse tasks, enabling MTL to be completed in a unified decoder without requiring additional design of specific task heads. To fully exploit the complementary information between BEV map segmentation and 3D object detection tasks in BEV space, we propose spatial modulation and scene-level context aggregation strategies. These strategies consider the inherent dependencies between BEV segmentation and 3D detection, naturally boosting MTL performance. Extensive experiments on nuScenes dataset show that compared with previous state-of-the-art MTL methods, MaskBEV achieves 1.3 NDS improvement in 3D object detection and 2.7 mIoU improvement in BEV map segmentation, while also demonstrating slightly leading inference speed.
Xukun Zhang, Dingkang Yang, Mingcheng Li, Shunli Wang 0001, Lihua Zhang 0002
ACM Multimedia3
2024 Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation Learning
abstract
Multimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases.
Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002
NeurIPS2
2024 PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
abstract
Developing intelligent pediatric consultation systems offers promising prospects for improving diagnostic efficiency, especially in China, where healthcare resources are scarce. Despite recent advances in Large Language Models (LLMs) for Chinese medicine, their performance is sub-optimal in pediatric applications due to inadequate instruction data and vulnerable training procedures. To address the above issues, this paper builds PedCorpus, a high-quality dataset of over 300,000 multi-task instructions from pediatric textbooks, guidelines, and knowledge graph resources to fulfil diverse diagnostic demands. Upon well-designed PedCorpus, we propose PediatricsGPT, the first Chinese pediatric LLM assistant built on a systematic and robust training pipeline. In the continuous pre-training phase, we introduce a hybrid instruction pre-training mechanism to mitigate the internal-injected knowledge inconsistency of LLMs for medical domain adaptation. Immediately, the full-parameter Supervised Fine-Tuning (SFT) is utilized to incorporate the general medical knowledge schema into the models. After that, we devise a direct following preference optimization to enhance the generation of pediatrician-like humanistic responses. In the parameter-efficient secondary SFT phase, a mixture of universal-specific experts strategy is presented to resolve the competency conflict between medical generalist and pediatric expertise mastery. Extensive results based on the metrics, GPT-4, and doctor evaluations on distinct downstream tasks show that PediatricsGPT consistently outperforms previous Chinese medical LLMs. The project and data will be released at https://github.com/ydk122024/PediatricsGPT.
Dingkang Yang, Jinjie Wei, Dongling Xiao, Shunli Wang 0001, Mingcheng Li, Shuaibing Wang, Jiawei Chen 0012, Qingyao Xu, Ke Li 0015, Peng Zhai, Lihua Zhang 0002
NeurIPS1
2024 A novel hierarchical distributed vehicular edge computing framework for supporting intelligent driving
Kun Yang 0010, Peng Sun 0007, Dingkang Yang, Jieyu Lin, Azzedine Boukerche
Ad Hoc Networks3
2024 Towards heart infarction detection via image-based dataset and three-stream fusion framework
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
Comput. Commun.2
2024 Expression guided medical condition detection via the Multi-Medical Condition Image Dataset
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002
Eng. Appl. Artif. Intell.2
2024 Memory-enhanced spatial-temporal encoding framework for industrial anomaly detection system
Yang Liu 0246, Bobo Ju, Dingkang Yang, Liyuan Peng, Peng Sun 0007, Chengfang Li, Hao Yang 0055, Jing Liu 0050
Expert Syst. Appl.3
2024 Boosting the transferability of adversarial attacks with global momentum initialization
Zhaoyu Chen 0001, Kaixun Jiang, Dingkang Yang, Lingyi Hong, Pinxue Guo, Haijing Guo
Expert Syst. Appl.4
2024 Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded Training
abstract
Understanding emotions from diverse contexts has received widespread attention in computer vision communities. The core philosophy of Context-Aware Emotion Recognition (CAER) is to provide valuable semantic cues for recognizing the emotions of target persons by leveraging rich contextual information. Current approaches invariably focus on designing sophisticated structures to extract perceptually critical representations from contexts. Nevertheless, a long-neglected dilemma is that a severe context bias in existing datasets results in an unbalanced distribution of emotional states among different contexts, causing biased visual representation learning. From a causal demystification perspective, the harmful bias is identified as a confounder that misleads existing models to learn spurious correlations based on likelihood estimation, limiting the models' performance. To address the issue, we embrace causal inference to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a customized causal graph. Subsequently, we present a Contextual Causal Intervention Module (CCIM) to de-confound the confounder, which is built upon backdoor adjustment theory to facilitate seeking approximate causal effects during model training. As a plug-and-play component, CCIM can easily integrate with existing approaches and bring significant improvements. Systematic experiments on three datasets demonstrate the effectiveness of our CCIM.
Dingkang Yang, Kun Yang 0010, Haopeng Kuang, Zhaoyu Chen 0001, Lihua Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Dual-stream framework for image-based heart infarction detection using convolutional neural networks
Chuyi Zhong, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
Soft Comput.2
2024 CPR-CLIP: Multimodal Pre-Training for Composite Error Recognition in CPR Training
abstract
The expensive cost of the medical skill training paradigm hinders the development of medical education, which has attracted widespread attention in the intelligent signal processing community. To address the issue of composite error action recognition in Cardiopulmonary Resuscitation (CPR) training, this letter proposes a multimodal pre-training framework named CPR-CLIP based on prompt engineering. Specifically, we design three prompts to fuse multiple errors naturally on the semantic level and then align linguistic and visual features via the contrastive pre-training loss. Extensive experiments verify the effectiveness of the CPR-CLIP. Ultimately, the CPR-CLIP is encapsulated to an electronic assistant, and four doctors are recruited for evaluation. Nearly four times efficiency improvement is observed in comparative experiments, which demonstrates the practicality of the system. We hope this work brings new insights to the intelligent medical skill training and signal processing communities simultaneously. Code is available onhttps://github.com/Shunli-Wang/CPR-CLIP.
Shunli Wang 0001, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
IEEE Signal Process. Lett.2
2024 Towards Asynchronous Multimodal Signal Interaction and Fusion via Tailored Transformers
abstract
The signals from human expressions are usually multimodal, including natural language, facial gestures, and acoustic behaviors. A key challenge is how to fuse multimodal time-series signals with temporal asynchrony. To this end, we present a Transformer-driven Signal Interaction and Fusion (TSIF) approach to effectively model asynchronous multimodal signal sequences. TSIF consists of linear and cross-modal transformer modules with different duties. The linear transformer module efficiently performs the global interaction for multimodal signals, and the vital philosophy is to replace the dot product similarity with the Exponential Kernel while achieving linear complexity by a low-rank matrix decomposition. By targeting the language modality, the cross-modal transformer module aims to capture reliable element correlations among distinct signals and mitigate noise interference in audio and visual modalities. Numerous experiments on two multimodal benchmarks show that our TSIF comparably outperforms previous state-of-the-art models with lower space-time complexities. The systematic analysis also proves the effectiveness of the proposed modules.
Dingkang Yang, Haopeng Kuang, Kun Yang 0010, Mingcheng Li, Lihua Zhang 0002
IEEE Signal Process. Lett.1
2024 Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
abstract
Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach.
Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2024 AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection System
abstract
As essential tools for industry safety protection, automatic video anomaly detection systems (AVADS) are designed to detect anomalous events of concern in surveillance videos. Existing VAD methods lack effective exploration of the prototypical appearance and motion features leading to poor performance in realistic scenarios. Specifically, they either misreport regular events as anomalies due to insufficient representation power, or lead to missed detections with over-power generalization. In this regard, we propose an appearance-motion prototype network (AMP-net) that uses external memories to record prototype features and augments the appearance-motion prototype with a spatial-temporal fusion. In addition, AMP-net sequentially fuses appearance features from deep to shallow to utilize multiscale spatial context. Additionally, we introduce temporal attention to capture important dynamics and enhance AMP-net for representing regular motion. The proposed method achieves a delicate balance of effective representation of normal events and limited generalization to anomalies. Experiments on three benchmark datasets demonstrate that our method can accurately detect anomalous events, achieving performance comparable to state-of-the-art methods with frame-level AUCs of 98.7%, 92.4%, and 78.8% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Moreover, we conducted a case study on the self-collected industrial dataset, and the results indicate that our AMP-net can cope with complex industrial scenarios and outperform existing methods.
Yang Liu 0246, Jing Liu 0050, Kun Yang 0010, Bobo Ju, Siao Liu, Dingkang Yang, Peng Sun 0007
IEEE Trans. Ind. Informatics7
2024 MGR3Net: Multigranularity Region Relation Representation Network for Facial Expression Recognition in Affective Robots
abstract
Automatic facial expression recognition (FER) based on face images is essential for affective robots, which are designed for interactive companions and intelligent healthcare. Although existing DL-based FERs have made significant progress, an accurate FER model in robots is challenging due to the subtle differences in facial expressions across various scenarios. To address this issue, we propose a multigranularity region relation representation network (MGR3Net) to improve the robustness and generalization of FER via attention-guided global-local fusion. The MGR3Net is composed of three modules: multigranularity attention (MGA), holistic-regional feature extractor (HRFE), and hybrid feature fusion. In the MGA module, we first process each holistic cropped face image into three granularity of face regions from coarse to fine, which are four region-cropped faces,$2^{2}$face partitions, and$4^{2}$face partitions. Then, we propose the region attention relation cell to model the relationship between each region and the aggregated representation while preserving the spatial information of the local features. In the HRFE module, we align multigranularity features from the coarse space to the finer space and extract one holistic embedding and multiple region embeddings for each granularity. Finally, we use a hybrid-level fusion strategy to combine global-local features from the three granularities for final classification. Extensive experiments demonstrate that the MGR3Net outperforms the state-of-the-art methods evaluated on the in-the-lab datasets, in-the-wild datasets, and occlusion/pose-based sets.
Yan Wang 0068, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jing Liu 0050, Dingkang Yang, Shuyong Gao
IEEE Trans. Ind. Informatics6
2023 Context De-Confounded Emotion Recognition
abstract
Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representations from subjects and contexts. However, a long-overlooked issue is that a context bias in existing datasets leads to a significantly unbalanced distribution of emotional states among different context scenarios. Concretely, the harmful bias is a confounder that misleads existing models to learn spurious correlations based on conventional likelihood estimation, significantly limiting the models' performance. To tackle the issue, this paper provides a causality-based perspective to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a tailored causal graph. Then, we propose a Contextual Causal Intervention Module (CCIM) based on the backdoor adjustment to de-confound the confounder and exploit the true causal effect for model training. CCIM is plug-in and model-agnostic, which improves diverse state-of-the-art approaches by considerable margins. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our CCIM and the significance of causal insight.
Dingkang Yang, Zhaoyu Chen 0001, Shunli Wang 0001, Mingcheng Li, Siao Liu, Zhiyan Dong, Peng Zhai, Lihua Zhang 0002
CVPR1
2023 Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention Mechanism
abstract
Accurate visualization of liver tumors and their surrounding blood vessels is essential for noninvasive diagnosis and prognosis prediction of tumors. In medical image segmentation, there is still a lack of in-depth research on the simultaneous segmentation of liver tumors and peritumoral blood vessels. To this end, we collect the first liver tumor, and vessel segmentation benchmark datasets containing 52 portal vein phase computed tomography images with liver, liver tumor, and vessel annotations. In this case, we propose a 3D U-shaped Cross-Attention Network (UCA-Net) that utilizes a tailored cross-attention mechanism instead of the traditional skip connection to effectively model the encoder and decoder feature. Specifically, the UCA-Net uses a channel-wise cross-attention module to reduce the semantic gap between encoder and decoder and a slice-wise cross-attention module to enhance the contextual semantic learning ability among distinct slices. Experimental results show that the proposed UCA-Net can accurately segment 3D medical images and achieve state-of-the-art performance on the liver tumor and intrahepatic vessel segmentation task.
Haopeng Kuang, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002
ICASSP2
2023 MSN-net: Multi-Scale Normality Network for Video Anomaly Detection
abstract
Existing unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories to learn multi-level prototypical spatial-temporal patterns of normal events. Specifically, the hierarchical memory module interacts with the encoder through the reading and writing operations during the training phase, preserving multi-scale normality in three separate memory pools. Then, the decoder decodes the features rewritten by the memorized normality to predict future frames so that its ability to predict anomalies is diminished. Experimental results show that MSN-net performs comparably to the state-of-the-art methods, and extension analysis demonstrates the effectiveness of multi-scale normality learning.
Yang Liu 0246, Dingkang Yang, Jing Liu 0050
ICASSP4
2023 Adversarial Contrastive Distillation with Adaptive Denoising
abstract
Adversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not always make correct predictions, interfering with the student’s robust performance. Besides, in the previous ARD methods, the robustness comes entirely from one-to-one imitation, ignoring the relationship between examples. To this end, we propose a novel structured ARD method called Contrastive Relationship DeNoise Distillation (CRDND). We design an adaptive compensation module to model the instability of the teacher. Moreover, we utilize the contrastive relationship to explore implicit robustness knowledge among multiple examples. Experimental results on multiple attack benchmarks show CRDND can transfer robust knowledge efficiently and achieves state-of-the-art performance.
Zhaoyu Chen 0001, Dingkang Yang, Yang Liu 0246, Siao Liu, Lizhe Qi
ICASSP3
2023 A Novel Efficient Multi-View Traffic-Related Object Detection Framework
abstract
With the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Accordingly, we propose a novel traffic-related framework named CEVAS to achieve efficient object detection using multi-view video data. Briefly, a fine-grained input filtering policy is introduced to produce a reasonable region of interest from the captured images. Also, we design a sharing object manager to manage the information of objects with spatial redundancy and share their results with other vehicles. We further derive a content-aware model selection policy to select detection methods adaptively. Experimental results show that our framework significantly reduces response latency while achieving the same detection accuracy as the state-of-the-art methods.
Kun Yang 0010, Jing Liu 0050, Dingkang Yang, Hanqi Wang, Peng Sun 0007
ICASSP3
2023 D-CONFORMER: Deformable Sparse Transformer Augmented Convolution for Voxel-Based 3D Object Detection
abstract
Although CNN-based and Transformer-based detectors have made impressive improvements in 3D object detection, these two network paradigms suffer from the interference of insufficient receptive field and local detail weakening, which significantly limits the feature extraction performance of the backbone. In this paper, we propose to fuse convolution and transformer, and simultaneously considering the different contributions of non-empty voxels at different positions in 3D space to object detection, it is not consistent with applying standard convolution and transformer directly on voxels. Specifically, we design a novel deformable sparse transformer to perform long-range information interaction on fine-grained local detail semantics aggregated by focal sparse convolution, termed D-Conformer. D-Conformer learns valuable voxels with position-wise in sparse space and can be applied to most voxel-based detectors as a backbone. Extensive experiments demonstrate that our method achieves satisfactory detection results and outperforms state-of-the-art 3D detection methods by a large margin.
Liuzhen Su, Xukun Zhang, Dingkang Yang, Shunli Wang 0001, Peng Zhai, Lihua Zhang 0002
ICASSP4
2023 Efficient Decision-based Black-box Patch Attacks on Video Recognition
abstract
Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have not been well investigated. Further, decision-based attacks, where attackers only access the predicted hard labels by querying threat models, have not been well explored on video models either, even if they are practical in real-world video recognition scenes. The absence of such studies leads to a huge gap in the robustness assessment for video models. To bridge this gap, this work first explores decision-based patch attacks on video models. We analyze that the huge parameter space brought by videos and the minimal information returned by decision-based models both greatly increase the attack difficulty and query burden. To achieve a query-efficient attack, we propose a spatial-temporal differential evolution (STDE) framework. First, STDE introduces target videos as patch textures and only adds patches on keyframes that are adaptively selected by temporal difference. Second, STDE takes minimizing the patch area as the optimization objective and adopts spatial-temporal mutation and crossover to search for the global optimum without falling into the local optimum. Experiments show STDE has demonstrated state-of-the-art performance in terms of threat, efficiency and imperceptibility. Hence, STDE has the potential to be a powerful tool for evaluating the robustness of video recognition models.
Kaixun Jiang, Zhaoyu Chen 0001, Dingkang Yang, Bo Li 0115, Yan Wang 0068
ICCV5
2023 Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation
abstract
Learning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficiency, suffering from serve performance degradation. In this paper, we first conduct qualitative analysis and illuminate the main causes: (i) high-variance gradient magnitudes and (ii) gradient conflicts existed in various augmentation methods. To alleviate these issues, we propose a general policy gradient optimization framework, named Conflict-aware Gradient Agreement Augmentation (CG2A), and better integrate augmentation combination into visual RL algorithms to address the generalization bias. In particular, CG2A develops a Gradient Agreement Solver to adaptively balance the varying gradient magnitudes, and introduces a Soft Gradient Surgery strategy to alleviate the gradient conflicts. Extensive experiments demonstrate that CG2A significantly improves the generalization performance and sample efficiency of visual RL algorithms.
Siao Liu, Zhaoyu Chen 0001, Yang Liu 0246, Dingkang Yang, Zhile Zhao, Ziqing Zhou, Xie Yi, Wei Li 0055, Zhongxue Gan 0001
ICCV5
2023 AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception
abstract
Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE.
Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002
ICCV1
2023 Spatio-Temporal Domain Awareness for Multi-Agent Collaborative Perception
abstract
Multi-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/.
Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007
ICCV2
2023 HandGCAT: Occlusion-Robust 3D Hand Mesh Reconstruction from Monocular Images
abstract
We propose a robust and accurate method for reconstructing 3D hand mesh from monocular images. This is a very challenging problem, as hands are often severely occluded by objects. Previous works often have disregarded 2D hand pose information, which contains hand prior knowledge that is strongly correlated with occluded regions. Thus, in this work, we propose a novel 3D hand mesh reconstruction network HandGCAT, that can fully exploit hand prior as compensation information to enhance occluded region features. Specifically, we designed the Knowledge-Guided Graph Convolution (KGC) module and the Cross-Attention Transformer (CAT) module. KGC extracts hand prior information from 2D hand pose by graph convolution. CAT fuses hand prior into occluded regions by considering their high correlation. Extensive experiments on popular datasets with challenging hand-object occlusions, such as HO3D v2, HO3D v3, and DexYCB demonstrate that our HandGCAT reaches state-of-the-art performance. The code is available at https://github.com/heartStrive/HandGCAT.
Shuaibing Wang, Shunli Wang 0001, Dingkang Yang, Mingcheng Li, Ziyun Qian, Liuzhen Su, Lihua Zhang 0002
ICME3
2023 FE-YOLOv5: Improved YOLOv5 Network for Multi-scale Drone-Captured Scene Detection
Zhiyan Dong, Dingkang Yang, Lihua Zhang 0002
ICONIP (2)4
2023 What2comm: Towards Communication-efficient Collaborative Perception via Feature Decoupling
abstract
Multi-agent collaborative perception has received increasing attention recently as an emerging application in driving scenarios. Despite advancements in previous approaches, challenges remain due to redundant communication patterns and vulnerable collaboration processes. To address these issues, we propose What2comm, an end-to-end collaborative perception framework to achieve a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we design an efficient communication mechanism based on feature decoupling to transmit exclusive and common feature maps among heterogeneous agents to provide perceptually holistic messages. Secondly, a spatio-temporal collaboration module is introduced to integrate complementary information from collaborators and temporal ego cues, leading to a robust collaboration procedure against transmission delay and localization errors. Ultimately, we propose a common-aware fusion strategy to refine final representations with informative common features. Comprehensive experiments in real-world and simulated scenarios demonstrate the effectiveness of What2comm.
Kun Yang 0010, Dingkang Yang, Hanqi Wang, Peng Sun 0007
ACM Multimedia2
2023 How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception
abstract
Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https://github.com/ydk122024/How2comm.
Dingkang Yang, Kun Yang 0010, Jing Liu 0050, Zhi Xu 0010, Rongbin Yin, Peng Zhai, Lihua Zhang 0002
NeurIPS1
2023 Stochastic video normality network for abnormal event detection in surveillance videos
Yang Liu 0246, Dingkang Yang, Gaoyun Fang, Donglai Wei 0002, Mengyang Zhao 0002, Kai Cheng 0001, Jing Liu 0050
Knowl. Based Syst.2
2023 Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002
Knowl. Based Syst.1
2023 Towards Robust Multimodal Sentiment Analysis Under Uncertain Signal Missing
abstract
Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of signal completeness. However, many inevitable factors in real applications lead to uncertain signal missing, causing significant degradation of model performance. To this end, we propose a Robust multimodal Missing Signal Framework (RMSF) to handle the problem of uncertain signal missing for MSA tasks and can be generalized to other multimodal patterns. Specifically, a hierarchical cross modal interaction module in RMSF exploits potential complementary semantics among modalities via coarse- and fine-grained cross modal attention. Furthermore, we design an adaptive feature refinement module to enhance the beneficial semantics of modalities and filter redundant features. Finally, we propose a knowledge integrated self-distillation module that enables dynamic knowledge integration and bidirectional knowledge transfer within a single network to precisely reconstruct missing semantics. Comprehensive experiments are conducted on two datasets, indicating that RMSF significantly improves MSA performance under both uncertain missing-signal and complete-signal cases.
Mingcheng Li, Dingkang Yang, Lihua Zhang 0002
IEEE Signal Process. Lett.2
2022 Robust Adversarial Reinforcement Learning with Dissipation Inequation Constraint
abstract
Robust adversarial reinforcement learning is an effective method to train agents to manage uncertain disturbance and modeling errors in real environments. However, for systems that are sensitive to disturbances or those that are difficult to stabilize, it is easier to learn a powerful adversary than establish a stable control policy. An improper strong adversary can destabilize the system, introduce biases in the sampling process, make the learning process unstable, and even reduce the robustness of the policy. In this study, we consider the problem of ensuring system stability during training in the adversarial reinforcement learning architecture. The dissipative principle of robust H-infinity control is extended to the Markov Decision Process, and robust stability constraints are obtained based on L2 gain performance in the reinforcement learning system. Thus, we propose a dissipation-inequation-constraint-based adversarial reinforcement learning architecture. This architecture ensures the stability of the system during training by imposing constraints on the normal and adversarial agents. Theoretically, this architecture can be applied to a large family of deep reinforcement learning algorithms. Results of experiments in MuJoCo and GymFc environments show that our architecture effectively improves the robustness of the controller against environmental changes and adapts to more powerful adversaries. Results of the flight experiments on a real quadcopter indicate that our method can directly deploy the policy trained in the simulation environment to the real environment, and our controller outperforms the PID controller based on hardware-in-the-loop. Both our theoretical and empirical results provide new and critical outlooks on the adversarial reinforcement learning architecture from a rigorous robust control perspective.
Peng Zhai, Zhiyan Dong, Lihua Zhang 0002, Shunli Wang 0001, Dingkang Yang
AAAI6
2022 Emotion Recognition for Multiple Context Awareness
Dingkang Yang, Shunli Wang 0001, Yang Liu 0246, Peng Zhai, Liuzhen Su, Mingcheng Li, Lihua Zhang 0002
ECCV (37)1
2022 Learning Appearance-Motion Normality for Video Anomaly Detection
abstract
Video anomaly detection is a challenging task in the Computer vision community. Most single task-based methods do not consider the independence of unique spatial and temporal patterns, while two-stream structures lack the exploration of the correlations. In this paper, we propose spatial-temporal memories augmented two-stream auto-encoder framework, which learns the appearance normality and motion normal-ity independently and explores the correlations via adversar-ial learning. Specifically, we first design two proxy tasks to train the two-stream structure to extract appearance and motion features in isolation. Then, the prototypical features are recorded in the corresponding spatial and temporal memory pools. Finally, the encoding-decoding network performs ad-versariallearning with the discriminator to explore the corre-lations between spatial and temporal patterns. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 98.1% and 89.8% on UCSD Ped2 and CUHK Avenue datasets.
Yang Liu 0246, Jing Liu 0050, Mengyang Zhao 0002, Dingkang Yang, Xiaoguang Zhu
ICME4
2022 CA-SpaceNet: Counterfactual Analysis for 6D Pose Estimation in Space
abstract
Reliable and stable 6D pose estimation of un-cooperative space objects plays an essential role in on-orbit servicing and debris removal missions. Considering that the pose estimator is sensitive to background interference, this paper proposes a counterfactual analysis framework named CA-SpaceNet to complete robust 6D pose estimation of the space-borne targets under complicated background. Specifically, conventional methods are adopted to extract the features of the whole image in the factual case. In the counterfactual case, a non-existent image without the target but only the background is imagined. Side effect caused by background interference is reduced by counterfactual analysis, which leads to unbiased prediction in final results. In addition, we also carry out low-bit-width quantization for CA-SpaceNet and deploy part of the framework to a Processing-In-Memory (PIM) accelerator on FPGA. Qualitative and quantitative results demonstrate the effectiveness and efficiency of our proposed method. To our best knowledge, this paper applies causal inference and network quantization to the 6D pose estimation of space-borne targets for the first time. The code is available at https://github.com/Shunli-Wang/CA-SpaceNet.
Shunli Wang 0001, Shuaibing Wang, Bo Jiao 0003, Dingkang Yang, Liuzhen Su, Peng Zhai, Chixiao Chen, Lihua Zhang 0002
IROS4
2022 Disentangled Representation Learning for Multimodal Emotion Recognition
abstract
Multimodal emotion recognition aims to identify human emotions from text, audio, and visual modalities. Previous methods either explore correlations between different modalities or design sophisticated fusion strategies. However, the serious problem is that the distribution gap and information redundancy often exist across heterogeneous modalities, resulting in learned multimodal representations that may be unrefined. Motivated by these observations, we propose a Feature-Disentangled Multimodal Emotion Recognition (FDMER) method, which learns the common and private feature representations for each modality. Specifically, we design the common and private encoders to project each modality into modality-invariant and modality-specific subspaces, respectively. The modality-invariant subspace aims to explore the commonality among different modalities and reduce the distribution gap sufficiently. The modality-specific subspaces attempt to enhance the diversity and capture the unique characteristics of each modality. After that, a modality discriminator is introduced to guide the parameter learning of the common and private encoders in an adversarial manner. We achieve the modality consistency and disparity constraints by designing tailored losses for the above subspaces. Furthermore, we present a cross-modal attention fusion module to learn adaptive weights for obtaining effective multimodal representations. The final representation is used for different downstream tasks. Experimental results show that the FDMER outperforms the state-of-the-art methods on two multimodal emotion recognition benchmarks. Moreover, we further verify the effectiveness of our model via experiments on the multimodal humor detection task.
Dingkang Yang, Haopeng Kuang, Yangtao Du, Lihua Zhang 0002
ACM Multimedia1
2022 Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language Sequences
abstract
Understanding human behaviors and intents from videos is a challenging task. Video flows usually involve time-series data from different modalities, such as natural language, facial gestures, and acoustic information. Due to the variable receiving frequency for sequences from each modality, the collected multimodal streams are usually unaligned. For multimodal fusion of asynchronous sequences, the existing methods focus on projecting multiple modalities into a common latent space and learning the hybrid representations, which neglects the diversity of each modality and the commonality across different modalities. Motivated by this observation, we propose a Multimodal Fusion approach for learning modality-Specific and modality-Agnostic representations (MFSA) to refine multimodal representations and leverage the complementarity across different modalities. Specifically, a predictive self-attention module is used to capture reliable contextual dependencies and enhance the unique features over the modality-specific spaces. Meanwhile, we propose a hierarchical cross-modal attention module to explore the correlations between cross-modal elements over the modality-agnostic space. In this case, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, the modality-specific and -agnostic multimodal representations are used together for downstream tasks. Comprehensive experiments on three multimodal datasets clearly demonstrate the superiority of our approach.
Dingkang Yang, Haopeng Kuang, Lihua Zhang 0002
ACM Multimedia1
2022 Contextual and Cross-Modal Interaction for Multi-Modal Speech Emotion Recognition
abstract
Speech emotion recognition combining linguistic content and audio signals in the dialog is a challenging task. Nevertheless, previous approaches have failed to explore emotion cues in contextual interactions and ignored the long-range dependencies between elements from different modalities. To tackle the above issues, this letter proposes a multimodal speech emotion recognition method using audio and text data. We first present a contextual transformer module to introduce contextual information via embedding the previous utterances between interlocutors, which enhances the emotion representation of the current utterance. Then, the proposed cross-modal transformer module focuses on the interactions between text and audio modalities, adaptively promoting the fusion from one modality to another. Furthermore, we construct associative topological relation over mini-batch and learn the association between deep fused features with graph convolutional network. Experimental results on the IEMOCAP and MELD datasets show that our method outperforms current state-of-the-art methods.
Dingkang Yang, Yang Liu 0246, Lihua Zhang 0002
IEEE Signal Process. Lett.1
2021 Learning Associative Representation for Facial Expression Recognition
abstract
The main inherent challenges with the Facial Expression Recognition (FER) are high intra-class variations and high inter-class similarities, while existing methods pay little attention to the association within inter- and intra-class expressions. This paper introduces a novel Expression Associative Network (EAN) to learn association of facial expression, specifically, from two aspects: 1) associative topological relation over mini-batch is constructed by similarity matrix with an adjacent regularization, and 2) learning association of expressions with Graph Convolutional Network (GCN). Besides, an auxiliary module as invariant feature generator based on Generative Adversarial Networks (GAN) is designed to suppress pose variations, illumination changes, and occlusions. Results on public benchmarks achieve comparable or better performance compared with current state-of-the-art methods, with 90.07% on FERPlus, 86.36% on RAF-DB, and improve by 3.92% over SOTA on synthetic wrong labeling datasets.
Yangtao Du, Dingkang Yang, Peng Zhai, Lihua Zhang 0002
ICIP2
2021 TSA-Net: Tube Self-Attention Network for Action Quality Assessment
abstract
In recent years, assessing action quality from videos has attracted growing attention in computer vision community and human-computer interaction. Most existing approaches usually tackle this problem by directly migrating the model from action recognition tasks, which ignores the intrinsic differences within the feature map such as foreground and background information. To address this issue, we propose a Tube Self-Attention Network (TSA-Net) for action quality assessment (AQA). Specifically, we introduce a single object tracker into AQA and propose the Tube Self-Attention Module (TSA), which can efficiently generate rich spatio-temporal contextual information by adopting sparse feature interactions. The TSA module is embedded in existing video networks to form TSA-Net. Overall, our TSA-Net is with the following merits: 1) High computational efficiency, 2) High flexibility, and 3) The state-of-the-art performance. Extensive experiments are conducted on popular action quality assessment datasets including AQA-7 and MTL-AQA. Besides, a dataset named Fall Recognition in Figure Skating (FR-FS) is proposed to explore the basic action assessment in the figure skating scene. Our TSA-Net achieves the Spearman's Rank Correlation of 0.8476 and 0.9393 on AQA-7 and MTL-AQA, respectively, which are the new state-of-the-art results. The results on FR-FS also verify the effectiveness of the TSA-Net. The code and FR-FS dataset are publicly available at https://github.com/Shunli-Wang/TSA-Net.
Shunli Wang 0001, Dingkang Yang, Peng Zhai, Chixiao Chen, Lihua Zhang 0002
ACM Multimedia2