EDBT 2026 Demo / reviewers in the wild / expert
Haopeng Kuang
dblp:320/1809
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-7951-2380ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards unified molecule-enhanced pathology image representation learning via integrating spatial transcriptomics
Minghao Han, Dingkang Yang, Jiabei Cheng, Xukun Zhang, Zizhi Chen, Haopeng Kuang, Lihua Zhang 0002 |
Pattern Recognit. | 6 |
| 2024 | CPR-Coach: Recognizing Composite Error Actions Based on Single-Class TrainingabstractFine- grained medical action analysis plays a vital role in improving medical skill training efficiency, but it faces the problems of data and algorithm shortage. Cardiopul-monary Resuscitation (CPR) is an essential skill in emer-gency treatment. Currently, the assessment of CPR skills mainly depends on dummies and trainers, leading to high training costs and low efficiency. For the first time, this pa-per constructs a vision-based system to complete error action recognition and skill assessment in CPR. Specifically, we define 13 types of single-error actions and 74 types of composite error actions during external cardiac compres-sion and then develop a video dataset named CPR-Coach. By taking the CPR-Coach as a benchmark, this paper in-vestigates and compares the performance of existing action recognition models based on different data modalities. To solve the unavoidable “Single-class Training & Multi-class Testing” problem, we propose a human-cognition-inspired framework named ImagineNet to improve the model's multi-error recognition performance under restricted supervision. Extensive comparison and actual deployment experiments verify the effectiveness of the framework. We hope this work could bring new inspiration to the computer vision and medical skills training communities simultaneously. The dataset and the code are publicly available on https://github.com/Shunli-Wang/CPR-Coach. Shunli Wang 0001, Shuaibing Wang, Dingkang Yang, Mingcheng Li, Haopeng Kuang, Liuzhen Su, Peng Zhai, Lihua Zhang 0002 |
CVPR | 5 |
| 2024 | Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide ImagesabstractSurvival prediction is a complex ordinal regression task that aims to predict the survival coefficient ranking among a cohort of patients, typically achieved by analyzing patients’ whole slide images. Existing deep learning approaches mainly adopt multiple instance learning or graph neural networks under weak supervision. Most of them are unable to uncover the diverse interactions between different types of biological entities(e.g., cell cluster and tissue block) across multiple scales, while such interactions are crucial for patient survival prediction. In light of this, we propose a novel multi-scale heterogeneity-aware hypergraph representation framework. Specifically, our framework first constructs a multi-scale heterogeneity-aware hypergraph and assigns each node with its biological entity type. It then mines diverse interactions between nodes on the graph structure to obtain a global representation. Experimental results demonstrate that our method outperforms state-of-the-art approaches on three benchmark datasets. Code is publicly available at https://github.com/Hanminghao/H2GT. Minghao Han, Xukun Zhang, Dingkang Yang, Tao Liu 0050, Haopeng Kuang, Jinghui Feng, Lihua Zhang 0002 |
ICME | 5 |
| 2024 | Dual knowledge-guided two-stage model for precise small organ segmentation in abdominal CT imagesabstractAbstract Multi‐organ segmentation from abdominal CT scans is crucial for various medical examinations and diagnoses. Despite the remarkable achievements of existing deep‐learning‐based methods, accurately segmenting small organs remains challenging due to their small size and low contrast. This article introduces a novel knowledge‐guided cascaded framework that utilizes two types of knowledge—image intrinsic (anatomy) and clinical expertise (radiology)—to improve the segmentation accuracy of small abdominal organs. Specifically, based on the anatomical similarities in abdominal CT scans, the approach employs entropy‐based registration techniques to map high‐quality segmentation results onto inaccurate results from the first stage, thereby guiding precise localization of small organs. Additionally, inspired by the practice of annotating images from multiple perspectives by radiologists, novel Multi‐View Fusion Convolution (MVFC) operator is developed, which can extract and adaptively fuse features from various directions of CT images to refine segmentation of small organs effectively. Simultaneously, the MVFC operator offers a seamless alternative to conventional convolutions within diverse model architectures. Extensive experiments on the Abdominal Multi‐Organ Segmentation (AMOS) dataset demonstrate the superiority of the method, setting a new benchmark in the segmentation of small organs. Tao Liu 0050, Xukun Zhang, Zhongwei Yang, Minghao Han, Haopeng Kuang, Shuwei Ma, Lihua Zhang 0002 |
IET Image Process. | 5 |
| 2024 | Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded TrainingabstractUnderstanding emotions from diverse contexts has received widespread attention in computer vision communities. The core philosophy of Context-Aware Emotion Recognition (CAER) is to provide valuable semantic cues for recognizing the emotions of target persons by leveraging rich contextual information. Current approaches invariably focus on designing sophisticated structures to extract perceptually critical representations from contexts. Nevertheless, a long-neglected dilemma is that a severe context bias in existing datasets results in an unbalanced distribution of emotional states among different contexts, causing biased visual representation learning. From a causal demystification perspective, the harmful bias is identified as a confounder that misleads existing models to learn spurious correlations based on likelihood estimation, limiting the models' performance. To address the issue, we embrace causal inference to disentangle the models from the impact of such bias, and formulate the causalities among variables in the CAER task via a customized causal graph. Subsequently, we present a Contextual Causal Intervention Module (CCIM) to de-confound the confounder, which is built upon backdoor adjustment theory to facilitate seeking approximate causal effects during model training. As a plug-and-play component, CCIM can easily integrate with existing approaches and bring significant improvements. Systematic experiments on three datasets demonstrate the effectiveness of our CCIM. Dingkang Yang, Kun Yang 0010, Haopeng Kuang, Zhaoyu Chen 0001, Lihua Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Adaptive Multiphase Liver Tumor Segmentation With Multiscale SupervisionabstractThe segmentation of liver tumors using multi-phase computed tomography (CT) images has garnered considerable attention in medical signal processing. However, existing multi-phase liver tumor segmentation methods primarily concentrate on feature integration across various phases, neglecting a comprehensive exploration of synergistic relationships among these phases and constraints on features across different scales. This limitation has led to performance bottlenecks in existing approaches. This article proposes a robust multi-phase liver tumor segmentation framework designed to address the aforementioned challenges. Specifically, we introduce a novel multi-phase and channel-stacked dual attention module, seamlessly integrated within a multi-scale architecture. This module adaptively captures essential semantic information among different phases, enhancing the segmentation network's feature extraction capabilities. A scale-weighted loss function for multi-scale supervision is also designed to mitigate false positives in the segmentation results. To facilitate a systematic evaluation of our model's performance on multi-phase data, we curate a new dataset comprising samples from four distinct phases. Our proposed framework is rigorously assessed through comprehensive quantitative and qualitative experiments, highlighting its compelling performance. Haopeng Kuang, Xue Yang 0013, Jingwei Wei, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Towards Asynchronous Multimodal Signal Interaction and Fusion via Tailored TransformersabstractThe signals from human expressions are usually multimodal, including natural language, facial gestures, and acoustic behaviors. A key challenge is how to fuse multimodal time-series signals with temporal asynchrony. To this end, we present a Transformer-driven Signal Interaction and Fusion (TSIF) approach to effectively model asynchronous multimodal signal sequences. TSIF consists of linear and cross-modal transformer modules with different duties. The linear transformer module efficiently performs the global interaction for multimodal signals, and the vital philosophy is to replace the dot product similarity with the Exponential Kernel while achieving linear complexity by a low-rank matrix decomposition. By targeting the language modality, the cross-modal transformer module aims to capture reliable element correlations among distinct signals and mitigate noise interference in audio and visual modalities. Numerous experiments on two multimodal benchmarks show that our TSIF comparably outperforms previous state-of-the-art models with lower space-time complexities. The systematic analysis also proves the effectiveness of the proposed modules. Dingkang Yang, Haopeng Kuang, Kun Yang 0010, Mingcheng Li, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Towards Simultaneous Segmentation Of Liver Tumors And Intrahepatic Vessels Via Cross-Attention MechanismabstractAccurate visualization of liver tumors and their surrounding blood vessels is essential for noninvasive diagnosis and prognosis prediction of tumors. In medical image segmentation, there is still a lack of in-depth research on the simultaneous segmentation of liver tumors and peritumoral blood vessels. To this end, we collect the first liver tumor, and vessel segmentation benchmark datasets containing 52 portal vein phase computed tomography images with liver, liver tumor, and vessel annotations. In this case, we propose a 3D U-shaped Cross-Attention Network (UCA-Net) that utilizes a tailored cross-attention mechanism instead of the traditional skip connection to effectively model the encoder and decoder feature. Specifically, the UCA-Net uses a channel-wise cross-attention module to reduce the semantic gap between encoder and decoder and a slice-wise cross-attention module to enhance the contextual semantic learning ability among distinct slices. Experimental results show that the proposed UCA-Net can accurately segment 3D medical images and achieve state-of-the-art performance on the liver tumor and intrahepatic vessel segmentation task. Haopeng Kuang, Dingkang Yang, Shunli Wang 0001, Lihua Zhang 0002 |
ICASSP | 1 |
| 2022 | Disentangled Representation Learning for Multimodal Emotion RecognitionabstractMultimodal emotion recognition aims to identify human emotions from text, audio, and visual modalities. Previous methods either explore correlations between different modalities or design sophisticated fusion strategies. However, the serious problem is that the distribution gap and information redundancy often exist across heterogeneous modalities, resulting in learned multimodal representations that may be unrefined. Motivated by these observations, we propose a Feature-Disentangled Multimodal Emotion Recognition (FDMER) method, which learns the common and private feature representations for each modality. Specifically, we design the common and private encoders to project each modality into modality-invariant and modality-specific subspaces, respectively. The modality-invariant subspace aims to explore the commonality among different modalities and reduce the distribution gap sufficiently. The modality-specific subspaces attempt to enhance the diversity and capture the unique characteristics of each modality. After that, a modality discriminator is introduced to guide the parameter learning of the common and private encoders in an adversarial manner. We achieve the modality consistency and disparity constraints by designing tailored losses for the above subspaces. Furthermore, we present a cross-modal attention fusion module to learn adaptive weights for obtaining effective multimodal representations. The final representation is used for different downstream tasks. Experimental results show that the FDMER outperforms the state-of-the-art methods on two multimodal emotion recognition benchmarks. Moreover, we further verify the effectiveness of our model via experiments on the multimodal humor detection task. Dingkang Yang, Haopeng Kuang, Yangtao Du, Lihua Zhang 0002 |
ACM Multimedia | 3 |
| 2022 | Learning Modality-Specific and -Agnostic Representations for Asynchronous Multimodal Language SequencesabstractUnderstanding human behaviors and intents from videos is a challenging task. Video flows usually involve time-series data from different modalities, such as natural language, facial gestures, and acoustic information. Due to the variable receiving frequency for sequences from each modality, the collected multimodal streams are usually unaligned. For multimodal fusion of asynchronous sequences, the existing methods focus on projecting multiple modalities into a common latent space and learning the hybrid representations, which neglects the diversity of each modality and the commonality across different modalities. Motivated by this observation, we propose a Multimodal Fusion approach for learning modality-Specific and modality-Agnostic representations (MFSA) to refine multimodal representations and leverage the complementarity across different modalities. Specifically, a predictive self-attention module is used to capture reliable contextual dependencies and enhance the unique features over the modality-specific spaces. Meanwhile, we propose a hierarchical cross-modal attention module to explore the correlations between cross-modal elements over the modality-agnostic space. In this case, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, the modality-specific and -agnostic multimodal representations are used together for downstream tasks. Comprehensive experiments on three multimodal datasets clearly demonstrate the superiority of our approach. Dingkang Yang, Haopeng Kuang, Lihua Zhang 0002 |
ACM Multimedia | 2 |