EDBT 2026 Demo / reviewers in the wild / expert
Wenhui Huang 0002
dblp:389/9131-2
· DBLP profile ↗
24ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0002-5435-8775ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ABE-Mamba: Few-shot medical image segmentation via adversarial bidirectional enhanced Mamba
Bingjie Guo, Wenhui Huang 0002 |
Expert Syst. Appl. | 2 |
| 2026 | LSDiff: Diffusion-Guided Level Set for Low-Contrast Lesion Boundary Segmentation
Wenhui Huang 0002, Jing Wang 0138, James C. Gee, Yuanjie Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Text-Image Co-Alignment for Weakly Supervised Polyp SegmentationabstractFully supervised polyp segmentation relies on costly pixel-level annotations. Although semi- and weakly supervised methods reduce annotation requirements, they still depend on partial mask supervision. Text-supervised segmentation is a promising alternative; however, for polyps, the key challenge is to ground instance-specific phrases to the correct lesion region under cluttered backgrounds and large appearance variations. Existing approaches often rely on coarse text-image alignment, limiting precise region-level semantic correspondence. In this paper, we propose Text-Image Co-Alignment (TICoA), a text-supervised framework for polyp segmentation. TICoA leverages large language models (LLMs)-generated structured clinical descriptions as weak supervision and formulates segmentation as a fine-grained phrase-region co-alignment problem. Through contrastive learning, TICoA explicitly associates query phrases with corresponding image regions to achieve robust semantic grounding under weak supervision. Architecturally, we adopt a State-Space Model (Mamba) to efficiently model long-range dependencies with linear computational complexity. To support effective cross-modal interaction, we further design a dedicated Mamba Fusion module with a Bi-Dimension Fusion (BiDF) strategy, which progressively propagates information along spatial and channel dimensions. Experiments on polyp datasets, with additional validation on skin lesion segmentation, demonstrate that TICoA is competitive with state-of-the-art weakly supervised methods. Our code and data are available at https://github.com/silentyuchen/TICoA. Wenhui Huang 0002, Zhen Pan, Yedi Zhang, Jingzhen He, James C. Gee, Yuanjie Zheng |
IEEE Trans. Medical Imaging | 1 |
| 2025 | MAMBA-Based Weakly Supervised Medical Image Segmentation with Cross-Modal Textual Information
Zhen Pan, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (8) | 2 |
| 2025 | MambaMER: Adaptive EEG-Guided Multimodal Emotion Recognition with Mamba
Xiangle Ping, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (1) | 2 |
| 2025 | Adversarial Bidirectional Enhanced Mamba for Few-Shot Medical Image Segmentation
Bingjie Guo, Wenhui Huang 0002 |
PRCV (13) | 2 |
| 2025 | Interactive EEG Emotion Recognition with Incremental Gaussian ProcessesabstractInteractivity is crucial for enabling models to adjust and optimize based on user feedback, thereby enhancing overall performance. However, existing electroencephalogram (EEG)-based emotion recognition models rely on static training paradigms, lack interactivity, and struggle to effectively handle uncertainty in predictions. To address this issue, we propose a novel paradigm for interactive emotion recognition based on incremental Gaussian processes (GP). Unlike existing methods, our approach introduces an expert interaction mechanism to correct samples with high predictive uncertainty and incrementally update the model accordingly, thereby optimizing its performance. First, we model the emotion recognition task as a GP-based framework, utilizing the variance of the GP to quantify the model's uncertainty, thereby guiding experts in targeted interactions. Second, within the GP framework, we propose a novel incremental update strategy that allows the GP to incrementally update prediction results and uncertainties based only on new data obtained through expert interactions, without reprocessing all existing data. This effectively overcomes the shortcomings of traditional GP in updating efficiency. Third, to address the high computational complexity of GP, we use a sparse approximation strategy, selecting inducing points and performing variational inference to efficiently approximate the GP posterior, thereby reducing computational complexity. Subject-dependent and subject-independent experiments conducted on the DEAP and DREAMER datasets demonstrate that the proposed method exhibits significant advantages over state-of-the-art (SOTA) methods. In subject-dependent experiments, our method achieved the highest improvement (1.73%) in the Dominance dimension on the DREAMER dataset. In subject-independent experiments, it attained the largest performance improvement (2.96%) in the Arousal dimension on the DEAP dataset. These results further validate the proposed method's effectiveness. Xiangle Ping, Wenhui Huang 0002 |
Int. J. Neural Syst. | 2 |
| 2025 | Rethinking interactive image matting as incremental Gaussian process regression problems
Bingjie Guo, Wenhui Huang 0002 |
Knowl. Based Syst. | 2 |
| 2024 | Uncertainty Guided Incremental Interactive Medical Image Segmentation with Sparse Variational Gaussian ProcessabstractMedical image segmentation often relies on fully automated methods, but achieving the necessary accuracy in clinical settings can be challenging. Interactive segmentation methods attempt to improve accuracy by incorporating user input in areas of inaccurate segmentation. However, these methods depend on the user’s subjective judgment, without considering the uncertainty in model predictions. To address these issues, we propose an incremental interactive medical image segmentation model guided by uncertainty. First, our proposed model combines medical image segmentation with a Gaussian process (GP), models medical image segmentation as a binary pixel classification problem based on a GP, and guides the selection of the interaction location by the magnitude of the variance value of the GP to reduce the uncertainty of the model prediction. Second, we introduce the sparse variational GP (SVGP), and we propose an incremental SVGP updating paradigm so that the SVGP can incrementally update its parameters by learning only the new interaction data without having to relearn all the old data. In addition, we propose a new induced point selection strategy and combine it with SVGP to reduce the complexity of the Gaussian process for calculating the covariance matrix. We trained and tested our model on the Medical Segmentation Decathlon’s three datasets of lung, pancreas, and colon, validating its generalization performance using the ISIC2018, CVC-ClinicDB, and KiTS19 datasets. Zhen Pan, Mingquan Jin, Maoling Qin, Wenhui Huang 0002 |
BIBM | 5 |
| 2024 | Uncertainty-Guided Incremental Interactive EEG Emotion Recognition with Gaussian ProcessabstractInteractivity is crucial for enabling models to adjust and optimize based on user feedback, thereby enhancing overall performance. However, existing methods for emotion recognition lack interactivity, failing to allow experts to provide feedback for interaction and updating prediction results, thus limiting the improvement of recognition accuracy. To address this problem, we propose a new paradigm for uncertainty-guided incremental interactive emotion recognition based on Gaussian process (GP). This paradigm involves interaction by having experts provide correction labels for samples with high prediction uncertainty, enabling incremental updates. Firstly, we model the emotion recognition task as a GP-based framework, utilizing the variance of GP to quantify the model’s uncertainty, thereby guiding experts in targeted interactions. Secondly, within the GP framework, we use an incremental update strategy to optimize prediction results. This method acquires new data through expert interaction, allowing for incremental updates based on new data without reprocessing all existing data. Thirdly, to tackle the computational complexity of inverting the covariance matrix in GP, we use a sparse approximation strategy, selecting inducing points and performing variational inference to efficiently approximate the GP posterior, thereby reducing computational complexity. Our method has been validated through extensive experiments on the DEAP and DREAMER datasets, demonstrating superior recognition performance compared to state-of-the-art methods. Xiangle Ping, Mingquan Jin, Maoling Qin, Wenhui Huang 0002 |
BIBM | 5 |
| 2024 | HiDiffSeg: A hierarchical diffusion model for blood vessel segmentation in retinal fundus images
Wenhui Huang 0002, Fengting Liu |
Expert Syst. Appl. | 1 |
| 2024 | Unsupervised EEG-Based Seizure Anomaly Detection with Denoising Diffusion Probabilistic ModelsabstractWhile many seizure detection methods have demonstrated great accuracy, their training necessitates a substantial volume of labeled data. To address this issue, we propose a novel method for unsupervised seizure anomaly detection called SAnoDDPM, which uses denoising diffusion probabilistic models (DDPM). We designed a novel pipeline that uses a variable lower bound on Markov chains to identify potential values that are unlikely to occur in anomalous data. The model is first trained on normal data, then anomalous data is input to the trained model. The model resamples the anomalous data and converts it to normal data. Finally, the presence of seizures can be determined by comparing the before and after data. Moreover, the input 2D spectrograms are encoded into vector-quantized representations, which enables powerful and efficient DDPM while maintaining its quality. Experimental comparisons on the publicly available datasets, CHB-MIT and TUH, show that our method delivers better results, significantly reduces inference time, and is suitable for deployment in a clinical environments. As far as we are aware, this is the first DDPM-based method for seizure anomaly detection. This novel approach significantly contributes to the progression of seizure detection algorithms, thereby augmenting their practicality in clinical settings. Mengxue Sun, Wenhui Huang 0002 |
Int. J. Neural Syst. | 3 |
| 2024 | Multi-spectral transformer with attention fusion for diabetic macular edema classification in multicolor image
Jingzhen He, Jingqi Song, Zeyu Han, Baojun Li, Qingtao Gong, Wenhui Huang 0002 |
Soft Comput. | 7 |
| 2024 | CorrDiff: Corrective Diffusion Model for Accurate MRI Brain Tumor SegmentationabstractAccurate segmentation of brain tumors in MRI images is imperative for precise clinical diagnosis and treatment. However, existing medical image segmentation methods exhibit errors, which can be categorized into two types: random errors and systematic errors. Random errors, arising from various unpredictable effects, pose challenges in terms of detection and correction. Conversely, systematic errors, attributable to systematic effects, can be effectively addressed through machine learning techniques. In this paper, we propose a corrective diffusion model for accurate MRI brain tumor segmentation by correcting systematic errors. This marks the first application of the diffusion model for correcting systematic segmentation errors. Additionally, we introduce the Vector Quantized Variational Autoencoder (VQ-VAE) to compress the original data into a discrete coding codebook. This not only reduces the dimensionality of the training data but also enhances the stability of the correction diffusion model. Furthermore, we propose the Multi-Fusion Attention Mechanism, which can effectively enhances the segmentation performance of brain tumor images, and enhance the flexibility and reliability of the corrective diffusion model. Our model is evaluated on the BRATS2019, BRATS2020, and Jun Cheng datasets. Experimental results demonstrate the effectiveness of our model over state-of-the-art methods in brain tumor segmentation. Wenhui Huang 0002, Yuanjie Zheng |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Prostate MRI Super-Resolution using Discrete Residual Diffusion ModelabstractProstate cancer (PCa) is one of the most common malignant tumors. High-resolution magnetic resonance imaging (HR MRI) is an effective tool for diagnosing PCa, but it requires patients to remain immobile for extended periods, increasing chances of image distortion due to motion. One solution is to utilize super-resolution (SR) techniques to create a higher-resolution MRI. However, existing medical SR models suffer from issues such as excessive smoothness and mode collapse. In this paper, we propose a novel generative model avoiding the problems, called Prostate MRI Super-Resolution using Discrete Residual Diffusion Model (DR-DM). First, the forward process of DR-DM gradually disrupts the input via a fixed Markov chain, producing a sequence of latent variables. The backward process optimizes a variant of the variational lower bound, training diffusion models effectively address the mode collapse. Second, to focus DR-DM on recovering high-frequency details, we synthesize residual images instead of synthesizing HR MRI directly. The residual image represents the difference between the HR and LR up-sampled MR image, and we convert residual image into discrete image tokens with a shorter sequence length by a vector quantized variational autoencoder (VQ-VAE), which reduced the computational complexity. Third, transformer architecture is integrated to model the relationship between LR MRI and residual image, which can capture the long-range dependencies between LR MRI and the synthesized imaging, thereby improving the fidelity of the reconstructed images. Our experiments on the Prostate-Diagnosis and PROSTATEx datasets demonstrate that the DR-DM model significantly improves image quality, resulting in greater clarity and improved diagnostic accuracy for patients. Zhitao Han, Wenhui Huang 0002 |
BIBM | 2 |
| 2023 | Joint Optic Disc and Cup Segmentation with Parallel Cooperative Diffusion ModelabstractGlaucoma is a common eye disease that can lead to permanent vision loss. Accurate segmentation of the optic disc (OD) and optic cup (OC) plays a crucial role in glaucoma screening and assessment. However, existing methods predominantly focus on separate segmentation of the OD and OC, neglecting the interrelationship between the two structures. Furthermore, due to the optic cup’s indistinct boundary, most existing methods fail to generate accurate OC region segmentation from fundus images, resulting in errors in cup-to-disc ratio (CDR) measurements. Therefore, we propose a parallel cooperative diffusion model, to improve the segmentation accuracy through joint segmentation of OD and OC. This is the first application of diffusion models in the joint segmentation of the OD and OC. We integrate a denoising autoencoder with coupling function into the diffusion model to enhance its performance and generalization ability. Additionally, we proposed a joint U-Net to simultaneously learn the joint distribution and semantic correlation of the OD and OC, thereby improving the segmentation accuracy and robustness of the results. Furthermore, we introduce a cross-attention block to bridge the two subnets of the OD and OC, achieving alignment of the segmentation results. The proposed model is evaluated on publicly available datasets including DrishtiGS, RIM-ONE (r3), and REFUGE. Experimental results demonstrate that our model outperforms state-of-the-art segmentation methods in OC and OD segmentation performance. Wenhui Huang 0002 |
BIBM | 2 |
| 2023 | A Diffusion Model-Based Joint Dual-Task Network for Low-Quality Retinal Image Enhancement and Vessel SegmentationabstractIn clinical screening, precise diagnosis of different diseases depends on extracting blood vessel information from fundus images. However, clinical fundus images often suffer from uneven illumination, blur, and artifacts caused by equipment or environmental factors. These problems can result in missed small blood vessels and vascular discontinuity, which hinder reliable diagnosis by optometrists or computer-aided systems. In this manuscript, we propose a joint framework, called ESDiff, to address these challenges by integrating retinal image enhancement and vessel segmentation. Specifically, we introduce a novel diffusion model-based framework for image enhancement, incorporating mask refinement as an auxiliary task via a image enhancement and vessel mask-aware diffusion model. Furthermore, we utilize low-quality retinal fundus images and their corresponding illumination maps as inputs to the modified UNet to obtain degradation factors that effectively preserve pathological features and pertinent information. This approach enhances the intermediate results within the iterative process of the diffusion model. To the best of our knowledge, this is the first utilization of a diffusion model to simultaneously achieve low-quality image enhancement and fundus vessel segmentation. Extensive experiments on publicly available fundus retinal datasets demonstrate the effectiveness of ESDiff compared to state-of-the-art methods. Fengting Liu, Wenhui Huang 0002 |
BIBM | 2 |
| 2023 | Instance-Aware Diffusion Model for Gland Segmentation in Colon Histology Images
Mengxue Sun, Wenhui Huang 0002, Yuanjie Zheng |
MICCAI (6) | 2 |
| 2023 | Dual-Modal Information Bottleneck Network for Seizure DetectionabstractIn recent years, deep learning has shown very competitive performance in seizure detection. However, most of the currently used methods either convert electroencephalogram (EEG) signals into spectral images and employ 2D-CNNs, or split the one-dimensional (1D) features of EEG signals into many segments and employ 1D-CNNs. Moreover, these investigations are further constrained by the absence of consideration for temporal links between time series segments or spectrogram images. Therefore, we propose a Dual-Modal Information Bottleneck (Dual-modal IB) network for EEG seizure detection. The network extracts EEG features from both time series and spectrogram dimensions, allowing information from different modalities to pass through the Dual-modal IB, requiring the model to gather and condense the most pertinent information in each modality and only share what is necessary. Specifically, we make full use of the information shared between the two modality representations to obtain key information for seizure detection and to remove irrelevant feature between the two modalities. In addition, to explore the intrinsic temporal dependencies, we further introduce a bidirectional long-short-term memory (BiLSTM) for Dual-modal IB model, which is used to model the temporal relationships between the information after each modality is extracted by convolutional neural network (CNN). For CHB-MIT dataset, the proposed framework can achieve an average segment-based sensitivity of 97.42%, specificity of 99.32%, accuracy of 98.29%, and an average event-based sensitivity of 96.02%, false detection rate (FDR) of 0.70/h. We release our code at https://github.com/LLLL1021/Dual-modal-IB. Xinting Ge, Yunfeng Shi, Mengxue Sun, Qingtao Gong, Wenhui Huang 0002 |
Int. J. Neural Syst. | 7 |
| 2023 | Information bottleneck-based interpretable multitask network for breast cancer classification and segmentation
Junxia Wang, Yuanjie Zheng, Xinmeng Li, Chongjing Wang, James Gee, Wenhui Huang 0002 |
Medical Image Anal. | 8 |
| 2023 | Image Matting With Deep Gaussian ProcessabstractWe observe a common characteristic between the classical propagation-based image matting and the Gaussian process (GP)-based regression. The former produces closer alpha matte values for pixels associated with a higher affinity, while the outputs regressed by the latter are more correlated for more similar inputs. Based on this observation, we reformulate image matting as GP and find that this novel matting-GP formulation results in a set of attractive properties. First, it offers an alternative view on and approach to propagation-based image matting. Second, an application of kernel learning in GP brings in a novel deep matting-GP technique, which is pretty powerful for encapsulating the expressive power of deep architecture on the image relative to its matting. Third, an existing scalable GP technique can be incorporated to further reduce the computational complexity to$\mathcal {O}(n)$from$\mathcal {O}(n^{3})$of many conventional matting propagation techniques. Our deep matting-GP provides an attractive strategy toward addressing the limit of widespread adoption of deep learning techniques to image matting for which a sufficiently large labeled dataset is lacking. A set of experiments on both synthetically composited images and real-world images show the superiority of the deep matting-GP to not only the classical propagation-based matting techniques but also modern deep learning-based approaches. Yuanjie Zheng, Yunshuai Yang, Tongtong Che, Sujuan Hou, Wenhui Huang 0002, Yue Gao 0002, Ping Tan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Exploiting Probabilistic Siamese Visual Tracking with a Conditional Variational AutoencoderabstractVisual tracking is a fundamental capability for robots tasked with humans and environment interaction. However, state-of-the-art visual tracking methods are still prone to failures and are imprecise when applied to challenging stereos, and their results are generally confidence agonistic. These methods depend on an embedded deep learning model to provide deterministic features or regression maps. A deterministic output with low confidence can result in disastrous consequences and lacks evidence needed for subsequent operations. Moreover, training data ambiguities or noise in the observations (so-called data uncertainty) can also lead to inherent uncertainty. In this paper, we focus on exploiting probabilistic Siamese visual tracking with a conditional variational autoencoder (CVAE). First, we build a bridge between the Siamese architecture and the CVAE and propose a novel Bayesian visual tracking method. Second, the proposed method generates a complete probability distribution that enables the production of multiple plausible tracking outputs. Third, CVAE conditioned by ground truth data encodes a low-dimensional latent space and conducts noise-injection training to prevent overfitting. Our proposed tracking method outperformed the state-of-the-art trackers on the VOT2016, VOT2018 and TColor-128 datasets. Wenhui Huang 0002, Jason Gu, Peiyong Duan, Sujuan Hou, Yuanjie Zheng |
ICRA | 1 |
| 2020 | End-to-end multitask Siamese network with residual hierarchical attention for real-time object tracking
Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001 |
Appl. Intell. | 1 |
| 2017 | Correlation filter-based self-paced object trackingabstractObject tracking is an important capability for robots tasked with interacting with humans and the environment, and it enables robots to manipulate objects. In object tracking, selecting samples to learn a robust and efficient appearance model is a challenging task. Model learning determines both the strategy and frequency of model updating, which concerns many details that can affect the tracking results. In this paper, we propose an object tracking approach by formulating a new objective function that integrates the learning paradigm of self-paced learning into object tracking such that reliable samples can be automatically selected for model learning. Sample weights and model parameters can be learned by minimizing this single objective function under the framework of kernelized correlation filters. Moreover, a real-valued error-tolerant self-paced function with a constraint vector is proposed to combine prior knowledge, i.e., the characteristics of object tracking, with information learned during tracking. We demonstrate the robustness and efficiency of our object tracking approach on a recent object tracking benchmark data set: OTB 2013. Wenhui Huang 0002, Jason Gu, Xin Ma 0001, Yibin Li 0001 |
ICRA | 1 |