Wendong Huang

dblp:47/5203 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Mask-Guided Proxy Mining Network for Few-Shot Medical Image Segmentation
abstract
Few-shot medical image segmentation (FSMIS) has attracted increasing attention as a promising technique for solving medical image segmentation tasks by relying on only a small amount of labeled data from new classes. Current FSMIS methods typically employ pixel-level semantic correlations between support-query image pairs to guide the segmentation of query images. However, the class information gap between support and query images may induce severe mismatches, leading to semantic ambiguity between foreground and background pixels. To address this issue, we propose a novel mask-guided proxy mining network (MPMNet), which mines a set of representative reference features (termed proxies) from support and query images to rectify foreground-background ambiguity. Specifically, to eliminate false pairwise matches caused by excessive intra-class variations, we design a mask-guided proxy mining module to adaptively learn representative proxies that can perceive visual differences between objects with different scales and shapes. Moreover, we integrate a hierarchical prior generation module and a context-aware feature enrichment module into MPMNet to obtain multi-scale information and enhance the discriminability of features. With these well-designed components and structures, our MPMNet can effectively overcome the adverse effects of false pixel matches by establishing proxy-level semantic correlations. Extensive experiments on three standard medical segmentation benchmarks demonstrate that our MPMNet significantly outperforms previous state-of-the-art methods, with a mean gain of 2.71% in DSC across all datasets. The code is available at: https://github.com/donglongzi/MPMNet.
Wendong Huang, Jinwu Hu, Yongchao Wang 0004, Xiuli Bi, Yucheng Shu, Xuezong Yang, Bin Xiao 0002
IEEE Trans. Image Process.1
2025 Who Controls the Authorization? Invertible Networks for Copyright Protection in Text-to-Image Synthesis
Baoyue Hu, Yang Wei 0002, Wendong Huang, Xiuli Bi, Bin Xiao 0002
ICCV4
2025 Spatially aligned graph transfer learning for characterizing spatial regulatory heterogeneity
abstract
Spatially resolved transcriptomics (SRT) technologies facilitate the exploration of cell fates or states within tissue microenvironments. Despite these advances, the field has not adequately addressed the regulatory heterogeneity influenced by microenvironmental factors. Here, we propose a novel Spatially Aligned Graph Transfer Learning (SpaGTL), pretrained on a large-scale multi-modal SRT data of about 100 million cells/spots to enable inference of context-specific spatial gene regulatory networks across multiple scales in data-limited settings. As a novel cross-dimensional transfer learning architecture, SpaGTL aligns spatial graph representations across gene-level graph transformers and cell/spot-level manifold-dominated variational autoencoder. This alignment facilitates the exploration of microenvironmental variations in cell types and functional domains from a molecular regulatory perspective, all within a self-supervised framework. We verified SpaGTL's precision, robustness, and speed over existing state-of-the-art algorithms and show SpaGTL's potential that facilitates the discovery of novel regulatory programs that exhibit strong associations with tissue functional regions and cell types. Importantly, SpaGTL could be extended to process multi-slice SRT data and map molecular regulatory landscape associated with three-dimensional spatial-temporal changes during development.
Wendong Huang, Yaofeng Hu, Lequn Wang, Guangsheng Wu, Chuanchao Zhang, Qianqian Shi 0004
Briefings Bioinform.1
2025 Neurocognitive Insights: Cognitive Comprehension Attention in Multi-Organ Segmentation
abstract
In multi-organ segmentation, attention mechanisms are frequently employed to enhance the focus on irregular organs, improving performance. However, current attention mechanisms exhibit notable limitations. On the one hand, their visual saliency-based attention bias results in incomplete region-of-interest coverage. On the other hand, their organ-specific cognitive deficiency exacerbates organ misclassification. Inspired by neurocognitive science, this paper proposes a Cognitive Comprehension Attention (CCA). Diverging from existing methods, CCA achieves refined attention allocation by decomposing visual representations into discrete visual stimuli. This fine-grained approach enables unbiased processing for each visual stimulus, preventing critical information omission and ensuring comprehensive organ region coverage. More importantly, CCA generates organ-specific attention representations by establishing distinct attention patterns across different organ regions, which empowers CCA with cognitive capacity, resolving organ misclassification. Extensive experiments across multiple datasets demonstrate that CCA significantly enhances backbone performance, achieving a max mDice improvement of 8.45% while surpassing state-of-the-art methods by 9% in Recall and 11.78% in Precision. Code is available at:https://github.com/robert1818118/CCA.
Yang Wei 0002, Wendong Huang, Xiuli Bi, Xuezong Yang, Bin Xiao 0002
IEEE Trans. Big Data3
2025 Prototype-Guided Graph Reasoning Network for Few-Shot Medical Image Segmentation
abstract
Few-shot semantic segmentation (FSS) is of tremendous potential for data-scarce scenarios, particularly in medical segmentation tasks with merely a few labeled data. Most of the existing FSS methods typically distinguish query objects with the guidance of support prototypes. However, the variances in appearance and scale between support and query objects from the same anatomical class are often exceedingly considerable in practical clinical scenarios, thus resulting in undesirable query segmentation masks. To tackle the aforementioned challenge, we propose a novel prototype-guided graph reasoning network (PGRNet) to explicitly explore potential contextual relationships in structured query images. Specifically, a prototype-guided graph reasoning module is proposed to perform information interaction on the query graph under the guidance of support prototypes to fully exploit the structural properties of query images to overcome intra-class variances. Moreover, instead of fixed support prototypes, a dynamic prototype generation mechanism is devised to yield a collection of dynamic support prototypes by mining rich contextual information from support images to further boost the efficiency of information interaction between support and query branches. Equipped with the proposed two components, PGRNet can learn abundant contextual representations for query images and is therefore more resilient to object variations. We validate our method on three publicly available medical segmentation datasets, namely CHAOS-T2, MS-CMRSeg, and Synapse. Experiments indicate that the proposed PGRNet outperforms previous FSS methods by a considerable margin and establishes a new state-of-the-art performance.
Wendong Huang, Jinwu Hu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002
IEEE Trans. Medical Imaging1
2024 Anatomical Prior Guided Spatial Contrastive Learning for Few-Shot Medical Image Segmentation
Wendong Huang, Jinwu Hu, Xiuli Bi, Bin Xiao 0002
ACM Multimedia1
2024 MCLEMCD: multimodal collaborative learning encoder for enhanced music classification from dances
Wenjuan Gong, Qingshuang Yu, Wendong Huang, Peng Cheng 0008, Jordi Gonzàlez 0001
Multim. Syst.4
2023 Location-Aware Transformer Network for Few-Shot Medical Image Segmentation
abstract
Automatic and precise organ segmentation plays a significant role in promoting the development of the diagnosis and treatment of the disease. Despite making enormous strides in medical image segmentation, conventional deep neural network-based methods are inherently massive data-driven techniques and are challenging to adapt to novel classes with a small number of labeled samples. Few-shot learning is a promising solution through learning novel classes from extremely limited annotated examples. However, existing few-shot segmentation methods focus excessively on targets in individual images while neglecting to model the global spatial correlation across images, which may cause severe performance degradation. To solve this issue, we propose a new Transformer-based few-shot segmentation framework for medical imaging, namely location-aware transformer network (LATNet), which establishes the spatial correlation between support and query objects, yielding location-aware prototypes, and then performs segmentation by computing the semantic similarity between query features and location-aware prototypes. Additionally, to further enhance the representativeness of the obtained location-aware prototypes in low-data regimes, we design a prediction iterative refinement module, which can iteratively exploit the query predictions output by each iteration to update the location-aware prototypes and progressively refine the query predictions. Extensive experiments on three challenging medical image datasets, i.e., Abd-MRI, Card-MRI, and Abd-CT, show that the proposed LATNet achieves remarkable improvements over current state-of-the-art methods by an average of 4.17%, 1.50%, and 4.63% in terms of the Dice Score, respectively.
Wendong Huang, Bin Xiao 0002, Jinwu Hu, Xiuli Bi
BIBM1
2023 Spatially aware self-representation learning for tissue structure characterization and spatial functional genes identification
abstract
Spatially resolved transcriptomics (SRT) enable the comprehensive characterization of transcriptomic profiles in the context of tissue microenvironments. Unveiling spatial transcriptional heterogeneity needs to effectively incorporate spatial information accounting for the substantial spatial correlation of expression measurements. Here, we develop a computational method, SpaSRL (spatially aware self-representation learning), which flexibly enhances and decodes spatial transcriptional signals to simultaneously achieve spatial domain detection and spatial functional genes identification. This novel tunable spatially aware strategy of SpaSRL not only balances spatial and transcriptional coherence for the two tasks, but also can transfer spatial correlation constraint between them based on a unified model. In addition, this joint analysis by SpaSRL deciphers accurate and fine-grained tissue structures and ensures the effective extraction of biologically informative genes underlying spatial architecture. We verified the superiority of SpaSRL on spatial domain detection, spatial functional genes identification and data denoising using multiple SRT datasets obtained by different platforms and tissue sections. Our results illustrate SpaSRL's utility in flexible integration of spatial information and novel discovery of biological insights from spatial transcriptomic datasets.
Chuanchao Zhang, Xinxing Li, Wendong Huang, Lequn Wang, Qianqian Shi 0004
Briefings Bioinform.3
2022 Classification of thermal image of clinical burn based on incremental reinforcement learning
Xianjun Wu, Wendong Huang, Shenghang Wu, Jinbo Huang
Neural Comput. Appl.2
2022 Correction to: Classification of thermal image of clinical burn based on incremental reinforcement learning
Xianjun Wu, Wendong Huang, Shenghang Wu, Jinbo Huang
Neural Comput. Appl.2
2009 A joint encoder-decoder framework for supporting energy efficient audio decoding
Wendong Huang, Ye Wang 0007
Multim. Syst.1
2009 An optimal speed control scheme supported by media servers for low-power multimedia applications
Wendong Huang, Ye Wang 0007
Multim. Syst.1
2006 Efficient Partial Spectrum Reconstruction using an Asymmetric PQMF Algorithm for MPEG-Coded Stereo Audio
abstract
This paper presents a novel algorithm of a scalable and efficient pseudo-quadrature mirror filters (PQMF), which is employed for partial decoding a single-layer audio bitstream such as MP3, typically coded in joint/MS mode. The proposed algorithm is a new extension to our previous work on scalable audio decoding and is designed for asymmetric partial spectrum reconstruction (APSR), where perceptually irrelevant computations are removed. Furthermore, an efficient up-sampling operation is introduced for right channel output. The slight distortions introduced by our simple up-sampling method are inaudible according to a set of perceptual evaluations. Simulation results show that 64.6% energy savings can be achieved for a typical configuration in comparison to the standard PQMF algorithm employed by MPEG-1 audio
Wendong Huang, Ye Wang 0007
ICME1
2005 Power-aware bandwidth and stereo-image scalable audio decoding
abstract
We propose a new workload-scalable audio decoding scheme that would enable users to control the tradeoff between playback quality and power consumption in battery-powered portable audio players. Our objective is to give users a control at the decoder side, similar to the Long Play (LP) recording mode at the encoder side in many media recording devices. The main contribution of this paper is a proposal for a Bandwidth and Stereo-image Scalable (BSS) decoding scheme for single-layer audio formats such as MP3. The proposed scheme is based on an analysis of the perceptual relevance of different audio components in the compressed bitstream. The bandwidth and stereo-image scalability directly translates into scalability in terms of the computational workload generated by the decoder. This can be exploited by a voltage/frequency scalable processor to save energy and prolong the battery life.
Wendong Huang, Ye Wang 0007, Samarjit Chakraborty
ACM Multimedia1
2004 A framework for robust and scalable audio streaming
abstract
We propose a framework to achieve bandwidth efficient, error robust and bitrate scalable audio streaming. Our approach is compatible with most audio compression format. The main contributions of this paper include: 1) the proposal of a Multi-Stage Interleaving (MSI) strategy which translates packet loss into loss of separate frequency components that are less perceptually significant; and 2) the design of a Layered Unequal-Sized Packetization (LUSP) scheme which enables bitrate scalability and prioritized packet transmission. The combination of the proposed MSI and LUSP allows the use of a set of simple yet effective methods of error concealment in the compressed domain. Our approach offers significant advantages over existing methods in terms of memory consumption (a savings of over 40 times in the sample MP3 implementation), and computational complexity, which are critical issues for battery-powered small devices.
Ye Wang 0007, Wendong Huang, Jari Korhonen
ACM Multimedia2
2003 Content-based UEP: a new scheme for packet loss recovery in music streaming
abstract
Bandwidth efficiency and error robustness are two essential and conflicting requirements for streaming media content over error-prone channels, such as wireless channels. This paper describes a new scheme called content-based unequal error protection (C-UEP), which aims to improve the user-perceived QoS in the case of packet loss. We use music streaming as an example to show the effectiveness of the new concept. C-UEP requires only a small fraction of the redundancy used in existing forward error correction (FEC) methods. C-UEP classifies every audio segment (e.g. an encoding frame) into different classes to improve encoding efficiency. Salient transients such as drumbeats and note onsets are encoded with more redundancy in a secondary bitstream used to recover lost packets by the receiver. Formal perceptual evaluations show that our scheme improves audio quality significantly over simple muting and packet repetition baselines. This improvement is achieved with a negligible amount of redundancy, which is transmitted to the receiver ahead of playback.
Ye Wang 0007, Ali Ahmaniemi, David Isherwood, Wendong Huang
ACM Multimedia4