VLDB 2026 Research / reviewers in the wild / expert
Huisi Wu
dblp:52/8869
· DBLP profile ↗
93ranked-venue papers
38as first author
81since 2021 · last 2026
0000-0002-0399-9089ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 18 first-author · 53 since 2021Artificial intelligence and machine learning · 39 · 15 first-author · 37 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 12 first-author · 21 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VPSentry: Semi-supervised Video Polyp Segmentation via Sentry-guided Long-term Prototype Fusion with Correlation Dynamic PropagationabstractAutomated polyp segmentation in colonoscopy videos is an essential computer-aided technology for early detection and removal of polyps. However, most existing video polyp segmentation methods are designed with pixel-level temporal learning mechanisms, at the cost of time-consuming frame-wise annotations. In this paper, we present VPSentry, a novel semi-supervised segmentation model with a sentry mechanism. Our model integrates a prototype memory to store the long-term spatiotemporal cues of colonoscopy videos. Moreover, we devise adaptive prototypes to capture and generalize critical representations from individual frames, enabling long-term temporal fusion across labeled and unlabeled frames. In addition, we propose a correlation dynamic propagation module that propagates information from prototypes to features while simultaneously extracting dynamic features to perceive variations in polyp details between adjacent frames. Since colonoscopy scenes may change among consecutive frames, we further employ a sentry mechanism to assess the inter-frame continuity. This mechanism guides the prototype memory updating and the correlation dynamic propagation, further facilitating robust temporal propagation and dynamic detail perception for semi-supervised learning of long-term colonoscopy video sequences. Extensive experiments on the large-scale SUN-SEG dataset demonstrate that our model achieves optimal segmentation performance with real-time inference efficiency. Guilian Chen, Xiaoling Luo 0001, Huisi Wu, Harry Qin |
AAAI | 3 |
| 2026 | Palimpsest: Reconciling the CISS Trilemma for Incremental Nuclei SegmentationabstractAdapting computational pathology models to evolving clinical diagnostics via Class-Incremental Semantic Segmentation (CISS) is critical. However, this task imposes a unique CISS Trilemma: a simultaneous failure to preserve the intricate tissue background (stability), distinguish morphologically similar new nuclei (plasticity), and maintain a constant model size (scalability), all under a strict exemplar-free constraint. To resolve this, we introduce Palimpsest, a novel framework that systematically decouples these conflicting demands. Palimpsest integrates three synergistic mechanisms: a Parameter-Conserving Synthesis (PCS) module merges lightweight adapters to ensure scalability; a novel Similarity-Aware Centroid Recalibration (SCR) module executes differentiated recalibration to counteract non-uniform foreground drift, securing plasticity; and an Adaptive Residual Shading (ARS) module performs logit-space decoupling to preserve background integrity, ensuring stability. Extensive experiments on two histopathology datasets demonstrate that Palimpsest significantly outperforms state-of-the-art methods, achieving a superior stability-plasticity balance, particularly in challenging long-term incremental scenarios. Huisi Wu |
AAAI | 2 |
| 2026 | CiNuSeg: Class Incremental Nuclei Segmentation via Anchor-driven Consistency Learning with Dual Region RegularizationabstractRecent advances in deep learning have led to significant improvements in nuclei segmentation from histological images, particularly when labels of all classes are available simultaneously during training. However, in clinical practice, real-world scenarios require a model to perform well in an incremental learning setting, where we anticipate the model to achieve satisfactory performance on previously unseen data while effectively mitigating catastrophic forgetting of old classes. Most previous methods alleviate forgetting by distilling old class knowledge through prototypes; however, they fail to adequately capture fine-grained details to address the challenge of high class similarity, which is particularly severe in histological images. To overcome these limitations, we propose a novel incremental learning method for nuclei segmentation (we call it CiNuSeg), which is composed of two key innovative modules. First, we propose a new Anchor-driven Consistency Learning (ACL) module to construct multi-level class anchors within each sample to effectively capture fine structural and textural details of nuclei, thereby significantly mitigating forgetting. Second, we develop a Dual Region Regularization (DRR) module to suppress new class representations within old class regions while enhancing new class representations within new class regions, strengthening the model's ability to discriminate between different nuclei types and improving inter-class separability. We further introduce an Adaptive Temperature Tuning (ATT) strategy to dynamically balance model stability and plasticity. Extensive experiments conducted on benchmarking MoNuSAC and CoNSeP pathological datasets demonstrate the effectiveness of our method, consistently achieving better performance than SOTAs in different settings. Codes will be available upon publication. Xuexin Wu, Zhenhui Ding, Huisi Wu, Harry Qin |
AAAI | 3 |
| 2026 | SSPF-Net: Learning Organ Shape-Aware and Self-Calibration Prototypes for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation is a promising solution for scarce data scenarios, especially for medical classes with limited labeled data. However, there are still several flaws in the current few-shot medical image segmentation methods. The diversity of the base classes in the training stage is insufficient, so it is still difficult to train a model with strong generalization through a large number of images and classes like natural images. Besides, due to the discrepancy between the query and support images, relying solely on representative prototypes from support features fails to handle intra-class variability. In this paper, we propose a novel network with an organ shape-aware module (OSA) to optimize the prototype-based network and enhance the model’s generalization ability. Specifically, the OSA is designed to obtain a variable number of prototypes that contain the shape and edge information of the organ. Compared with general models using fixed-shape prototypes, the prototypes obtained by our OSA are more precise and informative. In addition, we propose a new inferring prototype self-calibration (IPSC) module with a mechanism to mine the query potential information. Such a scheme can extract foreground and background inferring prototypes from query features to enhance support prototypes, reduce intra-class differences, and eliminate disturbing information around the target class. Furthermore, the self-calibration operation in IPSC can select reliable foreground inferring vectors to generate foreground inferring prototypes. Extensive experiments on two public datasets demonstrate the superiority of our method over state-of-the-art methods. Shihao Zheng, Huisi Wu |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | Echocardiography Video Segmentation via Mamba-Based Spatiotemporal Synergistic Network and Adaptive-Dynamic LearningabstractAutomatic echocardiography video segmentation is crucial for accurate diagnosis of cardiovascular diseases, as high-quality segmentation significantly improves automated lesion detection. However, deep learning methods still face challenges including speckle noise, dynamic ventricular changes, limited annotations, and the requirement for real-time inference in clinical practice. In this paper, we propose MSSNet, a novel semi-supervised method based on the efficient sequence modeling architecture Mamba, to address these challenges. To enhance noise robustness and foreground tracking, we design a flexible and efficient spatiotemporal synergistic guidance (SSG) module that leverages attention weights from historical frames to guide subsequent segmentation. By incorporating stable structural context and modeling inter-frame dependencies through weight propagation, SSG effectively mitigates segmentation errors caused by strong local noise and ventricular dynamics while maintaining low computational complexity. To alleviate limited annotation, we further introduce two semi-supervised modules: region-wise adaptive cross-mix (RAC) and dynamic offset correction (DOC). RAC simulates clinically plausible samples via regional mixing to enrich semantic details and strengthen feature learning, while DOC continuously integrates features from high-quality pseudo-labels during training. Experiments on the CAMUS and EchoNet-Dynamic datasets demonstrate that MSSNet outperforms existing SOTA methods in segmentation accuracy and achieves notable improvements in inference speed. The code is available at https://github.com/SSS666-klk/MSSNet. Yimu Sun, Guilian Chen, Jingxing Guo, Huisi Wu |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Semi-Supervised Breast Lesion Segmentation Using Confidence-Ranked Features and Bi-Level PrototypesabstractAutomated lesion segmentation through breast ultrasound (BUS) images is an essential prerequisite in computer-aided diagnosis. However, the task of breast segmentation remains challenging, due to the time-consuming and labor-intensive process of acquiring precise labeled data, as well as severely ambiguous lesion boundaries and low contrast in BUS images. In this article, we propose a novel semi-supervised breast segmentation framework based on confidence-ranked features and bi-level prototypes (CoBiNet) to alleviate these issues. Our outputs are derived from two branches: classifier and projector. In the projector branch, we first rank the features by multilevel sampling to obtain multiple feature sets with different confidence levels. Then, these sets are progressed in two directions. One is to acquire local prototypes at each level by local sampling and perform trans-confidence level (TCL) contrastive learning. This encourages the low-confidence features to converge to the high-confidence features, which enhances the model's ability to recognize ambiguous regions. The other process is to generate more representative global prototypes by global sampling, followed by generating more reliable predictions and performing cross-guidance (CG) consistency learning with the classifier output predictions, facilitating knowledge transfer between the structure-aware projector and the category-discriminative classifier branches. Extensive experiments on two well-known public datasets, BUSI and UDIAT, demonstrate the superiority of our method over state-of-the-art approaches. Codes will be released upon publication. Siyao Jiang, Huisi Wu, Yu Zhou 0027, Junyang Chen 0001, Harry Qin |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR), with its large patient population, has become a formidable threat to human visual health. In the clinical diagnosis of DR, multi-view fundus images are considered to be more suitable for DR diagnosis because of the wide coverage of the field of view. Therefore, different from most of the previous single-view DR grading methods, we design a dynamic selection-driven multi-view DR grading method to fit clinical scenarios better. Since lesion information plays a key role in DR diagnosis, previous methods usually boost the model performance by enhancing the lesion feature. However, during the actual diagnosis, ophthalmologists not only focus on the crucial parts, but also exclude irrelevant features to ensure the accuracy of judgment. To this end, we introduce the idea of dynamic selection and design a series of selection mechanisms from fine granularity to coarse granularity. In this work, we first introduce an Ophthalmic Image Reader (OIR) agent to provide the model with pixel-level prompts of suspected lesion areas. Moreover, a Multi-View Token Selection Module (MVTSM) is designed to prune redundant feature tokens and realize dynamic selection of key information. In the final decision stage, we dynamically fuse multi-view features through the novel Multi-View Mixture of Experts Module (MVMoEM), to enhance key views and reduce the impact of conflicting views. Extensive experiments on a large multi-view fundus image dataset with 34,452 images demonstrate that our method performs favorably against state-of-the-art models. Xiaoling Luo 0001, Qihao Xu, Huisi Wu, Chengliang Liu 0003, Zhihui Lai 0001, LinLin Shen |
AAAI | 3 |
| 2025 | Line Drawing Abstraction Based on Line Importance Evaluation
Shilong Deng, Xueting Liu 0001, Chengze Li, Ping Li 0016, Zhenkun Wen, Huisi Wu |
CGI (1) | 6 |
| 2025 | Cartoon Animation Shading Removal
Zhenhua Ou, Chengze Li, Xueting Liu 0001, Zhenkun Wen, Huisi Wu |
CGI (1) | 5 |
| 2025 | CSC-PA: Cross-image Semantic Correlation via Prototype Attentions for Single-network Semi-supervised Breast Tumor SegmentationabstractAccurate automatic breast ultrasound (BUS) image segmentation is essential for early breast cancer screening and diagnosis. However, it remains challenging owing to (1) breast lesions of various scale and shape, (2) ambiguous boundaries caused by speckle noise and artifacts, and (3) the scarcity of high-quality annotations. Most existing semi-supervised methods employ the mean-teacher architecture, which merely learns semantic information within a single image and heavily relies on the performance of the teacher model. Therefore, we present a novel cross-image semantic correlation semi-supervised framework, named CSC-PA, to improve the performance of BUS image segmentation. CSC-PA is trained based on a single network, which integrates a foreground prototype attention (FPA) and an edge prototype attention (EPA). Specifically, FPA transfers complementary foreground information for more stable and complete lesion segmentation. On the other hand, EPA enhances edge features of lesions by using edge prototype, where an adaptive edge container is proposed to store global edge features and generate the edge prototype. Additionally, we introduce a pixel affinity loss (PAL) to exploit previously ignored contextual correlation in supervision, which further improves performance on edges. Extensive experiments on two benchmark BUS datasets demonstrate that our model outperforms other state-of-the-art methods under different partition protocols. Codes are available at https://github.com/shdkdh/CSC-PA. Zhenhui Ding, Guilian Chen, Qin Zhang 0011, Huisi Wu, Harry Qin |
CVPR | 4 |
| 2025 | STDDNet: Harnessing Mamba for Video Polyp Segmentation via Spatial-aligned Temporal Modeling and Discriminative Dynamic Representation Learning
Guilian Chen, Huisi Wu, Harry Qin |
ICCV | 2 |
| 2025 | WeaveSeg: Iterative Contrast-weaving and Spectral Feature-refining for Nuclei Instance Segmentation
Huisi Wu, Harry Qin |
ICCV | 2 |
| 2025 | GDKVM: Echocardiography Video Segmentation via Spatiotemporal Key-Value Memory with Gated Delta RuleabstractAccurate segmentation of cardiac chambers in echocardiography sequences is crucial for the quantitative analysis of cardiac function, aiding in clinical diagnosis and treatment. The imaging noise, artifacts, and the deformation and motion of the heart pose challenges to segmentation algorithms. While existing methods based on convolutional neural networks, Transformers, and space-time memory networks have improved segmentation accuracy, they often struggle with the trade-off between capturing long-range spatiotemporal dependencies and maintaining computational efficiency with fine-grained feature representation. In this paper, we introduce GDKVM, a novel architecture for echocardiography video segmentation. The model employs Linear Key-Value Association (LKVA) to effectively model inter-frame correlations, and introduces Gated Delta Rule (GDR) to efficiently store intermediate memory states. Key-Pixel Feature Fusion (KPFF) module is designed to integrate local and global features at multiple scales, enhancing robustness against boundary blurring and noise interference. We validated GDKVM on two mainstream echocardiography video datasets (CAMUS and EchoNet-Dynamic) and compared it with various state-of-the-art methods. Experimental results show that GDKVM outperforms existing approaches in terms of segmentation accuracy and robustness, while ensuring real-time performance. Code is available at https://github.com/wangrui2025/GDKVM. Rui Wang 0186, Yimu Sun, Jingxing Guo, Huisi Wu, Harry Qin |
ICCV | 4 |
| 2025 | RA-BUSSeg: Relation-Aware Semi-Supervised Breast Ultrasound Image Segmentation via Adjacent Propagation and Cross-Layer Alignment
Wanting Zhang, Zhenhui Ding, Guilian Chen, Huisi Wu, Harry Qin |
ICCV | 4 |
| 2025 | Robust Character Stroke Segmentation For Diverse Fonts Via Contour Matching and Chain PropagationabstractStroke segmentation is a fundamental technique for various character analysis and synthesis applications. However, existing methods often face challenges such as over-segmentation, under-segmentation, low segmentation accuracy, and limited generalization capability when segmenting characters of diverse fonts. To address these issues, we propose a novel stroke segmentation method based on contour matching and utilize a similarity-based chain propagation strategy to tackle the challenges posed by fonts with significant structural and stylistic variations. Extensive visual and quantitative experiments on a newly created high-quality dataset demonstrate that our approach outperforms state-of-the-art methods and effectively handles a wide range of fonts. Xueting Liu 0001, Chengze Li, Zhenkun Wen, Huisi Wu |
ICIP | 5 |
| 2025 | Controllable Image Synthesis Workflow for Enhancing Cervical Cell Detection
Yihuang Hu, Qi Chen 0014, Linbo Liao, Weiping Lin, Huisi Wu, Liansheng Wang 0002 |
MICCAI (13) | 5 |
| 2025 | EchoVim: Making Vision Mamba Docile for Echocardiography Video Segmentation via Dynamic Interaction and Semantic Token-attentive RefinementabstractAutomatic echocardiography video segmentation is a powerful tool for improving the accuracy of cardiovascular function assessment. However, it remains a challenging task owing to (1) extensive speckle noise and blurred boundaries, (2) dramatic shape variations of targeting structures across frames, and (3) limited labeled data due to the high cost of annotation. In this paper, we present a novel semi-supervised segmentation model based on Vision Mamba (Vim) to comprehensively tackle these challenges; we call it EchoVim. Our framework introduces three technical innovations: First, a bidirectional inference mechanism (BIM) which can propagate label information bidirectionally from end-diastolic (ED) and end-systolic (ES) frames to generate pseudo-labels, coupled with confidence-aware dynamic updating to progressively refine supervision signals. Second, a dynamic interaction temporal alignment (DITA) module that establishes anatomical correspondence across frames by adaptively enhancing features near temporally stable regions while suppressing motion-irrelevant artifacts, effectively addressing variations in cardiac shape. Third, a semantic token-attentive refinement (STR) module that constructs low-rank semantic tokens to encode cardiac structure priors, utilizing attention-guided nonlinear transformations to disentangle speckle noise from true anatomical patterns. We conduct extensive experiments on two benchmarking echocardiography video datasets: CAMUS and EchoNet-Dynamic, and the results demonstrate that our method outperforms existing state-of-the-art approaches with real-time inference. Codes are available at https://github.com/guojx2255/EchoVim. Jingxing Guo, Guilian Chen, Yimu Sun, Huisi Wu, Harry Qin |
ACM Multimedia | 4 |
| 2025 | Hierarchical Spatiotemporal Context Aggregation and Speckle-aware Deformable Convolution for Echocardiography Video SegmentationabstractAutomatic segmentation of echocardiography videos is crucial for computer-aided cardiovascular function assessment in clinical practice. However, it is a challenging task owing to the existence of massive speckle noise, the large shape variations of heart structures between frames, and limited annotations. In this paper, we propose a novel semi-supervised video segmentation model to comprehensively meet these challenges. The proposed approach has two key techniques. First, we propose a dual-stream architecture that processes spatial and temporal features through separate pathways to capture structural details and motion patterns, then enhances spatiotemporal representations by interacting these decomposed features with query features generated from the original input. Second, as speckle noise primarily concentrates in high-frequency regions, we extend the traditional dilated convolution from a frequency perspective, enabling it to adaptively adjust the dilation rate and convolution kernel weights based on high frequency speckle noise information. This enables the network to focus on specific frequency bands, thereby enhancing its ability to capture both low-frequency context and high-frequency local details. Extensive experiments on the CAMUS and EchoNet-Dynamic datasets demonstrate that our method outperforms existing state-of-the-art methods in terms of both accuracy and inference speed. Codes are available at https://github.com/guojx2255/HSCA-SDC. Jingxing Guo, Guilian Chen, Yimu Sun, Huisi Wu, Harry Qin |
ACM Multimedia | 4 |
| 2025 | A Review of Few-Shot and Zero-Shot Learning for Node Classification in Social NetworksabstractNode classification tasks aim to assign labels or categories to entire graphs based on their structural properties or node attributes. It can be adopted for various types of graph systems, including but not limited to network traffic, biological networks, knowledge graphs, etc., especially to social networks. This problem is well-studied, and solutions have demonstrated significant success in numerous real-world applications. However, in the situation where emerging categories are scarce or even have no labeled data, classical methods perform poorly on the whole, which has attracted growing attention. Based on this, in this article, we divide researches for node classification in social networks into two broad categories: traditional methods and novel strategies (few-shot/zero-shot learning). In traditional node classification methods, we summarize some classical methods for both homogeneous and heterogeneous networks, which includes unsupervised classifier, matrix factorization techniques, supervised methods, random-walk, and meta-path. Meanwhile, we introduce novel methods in few-shot or zero-shot learning. The article outlines the technical principles of various methods and analyzes their performance across different classes. It further summarizes the benchmark datasets used for evaluating node classification tasks. Finally, the major opportunities, challenges, and future research directions in few-shot and zero-shot learning for node classification in graph scenarios are discussed. Junyang Chen 0001, Rui Mi, Huan Wang 0005, Huisi Wu, Jiqian Mo, Jingcai Guo, Zhihui Lai 0001, Liang-Jie Zhang, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Improving Video Moment Retrieval by Auxiliary Moment-Query Pairs With Hyper-InteractionabstractMost existing video moment retrieval (VMR) benchmark datasets face a common issue of sparse annotations-only a few moments being annotated. We argue that videos contain a broader range of meaningful moments that, if leveraged, could significantly enhance performance. Existing methods typically follow a generate-then-select paradigm, focusing primarily on generating moment-query pairs while neglecting the crucial aspect of selection. In this paper, we propose a new method, HyperAux, to yield auxiliary moment-query pairs by modeling the multi-modal hyper-interaction between video and language. Specifically, given a set of candidate moment-query pairs from a video, we construct a hypergraph with multiple hyperedges, each corresponding to a moment-query pair. Unlike traditional graphs where each edge connects only two nodes (frames or queries), each hyperedge connects multiple nodes, including all frames within a moment, semantically related frames outside the moment, and an input query. This design allows us to consider the frames within a moment as a whole, rather than modeling individual frame-query relationships separately. More importantly, constructing the relationships among all moment-query pairs within a video into a large hypergraph facilitates selecting higher-quality data from such pairs. On this hypergraph, we employ a hypergraph neural network to aggregate node information, update the hyperedge, and propagate video-language hyper-interactions to each connected node, resulting in context-aware node representations. This enables us to use node relevance to select high-quality moment-query pairs and refine the moments’ boundaries. We also exploit the discrepancy in semantic matching within and outside moments to construct a loss function for training the HGNN without human annotations. Our auxiliary data enhances the performance of twelve VMR models under fully-supervised, weakly-supervised, and zero-shot settings across three widely used VMR datasets: ActivityNet Captions, Charades-STA, and QVHighlights. We will release the source code and models publicly. Runhao Zeng, Yishen Zhuo, Yunjin Yang, Huisi Wu, Qi Chen 0014, Xiping Hu, Victor C. M. Leung |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Denoiser-Regulated Deep Unfolding Compressed Sensing With Learnable Fixed-Point ProjectionsabstractThe family of regularization by denoising (RED) methods introduce denoising operator as the regularization term to perform compressed sensing (CS) reconstruction, which shows higher flexibility and scalability. However, traditional RED framework has strict requirements on several properties of denoiser, making it hard to design the specific denoiser and limits the quality of reconstructed images. Although some relaxation for denoisers can be made by incorporating the fixed point projection during the iteration process, the involved parameters have great impact on the effectiveness and efficiency of the algorithm, which is non-trivial to set them properly. In this paper, we propose an innovative Deep Unfolding Network framework termed FP-DUN based on the iterative process of Regularization by Denoising via Fixed-Point Projection (RED-PRO). In FP-DUN, fix-point projection module is implemented with learnable weights of neural networks, where an effective denoiser based on dual attention mechanism (DAM) is developed to capture the details of the reconstructed image. Additionally, we propose a new loss function based on fixed point constraints, which is able to overcome the over-smoothness caused by multi-stage denoising and maintain the structural details to progressively improve the reconstruction quality. By training the DUN model, the parameters for the process of fix point projection and denoiser are learned automatically. Extensive experimental results comparing with state-of-the-art CS algorithms and traditional RED-PRO approach validate the effectiveness of FP-DUN, especially on some images with complex details. Yu Zhou 0027, Wei Xie 0020, Huisi Wu, Lei Huang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Echocardiography Video Segmentation via Neighborhood Correlation MiningabstractAccurate segmentation of the left ventricle in echocardiography is critical for diagnosing and treating cardiovascular diseases. However, accurate segmentation remains challenging due to the limitations of ultrasound imaging. Although numerous image and video segmentation methods have been proposed, existing methods still fail to effectively solve this task, which is limited by sparsity annotations. To address this problem, we propose a novel semi-supervised segmentation framework named NCM-Net for echocardiography. We first propose the neighborhood correlation mining (NCM) module, which sufficiently mines the correlations between query features and their spatiotemporal neighborhoods to resist noise influence. The module also captures cross-scale contextual correlations between pixels spatially to further refine features, thus alleviating the impact of noise on echocardiography segmentation. To further improve segmentation accuracy, we propose using unreliable-pixels masked attention (UMA). By masking reliable pixels, it pays extra attention to unreliable pixels to refine the boundary of segmentation. Further, we use cross-frame boundary constraints on the final predictions to optimize their temporal consistency. Through extensive experiments on two publicly available datasets, CAMUS and EchoNet-Dynamic, we demonstrate the effectiveness of the proposed, which achieves state-of-the-art performance and outstanding temporal consistency. Codes are available at https://github.com/dengxl0520/NCMNet. Huisi Wu |
IEEE Trans. Medical Imaging | 2 |
| 2025 | EPSegNet: Lightweight Semantic Recalibration and Assembly for Efficient Polyp SegmentationabstractColorectal cancer (CRC) is among the most common malignancies and the detection and removal of polyps at the early stage is of great importance to prevent it. However, current state-of-the-art high-accuracy methods for polyps segmentation have a large number of parameters and a stringent requirement for computational cost, while lightweight and fast models significantly sacrifice accuracy. Currently, medical semantic segmentation algorithms are mostly based on encoder-decoder architecture. Pixelwise spatial information has been proven to be very important to the quality of features extracted by encoders. However, almost all existing approaches capturing it suffer from high computational complexity. Furthermore, the capacity of the traditional decoder is limited by its limited receptive fields. To comprehensively address the above problems, we propose a novel efficient polyp segmentation network (EPSegNet) to simultaneously fulfill the requirements of accuracy, size, and speed. First, we propose a lightweight feature extraction and recalibration module (LFERM), which can efficiently extract dense multiscale features. Specifically, in LFERM, we propose a spatial information recalibration (SIR) block for efficiently refining spatial information. Based on LFERMs, we develop an encoder. Moreover, we propose a novel lightweight semantic assembly decoder (LSAD) that assembles both channelwise and pixelwise semantics from a global context view. Finally, we combine the encoder and LSAD to form the proposed EPSegNet. Experiments on Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB datasets demonstrate that the proposed EPSegNet achieves the best balance between accuracy and size among state-of-the-art models and obtains a fast speed for polyp segmentation. Without any pretraining and postprocessing, our method achieves 79.37% intersection over union (IoU) and 86.74% Dice on the Kvasie-SEG dataset with only 0.34 million parameters and a speed of 128 frames/s (FPS) at the input size of $3 { \times }384 \times 384$ on a single NVIDIA GEFORCE RTX 2080Ti card. Codes will be released upon publication. Huisi Wu, Zebin Zhao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Cartoon Animation Outpainting With Region-Guided Motion InferenceabstractCartoon animation video is a popular visual entertainment form worldwide, however many classic animations were produced in a 4:3 aspect ratio that is incompatible with modern widescreen displays. Existing methods like cropping lead to information loss while retargeting causes distortion. Animation companies still rely on manual labor to renovate classic cartoon animations, which is tedious and labor-intensive, but can yield higher-quality videos. Conventional extrapolation or inpainting methods tailored for natural videos struggle with cartoon animations due to the lack of textures in anime, which affects the motion estimation of the objects. In this article, we propose a novel framework designed to automatically outpaint 4:3 anime to 16:9 via region-guided motion inference. Our core concept is to identify the motion correspondences between frames within a sequence in order to reconstruct missing pixels. Initially, we estimate optical flow guided by region information to address challenges posed by exaggerated movements and solid-color regions in cartoon animations. Subsequently, frames are stitched to produce a pre-filled guide frame, offering structural clues for the extension of optical flow maps. Finally, a voting and fusion scheme utilizes learned fusion weights to blend the aligned neighboring reference frames, resulting in the final outpainting frame. Extensive experiments confirm the superiority of our approach over existing methods. Huisi Wu, Chengze Li, Xueting Liu 0001, Zhenkun Wen, Tong-Yee Lee |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | AdaMotif: Graph Simplification via Adaptive Motif DesignabstractWith the increase of graph size, it becomes difficult or even impossible to visualize graph structures clearly within the limited screen space. Consequently, it is crucial to design effective visual representations for large graphs. In this paper, we propose AdaMotif, a novel approach that can capture the essential structure patterns of large graphs and effectively reveal the overall structures via adaptive motif designs. Specifically, our approach involves partitioning a given large graph into multiple subgraphs, then clustering similar subgraphs and extracting similar structural information within each cluster. Subsequently, adaptive motifs representing each cluster are generated and utilized to replace the corresponding subgraphs, leading to a simplified visualization. Our approach aims to preserve as much information as possible from the subgraphs while simplifying the graph efficiently. Notably, our approach successfully visualizes crucial community information within a large graph. We conduct case studies and a user study using real-world graphs to validate the effectiveness of our proposed approach. The results demonstrate the capability of our approach in simplifying graphs while retaining important structural and community information. Peifeng Lai, Zhida Sun, Xiangyuan Chen, Huisi Wu, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | Few-shot medical image segmentation via query transformation learning
Shihao Zheng, Huisi Wu, Zhijian Gao, Ping Li 0016 |
Vis. Comput. | 2 |
| 2024 | An Embedding-Unleashing Video Polyp Segmentation Framework via Region Linking and Scale AlignmentabstractAutomatic polyp segmentation from colonoscopy videos is a critical task for the development of computer-aided screening and diagnosis systems. However, accurate and real-time video polyp segmentation (VPS) is a very challenging task due to low contrast between background and polyps and frame-to-frame dramatic variations in colonoscopy videos. We propose a novel embedding-unleashing framework consisting of a proposal-generative network (PGN) and an appearance-embedding network (AEN) to comprehensively address these challenges. Our framework, for the first time, models VPS as an appearance-level semantic embedding process to facilitate generate more global information to counteract background disturbances and dramatic variations. Specifically, PGN is a video segmentation network to obtain segmentation mask proposals, while AEN is a network we specially designed to produce appearance-level embedding semantics for PGN, thereby unleashing the capability of PGN in VPS. Our AEN consists of a cross-scale region linking (CRL) module and a cross-wise scale alignment (CSA) module. The former screens reliable background information against background disturbances by constructing linking of region semantics, while the latter performs the scale alignment to resist dramatic variations by modeling the center-perceived motion dependence with a cross-wise manner. We further introduce a parameter-free semantic interaction to embed the semantics of AEN into PGN to obtain the segmentation results. Extensive experiments on CVC-612 and SUN-SEG demonstrate that our approach achieves better performance than other state-of-the-art methods. Codes are available at https://github.com/zhixue-fang/EUVPS. Zhixue Fang, Xinrong Guo, Jingyin Lin, Huisi Wu, Harry Qin |
AAAI | 4 |
| 2024 | FedCD: Federated Semi-Supervised Learning with Class Awareness Balance via Dual TeachersabstractRecent advancements in deep learning have greatly improved the efficiency of auxiliary medical diagnostics. However, concerns over patient privacy and data annotation costs restrict the viability of centralized training models. In response, federated semi-supervised learning has garnered substantial attention from medical institutions. However, it faces challenges arising from knowledge discrepancies among local clients and class imbalance in non-independent and identically distributed data. Existing methods like class balance adaptation for addressing class imbalance often overlook low-confidence yet valuable rare samples in unlabeled data and may compromise client privacy. To address these issues, we propose a novel framework with class awareness balance and dual teacher distillation called FedCD. FedCD introduces a global-local framework to balance and purify global and local knowledge. Additionally, we introduce a novel class awareness balance module to effectively explore potential rare classes and encourage balanced learning in unlabeled clients. Importantly, our approach prioritizes privacy protection by only exchanging network parameters during communication. Experimental results on two medical datasets under various settings demonstrate the effectiveness of FedCD. The code is available at https://github.com/YunzZ-Liu/FedCD. Huisi Wu, Harry Qin |
AAAI | 2 |
| 2024 | MemSAM: Taming Segment Anything Model for Echocardiography Video Segmentation
Huisi Wu, Runhao Zeng, Harry Qin |
CVPR | 2 |
| 2024 | PH-Net: Semi-Supervised Breast Lesion Segmentation via Patch-Wise HardnessabstractWe present a novel semi-supervised framework for breast ultrasound (BUS) image segmentation, which is a very challenging task owing to (1) large scale and shape variations of breast lesions and (2) extremely ambiguous boundaries caused by massive speckle noise and artifacts in BUS images. While existing models achieved certain progress in this task, we believe the main bottleneck nowadays for further improvement is that we still cannot deal with hard cases well. Our framework aims to break through this bottleneck, which includes two innovative components: an adaptive patch augmentation scheme and a hard-patch contrastive learning module. We first identify hard patches by computing the average entropy of each patch and then shield hard patches to prevent them from being cropped out while performing random patch cutmix. Such a scheme is able to prevent hard regions from being inadequately trained under strong augmentation. We further develop a new hard-patch contrastive learning algorithm to direct model attention to hard regions by applying extra contrast to pixels in hard patches, further improving segmentation performance on hard cases. We demonstrate the superior-ity of our framework to state-of-the-art approaches on two famous BUS datasets, achieving better performance under different labeling conditions. The code is available at https://github.com/jjjsyyy/PH-Net. Siyao Jiang, Huisi Wu, Junyang Chen 0001, Qin Zhang 0011, Harry Qin |
CVPR | 2 |
| 2024 | Incremental Nuclei Segmentation from Histopathological Images via Future-class Awareness and Compatibility-inspired DistillationabstractWe present a novel semantic segmentation approach for incremental nuclei segmentation from histopathological images, which is a very challenging task as we have to in-crementally optimize existing models to make them perform well in both old and new classes without using training samples of old classes. Yet, it is an indispensable component of computer-aided diagnosis systems. The proposed approach has two key techniques. First, we propose a new future-class awareness mechanism by separating some potential regions for future classes from background based on their similari-ties to both old and new classes in the representation space. With this mechanism, we can not only reserve more parameter space for future updates but also enhance the repre-sentation capability of learned features. We further propose an innovative compatibility-inspired distillation scheme to make our model take full advantage of the knowledge learned by the old model. We conducted extensive experiments on two famous histopathological datasets and the results demonstrate the proposed approach achieves much better performance than state-of-the-art approaches. The code is available at https://github.com/why199911/nSeg. Huyong Wang, Huisi Wu, Harry Qin |
CVPR | 2 |
| 2024 | Benchmarking the Robustness of Temporal Action Detection Models Against Temporal CorruptionsabstractTemporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Although many methods have achieved promising results, their robustness has not been thoroughly studied. In practice, we observe that temporal information in videos can be occasionally corrupted, such as missing or blurred frames. Interestingly, existing methods often incur a significant performance drop even if only one frame is affected. To formally evaluate the robustness, we establish two temporal corruption robustness benchmarks, namely THUMOS14-C and ActivityNet-v1.3-C. In this paper, we extensively analyze the robustness of seven leading TAD methods and obtain some interesting findings: 1) Existing methods are particularly vulnerable to temporal corruptions, and end-to-end methods are often more susceptible than those with a pretrained feature extractor; 2) Vulnera-bility mainly comes from localization error rather than classification error; 3) When corruptions occur in the middle of an action instance, TAD models tend to yield the largest performance drop. Besides building a benchmark, we further develop a simple but effective robust training method to defend against temporal corruptions, through the Frame-Drop augmentation and Temporal-Robust Consistency loss. Remarkably, our approach not only improves robustness but also yields promising improvements on clean data. We believe that this study will serve as a benchmark for future research in robust video analysis. Source code and models are available at https://github.com/Alvin-Zeng/temporal-robustness-benchmark. Runhao Zeng, Xiaoyong Chen, Huisi Wu |
CVPR | 4 |
| 2024 | VP-SAM: Taming Segment Anything Model for Video Polyp Segmentation via Disentanglement and Spatio-Temporal Side Network
Zhixue Fang, Huisi Wu, Harry Qin |
ECCV (23) | 3 |
| 2024 | Domesticating SAM for Breast Ultrasound Image Segmentation via Spatial-Frequency Fusion and Uncertainty Correction
Wanting Zhang, Huisi Wu, Harry Qin |
ECCV (23) | 2 |
| 2024 | CONC: Complex-noise-resistant Open-set Node Classification with Adaptive Noise Detection
Qin Zhang 0011, Jiexin Lu, Huisi Wu, Shirui Pan, Junyang Chen 0001 |
IJCAI | 4 |
| 2024 | Boosting FFPE-to-HE Virtual Staining with Cell Semantics from Pretrained Segmentation Model
Yihuang Hu, Qiong Peng, Zhicheng Du, Huisi Wu, Jingxin Liu 0005, Hao Chen 0011, Liansheng Wang 0002 |
MICCAI (3) | 5 |
| 2024 | EGonc : Energy-based Open-Set Node Classification with substitute UnknownsabstractOpen-set Classification (OSC) is a critical requirement for safely deploying machine learning models in the open world, which aims to classify samples from known classes and reject samples from out-of-distribution (OOD).
Existing methods exploit the feature space of trained network and attempt at estimating the uncertainty in the predictions.
However, softmax-based neural networks are found to be overly confident in their predictions even on data they have never seen before and
the immense diversity of the OOD examples also makes such methods fragile.
To this end, we follow the idea of estimating the underlying density of the training data to decide whether a given input is close to the in-distribution (IND) data and adopt Energy-based models (EBMs) as density estimators.
A novel energy-based generative open-set node classification method, \textit{EGonc}, is proposed to achieve open-set graph learning.
Specifically, we generate substitute unknowns to mimic the distribution of real open-set samples firstly, based on the information of graph structures.
Then, an additional energy logit representing the virtual OOD class is learned from the residual of the feature against the principal space, and matched with the original logits by a constant scaling. This virtual logit serves as the indicator of OOD-ness.
EGonc has nice theoretical properties that guarantee an overall distinguishable margin between the detection scores for IND and OOD samples.
Comprehensive experimental evaluations of EGonc also demonstrate its superiority. Qin Zhang 0011, Zelin Shi, Shirui Pan, Junyang Chen 0001, Huisi Wu, Xiaojun Chen 0006 |
NeurIPS | 5 |
| 2024 | Multi-Level Object-Aware Guidance Network for Biomedical Image SegmentationabstractMost state-of-the-art models for biomedical image segmentation are developed based on U-shape architecture, which has two renowned, yet mutually affected, shortcomings: 1) difficulties in capturing global long-range dependencies, and 2) semantic information dilution in the decoding process. In this paper, we propose a novel network with a new object-aware module (OAM) to effectively establish global dependencies at multiple levels within the network and compensate high-level semantic information dilution when fusing the extracted multi-level features; we call the network MOG-Net. Specifically, the OAM is designed to figure out the relations between each pixel and targeting object region and recalibrate class-level semantic information according to the relations. Compared with non-local models, which construct pixel-wise global dependencies, our OAM is more efficient and target-specific, enabling us to achieve satisfactory results with less extra computational overhead. In addition, we embed a pyramid context encoder module (PCEM) in the proposed OAM to alleviate semantic information dilution; this scheme is able to bridge the spatial-semantic gap when fusing features extracted from different levels. We extensively evaluate the proposed MOG-Net on four diverse biomedical image segmentation tasks with different imaging modalities, achieving segmentation performance with 88.19%, 90.95% and 66.03% in Dice on three one-class datasets, as well as 88.83% and 87.11% in Dice for two classes on a multi-class dataset, respectively. Experimental results demonstrate the effectiveness of the proposed method, consistently outperforming state-of-the-art methods in most evaluation metrics.Note to Practitioners—Semantic segmentation of biomedical images is a critical prerequisite for subsequent diagnosis, treatment, and quantitative tasks in clinical practice. This article proposes a novel biomedical image segmentation network, namely MOG-Net, with a new object-aware module (OAM) to model global context dependencies from a category perspective and a pyramid context encoder module (PCEM) to enhance feature representation capabilities of spatial and channel dimensions. We experimentally demonstrate the effectiveness and generalization capability of proposed MOG-Net on diverse biomedical image segmentation tasks with different imaging modalities. We believe that our proposed method can serve as a practical clinical tool and has the potential to be applied to existing computer-aided medical systems and clinical measurement. Huisi Wu, Baiming Zhang, Junquan Pan, Harry Qin |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | META-Unet: Multi-Scale Efficient Transformer Attention Unet for Fast and High-Accuracy Polyp SegmentationabstractPolyp segmentation plays an important role in preventing Colorectal cancer. Although Vision Transformer has been widely introduced in medical image segmentation to compensate the limitations of traditional CNN in modeling global context, its shortcomings in learning the fine-detailed features and the heavy computation cost also hinder its application in challenging polyp segmentation due to the various shapes and sizes of polyps, the low-intensity contrast between polyps and surrounding tissues, and the inherent real-time requirement. In this paper, we propose a multi-scale efficient transformer attention (META) mechanism for fast and high-accuracy polyp segmentation, where efficient transformer blocks are employed to generate multi-scale element-wise attentions for adaptive feature fusion in the famous U-shape encoder-decoder architecture. Specifically, our META mechanism includes two branches to capture multi-scale long-term dependencies, which are implemented via two efficient transformer blocks with different resolutions. The local branch is used to capture a relatively smaller transform attention under a relatively lower resolution, while the global branch is used to capture high-resolution transform attention. The final poly segmentation results are progressively integrated based on the META mechanism in each layer of the decoder. Extensive experiments are conducted on four polyp segmentation datasets (CVC-ClinicDB, Endoscenestill, Kvasir-SEG and ETIS-Larib) to demonstrate its advantages, consistently outperforming different competitors. While using ResNet34 as backbones, it can achieve 85.78% IoU and 92.03% Dice, 88.99% IoU and 93.85% Dice, 86.42% IoU and 91.86% Dice respectively in CVC-ClinicDB, Endoscenestill, and Kvasir-SEG, and a speed of 98 FPS at the input size of$3 \times 512 \times 512$on a NVIDIA GeForce RTX 3090 card. The code is available at https://github.com/szuzzb/META-Unet.Note to Practitioners—Automatic polyp segmentation is a crucial step of polyp recognition and diagnostic of colonoscopy, which usually require both high-accuracy and real-time performance. This article proposes a novel polyp segmentation method, namely META-Unet, by modeling multi-scale attention maps effectively and efficiently based on a novel multi-scale efficient transformer attention (META) mechanism, for faster and higher-accuracy polyp segmentation. We evaluate our META-Unet on four public polyp image segmentation datasets (CVC-ClinicDB, Endoscenestill, Kvasir-SEG and ETIS-Larib). Comprehensive experimental results validate its outstanding performance with a better balance in both accuracy and inference speed. The proposed META mechanism is potentially to be embedded in various deep learning frameworks and facilitates more computer-aided applications in clinical practice. Huisi Wu, Zebin Zhao 0004, Zhaoze Wang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | 3DSN-Net: A 3-D Scale-Aware convNet With Nonlocal Context Guidance for Kidney and Tumor Segmentation From CT VolumesabstractAutomatic kidney and tumor segmentation from CT volumes is a critical prerequisite/tool for diagnosis and surgical treatment (such as partial nephrectomy). However, it remains a particularly challenging issue as kidneys and tumors often exhibit large-scale variations, irregular shapes, and blurring boundaries. We propose a novel 3-D network to comprehensively tackle these problems; we call it 3DSN-Net. Compared with existing solutions, it has two compelling characteristics. First, with a new scale-aware feature extraction (SAFE) module, the proposed 3DSN-Net is capable of adaptively selecting appropriate receptive fields according to the sizes of targets instead of indiscriminately enlarging them, which is particularly essential for improving the segmentation accuracy of the tumor with large scale variation. Second, we propose a novel yet efficient nonlocal context guidance (NCG) mechanism to capture global dependencies to tackle irregular shapes and blurring boundaries of kidneys and tumors. Instead of directly harnessing a 3-D NCG mechanism, which makes the number of parameters exponentially increase and hence the network difficult to be trained under limited training data, we develop a 2.5D NCG mechanism based on projections of feature cubes, which achieves a tradeoff between segmentation accuracy and network complexity. We extensively evaluate the proposed 3DSN-Net on the famous KiTS dataset with many challenging kidney and tumor cases. Experimental results demonstrate our solution consistently outperforms state-of-the-art 3-D networks after being equipped with scale aware and NCG mechanisms, particularly for tumor segmentation. Huisi Wu, Baiming Zhang, Zhuoying Li, Harry Qin, Tong-Yee Lee |
IEEE Trans. Cybern. | 1 |
| 2024 | Prior-Guided Adversarial Learning With Hypergraph for Predicting Abnormal Connections in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is characterized by alterations of the brain's structural and functional connectivity during its progressive degenerative processes. Existing auxiliary diagnostic methods have accomplished the classification task, but few of them can accurately evaluate the changing characteristics of brain connectivity. In this work, a prior-guided adversarial learning with hypergraph (PALH) model is proposed to predict abnormal brain connections using triple-modality medical images. Concretely, a prior distribution from anatomical knowledge is estimated to guide multimodal representation learning using an adversarial strategy. Also, the pairwise collaborative discriminator structure is further utilized to narrow the difference in representation distribution. Moreover, the hypergraph perceptual network is developed to effectively fuse the learned representations while establishing high-order relations within and between multimodal images. Experimental results demonstrate that the proposed model outperforms other related methods in analyzing and predicting AD progression. More importantly, the identified abnormal connections are partly consistent with previous neuroscience discoveries. The proposed model can evaluate the characteristics of abnormal brain connections at different stages of AD, which is helpful for cognitive disease study and early treatment. Qiankun Zuo, Huisi Wu, C. L. Philip Chen, Bai Ying Lei, Shuqiang Wang |
IEEE Trans. Cybern. | 2 |
| 2024 | Multi-Resolution Expansion of Analysis in Time-Frequency Domain for Time Series ForecastingabstractTime series forecasting plays a crucial role in various real-world applications, such as finance, energy, traffic, and healthcare, providing valuable insights for decision-making processes. The aggregation of information windows with different resolutions has proven effective in time series forecasting tasks and provides the model diverse contextual information. As a result, the network can better capture and model the heterogeneity present in the data, thereby improving performance. However, most of the current work focuses on extracting multilevel-resolution information without considering the possibility that important information can be supplemented. Meanwhile, these methods also tend to ignore the effect of resolution on frequency. To address these challenges, we introduce the Time-Frequency Domain Multi-Resolution Expansion Network (TFMRN) for long-series forecasting using multi-resolution time-frequency data. The proposed TFMRN aims to expand the data in both the time and frequency domains, enabling the model to capture finer details that may not be evident in the original data. In addition, we also propose an Information Gating Unit (IGU) to enhance the selection and guidance of rich information from the expanded time-frequency multi-resolution data. Experimental results demonstrate that the proposed method yields better performance compared with the state-of-the-art methods in both univariate and multivariate time forecasting tasks. Codes are available athttps://github.com/CurtainKevin/TFMRN-master. Kaiwen Yan, Chen Long, Huisi Wu, Zhenkun Wen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Dynamic-Guided Spatiotemporal Attention for Echocardiography Video SegmentationabstractLeft ventricle (LV) endocardium segmentation in echocardiography video has received much attention as an important step in quantifying LV ejection fraction. Most existing methods are dedicated to exploiting temporal information on top of 2D convolutional networks. In addition to single appearance semantic learning, some research attempted to introduce motion cues through the optical flow estimation (OFE) task to enhance temporal consistency modeling. However, OFE in these methods is tightly coupled to LV endocardium segmentation, resulting in noisy inter-frame flow prediction, and post-optimization based on these flows accumulates errors. To address these drawbacks, we propose dynamic-guided spatiotemporal attention (DSA) for semi-supervised echocardiography video segmentation. We first fine-tune the off-the-shelf OFE network RAFT on echocardiography data to provide dynamic information. Taking inter-frame flows as additional input, we use a dual-encoder structure to extract motion and appearance features separately. Based on the connection between dynamic continuity and semantic consistency, we propose a bilateral feature calibration module to enhance both features. For temporal consistency modeling, the DSA is proposed to aggregate neighboring frame context using deformable attention that is realized by offsets grid attention. Dynamic information is introduced into DSA through a bilateral offset estimation module to effectively combine with appearance semantics and predict attention offsets, thereby guiding semantic-based spatiotemporal attention. We evaluated our method on two popular echocardiography datasets, CAMUS and EchoNet-Dynamic, and achieved state-of-the-art. Jingyin Lin, Wende Xie, Huisi Wu |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Federated Semi-Supervised Medical Image Segmentation via Prototype-Based Pseudo-Labeling and Contrastive LearningabstractExisting federated learning works mainly focus on the fully supervised training setting. In realistic scenarios, however, most clinical sites can only provide data without annotations due to the lack of resources or expertise. In this work, we are concerned with the practical yet challenging federated semi-supervised segmentation (FSSS), where labeled data are only with several clients and other clients can just provide unlabeled data. We take an early attempt to tackle this problem and propose a novel FSSS method with prototype-based pseudo-labeling and contrastive learning. First, we transmit a labeled-aggregated model, which is obtained based on prototype similarity, to each unlabeled client, to work together with the global model for debiased pseudo labels generation via a consistency- and entropy-aware selection strategy. Second, we transfer image-level prototypes from labeled datasets to unlabeled clients and conduct prototypical contrastive learning on unlabeled models to enhance their discriminative power. Finally, we perform the dynamic model aggregation with a designed consistency-aware aggregation strategy to dynamically adjust the aggregation weights of each local model. We evaluate our method on COVID-19 X-ray infected region segmentation, COVID-19 CT infected region segmentation and colorectal polyp segmentation, and experimental results consistently demonstrate the effectiveness of our proposed method. Codes areavailable at https://github.com/zhangbaiming/FedSemiSeg. Huisi Wu, Baiming Zhang, Cheng Chen 0013, Harry Qin |
IEEE Trans. Medical Imaging | 1 |
| 2024 | MHW-GAN: Multidiscriminator Hierarchical Wavelet Generative Adversarial Network for Multimodal Image FusionabstractImage fusion technology aims to obtain a comprehensive image containing a specific target or detailed information by fusing data of different modalities. However, many deep learning-based algorithms consider edge texture information through loss functions instead of specifically constructing network modules. The influence of the middle layer features is ignored, which leads to the loss of detailed information between layers. In this article, we propose a multidiscriminator hierarchical wavelet generative adversarial network (MHW-GAN) for multimodal image fusion. First, we construct a hierarchical wavelet fusion (HWF) module as the generator of MHW-GAN to fuse feature information at different levels and scales, which avoids information loss in the middle layers of different modalities. Second, we design an edge perception module (EPM) to integrate edge information from different modalities to avoid the loss of edge information. Third, we leverage the adversarial learning relationship between the generator and three discriminators for constraining the generation of fusion images. The generator aims to generate a fusion image to fool the three discriminators, while the three discriminators aim to distinguish the fusion image and edge fusion image from two source images and the joint edge image, respectively. The final fusion image contains both intensity information and structure information via adversarial learning. Experiments on public and self-collected four types of multimodal image datasets show that the proposed algorithm is superior to the previous algorithms in terms of both subjective and objective evaluation. Cheng Zhao 0003, Peng Yang 0011, Feng Zhou 0003, Guanghui Yue 0001, Shuigen Wang, Huisi Wu, Guoliang Chen 0005, Tianfu Wang 0001, Bai Ying Lei |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Suitable and Style-Consistent Multi-Texture Recommendation for Cartoon IllustrationsabstractTexture plays an important role in cartoon illustrations to display object materials and enrich visual experiences. Unfortunately, manually designing and drawing an appropriate texture is not easy even for proficient artists, let alone novice or amateur people. While there exist tons of textures on the Internet, it is not easy to pick an appropriate one using traditional text-based search engines. Although several texture pickers have been proposed, they still require the users to browse the textures by themselves, which is still labor-intensive and time-consuming. In this article, an automatic texture recommendation system is proposed for recommending multiple textures to replace a set of user-specified regions in a cartoon illustration with visually pleasant look. Two measurements, the suitability measurement and the style-consistency measurement, are proposed to make sure that the recommended textures are suitable for cartoon illustration and at the same time mutually consistent in style. The suitability is measured based on the synthesizability, cartoonity, and region fitness of textures. The style-consistency is predicted using a learning-based solution since it is subjective to judge whether two textures are consistent in style. An optimization problem is formulated and solved via the genetic algorithm. Our method is validated on various cartoon illustrations, and convincing results are obtained. Huisi Wu, Zhaoze Wang, Xueting Liu 0001, Tong-Yee Lee |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Separating Shading and Reflectance From Cartoon IllustrationsabstractShading plays an important role in cartoon drawings to present the 3D lighting and depth information in a 2D image to improve the visual information and pleasantness. But it also introduces apparent challenges in analyzing and processing the cartoon drawings for different computer graphics and vision applications, such as segmentation, depth estimation, and relighting. Extensive research has been made in removing or separating the shading information to facilitate these applications. Unfortunately, the existing researches only focused on natural images, which are natively different from cartoons since the shading in natural images is physically correct and can be modeled based on physical priors. However, shading in cartoons is manually created by artists, which may be imprecise, abstract, and stylized. This makes it extremely difficult to model the shading in cartoon drawings. Without modeling the shading prior, in the paper, we propose a learning-based solution to separate the shading from the original colors using a two-branch system consisting of two subnetworks. To the best of our knowledge, our method is the first attempt in separating shading information from cartoon drawings. Our method significantly outperforms the methods tailored for natural images. Extensive evaluations have been performed with convincing results in all cases. Ziheng Ma, Chengze Li, Xueting Liu 0001, Huisi Wu, Zhenkun Wen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Shading-Guided Manga Screening From ReferenceabstractManga screening is a critical process in manga production, which still requires intensive labor and cost. Existing manga screening methods either generate simple dotted screentones only or rely on color information and manual hints during screentone selection. Due to the large domain gap between line drawings and screened manga, and the difficulties in generating high-quality, properly selected and shaded screentones, even state-of-the-art deep learning methods cannot convert line drawings to screened manga well. Besides, ambiguity exists in the screening process since different artists may screen differently for the same line drawing. In this article, we propose to introduce shaded line drawing as the intermediate counterpart of the screened manga so that the manga screening task can be decomposed into two sub-tasks, generating shading from a line drawing and replacing shading with proper screentones. The reference image is adopted to resolve the ambiguity issue and provides options and controls on the generated screened manga. We proposed a reference-based shading generation network and a reference-based screentone generation module to achieve the two sub-tasks individually. We conduct extensive visual and quantitative experiments to verify the effectiveness of our system. Results and statistics show that our method outperforms existing methods on the manga screening task. Huisi Wu, Ziheng Ma, Wenliang Wu, Xueting Liu 0001, Chengze Li, Zhenkun Wen |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Super-efficient Echocardiography Video Segmentation via Proxy- and Kernel-Based Semi-supervised LearningabstractAutomatic segmentation of left ventricular endocardium in echocardiography videos is critical for assessing various cardiac functions and improving the diagnosis of cardiac diseases. It is yet a challenging task due to heavy speckle noise, significant shape variability of cardiac structure, and limited labeled data. Particularly, the real-time demand in clinical practice makes this task even harder. In this paper, we propose a novel proxy- and kernel-based semi-supervised segmentation network (PKEcho-Net) to comprehensively address these challenges. We first propose a multi-scale region proxy (MRP) mechanism to model the region-wise contexts, in which a learnable region proxy with an arbitrary shape is developed in each layer of the encoder, allowing the network to identify homogeneous semantics and hence alleviate the influence of speckle noise on segmentation. To sufficiently and efficiently exploit temporal consistency, different from traditional methods which only utilize the temporal contexts of two neighboring frames via feature warping or self-attention mechanism, we formulate the semi-supervised segmentation with a group of learnable kernels, which can naturally and uniformly encode the appearances of left ventricular endocardium, as well as extracting the inter-frame contexts across the whole video to resist the fast shape variability of cardiac structures. Extensive experiments have been conducted on two famous public echocardiography video datasets, EchoNet-Dynamic and CAMUS. Our model achieves the best performance-efficiency trade-off when compared with other state-of-the-art approaches, attaining comparative accuracy with a much faster speed. The code is available at https://github.com/JingyinLin/PKEcho-Net. Huisi Wu, Jingyin Lin, Wende Xie, Harry Qin |
AAAI | 1 |
| 2023 | ACL-Net: Semi-supervised Polyp Segmentation via Affinity Contrastive LearningabstractAutomatic polyp segmentation from colonoscopy images is an essential prerequisite for the development of computer-assisted therapy. However, the complex semantic information and the blurred edges of polyps make segmentation extremely difficult. In this paper, we propose a novel semi-supervised polyp segmentation framework using affinity contrastive learning (ACL-Net), which is implemented between student and teacher networks to consistently refine the pseudo-labels for semi-supervised polyp segmentation. By aligning the affinity maps between the two branches, a better polyp region activation can be obtained to fully exploit the appearance-level context encoded in the feature maps, thereby improving the capability of capturing not only global localization and shape context, but also the local textural and boundary details. By utilizing the rich inter-image affinity context and establishing a global affinity context based on the memory bank, a cross-image affinity aggregation (CAA) module is also implemented to further refine the affinity aggregation between the two branches. By continuously and adaptively refining pseudo-labels with optimized affinity, we can improve the semi-supervised polyp segmentation based on the mutually reinforced knowledge interaction among contrastive learning and consistency learning iterations. Extensive experiments on five benchmark datasets, including Kvasir-SEG, CVC-ClinicDB, CVC-300, CVC-ColonDB and ETIS, demonstrate the effectiveness and superiority of our method. Codes are available at https://github.com/xiewende/ACL-Net. Huisi Wu, Wende Xie, Jingyin Lin, Xinrong Guo |
AAAI | 1 |
| 2023 | Zero-shot Micro-video Classification with Neural Variational Inference in Graph Prototype NetworkabstractMicro-video classification plays a central role in online content recommendation platforms, such as Kwai and Tik-Tok. Existing works on video classification largely exploit the interactions between users and items as well as the item labels to provide quality recommendation services. However, scarce or even no labeled data of emerging videos is a great challenge for existing classification methods. In this paper, we propose a zero-shot micro-video classification model (NVIGPN) by exploiting the hidden topics behind items to guide the representation learning in user-item interactions. Specifically, we study this zero-shot classification in two stages: (1) exploiting a generalized semantic hidden topic descriptions for transferable knowledge learning, and (2) designing a graph-based learning model for guiding the minor seen class information to the unseen ones. Through mining the transferable knowledge between the hidden topics and the small number of the seen classes, NVIGPN can achieves state-of-the-art performances in predicting the unseen classes of micro-videos. We conduct extensive experiments to demonstrate the effectiveness of our method. Junyang Chen 0001, Zhijiang Dai, Huisi Wu, Mengzhu Wang, Qin Zhang 0011, Huan Wang 0005 |
ACM Multimedia | 4 |
| 2023 | Interpolation Normalization for Contrast Domain GeneralizationabstractDomain generalization refers to the challenge of training a model from various source domains that can generalize well to unseen target domains. Contrastive learning is a promising solution that aims to learn domain-invariant representations by utilizing rich semantic relations among sample pairs from different domains. One simple approach is to bring positive sample pairs from different domains closer, while pushing negative pairs further apart. However, in this paper, we find that directly applying contrastive-based methods is not effective in domain generalization. To overcome this limitation, we propose to leverage a novel contrastive learning approach that promotes class-discriminative and class-balanced features from source domains. Essentially, clusters of sample representations from the same category are encouraged to cluster, while those from different categories are spread out, thus enhancing the model's generalization capability. Furthermore, most existing contrastive learning methods use batch normalization, which may prevent the model from learning domain-invariant features. Inspired by recent research on universal representations for neural networks, we propose a simple emulation of this mechanism by utilizing batch normalization layers to distinguish visual classes and formulating a way to combine them for domain generalization tasks. Our experiments demonstrate a significant improvement in classification accuracy over state-of-the-art techniques on popular domain generalization benchmarks, including Digits-DG, PACS, Office-Home and DomainNet. Mengzhu Wang, Junyang Chen 0001, Huan Wang 0005, Huisi Wu, Zhidan Liu 0001, Qin Zhang 0011 |
ACM Multimedia | 4 |
| 2023 | PolypSeg+: A Lightweight Context-Aware Network for Real-Time Polyp SegmentationabstractAutomatic polyp segmentation from colonoscopy videos is a prerequisite for the development of a computer-assisted colon cancer examination and diagnosis system. However, it remains a very challenging task owing to the large variation of polyps, the low contrast between polyps and background, and the blurring boundaries of polyps. More importantly, real-time performance is a necessity of this task, as it is anticipated that the segmented results can be immediately presented to the doctor during the colonoscopy intervention for his/her prompt decision and action. It is difficult to develop a model with powerful representation capability, yielding satisfactory segmentation results and, simultaneously, maintaining real-time performance. In this article, we present a novel lightweight context-aware network, namely, PolypSeg+, attempting to capture distinguishable features of polyps without increasing network complexity and sacrificing time performance. To achieve this, a set of novel lightweight techniques is developed and integrated into the proposed PolypSeg+, including an adaptive scale context (ASC) module equipped with a lightweight attention mechanism to tackle the large-scale variation of polyps, an efficient global context (EGC) module to promote the fusion of low-level and high-level features by excluding background noise and preserving boundary details, and a lightweight feature pyramid fusion (FPF) module to further refine the features extracted from the ASC and EGC. We extensively evaluate the proposed PolypSeg+ on two famous public available datasets for the polyp segmentation task: 1) Kvasir-SEG and 2) CVC-Endoscenestill. The experimental results demonstrate that our PolypSeg+ consistently outperforms other state-of-the-art networks by achieving better segmentation accuracy in much less running time. The code is available at https://github.com/szu-zzb/polypsegplus. Huisi Wu, Zebin Zhao 0004, Jiafu Zhong, Wei Wang 0117, Zhenkun Wen, Harry Qin |
IEEE Trans. Cybern. | 1 |
| 2023 | Cross-Image Dependency Modeling for Breast Ultrasound SegmentationabstractWe present a novel deep network (namely BUSSeg) equipped with both within- and cross-image long-range dependency modeling for automated lesions segmentation from breast ultrasound images, which is a quite daunting task due to (1) the large variation of breast lesions, (2) the ambiguous lesion boundaries, and (3) the existence of speckle noise and artifacts in ultrasound images. Our work is motivated by the fact that most existing methods only focus on modeling the within-image dependencies while neglecting the cross-image dependencies, which are essential for this task under limited training data and noise. We first propose a novel cross-image dependency module (CDM) with a cross-image contextual modeling scheme and a cross-image dependency loss (CDL) to capture more consistent feature expression and alleviate noise interference. Compared with existing cross-image methods, the proposed CDM has two merits. First, we utilize more complete spatial features instead of commonly used discrete pixel vectors to capture the semantic dependencies between images, mitigating the negative effects of speckle noise and making the acquired features more representative. Second, the proposed CDM includes both intra- and inter-class contextual modeling rather than just extracting homogeneous contextual dependencies. Furthermore, we develop a parallel bi-encoder architecture (PBA) to tame a Transformer and a convolutional neural network to enhance BUSSeg's capability in capturing within-image long-range dependencies and hence offer richer features for CDM. We conducted extensive experiments on two representative public breast ultrasound datasets, and the results demonstrate that the proposed BUSSeg consistently outperforms state-of-the-art approaches in most metrics. Huisi Wu, Xinrong Guo, Zhenkun Wen, Harry Qin |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Feature Masking on Non-Overlapping Regions for Detecting Dense Cells in Blood Smear ImageabstractDetecting cells in blood smear images is of great significance for automatic diagnosis of blood diseases. However, this task is rather challenging, mainly because there are dense cells that are often overlapping, making some of the occluded boundary parts invisible. In this paper, we propose a generic and effective detection framework that exploits non-overlapping regions (NOR) for providing discriminative and confident information to compensate the intensity deficiency. In particular, we propose a feature masking (FM) to exploit the NOR mask generated from the original annotation information, which can guide the network to extract NOR features as supplementary information. Furthermore, we exploit NOR features to directly predict the NOR bounding boxes (NOR BBoxes). NOR BBoxes are combined with the original BBoxes for generating one-to-one corresponding BBox-pairs that are used for further improving the detection performance. Different from the non-maximum suppression (NMS), our proposed non-overlapping regions NMS (NOR-NMS) uses the NOR BBoxes in the BBox-pairs to calculate intersection over union (IoU) for suppressing redundant BBoxes, and consequently retains the corresponding original BBoxes, circumventing the dilemma of NMS. We conducted extensive experiments on two publicly available datasets, with positive results demonstrating the effectiveness of the proposed method against existing methods. Huisi Wu, Canfeng Lin, Jiasheng Liu, Youyi Song, Zhenkun Wen, Harry Qin |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Continual Nuclei Segmentation via Prototype-Wise Relation Distillation and Contrastive LearningabstractDeep learning models have achieved remarkable success in multi-type nuclei segmentation. These models are mostly trained at once with the full annotation of all types of nuclei available, while lack the ability of continually learning new classes due to the problem of catastrophic forgetting. In this paper, we study the practical and important class-incremental continual learning problem, where the model is incrementally updated to new classes without accessing to previous data. We propose a novel continual nuclei segmentation method, to avoid forgetting knowledge of old classes and facilitate the learning of new classes, by achieving feature-level knowledge distillation with prototype-wise relation distillation and contrastive learning. Concretely, prototype-wise relation distillation imposes constraints on the inter-class relation similarity, encouraging the encoder to extract similar class distribution for old classes in the feature space. Prototype-wise contrastive learning with a hard sampling strategy enhances the intra-class compactness and inter-class separability of features, improving the performance on both old and new classes. Experiments on two multi-type nuclei segmentation benchmarks, i.e., MoNuSAC and CoNSeP, demonstrate the effectiveness of our method with superior performance over many competitive methods. Codes are available at https://github.com/zzw-szu/CoNuSeg. Huisi Wu, Zhaoze Wang, Zebin Zhao 0004, Cheng Chen 0013, Harry Qin |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Context Prior Guided Semantic Modeling for Biomedical Image SegmentationabstractMost state-of-the-art deep networks proposed for biomedical image segmentation are developed based on U-Net. While remarkable success has been achieved, its inherent limitations hinder it from yielding more precise segmentation. First, its receptive field is limited due to the fixed kernel size, which prevents the network from modeling global context information. Second, when spatial information captured by shallower layer is directly transmitted to higher layers by skip connections, the process inevitably introduces noise and irrelevant information to feature maps and blurs their semantic meanings. In this article, we propose a novel segmentation network equipped with a new context prior guidance (CPG) module to overcome these limitations for biomedical image segmentation, namely context prior guidance network (CPG-Net). Specifically, we first extract a set of context priors under the supervision of a coarse segmentation and then employ these context priors to model the global context information and bridge the spatial-semantic gap between high-level features and low-level features. The CPG module contains two major components: context prior representation (CPR) and semantic complement flow (SCF). CPR is used to extract pixels belonging to the same objects and hence produce more discriminative features to distinguish different objects. We further introduce deep semantic information for each CPR by the SCF mechanism to compensate the semantic information diluted during the decoding. We extensively evaluate the proposed CPG-Net on three famous biomedical image segmentation tasks with diverse imaging modalities and semantic environments. Experimental results demonstrate the effectiveness of our network, consistently outperforming state-of-the-art segmentation networks in all the three tasks. Codes are available at https://github.com/zzw-szu/CPGNet . Huisi Wu, Zhaoze Wang, Zhuoying Li, Zhenkun Wen, Harry Qin |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | AddCR: a data-driven cartoon remastering
Yinghua Liu, Chengze Li, Xueting Liu 0001, Huisi Wu, Zhenkun Wen |
Vis. Comput. | 4 |
| 2022 | Left and Right Ventricular Segmentation Based on 3D Region-Aware U-NetabstractThe cardiac is one of the essential organs, and the segmentation of the left and right ventricular of cardiac is essential in diagnosing various heart diseases. The most popular method for the segmentation of 3D MRI images is the nnUNet. However, the 3D MRI volume of the ventricular contains other organs which interfere with the segmentation of the ventricular. Hence, we proposed a novel region-aware U-Net segmentation method RegUNet for ventricular segmentation. RegUNet improves the ventricular's segmentation performance by first capturing the region of interest (RoI) of the ventricular and then segmenting the ventricular with the captured RoI features, which reduces the segmentation module's difficulty by keeping the cardiac's features and leaving others such that RegUNet can focus on ventricular segmentation. Besides, since the model segments the ventricular with the captured RoI features, it saves the model's computing resources from identifying the background of the volume. Since 3D cardiac MRI volumes scanned by the different devices have diverse statistical characteristics, which causes the model's performance in processing the multi-source cardiac volumes to be unstable. We stabilize the model's performance with a multi-sources feature normalization strategy, which normalizes the feature from a different source with different parameters. We validated the proposed method on the M&MS dataset, a multi-sources 3D MRI cardiac segmentation dataset. Experiments showed that RegUNet's segmentation ability reached the state-of-the-art. Xueting Liu 0001, Huisi Wu, Zhenkun Wen, LinLin Shen |
CBMS | 4 |
| 2022 | Authenticity Identification of Qi Baishi's Shrimp Painting with Dynamic Token Enhanced Visual Transformer
Xueting Liu 0001, Huisi Wu, Fu Qi |
CGI | 4 |
| 2022 | Cross-patch Dense Contrastive Learning for Semi-supervised Segmentation of Cellular Nuclei in Histopathologic ImagesabstractWe study the semi-supervised learning problem, using a few labeled data and a large amount of unlabeled data to train the network, by developing a cross-patch dense contrastive learning framework, to segment cellular nuclei in histopathologic images. This task is motivated by the expensive burden on collecting labeled data for histopathologic image segmentation tasks. The key idea of our method is to align features of teacher and student networks, sampled from cross-image in both patch- and pixel-levels, for enforcing the intra-class compactness and inter-class separability of features that as we shown is helpful for extracting valuable knowledge from unlabeled data. We also design a novel optimization framework that combines consistency regularization and entropy minimization techniques, showing good property in eviction of gradient vanishing. We assess the proposed method on two publicly available datasets, and obtain positive results on extensive experiments, outperforming the state-of-the-art methods. Codes are available at https://github.com/zzw-szu/CDCL. Huisi Wu, Zhaoze Wang, Youyi Song, Harry Qin |
CVPR | 1 |
| 2022 | Dual Contrastive Learning with Anatomical Auxiliary Supervision for Few-Shot Medical Image Segmentation
Huisi Wu, Fangyan Xiao, Chongxin Liang |
ECCV (20) | 1 |
| 2022 | Sample Hardness Based Gradient Loss for Long-Tailed Cervical Cell Detection
Minmin Liu, Xuechen Li 0001, Xiangbo Gao, Junliang Chen 0002, LinLin Shen, Huisi Wu |
MICCAI (2) | 6 |
| 2022 | Vectorizing Line Drawings of Arbitrary Thickness via Boundary-based Topology ReconstructionabstractAbstract Vectorization is a commonly used technique for converting raster images to vector format and has long been a research focus in computer graphics and vision. While a number of attempts have been made to extract the topology of line drawings and further convert them to vector representations, the existing methods commonly focused on resolving junctions composed of thin lines. They usually fail for line drawings composed of thick lines, especially at junctions. In this paper, we propose an automatic line drawing vectorization method that can reconstruct the topology of line drawings of arbitrary thickness. Our key observation is that no matter the lines are thin or thick, the boundaries of the lines always provide reliable hints for reconstructing the topology. For example, the boundaries of two continuous line segments at a junction are usually smoothly connected. By analyzing the continuity of boundaries, we can better analyze the topology at junctions. In particular, we first extract the skeleton of the input line drawing via thinning. Then we analyze the reliability of the skeleton points based on boundaries. Reliable skeleton points are preserved while unreliable skeleton points are reconstructed based on boundaries again. Finally, the skeleton after reconstruction is vectorized as the output. We apply our method on line drawings of various contents and styles. Satisfying results are obtained. Our method significantly outperforms existing methods for line drawings composed of thick lines. Xueting Liu 0001, Chengze Li, Huisi Wu, Zhenkun Wen |
Comput. Graph. Forum | 4 |
| 2022 | Reference-guided structure-aware deep sketch colorization for cartoonsabstractDigital cartoon production requires extensive manual labor to colorize sketches with visually pleasant color composition and color shading. During colorization, the artist usually takes an existing cartoon image as color guidance, particularly when colorizing related characters or an animation sequence. Reference-guided colorization is more intuitive than colorization with other hints, such as color points or scribbles, or textbased hints. Unfortunately, reference-guided colorization is challenging since the style of the colorized image should match the style of the reference image in terms of both global color composition and local color shading. In this paper, we propose a novel learning-based framework which colorizes a sketch based on a color style feature extracted from a reference color image. Our framework contains a color style extractor to extract the color feature from a color image, a colorization network to generate multi-scale output images by combining a sketch and a color feature, and a multi-scale discriminator to improve the reality of the output image. Extensive qualitative and quantitative evaluations show that our method outperforms existing methods, providing both superior visual quality and style reference consistency in the task of reference-based colorization. Xueting Liu 0001, Wenliang Wu, Chengze Li, Huisi Wu |
Comput. Vis. Media | 5 |
| 2022 | Semi-supervised segmentation of echocardiography videos via noise-resilient spatiotemporal semantic calibration and fusion
Huisi Wu, Jiasheng Liu, Fangyan Xiao, Zhenkun Wen, Harry Qin |
Neurocomputing | 1 |
| 2022 | FAT-Net: Feature adaptive transformers for automated skin lesion segmentation
Huisi Wu, Shihuai Chen, Guilian Chen, Wei Wang 0117, Bai Ying Lei, Zhenkun Wen |
Medical Image Anal. | 1 |
| 2022 | Semi-supervised segmentation of echocardiography videos via noise-resilient spatiotemporal semantic calibration and fusion
Huisi Wu, Jiasheng Liu, Fangyan Xiao, Zhenkun Wen, Harry Qin |
Medical Image Anal. | 1 |
| 2021 | Deep Style Transfer for Line DrawingsabstractLine drawings are frequently used to illustrate ideas and concepts in digital documents and presentations. To compose a line drawing, it is common for users to retrieve multiple line drawings from the Internet and combine them as one image. However, different line drawings may have different line styles and are visually inconsistent when put together. In order that the line drawings can have consistent looks, in this paper, we make the first attempt to perform style transfer for line drawings. The key of our design lies in the fact that centerline plays a very important role in preserving line topology and extracting style features. With this finding, we propose to formulate the style transfer problem as a centerline stylization problem and solve it via a novel style-guided image-to-image translation network. Results and statistics show that our method significantly outperforms the existing methods both visually and quantitatively. Xueting Liu 0001, Wenliang Wu, Huisi Wu, Zhenkun Wen |
AAAI | 3 |
| 2021 | Region-aware Global Context Modeling for Automatic Nerve Segmentation from Ultrasound ImagesabstractWe present a novel deep learning model equipped with a new region-aware global context modeling technique for automatic nerve segmentation from ultrasound images, which is a challenging task due to (1) the large variation and blurred boundaries of targets, (2) the large amount of speckle noise in ultrasound images, and (3) the inherent real-time requirement of this task. It is essential to efficiently capture long-range dependencies by global context modeling for a segmentation network to overcome these challenges. Traditional global context modeling techniques usually explore pixel-aware correlations to establish long-range dependencies, which are usually computation-intensive and greatly degrade time performance. In addition, in this application, pixel-aware modeling may inevitably introduce much speckle noise in the computation and potentially degrade segmentation performance. In this paper, we propose a novel region-aware modeling technique to establish long-range dependencies based on different regions to improve segmentation accuracy while maintaining real-time performance; we call it region-aware pyramid aggregation (RPA) module. In order to adaptively divide the feature maps into a set of semantic-independent regions, we develop an attention mechanism and integrate it into the spatial pyramid network to evaluate the semantic similarity of different regions. We further develop an adaptive pyramid fusion (APF) module to dynamically fuse the multi-level features generated from the decoder to refining the segmentation results. We conducted extensive experiments on a famous public ultrasound nerve image segmentation dataset. Experimental results demonstrate that our method consistently outperforms our rivals in terms of segmentation accuracy. The code is available at https://github.com/jsonliu-szu/RAGCM. Huisi Wu, Jiasheng Liu, Wei Wang 0117, Zhenkun Wen, Harry Qin |
AAAI | 1 |
| 2021 | Precise Yet Efficient Semantic Calibration and Refinement in ConvNets for Real-time Polyp Segmentation from Colonoscopy VideosabstractWe propose a novel convolutional neural network (ConvNet) equipped with two new semantic calibration and refinement approaches for automatic polyp segmentation from colonoscopy videos. While ConvNets set state-of-the-are performance for this task, it is still difficult to achieve satisfactory results in a real-time manner, which is a necessity in clinical practice. The main obstacle is the huge semantic gap between high-level features and low-level features, making it difficult to take full advantage of complementary semantic information contained in these hierarchical features. Compared with existing solutions, which either directly aggregate these features without considering the semantic gap or employ sophisticated non-local modeling techniques to refine semantic information by introduce many extra computational costs, the proposed ConvNet is able to more precisely yet efficiently calibrate and refine semantic information for better segmentation performance without increasing model complexity; we call the proposed ConvNet as SCR-Net, which has two key modules. We first propose a semantic calibration module (SCM) to effectively transmit the semantic information from high-level layers to low-level layers by learning the semantic-spatial relations during the training procedure. We then propose a semantic refinement module (SRM) to, based on the features calibrated by SCM, enhance the discrimination capability of the features for targeting objects. Extensive experiments on the Kvasir-SEG dataset demonstrate that the proposed SCR-Net is capable of achieving better segmentation accuracy than state-of-the-art approaches with a faster speed. The proposed techniques are general enough to be applied to similar applications where precise and efficient multi-level feature fusion is critical. The code is available at https://github.com/jiafuz/SCR-Net. Huisi Wu, Jiafu Zhong, Wei Wang 0117, Zhenkun Wen, Harry Qin |
AAAI | 1 |
| 2021 | Collaborative and Adversarial Learning of Focused and Dispersive Representations for Semi-supervised Polyp SegmentationabstractAutomatic polyp segmentation from colonoscopy images is an essential step in computer aided diagnosis for colorectal cancer. Most of polyp segmentation methods reported in recent years are based on fully supervised deep learning. However, annotation for polyp images by physicians during the diagnosis is time-consuming and costly. In this paper, we present a novel semi-supervised polyp segmentation via collaborative and adversarial learning of focused and dispersive representations learning model, where focused and dispersive extraction module are used to deal with the diversity of location and shape of polyps. In addition, confidence maps produced by a discriminator in an adversarial training framework shows the effectiveness of leveraging unlabeled data and improving the performance of segmentation network. Consistent regularization is further employed to optimize the segmentation networks to strengthen the representation of the outputs of focused and dispersive extraction module. We also propose an auxiliary adversarial learning method to better leverage unlabeled examples to further improve semantic segmentation accuracy. We conduct extensive experiments on two famous polyp datasets: Kvasir-SEG and CVC-Clinic DB. Experimental results demonstrate the effectiveness of the proposed model, consistently outperforming state-of-the-art semi-supervised segmentation models based on adversarial training and even some advanced fully supervised models. Huisi Wu, Guilian Chen, Zhenkun Wen, Harry Qin |
ICCV | 1 |
| 2021 | Automated Malaria Cells Detection from Blood Smears Under Severe Class Imbalance via Importance-Aware Balanced Group Softmax
Canfeng Lin, Huisi Wu, Zhenkun Wen, Harry Qin |
MICCAI (8) | 2 |
| 2021 | Deep texture cartoonization via unsupervised appearance regularization
Huisi Wu, Xueting Liu 0001, Chengze Li, Wenliang Wu |
Comput. Graph. | 1 |
| 2021 | Optimized HRNet for image semantic segmentation
Huisi Wu, Chongxin Liang, Meng-Shu Liu, Zhenkun Wen |
Expert Syst. Appl. | 1 |
| 2021 | Deep boundary-aware semantic image segmentationabstractAbstract While extensive research efforts have been made in semantic image segmentation, the state‐of‐the‐art methods still suffer from blurry boundaries and mismatched objects due to the insufficient multiscale adaptability. In this paper, we propose a two‐branch convolutional neural network (CNN) approach to capture the multiscale context and the boundary information with the two branches, respectively. To capture the multiscale context, we propose to embed self‐attention mechanism to the atrous spatial pyramid pooling network. To capture the boundary information, we propose to fuse the low‐level features in boundary feature extraction for refining the extracted boundaries via a feature fusion layer (FFL). With FFL, our method can improve the segmentation result with clearer boundaries. A new loss function is proposed which contains a segmentation loss and a boundary loss. Experiments show that our method can predict the boundaries of objects more clearly and have better performance for small‐scale objects. Huisi Wu, Xueting Liu 0001, Ping Li 0016 |
Comput. Animat. Virtual Worlds | 1 |
| 2021 | Automated left ventricular segmentation from cardiac magnetic resonance images via adversarial learning with multi-stage pose estimation network and co-discriminator
Huisi Wu, Xuheng Lu, Bai Ying Lei, Zhenkun Wen |
Medical Image Anal. | 1 |
| 2021 | SCS-Net: A Scale and Context Sensitive Network for Retinal Vessel Segmentation
Huisi Wu, Wei Wang 0117, Jiafu Zhong, Bai Ying Lei, Zhenkun Wen, Harry Qin |
Medical Image Anal. | 1 |
| 2021 | Automatic Symmetry Detection From Brain MRI Based on a 2-Channel Convolutional Neural NetworkabstractSymmetry detection is a method to extract the ideal mid-sagittal plane (MSP) from brain magnetic resonance (MR) images, which can significantly improve the diagnostic accuracy of brain diseases. In this article, we propose an automatic symmetry detection method for brain MR images in 2-D slices based on a 2-channel convolutional neural network (CNN). Different from the existing detection methods that mainly rely on the local image features (gradient, edge, etc.) to determine the MSP, we use a CNN-based model to implement the brain symmetry detection, which does not require any local feature detections and feature matchings. By training to learn a wide variety of benchmarks in the brain images, we can further use a 2-channel CNN to evaluate the similarity between the pairs of brain patches, which are randomly extracted from the whole brain slice based on a Poisson sampling. Finally, a scoring and ranking scheme is used to identify the optimal symmetry axis for each input brain MR slice. Our method was evaluated in 2166 artificial synthesized brain images and 3064 collected in vivo MR images, which included both healthy and pathological cases. The experimental results display that our method achieves excellent performance for symmetry detection. Comparisons with the state-of-the-art methods also demonstrate the effectiveness and advantages for our approach in achieving higher accuracy than the previous competitors. Huisi Wu, Xiujuan Chen, Ping Li 0016, Zhenkun Wen |
IEEE Trans. Cybern. | 1 |
| 2021 | Automated Skin Lesion Segmentation Via an Adaptive Dual Attention ModuleabstractWe present a convolutional neural network (CNN) equipped with a novel and efficient adaptive dual attention module (ADAM) for automated skin lesion segmentation from dermoscopic images, which is an essential yet challenging step for the development of a computer-assisted skin disease diagnosis system. The proposed ADAM has three compelling characteristics. First, we integrate two global context modeling mechanisms into the ADAM, one aiming at capturing the boundary continuity of skin lesion by global average pooling while the other dealing with the shape irregularity by pixel-wise correlation. In this regard, our network, thanks to the proposed ADAM, is capable of extracting more comprehensive and discriminative features for recognizing the boundary of skin lesions. Second, the proposed ADAM supports multi-scale resolution fusion, and hence can capture multi-scale features to further improve the segmentation accuracy. Third, as we harness a spatial information weighting method in the proposed network, our method can reduce a lot of redundancies compared with traditional CNNs. The proposed network is implemented based on a dual encoder architecture, which is able to enlarge the receptive field without greatly increasing the network parameters. In addition, we assign different dilation rates to different ADAMs so that it can adaptively capture distinguishing features according to the size of a lesion. We extensively evaluate the proposed method on both ISBI2017 and ISIC2018 datasets and the experimental results demonstrate that, without using network ensemble schemes, our method is capable of achieving better segmentation performance than state-of-the-art deep learning models, particularly those equipped with attention mechanisms. Huisi Wu, Junquan Pan, Zhuoying Li, Zhenkun Wen, Harry Qin |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Deep Texture Exemplar Extraction Based on Trimmed T-CNNabstractTexture exemplar has been widely used in synthesizing 3D movie scenes and appearances of virtual objects. Unfortunately, conventional texture synthesis methods usually only emphasized on generating optimal target textures with arbitrary sizes or diverse effects, and put little attention to automatic texture exemplar extraction. Obtaining texture exemplars is still a labor intensive task, which usually requires carefully cropping and post-processing. In this paper, we present an automatic texture exemplar extraction based on Trimmed Texture Convolutional Neural Network (Trimmed T-CNN). Specifically, our Trimmed T-CNN is filter banks for texture exemplar classification and recognition. Our Trimmed T-CNN is learned with a standard ideal exemplar dataset containing thousands of desired texture exemplars, which were collected and cropped by our invited artists. To efficiently identify the exemplar candidates from an input image, we employ a selective search algorithm to extract the potential texture exemplar patches. We then put all candidates into our Trimmed T-CNN for learning ideal texture exemplars based on our filter banks. Finally, optimal texture exemplars are identified with a scoring and ranking scheme. Our method is evaluated with various kinds of textures and user studies. Comparisons with different feature-based methods and different deep CNN architectures (AlexNet, VGG-M, Deep-TEN and FV-CNN) are also conducted to demonstrate its effectiveness. Huisi Wu, Wei Yan 0036, Ping Li 0016, Zhenkun Wen |
IEEE Trans. Multim. | 1 |
| 2020 | Memory-Efficient Automatic Kidney and Tumor Segmentation Based on Non-local Context Guided 3D U-Net
Zhuoying Li, Junquan Pan, Huisi Wu, Zhenkun Wen, Harry Qin |
MICCAI (4) | 3 |
| 2020 | RVSeg-Net: An Efficient Feature Pyramid Cascade Network for Retinal Vessel Segmentation
Wei Wang 0117, Jiafu Zhong, Huisi Wu, Zhenkun Wen, Harry Qin |
MICCAI (5) | 3 |
| 2020 | PolypSeg: An Efficient Context-Aware Network for Polyp Segmentation from Colonoscopy Videos
Jiafu Zhong, Wei Wang 0117, Huisi Wu, Zhenkun Wen, Harry Qin |
MICCAI (6) | 3 |
| 2020 | Automatic Video Segmentation Based on Information Centroid and Optimized SaliencyCut
Huisi Wu, Meng-Shu Liu, Lulu Yin, Ping Li 0016, Zhenkun Wen, Hon-Cheng Wong |
J. Comput. Sci. Technol. | 1 |
| 2019 | Video Tamper Detection Based on Convolutional Neural Network and Perceptual Hashing Learning
Huisi Wu, Yawen Zhou, Zhenkun Wen |
CGI | 1 |
| 2018 | A Second-Order Variational Framework for Joint Depth Map Estimation and Image DehazingabstractOutdoor images captured in poor weather conditions (e.g., fog or haze) commonly suffer from reduced contrast and visibility. Increasing attention has recently been paid to single image dehazing, i.e., improving image contrast and visibility. It is generally thought that the dehazing performance highly depends on the accurate depth information. In this work, we first obtain the initial depth map by using the popular dark channel prior. A unified second-order variational framework is then proposed to refine the depth map and restore the haze-free image. The introduced second-order framework has the capacity of preserving important structures in both depth map and haze-free image. Furthermore, the proposed framework performs well for several different types of haze situations. The resulting optimization problems related to depth map estimation and latent image restoration can be effectively handled using the primal-dual algorithm under a two-step numerical framework. The effectiveness of our proposed method has been demonstrated by comparing the imaging performance with several state-of-the-art dehazing methods. Ryan Wen Liu, Shengwu Xiong 0001, Huisi Wu |
ICASSP | 3 |
| 2018 | Automatic texture exemplar extraction based on global and local textureness measuresabstractTexture synthesis is widely used for modeling the appearance of virtual objects. However, traditional texture synthesis techniques emphasize creation of optimal target textures, and pay insufficient attention to choice of suitable input texture exemplars. Currently, obtaining texture exemplars from natural images is a labor intensive task for the artists, requiring careful photography and significant postprocessing. In this paper, we present an automatic texture exemplar extraction method based on global and local textureness measures. To improve the efficiency of dominant texture identification, we first perform Poisson disk sampling to randomly and uniformly crop patches from a natural image. For global textureness assessment, we use a GIST descriptor to distinguish textured patches from non-textured patches, in conjunction with SVM prediction. To identify real texture exemplars consisting solely of the dominant texture, we further measure the local textureness of a patch by extracting and matching the local structure (using binary Gabor pattern (BGP)) and dominant color features (using color histograms) between a patch and its sub-regions. Finally, we obtain optimal texture exemplars by scoring and ranking extracted patches using these global and local textureness measures. We evaluate our method on a variety of images with different kinds of textures. A convincing visual comparison with textures manually selected by an artist and a statistical study demonstrate its effectiveness. Huisi Wu, Xiaomeng Lyu, Zhenkun Wen |
Comput. Vis. Media | 1 |
| 2017 | Automatic Leaf Recognition Based on Deep Convolutional Networks
Huisi Wu, Yongkui Xiang, Zhenkun Wen |
ICONIP (3) | 1 |
| 2017 | Tensor Voting Guided Mesh DenoisingabstractMesh denoising is imperative for improving imperfect surfaces acquired by scanning devices. The main challenge is to faithfully retain geometric features and avoid introducing additional artifacts when removing noise. Unlike the existing mesh denoising techniques that focus only on either the first-order features or high-order differential properties, our approach exploits the synergy when facet normals and quadric surfaces are integrated to recover a piecewise smooth surface. In specific, we vote on surface normal tensors from robust statistics to guide the creation of consistent subneighborhoods subsequently used by moving least squares (MLS). This voting naturally leads to a conceptually simple way that gives a unified mesh-denoising framework for not only handling noise but also enabling the recovering of surfaces with both sharp and small-scale features. The effectiveness of our framework stems from: 1) the multiscale tensor voting that avoids the influence from noise; 2) the effective energy minimization strategy to searching the consistent subneighborhoods; and 3) the piecewise MLS that fully prevents the side effects from different subneighborhoods during surface fitting. Our framework is direct, practical, and easy to understand. Comparisons with the state-of-the-art methods demonstrate its outstanding performance on feature preservation and artifact suppression. Mingqiang Wei, Luming Liang, Wai-Man Pang, Jun Wang 0039, Huisi Wu |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2015 | Automatic Leaf Recognition from a Big Hierarchical Image DatabaseabstractAutomatic plant recognition has become a research focus and received more and more attentions recently. However, existing methods usually only focused on leaf recognition from small databases that usually only contain no more than hundreds of species, and none of them reported a stable performance in either recognition accuracy or recognition speed when compared with a big image database. In this paper, we present a novel method for leaf recognition from a big hierarchical image database. Unlike the existing approaches, our method combines the textural gradient histogram with the shape context to form a more distinctive feature for leaf recognition. To achieve efficient leaf image retrieval, we divided the big database into a set of subsets based on mean-shift clustering on the extracted features and build hierarchical k-dimensional trees (KD-trees) to index each cluster in parallel. Finally, the proposed parallel indexing and searching schemes are implemented with MapReduce architectures. Our method is evaluated with extensive experiments on different databases with different sizes. Comparisons to state-of-the-art techniques were also conducted to validate the proposed method. Both visual results and statistical results are shown to demonstrate its effectiveness. Huisi Wu, Zhenkun Wen |
Int. J. Intell. Syst. | 1 |
| 2013 | Parallel structure-aware halftoning
Huisi Wu, Tien-Tsin Wong, Pheng-Ann Heng |
Multim. Tools Appl. | 1 |
| 2010 | Resizing by symmetry-summarizationabstractImage resizing can be achieved more effectively if we have a better understanding of the image semantics. In this paper, we analyze the translational symmetry , which exists in many real-world images. By detecting the symmetric lattice in an image, we can summarize , instead of only distorting or cropping, the image content. This opens a new space for image resizing that allows us to manipulate, not only image pixels, but also the semantic cells in the lattice. As a general image contains both symmetry & non-symmetry regions and their natures are different, we propose to resize symmetry regions by summarization and non-symmetry region by warping. The difference in resizing strategy induces discontinuity at their shared boundary. We demonstrate how to reduce the artifact. To achieve practical resizing applications for general images, we developed a fast symmetry detection method that can detect multiple disjoint symmetry regions, even when the lattices are curved and perspectively viewed. Comparisons to state-of-the-art resizing techniques and a user study were conducted to validate the proposed method. Convincing visual results are shown to demonstrate its effectiveness. Huisi Wu, Yu-Shuen Wang, Kun-Chuan Feng, Tien-Tsin Wong, Tong-Yee Lee, Pheng-Ann Heng |
ACM Trans. Graph. | 1 |