Si Wu 0002

dblp:25/437-2 · DBLP profile ↗
← Back
127ranked-venue papers
13as first author
94since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 75 · 7 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 68 · 10 first-author · 51 since 2021Databases, data management, data science and information retrieval · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Refinement Contrastive Learning of Cell-Gene Associations for Unsupervised Cell Type Identification
abstract
Unsupervised cell type identification is crucial for uncovering and characterizing heterogeneous populations in single cell omics studies. Although a range of clustering methods have been developed, most focus exclusively on intrinsic cellular structure and ignore the pivotal role of cell-gene associations, which limits their ability to distinguish closely related cell types. To this end, we propose a Refinement Contrastive Learning framework (scRCL) that explicitly incorporates cell-gene interactions to derive more informative representations. Specifically, we introduce two contrastive distribution alignment components that reveal reliable intrinsic cellular structures by effectively exploiting cell-cell structural relationships. Additionally, we develop a refinement module that integrates gene-correlation structure learning to enhance cell embeddings by capturing underlying cell-gene associations. This module strengthens connections between cells and their associated genes, refining the representation learning to exploiting biologically meaningful relationships. Extensive experiments on several single-cell RNA-seq and spatial transcriptomics benchmark datasets demonstrate that our method consistently outperforms state-of-the-art baselines in cell-type identification accuracy. Moreover, downstream biological analyses confirm that the recovered cell populations exhibit coherent gene-expression signatures, further validating the biological relevance of our approach.
Yixuan Ye, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
AAAI6
2026 Hierarchical Cross-Modality Interaction for Unified Video-Text Retrieval Modeling
Tianshi Xu, Zhengzheng Sun, Yizheng Hu, Junyuan Shang, Si Wu 0002
MMM (2)5
2026 NicheDeSig: niche-aware deconvolution and adaptive signature analysis for spatial transcriptomics
abstract
MOTIVATION: For spot-based spatial transcriptomics (ST), accurate cell-type deconvolution is essential for downstream analysis since each spot captures mixtures of multiple cell types. Meanwhile, spatial niches define distinct micro-environmental contexts, also salient for biological interpretation. However, existing deconvolution methods usually rely on fixed reference signatures or mapping single cells onto ST spots, without incorporating niche priors or modeling niche-dependent shifts. Consequently, existing methods remain focused on spot-level proportion estimation, with limited ability to support functional analysis of niche-associated molecular programs. RESULTS: We present NicheDeSig for niche-aware deconvolution. NicheDeSig models each cell type through adaptive signatures, enabling spot deconvolution under context-dependent signatures and supporting niche-aware analysis of cell-state variation across spatial micro-environments. Our method achieves strong deconvolution performance across the simulated benchmark datasets and improves spatial fidelity in the simulated colon dataset. The learned signatures recover laminar and white-matter-associated programs in the human dorsolateral prefrontal cortex (DLPFC), domain-stratified tumor microenvironment patterns in breast cancer (BRCA), and region-associated signatures in pancreatic ductal adenocarcinoma sample A (PDAC-A) and colorectal liver metastasis analyses. AVAILABILITY AND IMPLEMENTATION: Source code and the archived code snapshot are available at https://github.com/Davidcoach/NicheDeSig and https://doi.org/10.5281/zenodo.20685597.
Juncheng Zhang, Jinjin Ma, Yong Xu 0007, Hau-San Wong, Si Wu 0002
Bioinform.8
2026 MVHDiff: Leveraging 3D priors for consistent multi-view human image generation with diffusion models
Yan Huang 0031, Hongxin Fu, Zhonghang Li, Yongcan Luo, Si Wu 0002
Neurocomputing6
2026 Diffusion-based blemish prior with statistical modulation for high-fidelity face retouching
Bingbing Zheng, Lianxin Xie, Si Wu 0002
Neurocomputing4
2026 Semantic-aware multi-view person image generation for re-identification
Si Wu 0002, Xin Li 0034, Yong Xu 0007, Yaowei Wang 0001
Image Vis. Comput.2
2026 HumanDiff: Leveraging body-part expert-assisted diffusion transformers for human image generation
Yan Huang 0031, Zhiren Wang, Hongzong Li, Yongcan Luo, Si Wu 0002
Knowl. Based Syst.7
2026 A new paradigm for multi-source sentiment analysis and adaptation with multiple pretrained language models
Rui Li 0045, Cheng Liu 0001, Dazhi Jiang, Si Wu 0002
Knowl. Based Syst.5
2026 ObjectDiff: An object-centric diffusion policy with modality-specific conditioning for robot manipulation
Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002
Knowl. Based Syst.5
2026 Learning region-aware style-content feature transformations for face image beautification
Si Wu 0002
Pattern Recognit.2
2026 SMART: Semantic Matching Contrastive Learning for Partially View-Aligned Clustering
abstract
Multi-view clustering has been empirically shown to improve learning performance by leveraging the inherent complementary information across multiple views of data. However, in real-world scenarios, collecting strictly aligned views is challenging, and learning from both aligned and unaligned data becomes a more practical solution. Partially View-aligned Clustering (PVC) aims to learn correspondences between misaligned view samples to better exploit the potential consistency and complementarity across views, including both aligned and unaligned data. However, most existing PVC methods fail to leverage unaligned data to capture the shared semantics among samples from the same cluster. Moreover, the inherent heterogeneity of multi-view data induces distributional shifts in representations, leading to inaccuracies in establishing meaningful correspondences between cross-view latent features and, consequently, impairing learning effectiveness. To address these challenges, we propose a Semantic MAtching contRasTive learning model (SMART) for PVC. The main idea of our approach is to alleviate the influence of cross-view distributional shifts, thereby facilitating semantic matching contrastive learning to fully exploit semantic relationships in both aligned and unaligned data. Specifically, we mitigate view distribution shifts by aligning cross-view covariance matrices, which enables the inference of a semantic graph for all data. Guided by the learned semantic graph, we further exploit semantic consistency across views through semantic matching contrastive learning. After the optimization of the above mechanisms, our model smoothly performs semantic matching for different view embeddings instead of the cumbersome view realignment, which enables the learned representations to enjoy richer category-level semantics and stronger robustness. Extensive experiments on eight benchmark datasets demonstrate that our method consistently outperforms existing approaches on the PVC problem. The code is available at https://github.com/THPengL/SMART.
Yixuan Ye, Cheng Liu 0001, Hangjun Che, Fei Wang 0056, Zhiwen Yu 0002, Si Wu 0002, Hau-San Wong
IEEE Trans. Circuits Syst. Video Technol.7
2026 Deep Self-Reinforced Multi-View Subspace Clustering for Cancer Subtyping
abstract
Identifying cancer subtypes is crucial for understanding disease progression. With advancements in high-throughput experimental technology, leveraging multiple types of omic data for subtype identification has become feasible. Various integrative cancer subtyping methods present a promising computational approach for identifying cancer subtypes from heterogeneous datasets. While existing integrative cancer subtyping methods have shown promising results in this task, efficiently integrating and clustering multi-omics datasets remains challenging due to high noise levels in omics data, which hinder accurate relationship capture among samples. To overcome this challenge, we propose a new deep multi-view subspace clustering model that introduces a self-reinforced learning strategy. This strategy iteratively enhances the quality of self-representation, crucial for capturing relationships among samples and for clustering. Specifically, during model training, our method is capable of learning a highly reliable self-representation by leveraging a good neighbor learning approach. This capability enables us to capture more accurate and robust relationships among samples. Subsequently, with the assistance of this highly reliable self-representation, we further develop a learnable view-graph fusion approach, which enables us to learn an accurate consensus for clustering and guides the overall model learning process. Additionally, we introduce a local graph-guided learning mechanism based on an initial graph learned from raw data. This mechanism helps prevent the model from converging to suboptimal solutions, thereby avoiding unsatisfactory and unstable results. Experimental results demonstrate that our method outperforms several state-of-the-art methods, verify the effectiveness of our approach in cancer subtype identification task.
Cheng Liu 0001, Baoyuan Zheng, Xibiao Wang, Hang Gao 0014, Fei Wang 0056, Si Wu 0002
IEEE J. Biomed. Health Informatics8
2026 Trustworthy Neighborhoods Mining: Homophily-Aware Neutral Contrastive Learning for Graph Clustering
Yixuan Ye, Cheng Liu 0001, Hangjun Che, Man-Fai Leung, Si Wu 0002, Hau-San Wong
IEEE Trans. Knowl. Data Eng.6
2026 ClassBooth: Boost Class Semantics With Bidirectional Feature Fusion in Text-to-Image Diffusion Models
abstract
Text-to-image (T2I) diffusion models aim to generate images that are both visually realistic and aligned with open-domain textual prompts. However, the leading T2I models often fall short in capturing semantic details, especially for fine-grained object categories. We find that simply expanding original prompts cannot effectively guide the model to generate the desired content, since T2I diffusion models pre-trained on generic datasets lack a mechanism to boost class semantics. In this work, we present ClassBooth, a flexible framework that improves pre-trained T2I diffusion model in rendering class-specific content, while preserving the open-domain generation capability. Toward this end, we introduce an auxiliary class-specific semantic booster conditioned on learnable prompts, which are associated with specific classes to encode fine-grained conditioning information. To enable the T2I model to synthesize fine-grained details of specific categories, we perform bidirectional fusion on the features conditioned on different information, and this design is beneficial for class-specific semantic expression, thereby synthesizing high-fidelity data encapsulating precise class-specific details. We validate ClassBooth across multiple benchmarks, demonstrating its superiority over existing methods through comprehensive quantitative and qualitative evaluations.
Yan Huang 0031, Hau-San Wong, Si Wu 0002
IEEE Trans. Multim.6
2025 SpotDiff: Spatial Gene Expression Imputation Diffusion with Single-Cell RNA Sequencing Data Integration
abstract
The advent of Spatial Transcriptomics (ST) has revolutionized understanding of tissue architecture by creating high-resolution maps of gene expression patterns. However, the low capture rate of ST leads to significant sparsity. The aim of imputation is to recover biological signals by imputing the dropouts in ST data to approximate the true expression values. In this paper, we introduce a Spatial Gene Expression Imputation Diffusion model to facilitate ST data imputation, and our model is referred to as SpotDiff. Specifically, we incorporate a spot-gene prompt learning module to capture the association between spots and genes. Further, SpotDiff integrates single-cell RNA sequencing data to impute gene expression at each spot. The proposed approach is able to reduce the uncertainty in the imputation process, since the aggregation of multiple single-cell measurements yield a stable representation of the corresponding spot expression profile. Extensive experiments have been performed to demonstrate that SpotDiff outperforms existing imputation methods across multiple benchmarks in terms of yielding more accurate and biologically relevant gene expression profiles, particularly in highly sparse scenarios.
Lianxin Xie, Si Wu 0002, Hau-San Wong
AAAI5
2025 Self-Correcting Robot Manipulation via Gaussian-Splatted Foresight
abstract
Language-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the robotic arm entering an untrained state, potentially causing destructive results. Consequently, the ability to detect and self-correct failures is crucial for the development of practical robotic systems. To address this challenge, we propose a foresight-driven failure detection and self-correction module for robot manipulation. By leveraging 3D Gaussian Splatting, we represent the current scene with multiple Gaussians. Subsequently, we train a prediction network to forecast the Gaussian representation of future scenes conditioned on planned actions. Failure is detected when the predicted future significantly deviates from the real observation after action execution. In such cases, the end-effector rolls back to the previous action to avoid an untrained state. Integrating this approach with the PerACT framework, we develop a self-correcting robot manipulation policy. Evaluations on ten RLBench tasks with 166 variations demonstrate the superior performance of the proposed method, which outperforms state-of-the-art methods by 12.0% success rate on average.
Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu
AAAI5
2025 Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video Restoration
abstract
Blind face video restoration aims to restore high-fidelity details from videos subjected to complex and unknown degradations. This task poses a significant challenge of managing temporal heterogeneity while at the same time maintaining stable face attributes. In this paper, we introduce a Discrete Prior-based Temporal-Coherent content prediction transformer to address the challenge, and our model is referred to as DP-TempCoh. Specifically, we incorporate a spatial-temporal-aware content prediction module to synthesize high-quality content from discrete visual priors, conditioned on degraded video tokens. To further enhance the temporal coherence of the predicted content, a motion statistics modulation module is designed to adjust the content, based on discrete motion priors in terms of cross-frame mean and variance. As a result, the statistics of the predicted content can match with that of real videos over time. By performing extensive experiments, we verify the effectiveness of the design elements and demonstrate the superior performance of our DP-TempCoh in both synthetically and naturally degraded video restoration.
Lianxin Xie, Bingbing Zheng, Ruotao Xu, Si Wu 0002, Hau-San Wong
AAAI7
2025 3DHumanEdit: Multi-modal Body Part-aware Conditioning Information Integration for 3D Human Manipulation
abstract
The rapid advancement of 3D Generative Adversarial Networks (GANs) has significantly enhanced the diversity and quality of generated 3D images. Despite these breakthroughs, the manipulation capabilities of 3D GANs remain unexplored, presenting substantial challenges for practical applications where user interaction and modification are essential. Current manipulation methods often lack the precision needed for fine-grained attribute manipulation, and struggle to maintain multi-view consistency during the editing process. To address these limitations, we propose 3DHumanEdit, a novel approach for 3D human body part-aware manipulation. 3DHumanEdit leverages multi-modal feature fusion and body part-aware feature alignment to achieve precise manipulation of individual body parts based on detailed text inputs and segmentation images. By exploring 3D prior for accurate editing and enforcing correspondence in latent space, 3DHumanEdit ensures coherence across multiple views. Experiments demonstrate that 3DHumanEdit outperforms existing methods in both editing fidelity and multi-view consistency, offering a robust solution for fine-grained 3D manipulation.
Fan Yang 0103, Si Wu 0002
AAAI5
2025 RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection Prompting
abstract
Face retouching aims to remove facial imperfections from image and videos while at the same time preserving face attributes. The existing methods are designed to perform non-interactive end-to-end retouching, while the ability to interact with users is highly demanded in downstream applications. In this paper, we propose RetouchGPT, a novel framework that leverages Large Language Models (LLMs) to guide the interactive retouching process. Towards this end, we design an instruction-driven imperfection prediction module to accurately identify imperfections by integrating textual and visual features. To learn imperfection prompts, we further incorporate a LLM-based embedding module to fuse multi-modal conditioning information. The prompt-based feature modification is performed in each transformer block, such that the imperfection features are suppressed and replaced with the features of normal skin progressively. Extensive experiments have been performed to verify effectiveness of our design elements and demonstrate that RetouchGPT is a useful tool for interactive face retouching and achieves superior performance over state-of-the-arts.
Chun Ding, Ruotao Xu, Si Wu 0002, Yong Xu 0007, Hau-San Wong
AAAI4
2025 Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding
abstract
The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks. To establish precise associations between the text and the corresponding visual region, we propose a Task-aware Crossmodal feature Refinement Transformer with LLMs for visual grounding, and our model is referred to as TCRT. To enable the LLM trained solely on text to understand images, we introduce an LLM adaptation module that extracts textrelated visual features to bridge the domain discrepancy between the textual and visual modalities. We feed the text and visual features into the LLM to obtain task-aware priors. To enable the priors to guide the feature fusion process, we further incorporate a cross-modal feature fusion module, which allows task-aware embeddings to refine visual features and facilitate information interaction between the Referring Expression Comprehension (REC) and Referring Expression Segmentation (RES) tasks. We have performed extensive experiments to verify the effectiveness of the main components and demonstrate the superior performance of the proposed TCRT over state-of-the-art end-to-end visual grounding methods on RefCOCO, RefCOCOg, RefCOCO+ and ReferItGame.
Ruotao Xu, Si Wu 0002, Hau-San Wong
CVPR4
2025 Dynamic Content Prediction with Motion-aware Priors for Blind Face Video Restoration
abstract
Blind Face Video Restoration (BFVR) focuses on reconstructing high-quality facial image sequences from degraded video inputs. The main challenge is address unknown degradations, while maintaining temporal consistency across frames. Current blind face restoration methods are primarily designed for images, and directly applying these approaches to BFVR will encounter a significant drop in restoration performance. In this work, we proposed Dynamic Content Prediction with Motion-aware Priors, referred to as DCP-MP. We develop a motion-aware semantic dictionary by encoding the semantic information of high-quality videos into discrete elements, and capturing the motion information in terms of element relationships, which are derived from the dynamic temporal changes within videos. For the purpose of utilizing dictionary to represent the degraded video, we train a temporal-aware element predictor, conditioned on degraded content, to learn the prediction of discrete elements in dictionary. The predicted elements will be refined, conditioned on motion information captured by the motion-aware semantic dictionary, to enhance temporal coherence. To alleviate deviation from the original structure information, we propose a conditional structure feature correction module that corrects the features flowing from the encoder to the generator. Through extensive experiments, we validate the effectiveness of our design components and demonstrate the superior performance of DCP-MP in synthesizing high-quality video.
Lianxin Xie, Bingbing Zheng, Si Wu 0002, Hau-San Wong
CVPR3
2025 Prompt-augmented Feature with Cross-domain Contrastive Learning for Efficient Multi-domain Sentiment Analysis
abstract
Pre-trained language models (PrLMs) demonstrate impressive performance on the sentiment analysis task. However, the large number of trainable parameters brings about heavy computational costs, which become more serious in multi-domain scenarios. In this paper, we propose to extract multi-layer features from the PrLM for efficient training since the training process is independent to its large backbone. Meanwhile, compared with the conventional feature extraction, we leverage prompts to induce PrLM for generating sentiment-aware features which lead to significant improvement on the sentiment analysis. In addition, most previous methods adopted a domain alignment paradigm for multi-domain learning, which becomes cumbersome when the number of domains is large. Therefore, we propose a novel prompt-augmented cross-domain contrastive learning for generalizable performance, which clusters samples with the same label under different prompts or domains. Our method is evaluated on two public multi-domain sentiment analysis benchmarks, which significantly outperforms recent state-of-the-art methods. Extensive ablation studies also verify the effectiveness of each proposed component.
Rui Li 0045, Cheng Liu 0001, Dazhi Jiang, Hau-San Wong, Si Wu 0002
ICASSP6
2025 Facilitating Semi-Supervised Pedestrian Detection with Structurally Controllable Instance Synthesis
abstract
The performance of pedestrian detectors typically relies on sufficient labeled data, and semi-supervised learning is a promising way to address the deficiency in manual annotations by utilizing sufficient unlabeled images. In this work, we design a Structure-Controllable Pedestrian Instance Generation approach (SCPIG), which is tailored to semi-supervised pedestrian detection. Specifically, we adopt a mask encoder to transform mask images into the embeddings encapsulating structure knowledge. In addition, we incorporate a mapping network to transform random latent code and a conditional generation network to synthesize diverse pedestrian instances, where the transformed code and mask embedding control pedestrian appearance and structure, respectively. The synthesized pedestrian instances are used to construct high-quality pseudo-labeled images for training pedestrian detectors. Extensive experiments validate the effectiveness of SCPIG in controllable pedestrian instance synthesizing and semi-supervised pedestrian detection.
Tianyou Zhang, Si Wu 0002, Rui Li 0045
ICASSP3
2025 IP-KGQA: Intent-Aware Prompt Learning for Knowledge Graph Question Answering
abstract
Knowledge Graph Question Answering (KGQA) addresses natural language questions by leveraging structured information stored in knowledge graphs. However, existing KGQA methods are overly concerned with improving the quality of responses by retrieving information, neglecting to identify which type of knowledge is truly useful to optimize the performance of the KGQA system, resulting in redundant retrieval. At the same time, these methods have limitations in aligning user intent and insufficient semantic richness in responses. In this work, we propose IP-KGQA, which introduce a intent-aware prompt learning scheme for KGQA framework. A Selection-Driven Efficient Retrieval (SER) module is incorporated in the framework, which classifies user questions to ensure that only long-tail questions are directed to the knowledge graph retrieval to enhance system efficiency. To filter and select the most relevant triplets, aligning retrieved information more closely with user intent, we introduce the User Intent-aware Filtering (UIF) module, where Monte Carlo sampling is applied to obtain the optimal triplets. The Domain-specific Context Prompt Extension (DCPE) module is utilized in collaboration with a fine-tuned large language model (LLM) to integrate domain-specific knowledge into the responses, ensuring that the answers are enriched in terms of semantic quality. Extensive experiments have been conducted on the CommonSenseQA and TriviaQA datasets, which demonstrate that IP-KGQA outperforms the existing methods in terms of retrieval efficiency, answer accuracy and user intent alignment.
Zheng Dai, Chun Ding, Si Wu 0002, Yong Xu 0007, Runzhe Liang, Tianshi Xu, Yedong Li, Dapeng Oliver Wu
ICME4
2025 InpaintFormer: Prompt-guided High-Quality Face Inpainting with Mask-Aware Self-Attention
abstract
Face image inpainting, especially with user-controllable customization, aims to restore degraded facial regions while adhering to user-provided instructions. Traditional inpainting methods often focus solely on restoring visual fidelity, lacking the ability to incorporate user prompts or semantic guidance. In this work, we present InpaintFormer, a novel framework for user-controlled face image inpainting guided by textual prompts. Specifically, we propose a Prompt-guided Feature Modulation (PGFM) module to align visual features with user instructions by utilizing a pre-trained CLIP model to extract text and image embeddings. These embeddings are fused to modulate the encoded image features, ensuring semantic consistency with the prompt. Additionally, a Degradation Mask Predictor (DMP) is introduced to identify degraded regions requiring inpainting, while a Mask-Aware Self-Attention (MASA) mechanism within the Transformer refines the inpainting process by selectively attending to non-degraded regions for generating realistic results. By combining PGFM, DMP, and MASA, InpaintFormer enables controllable face image inpainting with high fidelity and semantic alignment. Extensive experiments demonstrate that InpaintFormer outperforms state-of-the-art inpainting methods in terms of controllability and naturalness.
Zhouhao Ouyang, Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet, Dapeng Oliver Wu
ICME5
2025 Rethinking 3D Robotic Perception: Elastic Voxel Representation with Splatting Distillation
abstract
Language-guided robotic manipulation is advancing rapidly with Vision-Language-Action (VLA) models, yet faces fundamental challenges in 3D perception. This paper addresses two critical challenges: the scale elasticity requirement for simultaneously processing coarse environmental context and fine manipulation details, and the scarcity of action-annotated training data. We present Splat-Actor, a novel robotic manipulation framework that introduces two key innovations. First, we develop an elastic voxel encoder that combines multi-scale processing with selective tokenization, enabling efficient 3D spatial reasoning while adaptively focusing on informative regions. Second, we propose a depth-constrained feature distillation framework that leverages Gaussian Splatting to bridge 2D and 3D representations, transferring rich semantic features from pre-trained vision models to enhance 3D understanding. Extensive experiments across 10 manipulation tasks with 166 variations demonstrate that Splat-Actor achieves a 6.8% improvement over state-of-the-art methods while maintaining the computational efficiency.
Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu, Patrick Le Callet
ICME5
2025 Text to Trajectory: Enhancing and Evaluating LLMs for Embodied Task Planning
abstract
The increasing demand for effective human-machine interaction highlights the importance of integrating natural language processing with robotics technology. This paper addresses the challenges of using Large Language Models (LLMs) for embodied task planning in complex environments. We propose a comprehensive framework that combines environmental-aware LLM fine-tuning with a novel Stepwise Beam Search (SBS) strategy. In conjunction with the environmentally enhanced LLM, the SBS strategy facilitates comprehensive exploration of both token-level and step-level search spaces, overcoming the limitations of conventional greedy search methods. Additionally, to evaluate the effectiveness of embodied task planning, we introduce the Trajectory Match Score (TMS), a robust evaluation metric that leverages state-based simulation to assess plan success. Through extensive experiments on standard benchmarks, our framework demonstrates substantial improvements in both plan generation quality and task success rates, advancing the state-of-the-art in embodied task planning.
Yihan Tang, Yong Xu 0007, Ruotao Xu, Yan Huang 0031, Si Wu 0002, Patrick Le Callet
ICME5
2025 SemanticLoom: Category-aware Dynamic Fusion for Multi-class Few-shot Image Synthesis
abstract
Few-shot text-to-image (T2I) generation seeks to efficiently integrate new semantics into existing pre-trained models while preserving their capacity to generate diverse, high-quality images. However, existing methods often suffer from inefficiency and poor scalability due to the need for separate training processes for each new concept. These challenges hinder their practical application in multi-class few-shot scenarios. To overcome these issues, we propose SemanticLoom that dynamically incorporates novel concepts into pre-trained diffusion models through category-aware dynamic feature fusion. Our approach introduces a lightweight semantic expander that captures fine-grained semantics, guided by learnable identifier to ensure precise semantic integration. By dynamically adjusting feature fusion coefficients based on category diversity and training progress, our method harmonizes the integration of new semantic features with the original model’s capabilities, ensuring consistency and generalization. Experiments demonstrate that our method successfully integrates new semantics without compromising the generative diversity and versatility of the pre-trained model.
Yan Huang 0031, Si Wu 0002, Yong Xu 0007, Patrick Le Callet
ICME5
2025 Adaptive Illumination Transfer Network for Shadow Removal
abstract
Shadow removal aims to harmonize illumination between shadow and non-shadow regions. However, existing methods often struggle to achieve this goal due to inadequate modeling of illumination relationships between these two regions. Moreover, the prevalent reliance on binary shadow masks hinders their capability to address non-uniform shadows. To address these limitations, we propose an adaptive illumination transfer network (AITNet), which incorporates two complementary shadow enhancement strategies. First, a global illumination transfer strategy is designed to model the illumination relationship between shadow and non-shadow regions, enabling the holistic enhancement of shadow regions. Second, an illumination-adaptive strategy is developed to estimate an illumination degradation map, which guides the adaptive enhancement of shadow regions. Furthermore, to preserve the original structural details, the enhancement process is applied exclusively to the illumination map obtained after Retinex decomposition. Extensive experiments have demonstrated the superiority of our method over existing approaches on public datasets.
Si Wu 0002, Yong Xu 0007, Yan Huang 0031, Patrick Le Callet
ICME2
2025 Cross-View Neighborhood Contrastive Multi-View Clustering with View Mixup Feature Learning
abstract
Multi-view clustering (MVC) has shown that leveraging both consistency and complementary information across views enhances clustering performance. However, most existing methods focus on aligning features into the same dimension, often neglecting cross-view heterogeneity and introducing discrepancies. To address this, we propose a novel multi-view clustering framework that combines cross-view neighborhood contrastive learning with a cross-attention view-mixup feature learning mechanism. Specifically, the cross-attention view-mixup module learns view-invariant feature representations by capturing complementary and consistent information, while the neighborhood contrastive learning module uncovers semantic structures across views based on the learned mixup features. By implicitly performing feature mixup across views and effectively integrating cross-view neighborhood contrastive learning, our method alleviates cross-view discrepancies and enables more effective integration of complementary and consistent information, ultimately enhancing clustering performance. Experiments conducted on several real datasets demonstrate the effectiveness of our proposed method in comparision with several representative MVC approaches.
Yixuan Ye, Yang Zhang 0073, Rui Li 0045, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
ICME6
2025 Adversarial Style Interpolation for High-fidelity Stroke-aware Chinese Character Image Synthesis
abstract
As a specific image-to-image translation task, Chinese character style transfer aims to synthesize new character images with the style of reference characters and the content of target ones. This task is still challenging due to the significant number of characters with complex structures. We present an Adversarial Style Interpolation-based approach for reliable Style Transfer (ASIST). To capture style information from reference images, we adopt a style predictor to compete with a style encoder and a generator by predicting the interpolation coefficients of the character images synthesized from the interpolated style features. Further, we incorporate a stroke-aware content predictor, which applies attention on local strokes regardless of the variance in style. By jointly optimized with the content predictor, the generator is able to attend on the structural integrity of synthesized characters. Extensive experiments are performed to verify that the proposed approach can synthesize high-quality stylized Chinese characters over the leading methods.
Chun Ding, Si Wu 0002
IJCNN3
2025 Prior-Free Augmentation for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CCReID) aims to recognize individuals across clothing variations by learning clothing-invariant representations. However, obtaining sufficient samples of the same person in diverse outfits is often impractical. While synthesizing realistic person images provides an effective solution, existing augmentation methods require labeled data and external priors (e.g., pose skeletons, semantic maps), resulting in high costs and limited generalization. To this end, we propose a Prior-Free Augmentation method for Cloth-changing person re-identification (PFAC), which leverages text guidance to synthesize images with clothing variations while maintaining identity consistency. Our approach features: (1) a truncated diffusion model that preserves clothing-invariant structural cues from intermediate noisy images, (2) a dual-branch denoising network that decouples text-guided clothing synthesis from identity consistency via cross-modal alignment, and (3) a joint optimization strategy with identity-focused losses and image filtering to enhance realism and discriminability. Experimental results on PRCC, LTCC, and Celeb-reID datasets demonstrate that PFAC achieves state-of-the-art CCReID performance, effectively generating high-fidelity, identity-consistent images for robust augmentation without external priors.
Xin Li 0034, Si Wu 0002, Yong Xu 0007, Yaowei Wang 0001
ACM Multimedia3
2025 COME: contrastive mapping learning for spatial reconstruction of single-cell RNA sequencing data
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) enables high-throughput transcriptomic profiling at single-cell resolution. The inherent spatial location is crucial for understanding how single cells orchestrate multicellular functions and drive diseases. However, spatial information is often lost during tissue dissociation. Spatial transcriptomic (ST) technologies can provide precise spatial gene expression atlas, while their practicality is constrained by the number of genes they can assay or the associated costs at a larger scale and the fine-grained cell-type annotation. By transferring knowledge between scRNA-seq and ST data through cell correspondence learning, it is possible to recover the spatial properties inherent in scRNA-seq datasets. RESULTS: In this study, we introduce COME, a COntrastive Mapping lEarning approach that learns mapping between ST and scRNA-seq data to recover the spatial information of scRNA-seq data. Extensive experiments demonstrate that the proposed COME method effectively captures precise cell-spot relationships and outperforms previous methods in recovering spatial location for scRNA-seq data. More importantly, our method is capable of precisely identifying biologically meaningful information within the data, such as the spatial structure of missing genes, spatial hierarchical patterns, and the cell-type compositions for each spot. These results indicate that the proposed COME method can help to understand the heterogeneity and activities among cells within tissue environments. AVAILABILITY AND IMPLEMENTATION: The COME is freely available in GitHub (https://github.com/cindyway/COME).
Xindian Wei, Xibiao Wang, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
Bioinform.6
2025 Robust subspace structure discovery for cell type identification in scRNA-seq data
abstract
Single-cell RNA sequencing (scRNA-seq) technology has transformed gene expression studies by enabling analysis at the individual cell level, offering unprecedented insights into cellular heterogeneity. A key challenge in scRNA-seq data analysis is cell type identification, which requires grouping cells with similar gene expression profiles using unsupervised clustering methods. However, the high dimensionality, inherent noise, and significant sparsity of scRNA-seq data present substantial obstacles to accurately determining relationships among cell samples. To address these challenges, we propose a novel deep subspace clustering approach for cell type identification that captures a more reliable subspace structure from scRNA-seq data. Our method leverages a robust self-representation learning framework to effectively characterize and learn the underlying cluster structure. This framework is optimized through an integrated strategy combining a structure-guided approach with an optimal transport algorithm, enhancing the robustness of the subspace clustering process. By mitigating the effects of noise and sparsity in scRNA-seq data, this approach enables more accurate cell clustering. Experimental results on 18 real scRNA-seq datasets demonstrate that our method outperforms several state-of-the-art clustering approaches tailored for scRNA-seq data, excelling in both accuracy and interpretability.
Xianyong Zhou, Xindian Wei, Cheng Liu 0001, Ping Xuan, Si Wu 0002, Hau-San Wong
BMC Bioinform.6
2025 Controllable instance synthesis with hierarchical regularization for semi-supervised pedestrian detection
Gaozhe Li, Xiangyu Sai, Lianxin Xie, Si Wu 0002
Neurocomputing6
2025 E-Net for pansharpening: A super-resolution perspective
Si Wu 0002, Yong Xu 0007, Yan Huang 0031
Image Vis. Comput.2
2025 Diverse Semantic Image Synthesis with various conditioning modalities
Chaoyue Wu, Rui Li 0045, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
Knowl. Based Syst.4
2025 Class-conditional image synthesis with intra-class relation preservation
Xiaoyang Huo, Si Wu 0002, Hau-San Wong
Knowl. Based Syst.4
2025 DA-GAN: Dual-attention generative adversarial networks for real-world exquisite makeup transfer
Qianfen Jiao, Si Wu 0002, Hau-San Wong
Pattern Recognit.3
2025 EviD-GAN: Improving GAN With an Infinite Set of Discriminators at Negligible Cost
abstract
Ensemble learning improves the capability of convolutional neural network (CNN)-based discriminators, whose performance is crucial to the quality of generated samples in generative adversarial network (GAN). However, this learning strategy results in a significant increase in the number of parameters along with computational overhead. Meanwhile, the suitable number of discriminators required to enhance GAN performance is still being investigated. To mitigate these issues, we propose an evidential discriminator for GAN (EviD-GAN)-code is available at https://github.com/Tohokantche/EviD-GAN-to learn both the model (epistemic) and data (aleatoric) uncertainties. Specifically, by analyzing three GAN models, the relation between the distribution of discriminator's output and the generator performance has been discovered yielding a general formulation of GAN framework. With the above analysis, the evidential discriminator learns the degree of aleatoric and epistemic uncertainties via imposing a higher order distribution constraint over the likelihood as expressed in the discriminator's output. This constraint can learn an ensemble of likelihood functions corresponding to an infinite set of discriminators. Thus, EviD-GAN aggregates knowledge through the ensemble learning of discriminator that allows the generator to benefit from an informative gradient flow at a negligible computational cost. Furthermore, inspired by the gradient direction in maximum mean discrepancy (MMD)-repulsive GAN, we design an asymmetric regularization scheme for EviD-GAN. Unlike MMD-repulsive GAN that performs at the distribution level, our regularization scheme is based on a pairwise loss function, performs at the sample level, and is characterized by an asymmetric behavior during the training of generator and discriminator. Experimental results show that the proposed evidential discriminator is cost-effective, consistently improves GAN in terms of Frechet inception distance (FID) and inception score (IS), and performs better than other competing models that use multiple discriminators.
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Beyond Euclidean Structures: Collaborative Topological Graph Learning for Multiview Clustering
abstract
Graph-based multiview clustering (MVC) approaches have demonstrated impressive performance by leveraging the consistency properties of multiview data in an unsupervised manner. However, existing methods for graph learning heavily rely on either Euclidean structures or the manifold topological structures derived from fixed view-specific graphs. Unfortunately, these approaches may not accurately reflect the consensus topological structure in a multiview setting. To address this limitation and enhance the intrinsic graph learning process, an adaptive exploration of a more appropriate consistency topological structure is required. Toward this end, we propose a novel approach called collaborative topological graph learning (CTGL) for MVC. The key idea is to adaptively discover the consistent topological structure to guide intrinsic graph learning. We achieve this by introducing an auxiliary consistency graph that formulates the topological relevance learning function. However, estimating the auxiliary consistency graph is not straightforward, as it is based on the learned view-specific graphs and requires prior availability. To overcome this challenge, we develop a collaborative learning strategy that simultaneously learns both the auxiliary consistency graph and view-specific graphs using tensor learning techniques. This strategy enables the adaptive exploration of the consistency topological structure during graph learning, resulting in more accurate clustering outcomes. Extensive experiments are provided to show the effectiveness of the proposed method. The source code can be found at https://github.com/CLiu272/CTGL.
Cheng Liu 0001, Rui Li 0045, Hangjun Che, Man-Fai Leung, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Neural Networks Learn. Syst.5
2024 RetouchFormer: Semi-supervised High-Quality Face Retouching Transformer with Prior-Based Selective Self-Attention
abstract
Face retouching is to beautify a face image, while preserving the image content as much as possible. It is a promising yet challenging task to remove face imperfections and fill with normal skin. Generic image enhancement methods are hampered by the lack of imperfection localization, which often results in incomplete removal of blemishes at large scales. To address this issue, we propose a transformer-based approach, RetouchFormer, which simultaneously identify imperfections and synthesize realistic content in the corresponding regions. Specifically, we learn a latent dictionary to capture the clean face priors, and predict the imperfection regions via a reconstruction-oriented localization module. Also based on this, we can realize face retouching by explicitly suppressing imperfections in our selective self-attention computation, such that local content will be synthesized from normal skin. On the other hand, multi-scale feature tokens lead to increased flexibility in dealing with the imperfections at various scales. The design elements bring greater effectiveness and efficiency. RetouchFormer outperforms the advanced face retouching methods and synthesizes clean face images with high fidelity in our list of extensive experiments performed.
Lianxin Xie, Si Wu 0002, Cheng Liu 0001, Hau-San Wong
AAAI5
2024 Relational Matching for Weakly Semi-Supervised Oriented Object Detection
abstract
Oriented object detection has witnessed significant progress in recent years. However, the impressive performance of oriented object detectors is at the huge cost of labor-intensive annotations, and deteriorates once the an-notated data becomes limited. Semi-supervised learning, in which sufficient unannotated data are utilized to enhance the base detector, is a promising method to address the annotation deficiency problem. Motivated by weakly supervised learning, we introduce annotation-efficient point annotations for unannotated images and propose a weakly semi-supervised method for oriented object detection to balance the detection performance and annotation cost. Specifically, we propose a Rotation-Modulated Relational Graph Matching method to match relations of proposals centered on an-notated points between the teacher and student models to alleviate the ambiguity of point annotations in depicting the oriented object. In addition, we further propose a Relational Rank Distribution Matching method to align the rank distribution on classification and regression between different models. Finally, to handle the difficult annotated points that both models are confused about, we introduce weakly supervised learning to impose positive signals for difficult point-induced clusters to the base model, and focus the base model on the occupancy between the predictions and an-notated points. We perform extensive experiments on chal-lenging datasets to demonstrate the effectiveness of our proposed weakly semi-supervised method in leveraging point-annotated data for significant performance improvement.
Hau-San Wong, Si Wu 0002, Tianyou Zhang
CVPR3
2024 Learning Degradation-Unaware Representation with Prior-Based Latent Transformations for Blind Face Restoration
abstract
Blind face restoration focuses on restoring high-fidelity details from images subjected to complex and unknown degradations, while preserving identity information. In this paper, we present a Prior-based Latent Transformation approach (PLTrans), which is specifically designed to learn a degradation-unaware representation, thereby allowing the restoration network to effectively generalize to real-world degradation. Toward this end, PLTrans learns a degradation-unaware query via a latent diffusion-based regularization module. Furthermore, conditioned on the features of a degraded face image, a latent dictionary that captures the priors of HQ face images is leveraged to refine the features by mapping the top-d nearest elements. The refined version will be used to build key and value for the cross-attention computation, which is tailored to each degraded image and exhibits reduced sensitivity to different degradation factors. Conditioned on the resulting representation, we train a decoding network that synthesizes face images with authentic details and identity preservation. Through extensive experiments, we verify the effectiveness of the design elements and demonstrate the generalization ability of our proposed approach for both synthetic and unknown degradations. We finally demonstrate the applicability of PLTrans in other vision tasks.
Lianxin Xie, Bingbing Zheng, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
CVPR6
2024 Text-Conditional Attribute Alignment Across Latent Spaces for 3D Controllable Face Image Synthesis
abstract
With the advent of generative models and vision-language pre-training, significant improvement has been made in text-driven face manipulation. The text embedding can be used as target supervision for expression control. However, it is non-trivial to associate with its 3D attributes, i.e., pose and illumination. To address these issues, we propose a Text-conditional Attribute aLignment approach for 3D controllable face image synthesis, and our model is referred to as TcALign. Specifically, since the 3D rendered image can be precisely controlled with the 3D face representation, we first propose a Text-conditional 3D Editor to produce the target face representation to realize text-driven manipulation in the 3D space. An attribute embedding space spanned by the target-related attributes embeddings is also introduced to infer the disentangled task-specific direction. Next, we train a cross-modal latent mapping network conditioned on the derived difference of 3D representation to infer a correct vector in the latent space of Style-GAN. This correction vector learning design can accurately transfer the attribute manipulation on 3D images to 2D images. We show that the proposed method delivers more precise text-driven multi-attribute manipulation for 3D controllable face image synthesis. Extensive qualitative and quantitative experiments verify the effectiveness and superiority of our method over the other competing methods.
Rui Li 0045, Si Wu 0002, Yong Xu 0007, Hau-San Wong
CVPR3
2024 VRetouchEr: Learning Cross-Frame Feature Interdependence with Imperfection Flow for Face Retouching in Videos
abstract
Face Video Retouching is a complex task that often requires labor-intensive manual editing. Conventional image retouching methods perform less satisfactorily in terms of generalization performance and stability when applied to videos without exploiting the correlation among frames. To address this issue, we propose a Video Retouching transformEr to remove facial imperfections in videos, which is referred to as VRetouchEr. Specifically, we estimate the apparent motion of imperfections between two consecutive frames, and the resulting displacement vectors are used to refine the imperfection map, which is synthesized from the current frame together with the corresponding encoder features. The flow-based imperfection refinement is critical for precise and stable retouching across frames. To leverage the temporal contextual information, we inject the refined imperfection map into each transformer block for multi-frame masked attention computation, such that we can capture the interdependence between the current frame and multiple reference frames. As a result, the imperfection regions can be replaced with normal skin with high fidelity, while at the same time keeping the other regions unchanged. Extensive experiments are performed to verify the superiority of VRetouchEr over state-of-the-art image retouching methods in terms of fidelity and stability.
Lianxin Xie, Si Wu 0002, Yong Xu 0007, Hau-San Wong
CVPR4
2024 AttriHuman-3D: Editable 3D Human Avatar Generation with Attribute Decomposition and Indexing
abstract
Editable 3D-aware generation, which supports user-interacted editing, has witnessed rapid development re-cently. However, existing editable 3D GANs either fail to achieve high-accuracy local editing or suffer from huge computational costs. We propose AttriHuman-3D, an ed-itable 3D human generation model, which address the aforementioned problems with attribute decomposition and indexing. The core idea of the proposed model is to generate all attributes (e.g. human body, hair, clothes and so on) in an overall attribute space with six feature planes, which are then decomposed and manipulated with different attribute indexes. To precisely extract features of different attributes from the generated feature planes, we propose a novel at-tribute indexing method as well as an orthogonal projection regularization to enhance the disentanglement. We also introduce a hyper-latent training strategy and an attribute-specific sampling strategy to avoid style entanglement and misleading punishment from the discriminator. Our method allows users to interactively edit selected attributes in the generated 3D human avatars while keeping others fixed. Both qualitative and quantitative experiments demonstrate that our model provides a strong disentanglement between different attributes, allows fine-grained image editing and generates high-quality 3D human avatars.
Fan Yang 0103, Xiaosheng He, Zhongang Cai, Lei Yang 0045, Si Wu 0002, Guosheng Lin
CVPR6
2024 Reference-conditional Makeup-aware Discrimination for Face Image Beautification
abstract
Facial makeup transfer aims to replicate reference makeup on target face, and the existing methods are mainly based on a generic adversarial training process. In this work, we design a Reference-conditional Makeup-aware Discrimination approach (RcMD) to facilitate makeup transfer. Specifically, we perform region-wise semantic feature extraction from a reference makeup image and a source image without makeup. A generator learns to capture and render the reference makeup by modulating the region-wise intermediate features. To ensure precise makeup on target face, we incorporate a reference-conditional discrimination network, which learns to measure the regional makeup consistency between reference and synthesized images. Considering the discrepancy between reference and target faces, an alignment module is trained to fuse the extracted features, conditioned on the reference style. Based on the feature statistics, we perform regional real-synthesized makeup discrimination to ensure precise makeup rendering. Extensive experiments are performed to demonstrate the effectiveness of our designed modules and the superior performance of RcMD in transferring diverse real-world facial makeup.
Si Wu 0002, Xindian Wei, Qianfen Jiao, Cheng Liu 0001, Rui Li 0045
ICME2
2024 SCTrans: Multi-scale scRNA-seq Sub-vector Completion Transformer for Gene-selective Cell Type Annotation
Xindian Wei, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
IJCAI6
2024 Hunting Blemishes: Language-guided High-fidelity Face Retouching Transformer with Limited Paired Data
abstract
The prevalence of multimedia applications has led to increased concerns and demand for auto face retouching. Face retouching aims to enhance portrait quality by removing blemishes. However, the existing auto-retouching methods rely heavily on a large amount of paired training samples, and perform less satisfactorily when handling complex and unusual blemishes. To address this issue, we propose a Language-guided Blemish Removal Transformer for automatically retouching face images, while at the same time reducing the dependency of the model on paired training data. Our model is referred to as LangBRT, which leverages vision-language pre-training for precise facial blemish removal. Specifically, we design a text-prompted blemish detection module that indicates the regions to be edited. The priors not only enable the transformer network to handle specific blemishes in certain areas, but also reduce the reliance on retouching training data. Further, we adopt a target-aware cross attention mechanism, such that the blemish-like regions are edited accurately while at the same time maintaining the normal skin regions unchanged. Finally, we adopt a regularization approach to encourage the semantic consistency between the synthesized image and the text description of the desired retouching outcome. Extensive experiments are performed to demonstrate the superior performance of LangBRT over competing auto-retouching methods in terms of dependency on training data, blemish detection accuracy and synthesis quality.
Yan Huang 0031, Lianxin Xie, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
ACM Multimedia6
2024 Contrastive Graph Distribution Alignment for Partially View-Aligned Clustering
abstract
Partially View-aligned Clustering (PVC) presents a challenge as it requires a comprehensive exploration of complementary and consistent information in the presence of partial alignment of view data. Existing PVC methods typically learn view correspondence based on latent features that are expected to contain common semantic information. However, latent features obtained from heterogeneous spaces, along with the enforcement of alignment into the same feature dimension, can introduce cross-view discrepancies. In particular, partially view-aligned data lacks sufficient shared correspondences for the critical common semantic feature learning, resulting in inaccuracies in establishing meaningful correspondences between latent features across different views. While feature representations may differ across views, instance relationships within each view could potentially encode consistent common semantics across views. Motivated by this, our aim is to learn view correspondence based on graph distribution metrics that capture semantic view-invariant instance relationships. To achieve this, we utilize similarity graphs to depict instance relationships and learn view correspondence by aligning semantic similarity graphs through optimal transport with graph distribution. This facilitates the precise learning of view alignments, even in the presence of heterogeneous view-specific feature distortions. Furthermore, leveraging well-established cross-view correspondence, we introduce a cross-view contrastive learning to learn semantic features by exploiting consistency information. The resulting meaningful semantic features effectively isolate shared latent patterns, avoiding the inclusion of irrelevant private information. We conduct extensive experiments on several real datasets, demonstrating the effectiveness of our proposed method for the PVC task.
Xibiao Wang, Hang Gao 0014, Xindian Wei, Rui Li 0045, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
ACM Multimedia7
2024 SELF-Former: multi-scale gene filtration transformer for single-cell spatial reconstruction
abstract
The spatial reconstruction of single-cell RNA sequencing (scRNA-seq) data into spatial transcriptomics (ST) is a rapidly evolving field that addresses the significant challenge of aligning gene expression profiles to their spatial origins within tissues. This task is complicated by the inherent batch effects and the need for precise gene expression characterization to accurately reflect spatial information. To address these challenges, we developed SELF-Former, a transformer-based framework that utilizes multi-scale structures to learn gene representations, while designing spatial correlation constraints for the reconstruction of corresponding ST data. SELF-Former excels in recovering the spatial information of ST data and effectively mitigates batch effects between scRNA-seq and ST data. A novel aspect of SELF-Former is the introduction of a gene filtration module, which significantly enhances the spatial reconstruction task by selecting genes that are crucial for accurate spatial positioning and reconstruction. The superior performance and effectiveness of SELF-Former's modules have been validated across four benchmark datasets, establishing it as a robust and effective method for spatial reconstruction tasks. SELF-Former demonstrates its capability to extract meaningful gene expression information from scRNA-seq data and accurately map it to the spatial context of real ST data. Our method represents a significant advancement in the field, offering a reliable approach for spatial reconstruction.
Xindian Wei, Lianxin Xie, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
Briefings Bioinform.7
2024 Cluster-based Adversarial Decision Boundary for domain-adaptive open set recognition
Qianfen Jiao, Si Wu 0002, Cheng Liu 0001, Hau-San Wong
Knowl. Based Syst.3
2024 Semi-supervised class-conditional image synthesis with Semantics-guided Adaptive Feature Transforms
Xiaoyang Huo, Si Wu 0002
Pattern Recognit.3
2024 Collaborative Structure-Preserved Missing Data Imputation for Single-Cell RNA-Seq Clustering
abstract
Clustering of the single-cell RNA-seq (scRNA-seq) transcriptome profiles is able to identify cell types, which is beneficial to improve the understanding of disease progression. However, in practice, the single-cell expression data often contains a significant number of missing values as a result of technical variability. Missing data is a critical challenge in scRNA-seq clustering analysis since the unknown value does not reflect the underlying true expression level and makes it difficult to discovering cell types by applying clustering algorithms directly. Various approaches have been developed to overcome missing data issue in scRNA-seq clustering. Most of them recover missing expression values by borrowing observed data from similar cells or synthesizing data via generative adversarial networks. Such that the biologically meaningful cluster structure has not been sufficiently exploited. In this work, we introduce ColImpute, a collaborative structure-preserved missing data imputation approach for the scRNA-seq clustering. Specifically, a cluster structure-preserved imputation module and a subspace clustering module, which respectively perform missing data imputation and cell subtypes identification, are integrated into a unified optimization framework to train the two networks in a collaborative manner. Consequently, the clustering module effectively contributes cluster-structure information to guide the trainning process of the missing data imputation module. Simultaneously, the cluster structure-preserved imputation module reciprocally enhances the performance of the clustering module by generating more precise recovered samples. Promising experimental results show that the proposed method is effective for both the data imputation and the cell types identification.
Hang Gao 0014, Rui Li 0045, Cheng Liu 0001, Si Wu 0002
IEEE ACM Trans. Comput. Biol. Bioinform.5
2024 Pseudo-Siamese Teacher for Semi-Supervised Oriented Object Detection
abstract
Oriented object detection, which aims at detecting objects with orientation property, shows great potential for visual analysis in complex scenarios, such as aerial images. However, the powerful detection performance relies on abundant and accurate annotations, and deteriorates once the annotations become insufficient. Semi-supervised learning, which utilizes unannotated data to improve the target model, is a promising method to address the problem of annotation deficiency. In this work, we propose Pseudo-Siamese Teacher (PST), a new semi-supervised learning framework for oriented object detection. In this architecture, two teacher models, updated from the same student model with different optimizations, inspect the predictions of each other and collaborate to generate high-quality pseudo annotations. To reduce the unreliability of pseudo annotations on the localization, scale and orientation, we propose to model the oriented object as a Gaussian distribution, and apply a symmetric and bounded Jensen–Shannon divergence (JSD) to evaluate the divergence between predictions of different teacher models, the results of which serve as an indicator to remove confusing pseudo annotations without consistent regression estimation of teacher models. Scale invariance is also an important challenge in oriented object detection, which we address by proposing a scale-adaptive knowledge distillation to align information between the feature maps from the student model on images with flexible scales and the feature maps, interpolated from adjacent feature maps with scales closest to that of the down-sampled images, from the teacher models. We perform extensive experiments to demonstrate the effectiveness of our proposed method in leveraging unannotated data for performance improvement.
Hau-San Wong, Si Wu 0002
IEEE Trans. Geosci. Remote. Sens.3
2024 Latent Structure-Aware View Recovery for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) presents a significant challenge due to the need for effectively exploring complementary and consistent information within the context of missing views. One promising strategy to tackle this challenge is to recover missing views by inferring the missing samples. However, such approaches often fail to fully utilize discriminative structural information or adequately address consistency, as it requires such information to be known or learnable in advance, which contradicts the incomplete data setting. In this study, we propose a novel approach calledLatentStructure-Aware view recovery (LaSA) for the IMVC task. Our objective is to recover missing views through discriminative latent representations by leveraging structural information. Specifically, our method offers a unified closed-form formulation that simultaneously performs missing data inference and latent representation learning, using a learned intrinsic graph as structural information. This formulation, incorporating graph structure information, enhances the inference of missing data while facilitating discriminative feature learning. Even when intrinsic graph is initially unknown due to incomplete data, our formulation allows for effective view recovery and intrinsic graph learning through an iterative optimization process. To further enhance performance, we introduce an iterative consistency diffusion process, which effectively leverages the consistency and complementary information across multiple views. Extensive experiments demonstrate the effectiveness of the proposed method compared to state-of-the-art approaches.
Cheng Liu 0001, Rui Li 0045, Hangjun Che, Man-Fai Leung, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Knowl. Data Eng.5
2024 A Graph-Based Discriminator Architecture for Multi-Attribute Facial Image Editing
abstract
Multi-attribute editing aims to synthesize new facial images with multiple desired attributes while at the same time preserving other contents. Generative Adversarial Networks (GANs) with encoder-decoder-based generators are typically applied to this task, while the co-occurrence nature of attributes is overlooked by generic discriminators when identifying real and synthesized instances. To address the issue, we focus on precisely capturing semantics associated with target attributes in this work, and propose a Graph-based Discriminator architecture for a GAN model, which is referred to as GD-GAN, for explicitly modeling and leveraging the attribute dependencies. Specifically, the co-occurrence ratio between attributes is used to build a correlation matrix, which captures inter-attribute relationships. We design a discriminator with a Graph Convolutional Network (GCN) to integrate knowledge about the attribute dependencies into the adversarial training process. Different from the existing methods that identify the synthesized data conditioned on the attributes individually, we leverage the attribute correlations by performing feature propagation over the graph of attributes, which leads to interdependent representations for real-fake instance identification. Incorporating the relationships of attributes eventually induces the generator to capture precise semantics associated with the attributes. Empirical results on multiple benchmarks demonstrate the superior performance of GD-GAN in high-quality semantic manipulation.
Quanpeng Song, Si Wu 0002, Hau-San Wong
IEEE Trans. Multim.3
2024 Self-Guided Partial Graph Propagation for Incomplete Multiview Clustering
abstract
In this work, we study a more realistic challenging scenario in multiview clustering (MVC), referred to as incomplete MVC (IMVC) where some instances in certain views are missing. The key to IMVC is how to adequately exploit complementary and consistency information under the incompleteness of data. However, most existing methods address the incompleteness problem at the instance level and they require sufficient information to perform data recovery. In this work, we develop a new approach to facilitate IMVC based on the graph propagation perspective. Specifically, a partial graph is used to describe the similarity of samples for incomplete views, such that the issue of missing instances can be translated into the missing entries of the partial graph. In this way, a common graph can be adaptively learned to self-guide the propagation process by exploiting the consistency information, and the propagated graph of each view is in turn used to refine the common self-guided graph in an iterative manner. Thus, the associated missing entries can be inferred through graph propagation by exploiting the consistency information across all views. On the other hand, existing approaches focus on the consistency structure only, and the complementary information has not been sufficiently exploited due to the data incompleteness issue. By contrast, under the proposed graph propagation framework, an exclusive regularization term can be naturally adopted to exploit the complementary information in our method. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods. The source code of our method is available at the https://github.com/CLiu272/TNNLS-PGP.
Cheng Liu 0001, Rui Li 0045, Si Wu 0002, Hangjun Che, Dazhi Jiang, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Neural Networks Learn. Syst.3
2023 Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image Manipulation
abstract
Great progress has been made in StyleGAN-based image editing. To associate with preset attributes, most existing approaches focus on supervised learning for semantically meaningful latent space traversal directions, and each manipulation step is typically determined for an individual attribute. To address this limitation, we propose a Text-guided Unsupervised StyleGAN Latent Transformation (TUSLT) model, which adaptively infers a single transformation step in the latent space of StyleGAN to simultaneously manipulate multiple attributes on a given input image. Specifically, we adopt a two-stage architecture for a latent mapping network to break down the transformation process into two manageable steps. Our network first learns a diverse set of semantic directions tailored to an input image, and later nonlinearly fuses the ones associated with the target attributes to infer a residual vector. The resulting tightly interlinked two-stage architecture delivers the flexibility to handle diverse attribute combinations. By leveraging the cross-modal text-image representation of CLIP, we can perform pseudo annotations based on the semantic similarity between preset attribute text descriptions and training images, and further jointly train an auxiliary attribute classifier with the latent mapping network to provide semantic guidance. We perform extensive experiments to demonstrate that the adopted strategies contribute to the superior performance of TUSLT.
Xiwen Wei, Cheng Liu 0001, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
CVPR4
2023 Semi-Supervised Stereo-Based 3D Object Detection via Cross-View Consensus
abstract
Stereo-based 3D object detection, which aims at detecting 3D objects with stereo cameras, shows great potential in low-cost deployment compared to LiDAR-based methods and excellent performance compared to monocular-based algorithms. However, the impressive performance of stereo-based 3D object detection is at the huge cost of high-quality manual annotations, which are hardly attainable for any given scene. Semi-supervised learning, in which limited annotated data and numerous unannotated data are required to achieve a satisfactory model, is a promising method to address the problem of data deficiency. In this work, we propose to achieve semi-supervised learning for stereo-based 3D object detection through pseudo annotation generation from a temporal-aggregated teacher model, which temporally accumulates knowledge from a student model. To facilitate a more stable and accurate depth estimation, we introduce Temporal-Aggregation-Guided (TAG) disparity consistency, a cross-view disparity consistency constraint between the teacher model and the student model for robust and improved depth estimation. To mitigate noise in pseudo annotation generation, we propose a cross-view agreement strategy, in which pseudo annotations should attain high degree of agreements between 3D and 2D views, as well as between binocular views. We perform extensive experiments on the KITTI 3D dataset to demonstrate our proposed method's capability in leveraging a huge amount of unannotated stereo images to attain significantly improved detection results.
Hau-San Wong, Si Wu 0002
CVPR3
2023 Blemish-aware and Progressive Face Retouching with Limited Paired Data
abstract
Face retouching aims to remove facial blemishes, while at the same time maintaining the textual details of a given input image. The main challenge lies in distinguishing blemishes from the facial characteristics, such as moles. Training an image-to-image translation network with pixel-wise supervision suffers from the problem of expensive paired training data, since professional retouching needs specialized experience and is time-consuming. In this paper, we propose a Blemish-aware and Progressive Face Retouching model, which is referred to as BPFRe. Our framework can be partitioned into two manageable stages to perform progressive blemish removal. Specifically, an encoder-decoder-based module learns to coarsely remove the blemishes at the first stage, and the resulting intermediate features are injected into a generator to enrich local detail at the second stage. We find that explicitly suppressing the blemishes can contribute to an effective collaboration among the components. Toward this end, we incorporate an attention module, which learns to infer a blemish-aware map and further determine the corresponding weights, which are then used to refine the intermediate features transferred from the encoder to the decoder, and from the decoder to the generator. Therefore, BPFRe is able to deliver significant performance gains on a wide range of face retouching tasks. It is worth noting that we reduce the dependence of BPFRe on paired training samples by imposing effective regularization on unpaired ones.
Lianxin Xie, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
CVPR4
2023 Exploring Intra-class Variation Factors with Learnable Cluster Prompts for Semi-supervised Image Synthesis
abstract
Semi-supervised class-conditional image synthesis is typically performed by inferring and injecting class labels into a conditional Generative Adversarial Network (GAN). The supervision in the form of class identity may be inadequate to model classes with diverse visual appearances. In this paper, we propose a Learnable Cluster Prompt-based GAN (LCP-GAN) to capture class-wise characteristics and intra-class variation factors with a broader source of supervision. To exploit partially labeled data, we perform soft partitioning on each class, and explore the possibility of associating intra-class clusters with learnable visual concepts in the feature space of a pre-trained language-vision model, e.g., CLIP. For class-conditional image generation, we design a cluster-conditional generator by injecting a combination of intra-class cluster label embeddings, and further incorporate a real-fake classification head on top of CLIP to distinguish real instances from the synthesized ones, conditioned on the learnable cluster prompts. This significantly strengthens the generator with more semantic language supervision. LCP-GAN not only possesses superior generation capability but also matches the performance of the fully supervised version of the base models: BigGAN and StyleGAN2-ADA, on multiple standard benchmarks.
Xiaoyang Huo, Si Wu 0002, Hau-San Wong
CVPR4
2023 Unknown Class Feature Transformation for Open Set Domain Adaptation Without Source Data
abstract
Utilizing deep neural network for Domain Adaptation (DA) has made great progress on learning knowledge from source domain to solve tasks in other relevant target domains. However, conventional DA methods have several restrictions: first, the category set from the source and target domain should be identical; second, during the training process data from different domains are fed into the network simultaneously, which may not be practical in real-world applications. Consequently, we aim at tackling the Source Free Open Set Domain Adaptation scenario: source and target data cannot meet with each other and the target domain contains exclusive unknown classes. Specifically, we propose a method enhancing the unknown class identification ability by synthesizing unknown class data: a feature modifier with multiple Gated Recurrent Units (GRU) is designed to learn and modify key features from the input data in order that the modified data ’looks like’ the real unknown class class data. Both real and synthesis data are used to train the classifier such that it can identify each class correctly. We evaluate our method in multiple benchmarks and the proposed framework outperforms other methods in comparison.
Si Wu 0002, Hau-San Wong
ICIP2
2023 Collaborative learning-based unknown-class instance identification for open-set domain adaptation
Haohong Zhou, Si Wu 0002, Cheng Liu 0001, Hau-San Wong
Inf. Sci.3
2023 Discriminator feature-based progressive GAN inversion
Quanpeng Song, Guanyue Li, Si Wu 0002, Hau-San Wong
Knowl. Based Syst.3
2023 Collaborative Learning with Unreliability Adaptation for Semi-Supervised Image Classification
Xiaoyang Huo, Xiangping Zeng, Si Wu 0002, Hau-San Wong
Pattern Recognit.3
2023 Self-Supervised Graph Completion for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) is challenging, as it requires adequately exploring complementary and consistency information under the incompleteness of data. Most existing approaches attempt to overcome the incompleteness at instance-level. In this work, we develop a new approach to facilitate IMVC from a new perspective. Specifically, we transfer the issue of missing instances to a similarity graph completion problem for incomplete views, and propose a self-supervised multi-view graph completion algorithm to infer the associated missing entries. Further, by incorporating constrained feature learning, the inferred graph can be naturally leveraged in representation learning. We theoretically show that our feature learning process performs an Auto-Regressive filter function by encoding the learned similarity graph, which could yield discriminative representation for a clustering task. Extensive experiments demonstrate the effectiveness of the proposed method in comparison with state-of-the-art methods.
Cheng Liu 0001, Si Wu 0002, Rui Li 0045, Dazhi Jiang, Hau-San Wong
IEEE Trans. Knowl. Data Eng.2
2022 Unreliability-Aware Disentangling for Cross-Domain Semi-supervised Pedestrian Detection
Si Wu 0002, Hau-San Wong
ACCV (2)2
2022 SphericGAN: Semi-supervised Hyper-spherical Generative Adversarial Networks for Fine-grained Image Synthesis
abstract
Generative Adversarial Network (GAN)-based models have greatly facilitated image synthesis. However, the model performance may be degraded when applied to finegrained data, due to limited training samples and subtle distinction among categories. Different from generic GAN-s, we address the issue from a new perspective of discovering and utilizing the underlying structure of real data to explicitly regularize the spatial organization of latent space. To reduce the dependence of generative models on labeled data, we propose a semi-supervised hyper-spherical GAN for class-conditional fine-grained image generation, and our model is referred to as SphericGAN. By projecting random vectors drawn from a prior distribution onto a hyper-sphere, we can model more complex distributions, while at the same time the similarity between the resulting latent vectors depends only on the angle, but not on their magnitudes. On the other hand, we also incorporate a mapping network to map real images onto the hyper-sphere, and match latent vectors with the underlying structure of real data via real-fake cluster alignment. As a result, we obtain a spatially organized latent space, which is useful for capturing class-independent variation factors. The experi-mental results suggest that our SphericGAN achieves state-of-the-art performance in synthesizing high-fidelity images with precise class semantics.
Xiaoyang Huo, Si Wu 0002, Yong Xu 0007, Hau-San Wong
CVPR4
2022 Semi-Supervised Generative Learning with Extended Distribution Matching for Class-Conditional Image Synthesis
abstract
Generative Adversarial Network (GAN)-based models have made remarkable progress in high-fidelity image synthesis. However, the performance of class-conditional image synthesis may significantly deteriorate for the case where a limited number of labeled training samples are available. To reduce the dependence on labeled data, we propose a semi-supervised GAN with Extended Distribution Matching, and our model is referred to as EDM-GAN. To prevent a class-conditional discriminator from overfitting the limited labeled data, we perform a transformation of random regional replacement on both real and synthesized samples. By matching the extended distributions, the discriminator is encouraged to focus more on the spatial regions that contain certain objects, while at the same time a class-conditional generator is induced to capture precise class semantics. The adversarial training process can be effectively stabilized and converges to a better solution. Our experimental results on multiple standard benchmarks demonstrate consistent performance gains in synthesis quality and class-semantic accuracy.
Xiaoyang Huo, Guangchang Deng, Si Wu 0002, Zhiwen Yu 0002
ICME3
2022 Source-Free Unsupervised Cross-Domain Pedestrian Detection via Pseudo Label Mining and Screening
abstract
Although current cross-domain pedestrian detection frame-works have obtained certain positive results, the performance is still source data dependent, which is cumbersome and im-practical in practical applications. To address this issue, we propose a source-free unsupervised pedestrian detection with pseudo label mining and screening. First, a modified CSP de-tector with DropBlock and three detection heads is presented. Then, a multi-expert method is proposed to fuse pseudo la-bels from three detection heads. Finally, a clustering-based self-supervised learning is adopted to categorize pseudo la-bels into positive and negative classes, which forms a set of clusters via similarity of pseudo labels and give classification results based on two confidence scores of each label from the detector backbone and multi-expert fusion. Experimental re-sults on three benchmark datasets show that the proposed approach can achieve state-of-the-art performance and be even comparable with other latest works using source data.
Qianfen Jiao, Si Wu 0002, Hau-San Wong
ICME4
2022 αβ-GAN: Robust generative adversarial networks
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
Inf. Sci.4
2022 Perturbation-insensitive cross-domain image enhancement for low-quality face verification
Qianfen Jiao, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
Inf. Sci.4
2022 Learning scene-adaptive pseudo annotations for pedestrian detection in semi-supervised scenarios
Qianfen Jiao, Hau-San Wong, Gaozhe Li, Si Wu 0002
Knowl. Based Syst.5
2022 TSEV-GAN: Generative Adversarial Networks with Target-aware Style Encoding and Verification for facial makeup transfer
Si Wu 0002, Qianfen Jiao, Hau-San Wong
Knowl. Based Syst.2
2022 The residual generator: An improved divergence minimization framework for GAN
Aurele Tohokantche Gnanha, Wenming Cao 0002, Xudong Mao, Si Wu 0002, Hau-San Wong, Qing Li 0001
Pattern Recognit.4
2022 Attention regularized semi-supervised learning with class-ambiguous data for image classification
Xiaoyang Huo, Xiangping Zeng, Si Wu 0002, Hau-San Wong
Pattern Recognit.3
2022 Supervised Graph Clustering for Cancer Subtyping Based on Survival Analysis and Integration of Multi-Omic Tumor Data
abstract
Identifying cancer subtypes by integration of multi-omic data is beneficial to improve the understanding of disease progression, and provides more precise treatment for patients. Cancer subtypes identification is usually accomplished by clustering patients with unsupervised learning approaches. Thus, most existing integrative cancer subtyping methods are performed in an entirely unsupervised way. An integrative cancer subtyping approach can be improved to discover clinically more relevant cancer subtypes when considering the clinical survival response variables. In this study, we propose a Survival Supervised Graph Clustering (S2GC)for cancer subtyping by taking into consideration survival information. Specifically, we use a graph to represent similarity of patients, and develop a multi-omic survival analysis embedding with patient-to-patient similarity graph learning for cancer subtype identification. The multi-view (omic)survival analysis model and graph of patients are jointly learned in a unified way. The learned optimal graph can be unitized to cluster cancer subtypes directly. In the proposed model, the survival analysis model and adaptive graph learning could positively reinforce each other. Consequently, the survival time can be considered as supervised information to improve the quality of the similarity graph and explore clinically more relevant subgroups of patients. Experiments on several representative multi-omic cancer datasets demonstrate that the proposed method achieves better results than a number of state-of-the-art methods. The results also suggest that our method is able to identify biologically meaningful subgroups for different cancer types. (Our Matlab source code is available online at github: https://github.com/CLiu272/S2GC).
Cheng Liu 0001, Wenming Cao 0002, Si Wu 0002, Dazhi Jiang, Zhiwen Yu 0002, Hau-San Wong
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Semisupervised Multiple Choice Learning for Ensemble Classification
abstract
Ensemble learning has many successful applications because of its effectiveness in boosting the predictive performance of classification models. In this article, we propose a semisupervised multiple choice learning (SemiMCL) approach to jointly train a network ensemble on partially labeled data. Our model mainly focuses on improving a labeled data assignment among the constituent networks and exploiting unlabeled data to capture domain-specific information, such that semisupervised classification can be effectively facilitated. Different from conventional multiple choice learning models, the constituent networks learn multiple tasks in the training process. Specifically, an auxiliary reconstruction task is included to learn domain-specific representation. For the purpose of performing implicit labeling on reliable unlabeled samples, we adopt a negative$\ell _{1}$-norm regularization when minimizing the conditional entropy with respect to the posterior probability distribution. Extensive experiments on multiple real-world datasets are conducted to verify the effectiveness and superiority of the proposed SemiMCL model.
Xiangping Zeng, Wenming Cao 0002, Si Wu 0002, Cheng Liu 0001, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Cybern.4
2022 Semantic Regularized Class-Conditional GANs for Semi-Supervised Fine-Grained Image Synthesis
abstract
Learning effective generative models for natural image synthesis is a promising way to reduce the dependence of deep models on massive training data. This work focuses on Fine-Grained Image Synthesis (FGIS) in the semi-supervised setting where a small number of training instances are labeled. Different from generic image synthesis tasks, the available fine-grained data may be inadequate, and the differences among the object categories are typically subtle. To address these issues, we propose a Semantic Regularized class-conditional Generative Adversarial Network, which is referred to as SReGAN. We incorporate an additional discriminator and classifier into the generator-discriminator minimax game. Competing with two discriminators enforces the generator to model both marginal and class-conditional data distributions, which alleviates the problem of limited training data and labels. However, the discriminators may overlook the class separability. To induce the generator to discover the distinctions between classes, we construct semantically congruent and incongruent pairs in the generation process, and further regularize the generator by encouraging high similarities of congruent pairs, while penalizing that of incongruent ones in the classifier's feature space. We have conducted extensive experiments to verify the capability of SReGAN in generating high-fidelity images on a variety of FGIS benchmarks.
Si Wu 0002, Xuhui Yang, Yong Xu 0007, Hau-San Wong
IEEE Trans. Multim.2
2022 Unreliable-to-Reliable Instance Translation for Semi-Supervised Pedestrian Detection
abstract
Generating realistic pedestrian instances in a semi-supervised setting is promising but challenging due to the limited labeled data. We propose an unreliable-to-reliable instance translation model (Un2Reliab) conditioned on unreliable instances which poorly align with pedestrians. Un2Reliab mainly consists of an encoder-decoder-like generative network and a discriminative network, which are jointly trained in a minimax game. We adopt regularization to ensure that the synthesized instances are semantically similar to the corresponding ground truth. Furthermore, to preserve the identities of persons, we propose another regularization to ensure that the synthesized instances associated with the same person should be consistent in appearance. As a result, Un2Reliab learns to restore the missing parts of the original instances. As a side benefit, the synthesized instances are brought into better alignment. Inclusion of the synthesized data improves both the diversity and quality of training data, which eventually leads to better generalization performance. Extensive experiments indicate that Un2Reliab is able to synthesize high-fidelity pedestrian instances and improve the previous state-of-the-art results on multiple semi-supervised pedestrian detection benchmarks.
Sihao Lin, Si Wu 0002, Yong Xu 0007, Hau-San Wong
IEEE Trans. Multim.3
2022 Asymmetric Graph-Guided Multitask Survival Analysis With Self-Paced Learning
abstract
Recently, multitask learning has been successfully applied to survival analysis problems. A critical challenge in real-world survival analysis tasks is that not all instances and tasks are equally learnable. A survival analysis model can be improved when considering the complexities of instances and tasks during the model training. To this end, we propose an asymmetric graph-guided multitask learning approach with self-paced learning for survival analysis applications. The proposed model is able to improve the learning performance by identifying the complex structure among tasks and considering the complexities of training instances and tasks during the model training. Especially, by incorporating the self-paced learning strategy and asymmetric graph-guided regularization, the proposed model is able to learn the model in a progressive way from "easy" to "hard" loss function items. In addition, together with the self-paced learning function, the asymmetric graph-guided regularization allows the related knowledge transfer from one task to another in an asymmetric way. Consequently, the knowledge acquired from those earlier learned tasks can help to solve complex tasks effectively. The experimental results on both synthetic and real-world TCGA data suggest that the proposed method is indeed useful for improving survival analysis and achieves higher prediction accuracies than the previous state-of-the-art methods.
Cheng Liu 0001, Wenming Cao 0002, Si Wu 0002, Dazhi Jiang, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Neural Networks Learn. Syst.3
2021 High Fidelity GAN Inversion via Prior Multi-Subspace Feature Composition
abstract
Generative Adversarial Networks (GANs) have shown impressive gains in image synthesis. GAN inversion was recently studied to understand and utilize the knowledge it learns, where a real image is inverted back to a latent code and can thus be reconstructed by the generator. Although increasing the number of latent codes can improve inversion quality to a certain extent, we find that important details may still be neglected when performing feature composition over all the intermediate feature channels. To address this issue, we propose a Prior multi-Subspace Feature Composition (PmSFC) approach for high-fidelity inversion. Considering that the intermediate features are highly correlated with each other, we incorporate a self-expressive layer in the generator to discover meaningful subspaces. In this case, the features at a channel can be expressed as a linear combination of those at other channels in the same subspace. We perform feature composition separately in the subspaces. The semantic differences between them benefit the inversion quality, since the inversion process is regularized based on different aspects of semantics. In the experiments, the superior performance of PmSFC demonstrates the effectiveness of prior subspaces in facilitating GAN inversion together with extended applications in visual manipulation.
Guanyue Li, Qianfen Jiao, Sheng Qian, Si Wu 0002, Hau-San Wong
AAAI4
2021 Mask-Embedded Discriminator With Region-Based Semantic Regularization for Semi-Supervised Class-Conditional Image Synthesis
abstract
Semi-supervised generative learning (SSGL) makes use of unlabeled data to achieve a trade-off between the data collection/annotation effort and generation performance, when adequate labeled data are not available. Learning precise class semantics is crucial for class-conditional image synthesis with limited supervision. Toward this end, we propose a semi-supervised Generative Adversarial Network with a Mask-Embedded Discriminator, which is referred to as MED-GAN. By incorporating a mask embedding module, the discriminator features are associated with spatial information, such that the focus of the discriminator can be limited in the specified regions when distinguishing between real and synthesized images. A generator is enforced to synthesize the instances holding more precise class semantics in order to deceive the enhanced discriminator. Also benefiting from mask embedding, region-based semantic regularization is imposed on the discriminator feature space, and the degree of separation between real and fake classes and among object categories can thus be increased. This eventually improves class-conditional distribution matching between real and synthesized data. In the experiments, the superior performance of MED-GAN demonstrates the effectiveness of mask embedding and associated regularizers in facilitating SSGL.
Xiaoyang Huo, Xiangping Zeng, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
CVPR5
2021 Semi-Supervised Single-Stage Controllable GANs for Conditional Fine-Grained Image Generation
abstract
Previous state-of-the-art deep generative models improve fine-grained image generation quality by designing hierarchical model structures and synthesizing images across multiple stages. The learning process is typically performed without any supervision in object categories. To address this issue, while at the same time to alleviate the level of complexity of both model design and training, we propose a Single-Stage Controllable GAN (SSCGAN) for conditional fine-grained image synthesis in a semi-supervised setting. Considering the fact that fine-grained object categories may have subtle distinctions and shared attributes, we take into account three factors of variation for generative modeling: class-independent content, cross-class attributes and class semantics, and associate them with different variables. To ensure disentanglement among the variables, we maximize mutual information between the class-independent variable and synthesized images, map real data to the latent space of a generator to perform consistency regularization of cross-class attributes, and incorporate class semantic-based regularization into a discriminator’s feature space. We show that the proposed approach delivers a single-stage controllable generator and high-fidelity synthesized images of fine-grained categories. SSC-GAN establishes state-of-the-art semi-supervised image synthesis results across multiple fine-grained datasets.
Si Wu 0002, Yong Xu 0007, Liangbing Feng, Hau-San Wong
ICCV4
2021 Unsupervised Domain Adaptation VIA Cluster Alignment with Maximum Classifier Discrepancy
abstract
One way of addressing the problem of unsupervised domain adaptation (UDA) is to perform adversarial training between two classifiers and their shared feature extractor. The two classifiers are enforced to detect the misaligned regions between the source and target domains, while the feature extractor aligns the features by confusing the classifiers. Although this method yields improvement, it ignores the relationship among target neighbors, which may consequently limit the model performance. In this work, we propose a new alignment strategy based on the "cluster assumption" to ensure the aligned target features preserve their clusters by avoiding overlap with decision boundaries. Furthermore, to make the aligned features more compact, we constrain them to be ro-bust against adversarial perturbation using the different views of the classifiers. Extensive experiments demonstrate the effectiveness of our solution on various datasets.
Mohamed Azzam, Si Wu 0002, Aurele Tohokantche Gnanha, Qianfen Jiao, Hau-San Wong
ICME2
2021 Adversarial Adaptive Interpolation for Regularizing Representation Learning and Image Synthesis in Autoencoders
abstract
Data interpolation is typically used to explore and understand the latent representation learnt by a deep network. Naive linear interpolation may induce mismatch between the interpolated data and the underlying manifold of the original data. In this paper, we propose an Adversarial Adaptive Interpolation (AdvAI) approach for facilitating representation learning and image synthesis in autoencoders. To determine an interpolation path that stays on the manifold, we incorporate an interpolation correction module, which learns to offset the deviation from the manifold. Further, we perform matching with a prior distribution to control the characteristics of the representation. The data synthesized from random codes along with interpolation-based regularization are in turn used to constrain the representation learning process. In the experiments, the superior performance of the proposed approach demonstrates the effectiveness of AdvAI and associated regularizers in a variety of downstream tasks.
Guanyue Li, Xiwen Wei, Sheng Qian, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
ICME4
2021 Adversarially Smoothed Feature Alignment for Visual Domain Adaptation
abstract
Recent approaches to unsupervised domain adaptation focus on transferring knowledge from source (labeled) data to target (unlabeled) data. Both data types share the same class space but originate from different domains. One way to achieve this transfer is tasking two classifiers to detect target features that diverge from source features. Meanwhile, a feature generator is adversarially refined to match the detected features with the source ones. However, the aligned features are still subject to ambiguity. Samples that are not smoothly distributed on the latent manifold are often missed in training. Moreover, target data may not be sufficient for adversarial learning. To overcome these problems, our proposed Adversarially Smoothed Feature Alignment (AdvSFA) model is designed to identify ambiguous target inputs by maximizing classifiers discrepancy in an extended class space. This enables the generator to receive valuable feedback from the classifiers and consequently learn more discriminative and smooth representation. Imposing smoothness on the latent manifold is a desirable property to improve model generalization and avoid having neighboring samples of different classes. To further promote such property, we not only task the generator to conduct feature alignment on target examples and but also in-between them. By adopting these constraints, our method shows a remarkable improvement across different adaptation tasks using two benchmark datasets.
Mohamed Azzam, Si Wu 0002, Yang Zhang 0073, Aurele Tohokantche Gnanha, Hau-San Wong
IJCNN2
2021 Discovering Density-Preserving Latent Space Walks in GANs for Semantic Image Transformations
abstract
Generative adversarial network (GAN)-based models possess superior capability of high-fidelity image synthesis. There are a wide range of semantically meaningful directions in the latent representation space of well-trained GANs, and the corresponding latent space walks are meaningful for semantic controllability in the synthesized images. To explore the underlying organization of a latent space, we propose an unsupervised Density-Preserving Latent Semantics Exploration model (DP-LaSE). The important latent directions are determined by maximizing the variations in intermediate features, while the correlation between the directions is minimized. Considering that latent codes are sampled from a prior distribution, we adopt a density-preserving regularization approach to ensure latent space walks are maintained in iso-density regions, since moving to a higher/lower density region tends to cause unexpected transformations. To further refine semantics-specific transformations, we perform subspace learning over intermediate feature channels, such that the transformations are limited to the most relevant subspaces. Extensive experiments on a variety of benchmark datasets demonstrate that DP-LaSE is able to discover interpretable latent space walks, and specific properties of synthesized images can thus be precisely controlled.
Guanyue Li, Xiwen Wei, Yang Zhang 0073, Si Wu 0002, Yong Xu 0007, Hau-San Wong
ACM Multimedia5
2021 Behavior regularized prototypical networks for semi-supervised few-shot image classification
Shixin Huang, Xiangping Zeng, Si Wu 0002, Zhiwen Yu 0002, Mohamed Azzam, Hau-San Wong
Pattern Recognit.3
2021 DDAT: Dual domain adaptive translation for low-resolution face verification in the wild
Qianfen Jiao, Rui Li 0045, Wenming Cao 0002, Si Wu 0002, Hau-San Wong
Pattern Recognit.5
2021 Knowledge Exchange Between Domain-Adversarial and Private Networks Improves Open Set Image Classification
abstract
Both target-specific and domain-invariant features can facilitate Open Set Domain Adaptation (OSDA). To exploit these features, we propose a Knowledge Exchange (KnowEx) model which jointly trains two complementary constituent networks: (1) a Domain-Adversarial Network (DAdvNet) learning the domain-invariant representation, through which the supervision in source domain can be exploited to infer the class information of unlabeled target data; (2) a Private Network (PrivNet) exclusive for target domain, which is beneficial for discriminating between instances from known and unknown classes. The two constituent networks exchange training experience in the learning process. Toward this end, we exploit an adversarial perturbation process against DAdvNet to regularize PrivNet. This enhances the complementarity between the two networks. At the same time, we incorporate an adaptation layer into DAdvNet to address the unreliability of the PrivNet's experience. Therefore, DAdvNet and PrivNet are able to mutually reinforce each other during training. We have conducted thorough experiments on multiple standard benchmarks to verify the effectiveness and superiority of KnowEx in OSDA.
Haohong Zhou, Mohamed Azzam, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
IEEE Trans. Image Process.5
2021 KTransGAN: Variational Inference-Based Knowledge Transfer for Unsupervised Conditional Generative Learning
abstract
Class-conditional generative models have gained popularity due to their characteristics of learning disentangled representations. However, these models typically require labeled examples in training. In this paper, we explore the feasibility of training these models on completely unlabeled data, under the assumption that we have access to other labeled data. The labeled data share the same label space, while their domain is shifted. Our model, which we refer to as KTransGAN, incorporates a classifier to transfer knowledge from the labeled data and performs collaborative learning with the conditional generator. By adopting these measures, KTransGAN is able to approximate the conditional distribution of the unlabeled data and simultaneously introduces a new solution to the unsupervised domain adaptation problem. To mitigate the training difficulty of our generative adversarial networks-based model, variational encoding and feature matching are also considered. From the empirical results, KTransGAN exhibits outstanding performance on a number of synthetic datasets and multiple real-world benchmarks. The quality of the synthesized instances is far superior to the pure variational autoencoding model. For example, on the CIFAR-10 dataset, our model scores 35.3 in FID, while the other model scores 128.45. In addition, the synthesis quality is close to the case when the model is trained in a fully supervised setting over the same number of training iterations. Regarding the classification performance, for instance, our model surpasses the highest state-of-the-art results (89.19%) by a large margin and achieves a test accuracy of 95.31% on the unlabeled data SVHN, while MNIST represents the labeled data. These results highlight the effectiveness of our proposed framework.
Mohamed Azzam, Wenming Cao 0002, Si Wu 0002, Hau-San Wong
IEEE Trans. Multim.4
2020 Model Adaptation: Unsupervised Domain Adaptation Without Source Data
abstract
In this paper, we investigate a challenging unsupervised domain adaptation setting --- unsupervised model adaptation. We aim to explore how to rely only on unlabeled target data to improve performance of an existing source prediction model on the target domain, since labeled source data may not be available in some real-world scenarios due to data privacy issues. For this purpose, we propose a new framework, which is referred to as collaborative class conditional generative adversarial net to bypass the dependence on the source data. Specifically, the prediction model is to be improved through generated target-style data, which provides more accurate guidance for the generator. As a result, the generator and the prediction model can collaborate with each other without source data. Furthermore, due to the lack of supervision from source data, we propose a weight constraint that encourages similarity to the source model. A clustering-based regularization is also introduced to produce more discriminative features in the target domain. Compared to conventional domain adaptation methods, our model achieves superior performance on multiple adaptation tasks with only unlabeled target data, which verifies its effectiveness in this challenging setting.
Rui Li 0045, Qianfen Jiao, Wenming Cao 0002, Hau-San Wong, Si Wu 0002
CVPR5
2020 Regularizing Discriminative Capability of CGANs for Semi-Supervised Generative Learning
abstract
Semi-supervised generative learning aims to learn the underlying class-conditional distribution of partially labeled data. Generative Adversarial Networks (GANs) have led to promising progress in this task. However, it still needs to further explore the issue of imbalance between real labeled data and fake data in the adversarial learning process. To address this issue, we propose a regularization technique based on Random Regional Replacement (R3-regularization) to facilitate the generative learning process. Specifically, we construct two types of between-class instances: cross-category ones and real-fake ones. These instances could be closer to the decision boundaries and are important for regularizing the classification and discriminative networks in our class-conditional GANs, which we refer to as R3-CGAN. Better guidance from these two networks makes the generative network produce instances with class-specific information and high fidelity. We experiment with multiple standard benchmarks, and demonstrate that the R3-regularization can lead to significant improvement in both classification and class-conditional image synthesis.
Guangchang Deng, Xiangping Zeng, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
CVPR4
2020 Adversarially Constrained Interpolation for Unsupervised Domain Adaptation
abstract
We address the problem of unsupervised domain adaptation (UDA) which aims at adapting models trained on a labeled domain to a completely unlabeled domain. One way to achieve this goal is to learn a domain-invariant representation. However, this approach is subject to two challenges: samples from two domains are insufficient to guarantee domain-invariance at most part of the latent space, and neighboring samples from the target domain may not belong to the same class on the low-dimensional manifold. To mitigate these shortcomings, we propose two strategies. First, we incorporate a domain mixup strategy in domain adversarial learning model by linearly interpolating between the source and target domain samples. This allows the latent space to be continuous and yields an improvement of the domain matching. Second, the domain discriminator is regularized via judging the relative difference between both domains for the input mixup features, which speeds up the domain matching. Experiment results show that our proposed model achieves a superior performance on different tasks under various domain shifts and data complexity.
Mohamed Azzam, Aurele Tohokantche Gnanha, Hau-San Wong, Si Wu 0002
ICPR4
2020 Joint subspace and discriminative learning for self-paced domain adaptation
Cheng Liu 0001, Si Wu 0002, Wenming Cao 0002, Dazhi Jiang, Zhiwen Yu 0002, Hau-San Wong
Knowl. Based Syst.2
2020 Simplified unsupervised image translation for semantic segmentation adaptation
Rui Li 0045, Wenming Cao 0002, Qianfen Jiao, Si Wu 0002, Hau-San Wong
Pattern Recognit.4
2020 Multitask Feature Selection by Graph-Clustered Feature Sharing
abstract
Multitask feature selection (MTFS) methods have become more important for many real world applications, especially in a high-dimensional setting. The most widely used assumption is that all tasks share the same features, and the l2,1 regularization method is usually applied. However, this assumption may not hold when the correlations among tasks are not obvious. Learning with unrelated tasks together may result in negative transfer and degrade the performance. In this paper, we present a flexible MTFS by graph-clustered feature sharing approach. To avoid the above limitation, we adopt a graph to represent the relevance among tasks instead of adopting a hard task set partition. Furthermore, we propose a graph-guided regularization approach such that the sparsity of the solution can be achieved on both the task level and the feature level, and a variant of the smooth proximal gradient method is developed to solve the corresponding optimization problem. An evaluation of the proposed method on multitask regression and multitask binary classification problem has been performed. Extensive experiments on synthetic datasets and real-world datasets demonstrate the effectiveness of the proposed approach to capture task structure.
Cheng Liu 0001, Chutao Zheng, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Cybern.3
2020 Generating Target Image-Label Pairs for Unsupervised Domain Adaptation
abstract
Deep learning demonstrates its impressive success across various machine learning problems. However, its performance often suffers in the case where the training and test data sets follow different distributions, due to the domain shift. Most current domain adaptation methods minimize the discrepancy between the source and target domains by enforcing the alignment of their marginal distributions without considering the class-level matching. Consequently, data from different classes may become close together after mapping. To address this issue, we propose an unsupervised domain adaptation method by generating image-label pairs in the target domain, in which the model is augmented with the generated target pairs and achieve class-level transfer. Specifically, we integrate generative adversarial networks (GAN) into the model predictor, where the generator fed with labels aims to produce corresponding target domain images with a well-designed semantic loss. Meanwhile, compared to previous methods which focus on discrepancy reduction across domains, i.e., image to image translation, our model focuses on semantic preservation during image generation. Our model is straightforward yet effective for unsupervised domain adaptation problems. Without any labels in the target domain in all the experiments, we demonstrate the validity of our approach by presenting the plausible generated target image-label pairs. In addition, our proposed method achieves the best or comparable performance on multiple unsupervised domain adaptation datasets which include image classification and semantic segmentation.
Rui Li 0045, Wenming Cao 0002, Si Wu 0002, Hau-San Wong
IEEE Trans. Image Process.3
2020 Semi-Supervised Deep Coupled Ensemble Learning With Classification Landmark Exploration
abstract
Using an ensemble of neural networks with consistency regularization is effective for improving performance and stability of deep learning, compared to the case of a single network. In this paper, we present a semi-supervised Deep Coupled Ensemble (DCE) model, which contributes to ensemble learning and classification landmark exploration for better locating the final decision boundaries in the learnt latent space. First, multiple complementary consistency regularizations are integrated into our DCE model to enable the ensemble members to learn from each other and themselves, such that training experience from different sources can be shared and utilized during training. Second, in view of the possibility of producing incorrect predictions on a number of difficult instances, we adopt class-wise mean feature matching to explore important unlabeled instances as classification landmarks, on which the model predictions are more reliable. Minimizing the weighted conditional entropy on unlabeled data is able to force the final decision boundaries to move away from important training data points, which facilitates semi-supervised learning. Ensemble members could eventually have similar performance due to consistency regularization, and thus only one of these members is needed during the test stage, such that the efficiency of our model is the same as the non-ensemble case. Extensive experimental results demonstrate the superiority of our proposed DCE model over existing state-of-the-art semi-supervised learning methods.
Jichang Li, Si Wu 0002, Cheng Liu 0001, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Image Process.2
2020 Semi-Supervised Human Detection via Region Proposal Networks Aided by Verification
abstract
In this paper, we explore how to leverage readily available unlabeled data to improve semi-supervised human detection performance. For this purpose, we specifically modify the region proposal network (RPN) for learning on a partially labeled dataset. Based on commonly observed false positive types, a verification module is developed to assess foreground human objects in the candidate regions to provide an important cue for filtering the RPN's proposals. The remaining proposals with high confidence scores are then used as pseudo annotations for re-training our detection model. To reduce the risk of error propagation in the training process, we adopt a self-paced training strategy to progressively include more pseudo annotations generated by the previous model over multiple training rounds. The resulting detector re-trained on the augmented data can be expected to have better detection performance. The effectiveness of the main components of this framework is verified through extensive experiments, and the proposed approach achieves state-of-the-art detection results on multiple scene-specific human detection benchmarks in the semi-supervised setting.
Si Wu 0002, Shiyao Lei, Sihao Lin, Rui Li 0045, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Image Process.1
2019 Improving Domain-Specific Classification by Collaborative Learning with Adaptation Networks
abstract
For unsupervised domain adaptation, the process of learning domain-invariant representations could be dominated by the labeled source data, such that the specific characteristics of the target domain may be ignored. In order to improve the performance in inferring target labels, we propose a targetspecific network which is capable of learning collaboratively with a domain adaptation network, instead of directly minimizing domain discrepancy. A clustering regularization is also utilized to improve the generalization capability of the target-specific network by forcing target data points to be close to accumulated class centers. As this network learns and specializes to the target domain, its performance in inferring target labels improves, which in turn facilitates the learning process of the adaptation network. Therefore, there is a mutually beneficial relationship between these two networks. We perform extensive experiments on multiple digit and object datasets, and the effectiveness and superiority of the proposed approach is presented and verified on multiple visual adaptation benchmarks, e.g., we improve the state-ofthe-art on the task of MNIST→SVHN from 76.5% to 84.9% without specific augmentation.
Si Wu 0002, Wenming Cao 0002, Rui Li 0045, Zhiwen Yu 0002, Hau-San Wong
AAAI1
2019 Enhancing TripleGAN for Semi-Supervised Conditional Instance Synthesis and Classification
abstract
Learning class-conditional data distributions is crucial for Generative Adversarial Networks (GAN) in semi-supervised learning. To improve both instance synthesis and classification in this setting, we propose an enhanced TripleGAN (EnhancedTGAN) model in this work. We follow the adversarial training scheme of the original TripleGAN, but completely re-design the training targets of the generator and classifier. Specifically, we adopt feature-semantics matching to enhance the generator in learning class-conditional distributions from both the aspects of statistics in the latent space and semantics consistency with respect to the generator and classifier. Since a limited amount of labeled data is not sufficient to determine satisfactory decision boundaries, we include two classifiers, and incorporate collaborative learning into our model to provide better guidance for generator training. The synthesized high-fidelity data can in turn be used for improving classifier training. In the experiments, the superior performance of our approach on multiple benchmark datasets demonstrates the effectiveness of the mutual reinforcement between the generator and classifiers in facilitating semi-supervised instance synthesis and classification.
Si Wu 0002, Guangchang Deng, Jichang Li, Rui Li 0045, Zhiwen Yu 0002, Hau-San Wong
CVPR1
2019 Mutual Learning of Complementary Networks via Residual Correction for Improving Semi-Supervised Classification
abstract
Deep mutual learning jointly trains multiple essential networks having similar properties to improve semi-supervised classification. However, the commonly used consistency regularization between the outputs of the networks may not fully leverage the difference between them. In this paper, we explore how to capture the complementary information to enhance mutual learning. For this purpose, we propose a complementary correction network (CCN), built on top of the essential networks, to learn the mapping from the output of one essential network to the ground truth label, conditioned on the features learnt by another. To make the second essential network increasingly complementary to the first one, this network is supervised by the corrected predictions. As a result, minimizing the prediction divergence between the two complementary networks can lead to significant performance gains in semi-supervised learning. Our experimental results demonstrate that the proposed approach clearly improves mutual learning between essential networks, and achieves state-of-the-art results on multiple semi-supervised classification benchmarks. In particular, the test error rates are reduced from previous 21.23% and 14.65% to 12.05% and 10.37% on CIFAR-10 with 1000 and 2000 labels, respectively.
Si Wu 0002, Jichang Li, Cheng Liu 0001, Zhiwen Yu 0002, Hau-San Wong
CVPR1
2019 Semi-Supervised Pedestrian Instance Synthesis and Detection With Mutual Reinforcement
abstract
We propose a GAN-based scene-specific instance synthesis and classification model for semi-supervised pedestrian detection. Instead of collecting unreliable detections from unlabeled data, we adopt a class-conditional GAN for synthesizing pedestrian instances to alleviate the problem of insufficient labeled data. With the help of a base detector, we integrate pedestrian instance synthesis and detection by including a post-refinement classifier (PRC) into a minimax game. A generator and the PRC can mutually reinforce each other by synthesizing high-fidelity pedestrian instances and providing more accurate categorical information. Both of them compete with a class-conditional discriminator and a class-specific discriminator, such that the four fundamental networks in our model can be jointly trained. In our experiments, we validate that the proposed model significantly improves the performance of the base detector and achieves state-of-the-art results on multiple benchmarks. As shown in Figure 1, the result indicates the possibility of using inexpensively synthesized instances for improving semi-supervised detection models.
Si Wu 0002, Sihao Lin, Mohamed Azzam, Hau-San Wong
ICCV1
2019 Improving representation learning in autoencoders via multidimensional interpolation and dual regularizations
abstract
Autoencoders enjoy a remarkable ability to learn data representations. Research on autoencoders shows that the effectiveness of data interpolation can reflect the performance of representation learning. However, existing interpolation methods in autoencoders do not have enough capability of traversing a possible region between two datapoints on a data manifold, and the distribution of interpolated latent representations is not considered.To address these issues, we aim to fully exert the potential of data interpolation and further improve representation learning in autoencoders. Specifically, we propose the multidimensional interpolation to increase the capability of data interpolation by randomly setting interpolation coefficients for each dimension of latent representations. In addition, we regularize autoencoders in both the latent and the data spaces by imposing a prior on latent representations in the Maximum Mean Discrepancy (MMD) framework and encouraging generated datapoints to be realistic in the Generative Adversarial Network (GAN) framework. Compared to representative models, our proposed model has empirically shown that representation learning exhibits better performance on downstream tasks on multiple benchmarks.
Sheng Qian, Guanyue Li, Wenming Cao 0002, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
IJCAI5
2019 Unsupervised Multi-task Learning with Hierarchical Data Structure
Wenming Cao 0002, Sheng Qian, Si Wu 0002, Hau-San Wong
Pattern Recognit.3
2019 Encoding sparse and competitive structures among tasks in multi-task learning
Cheng Liu 0001, Chutao Zheng, Sheng Qian, Si Wu 0002, Hau-San Wong
Pattern Recognit.4
2019 Multiobjective Semisupervised Classifier Ensemble
abstract
Classification of high-dimensional data with very limited labels is a challenging task in the field of data mining and machine learning. In this paper, we propose the multiobjective semisupervised classifier ensemble (MOSSCE) approach to address this challenge. Specifically, a multiobjective subspace selection process (MOSSP) in MOSSCE is first designed to generate the optimal combination of feature subspaces. Three objective functions are then proposed for MOSSP, which include the relevance of features, the redundancy between features, and the data reconstruction error. Then, MOSSCE generates an auxiliary training set based on the sample confidence to improve the performance of the classifier ensemble. Finally, the training set, combined with the auxiliary training set, is used to select the optimal combination of basic classifiers in the ensemble, train the classifier ensemble, and generate the final result. In addition, diversity analysis of the ensemble learning process is applied, and a set of nonparametric statistical tests is adopted for the comparison of semisupervised classification approaches on multiple datasets. The experiments on 12 gene expression datasets and two large image datasets show that MOSSCE has a better performance than other state-of-the-art semisupervised classifiers on high-dimensional data.
Zhiwen Yu 0002, C. L. Philip Chen, Jane You, Hau-San Wong, Dan Dai, Si Wu 0002, Jun Zhang 0003
IEEE Trans. Cybern.7
2019 Exploring Correlations Among Tasks, Clusters, and Features for Multitask Clustering
abstract
Multitask clustering methods are proposed to improve performances of related tasks concurrently, because they explore the relationship among tasks via exploiting the coefficient matrix or the shared feature matrix. However, divergent effects of features in learning this relationship are seldom considered. To further improve performances, we propose a new multitask clustering approach through exploring correlations among tasks, clusters, and features based on effects of features on clusters. First, a Feature-Cluster (FeaCluster) matrix is introduced to capture the similarity and the distinct task-feature information simultaneously for each task. With the FeaCluster matrix, two affinities are calculated to constitute the interdependencies among tasks: the former is the graphical affinity based on feature-task and task-cluster correlations, while the latter is the reconstructive affinity. Here, the feature-task correlation considers effects of features on tasks, and the task-cluster correlation considers the overall effects of features on clusters. The reconstructive affinity is obtained by minimizing the reconstruction error when representing the FeaCluster matrix for a given task with a linear combination of others. The interdependencies among tasks allow transferring asymmetric shared information, exploring significant features and preserving key information when mapping data into the subspace. The experimental results on multiple data sets reveal that the proposed approach outperforms the state-of-the-art clustering methods in terms of accuracy and normal mutual information.
Wenming Cao 0002, Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
IEEE Trans. Neural Networks Learn. Syst.2
2018 Efficient Direct Structured Subspace Clustering
Wenming Cao 0002, Rui Li 0045, Sheng Qian, Si Wu 0002, Hau-San Wong
ICONIP (4)4
2018 Cross-domain Semantic Feature Learning via Adversarial Adaptation Networks
abstract
Existing domain adaptation approaches generalize models trained on the labeled source domain data to the unlabeled target domain data by forcing feature distributions of two domains closer. However, these approaches are likely to ignore the semantic information during the feature alignment between source and target domain. In this paper, we propose a new unsupervised domain adaptation framework to learn the cross-domain features and disentangle the semantic information concurrently. Specifically, we firstly combine the task-specific classification and domain adversarial learning to obtain the cross-domain features by mapping the data of both domains with the shared feature extractor. Secondly, we integrate the domain adversarial learning and the within-domain reconstruction to disentangle the semantic information from the domain information. Thirdly, we include a cross-domain transformation to further refine the feature extractor, which in turn improves the performances of the task classifier. We compare our proposed model to previous state-of-the-art methods on domain adaptation digit classification tasks. Experimental results show that our model achieves better performances than the other counterparts, which demonstrates the superiority and effectiveness of our model.
Rui Li 0045, Wenming Cao 0002, Sheng Qian, Hau-San Wong, Si Wu 0002
ICPR5
2018 Adaptive activation functions in convolutional neural networks
Sheng Qian, Cheng Liu 0001, Si Wu 0002, Hau-San Wong
Neurocomputing4
2018 Variant SemiBoost for Improving Human Detection in Application Scenes
abstract
Generic human detectors perform poorly in application scenes in which conditions are significantly different from those of the benchmark data sets. Based on the assumption that only a limited number of labeled examples are available, we propose a variant semi-supervised boosting approach for improving scene adaptiveness by utilizing unlabeled data. Specifically, we train a max-margin-based model as an initial detector for new example collection, instead of using generic detectors, and then a better model is trained via boosting in which the newly obtained examples influence the training process through their similarities to the labeled examples. Since the widely used human descriptors are usually high dimensional and redundant, we employ a graph-based method to determine the weight representing the importance of each feature, such that the weighted similarity measurement leads to a performance gain. In the experiments, the effectiveness of the proposed approach “Variant SemiBoost” is demonstrated and state-of-the-art performance on challenging data sets is achieved.
Si Wu 0002, Hau-San Wong, Shufeng Wang
IEEE Trans. Circuits Syst. Video Technol.1
2018 Exploiting Target Data to Learn Deep Convolutional Networks for Scene-Adapted Human Detection
abstract
The difference between sample distributions of public data sets and specific scenes can be very significant. As a result, the deployment of generic human detectors in real-world scenes most often leads to sub-optimal detection performance. To avoid the labor-intensive task of manual annotations, we propose a semi-supervised approach for training deep convolutional networks on partially labeled data. To exploit a large amount of unlabeled target data, the knowledge learnt from public data sets is transferred to new model training by adapting an auxiliary detector to the target scene. We hypothesize that the components of the auxiliary detector capture essential human characteristics useful for constructing a scene-adapted detector. A selective ensemble algorithm is proposed to select a subset of the components relevant to the target scene for recombination. The resulting model is applied for collecting high-confidence samples from unlabeled target data. Furthermore, a deep convolutional network is trained by progressively labeling and selecting new training samples in a self-paced way. The detailed experimental evaluation verifies the effectiveness and superiority of the proposed approach in scene-specific human detection.
Si Wu 0002, Shufeng Wang, Robert Laganière, Cheng Liu 0001, Hau-San Wong, Yong Xu 0007
IEEE Trans. Image Process.1
2018 Semi-Supervised Image Classification With Self-Paced Cross-Task Networks
abstract
In a semi-supervised setting, direct training of a deep discriminative model on partially labeled images often suffers from overfitting and poor performance, because only a small number of labeled images are available, and errors in label propagation are, in many cases, inevitable. In this paper, we introduce an auxiliary clustering task to explore the structure of the image data, and judiciously weigh unlabeled data to alleviate the influence of ambiguous data on model training. For this purpose, we propose a cross-task network composed of two streams to jointly learn two tasks: classification and clustering. Based on the model predictions, a large number of pairwise constraints can be generated from unlabeled images, and are fed to the clustering stream. Since pairwise constraints encode weak supervision information, the clustering is tolerant of errors in labeling. Unlabeled images are weighted according to the distances to the clusters discovered, and a better discriminative model is trained on the classification stream associated with a weighted softmax loss. Furthermore, a self-paced learning paradigm is adopted to gradually train our deep model from easy examples to difficult ones. Experimental results on widely used image classification datasets confirm the effectiveness and superiority of the proposed approach.
Si Wu 0002, Qiujia Ji, Shufeng Wang, Hau-San Wong, Zhiwen Yu 0002, Yong Xu 0007
IEEE Trans. Multim.1
2016 Incremental semi-supervised clustering ensemble for high dimensional data clustering
abstract
Recently, cluster ensemble approaches have gained more and more attention [1]–[2], due to useful applications in the areas of pattern recognition, data mining, bioinformatics, and so on. When compared with traditional single clustering algorithms, cluster ensemble approaches are able to integrate multiple clustering solutions obtained from different data sources into a unified solution, and provide a more robust, stable and accurate final result.
Zhiwen Yu 0002, Peinan Luo, Si Wu 0002, Guoqiang Han 0002, Jane You, Hareton K. N. Leung, Hau-San Wong, Jun Zhang 0003
ICDE3
2016 Progressive subspace ensemble learning
Zhiwen Yu 0002, Daxing Wang, Jane You, Hau-San Wong, Si Wu 0002, Jun Zhang 0003, Guoqiang Han 0002
Pattern Recognit.5
2016 Incremental Semi-Supervised Clustering Ensemble for High Dimensional Data Clustering
abstract
Traditional cluster ensemble approaches have three limitations: (1) They do not make use of prior knowledge of the datasets given by experts. (2) Most of the conventional cluster ensemble methods cannot obtain satisfactory results when handling high dimensional data. (3) All the ensemble members are considered, even the ones without positive contributions. In order to address the limitations of conventional cluster ensemble approaches, we first propose an incremental semi-supervised clustering ensemble framework (ISSCE) which makes use of the advantage of the random subspace technique, the constraint propagation approach, the proposed incremental ensemble member selection process, and the normalized cut algorithm to perform high dimensional data clustering. The random subspace technique is effective for handling high dimensional data, while the constraint propagation approach is useful for incorporating prior knowledge. The incremental ensemble member selection process is newly designed to judiciously remove redundant ensemble members based on a newly proposed local cost function and a global cost function, and the normalized cut algorithm is adopted to serve as the consensus function for providing more stable, robust, and accurate results. Then, a measure is proposed to quantify the similarity between two sets of attributes, and is used for computing the local cost function in ISSCE. Next, we analyze the time complexity of ISSCE theoretically. Finally, a set of nonparametric tests are adopted to compare multiple semisupervised clustering ensemble approaches over different datasets. The experiments on 18 real-world datasets, which include six UCI datasets and 12 cancer gene expression profiles, confirm that ISSCE works well on datasets with very high dimensionality, and outperforms the state-of-the-art semi-supervised clustering ensemble approaches.
Zhiwen Yu 0002, Peinan Luo, Jane You, Hau-San Wong, Hareton K. N. Leung, Si Wu 0002, Jun Zhang 0003, Guoqiang Han 0002
IEEE Trans. Knowl. Data Eng.6
2015 Improving pedestrian detection with selective gradient self-similarity feature
Si Wu 0002, Robert Laganière, Pierre Payeur
Pattern Recognit.1
2014 A Bayesian Model for Crowd Escape Behavior Detection
abstract
People naturally escape from a place when unexpected events happen. Based on this observation, efficient detection of crowd escape behavior in surveillance videos is a promising way to perform timely detection of anomalous situations. In this paper, we propose a Bayesian framework for escape detection by directly modeling crowd motion in both the presence and absence of escape events. Specifically, we introduce the concepts of potential destinations and divergent centers to characterize crowd motion in the above two cases respectively, and construct the corresponding class-conditional probability density functions of optical flow. Escape detection is finally performed based on the proposed Bayesian framework. Although only data associated with nonescape behavior are included in the training set, the density functions associated with the case of escape can also be adaptively updated using observed data. In addition, the identified divergent centers indicate possible locations at which the unexpected events occur. The performance of our proposed method is validated in a number of experiments on crowd escape detection in various scenarios.
Si Wu 0002, Hau-San Wong, Zhiwen Yu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2012 A fuzzy minimax clustering model and its applications
Xiang Li 0006, Hau-San Wong, Si Wu 0002
Inf. Sci.3
2012 Joint segmentation of collectively moving objects using a bag-of-words model and level set evolution
Si Wu 0002, Hau-San Wong
Pattern Recognit.1
2012 Crowd Motion Partitioning in a Scattered Motion Field
abstract
In this paper, we propose a crowd motion partitioning approach based on local-translational motion approximation in a scattered motion field. To represent crowd motion in an accurate and parsimonious way, we compute optical flow at the salient locations instead of at all the pixel locations. We then transform the problem of crowd motion partitioning into a problem of scattered motion field segmentation. Based on our assumption that local crowd motion can be approximated by a translational motion field, we develop a local-translation domain segmentation (LTDS) model in which the evolution of domain boundaries is derived from the Gâteaux derivative of an objective functional and further extend LTDS to the case of scattered motion field. The experiment results on a set of synthetic vector fields and a set of videos depicting real-world crowd scenes indicate that the proposed approach is effective in identifying the homogeneous crowd motion components under different scenarios.
Si Wu 0002, Hau-San Wong
IEEE Trans. Syst. Man Cybern. Part B1
2009 A Shape Derivative Based Approach for Crowd Flow Segmentation
Si Wu 0002, Zhiwen Yu 0002, Hau-San Wong
ACCV (1)1