Zijian Wang 0009

dblp:03/4540-9 · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
27since 2021 · last 2026
0000-0002-7190-9620ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCORE: Soft Label Compression-Centric Dataset Condensation via Coding Rate Optimization
abstract
Dataset Condensation (DC) aims to obtain a condensed dataset that allows models trained on the condensed dataset to achieve performance comparable to those trained on the full dataset. Recent DC approaches increasingly focus on encoding knowledge into realistic images with soft labeling, for their scalability to ImageNet-scale datasets and strong capability of cross-domain generalization. However, this strong performance comes at a substantial storage cost which could significantly exceed the storage cost of the original dataset. We argue that the three key properties to alleviate this performance-storage dilemma are informativeness, discriminativeness, and compressibility of the condensed data. Towards this end, this paper proposes a Soft label compression-centric dataset condensation framework using COding RatE (SCORE). SCORE formulates dataset condensation as a min-max optimization problem, which aims to balance the three key properties from an information-theoretic perspective. In particular, we theoretically demonstrate that our coding rate-inspired objective function is sub-modular, and its optimization naturally enforces low-rank structure in the soft label set corresponding to each condensed data. Extensive experiments on large-scale datasets, including ImageNet-1K and Tiny-ImageNet, demonstrate that SCORE outperforms existing methods in most cases. Even with a 30× compression of soft labels, performance decreases by only 5.5% and 2.7% for ImageNet-1K with IPC 10 and 50, respectively. The code is available at https://github.com/KeViNYuAn0314/SCORE.
Yuxia Fu, Zijian Wang 0009, Yadan Luo, Zi Huang
WACV3
2025 PEFTDiff: Diffusion-Guided Transferability Estimation for Parameter-Efficient Fine-Tuning
Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash
ICCV2
2025 Improving Out-of-Distribution Detection via Dynamic Covariance Calibration
abstract
Out-of-Distribution (OOD) detection is essential for the trustworthiness of AI systems. Methods using prior information (i.e., subspace-based methods) have shown effective performance by extracting information geometry to detect OOD data with a more appropriate distance metric. However, these methods fail to address the geometry distorted by ill-distributed samples, due to the limitation of statically extracting information geometry from the training distribution. In this paper, we argue that the influence of ill-distributed samples can be corrected by dynamically adjusting the prior geometry in response to new data. Based on this insight, we propose a novel approach that dynamically updates the prior covariance matrix using real-time input features, refining its information. Specifically, we reduce the covariance along the direction of real-time input features and constrain adjustments to the residual space, thus preserving essential data characteristics and avoiding effects on unintended directions in the principal space. We evaluate our method on two pre-trained models for the CIFAR dataset and five pre-trained models for ImageNet-1k, including the self-supervised DINO model. Extensive experiments demonstrate that our approach significantly enhances OOD detection across various models. The code is released at https://github.com/workerbcd/ooddcc.
Kaiyu Guo, Zijian Wang 0009, Tan Pan, Brian C. Lovell, Mahsa Baktash
ICML2
2025 WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management
abstract
Wheat management strategies play a critical role in determining yield. Traditional management decisions often rely on labour-intensive expert inspections, which are expensive, subjective and difficult to scale. Recently, Vision-Language Models (VLMs) have emerged as a promising solution to enable scalable, data-driven management support. However, due to a lack of domain-specific knowledge, directly applying VLMs to wheat management tasks results in poor quantification and reasoning capabilities, ultimately producing vague or even misleading management recommendations. In response, we propose WisWheat, a wheat-specific dataset with a three-layered design to enhance VLM performance on wheat management tasks: (1) a foundational pretraining dataset of 47,871 image-caption pairs for coarsely adapting VLMs to wheat morphology; (2) a quantitative dataset comprising 7,263 VQA-style image-question-answer triplets for quantitative trait measuring tasks; and (3) an Instruction Fine-tuning dataset with 4,888 samples targeting biotic and abiotic stress diagnosis and management plan for different phenological stages. Extensive experimental results demonstrate that fine-tuning open-source VLMs (e.g., Qwen2.5 7B) on our dataset leads to significant performance improvements. Specifically, the Qwen2.5 VL 7B fine-tuned on our wheat instruction dataset achieves accuracy scores of 79.2% and 84.6% on wheat stress and growth stage conversation tasks respectively, surpassing even general-purpose commercial models such as GPT-4o by a margin of 11.9% and 34.6%.
Selena Song, Javier Fernandez, Yadan Luo, Mahsa Baktash, Zijian Wang 0009
ACM Multimedia6
2025 Spectral Distribution Alignment for Enhanced Generalization in Regression
Kaiyu Guo, Zijian Wang 0009, Brian C. Lovell, Mahsa Baktash
ECML/PKDD (6)2
2025 Feature Space Perturbation: A Panacea to Enhanced Transferability Estimation
abstract
Leveraging a transferability estimation metric facilitates the non-trivial challenge of selecting the optimal model for the downstream task from a pool of pre-trained models. Most existing metrics primarily focus on identifying the statistical relationship between feature embeddings and the corresponding labels within the target dataset, but over-look crucial aspect of model robustness. This oversight may limit their effectiveness in accurately ranking pre-trained models. To address this limitation, we introduce a feature perturbation method that enhances the transferability estimation process by systematically altering the feature space. Our method includes a Spread operation that increases intra-class variability, adding complexity within classes, and an Attract operation that minimizes the distances between different classes, thereby blurring the class boundaries. Through extensive experimentation, we demonstrate the efficacy of our feature perturbation method in providing a more precise and robust estimation of model transferability. Notably, the existing LogMe method exhibited a significant improvement, showing a 28.84% increase in performance after applying our feature perturbation method. The implementation is available at https://github.com/prafful-kumar/enhancing_TE.git
Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash
WACV2
2025 Open-CRB: Toward Open World Active Learning for 3D Object Detection
abstract
LiDAR-based 3D object detection has recently seen significant advancements through active learning (AL), attaining satisfactory performance by training on a small fraction of strategically selected point clouds. However, in real-world deployments where streaming point clouds may include unknown or novel objects, the ability of current AL methods to capture such objects remains unexplored. This paper investigates a more practical and challenging research task: Open World Active Learning for 3D Object Detection (OWAL-3D), aimed at acquiring informative point clouds with new concepts. To tackle this challenge, we propose a simple yet effective strategy called Open Label Conciseness (OLC), which mines novel 3D objects with minimal annotation costs. Our empirical results show that OLC successfully adapts the 3D detection model to the open world scenario with just a single round of selection. Any generic AL policy can then be integrated with the proposed OLC to efficiently address the OWAL-3D problem. Based on this, we introduce the Open-CRB framework, which seamlessly integrates OLC with our preliminary AL method, CRB, designed specifically for 3D object detection. We develop a comprehensive codebase for easy reproducing and future research, supporting 15 baseline methods (i.e., active learning, out-of-distribution detection and open world detection), 2 types of modern 3D detectors (i.e., one-stage SECOND and two-stage PV-RCNN) and 3 benchmark 3D datasets (i.e., KITTI, nuScenes and Waymo). Extensive experiments evidence that the proposed Open-CRB demonstrates superiority and flexibility in recognizing both novel and known classes with very limited labeling costs, compared to state-of-the-art baselines.
Zhuoxiao Chen, Yadan Luo, Zijian Wang 0009, Zi Huang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 CASIA-PR-V1: A Multi-Ethnic, Multi-Device and Cross-Spectral Dataset and a Multiscale Disentangled Model for Periocular Recognition
abstract
Periocular recognition is regarded as an alternative trait for biometric recognition that can effectively solve the identification problem under large occlusions. However, few datasets are tailored for periocular recognition. For most compromises, iris datasets at near-infrared wavelengths, miss information about the eyebrows or eyelids. In this paper, a challenging dataset for real scenarios named CASIA-PR-V1 with evaluation protocols is released for periocular recognition. It is collected from multiple types of mobile devices with different resolutions or wavelengths. A rich set of attributes, e.g., ethnicities, is tagged to support fine-grained classification tasks. Moreover, we consider a wide range of noisy data in unconstrained environment, especially for glasses. Superior to its counterparts, this periocular dataset is highly valuable for studying cross-device and cross-spectral periocular recognition with occlusions, as well as fine-grained attribute classification. Additionally, a multiscale disentangled model is proposed to extract discriminating representations for periocular recognition with severe occlusions. Extensive experiments are conducted on CASIA-PR-V1, and the results indicate the superiority of our model for unconstraint periocular recognition.
Yiwei Ru, Yushan Han, Longteng Kong, Zijian Wang 0009, Yong He 0009, Zhenan Sun
IEEE Trans. Multim.5
2025 Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation Learning
abstract
Biphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets.
Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan
IEEE Trans. Neural Networks Learn. Syst.3
2024 CIFAR-10-Warehouse: Broad and More Realistic Testbeds in Model Generalization Analysis
abstract
Analyzing model performance in various unseen environments is a critical research problem in the machine learning community. To study this problem, it is important to construct a testbed with out-of-distribution test sets that have broad coverage of environmental discrepancies. However, existing testbeds typically either have a small number of domains or are synthesized by image corruptions, hindering algorithm design that demonstrates real-world effectiveness. In this paper, we introduce CIFAR-10-Warehouse, consisting of 180 datasets collected by prompting image search engines and diffusion models in various ways. Generally sized between 300 and 8,000 images, the datasets contain natural images, cartoons, certain colors, or objects that do not naturally appear. With CIFAR-10-W, we aim to enhance the evaluation and deepen the understanding of two generalization tasks: domain generalization and model accuracy prediction in various out-of-distribution environments. We conduct extensive benchmarking and comparison experiments and show that CIFAR-10-W offers new and interesting insights inherent to these tasks. We also discuss other fields that would benefit from CIFAR-10-W. Data and code are available at https://sites.google.com/view/CIFAR-10-warehouse/.
Xiaoxiao Sun 0002, Xingjian Leng, Zijian Wang 0009, Yang Yang 0223, Zi Huang, Liang Zheng 0001
ICLR3
2024 Color-Oriented Redundancy Reduction in Dataset Distillation
abstract
Dataset Distillation (DD) is designed to generate condensed representations of extensive image datasets, enhancing training efficiency. Despite recent advances, there remains considerable potential for improvement, particularly in addressing the notable redundancy within the color space of distilled images. In this paper, we propose a two-fold optimization strategy to minimize color redundancy at the individual image and overall dataset levels, respectively. At the image level, we employ a palette network, a specialized neural network, to dynamically allocate colors from a reduced color space to each pixel. The palette network identifies essential areas in synthetic images for model training, and consequently assigns more unique colors to them. At the dataset level, we develop a color-guided initialization strategy to minimize redundancy among images. Representative images with the least replicated color patterns are selected based on the information gain. A comprehensive performance study involving various datasets and evaluation scenarios is conducted, demonstrating the superior performance of our proposed color-aware DD compared to existing DD methods.
Zijian Wang 0009, Mahsa Baktash, Yadan Luo, Zi Huang
NeurIPS2
2023 How Far Pre-trained Models Are from Neural Collapse on the Target Dataset Informs their Transferability
abstract
This paper focuses on model transferability estimation, i.e., assessing the performance of pre-trained models on a downstream task without performing fine-tuning. Motivated by neural collapse (NC) [25] that reveals particular feature geometry at the terminal stage of training, we consider model transferability as how far the target activations obtained by pre-trained models are from their hypothetical state in the terminal phase of the model fine-tuned on the target domain. We propose a metric that measures this proximity based on three phenomena of NC: within-class variability collapse, simplex encoded label interpolation geometry structure is formed, and the nearest center classifier becomes optimal on training data. Through experiments on 11 datasets, we confirm none of the three NC proxies are dispensable, which allows us to obtain very competitive transferability estimation accuracy with approximately 10× wall-clock time speed up compared to state-of-the-art approaches.
Zijian Wang 0009, Yadan Luo, Liang Zheng 0001, Zi Huang, Mahsa Baktash
ICCV1
2023 Learning Efficient Unsupervised Satellite Image-based Building Damage Detection
abstract
Existing Building Damage Detection (BDD) methods always require labour-intensive pixel-level annotations of buildings and their conditions, hence largely limiting their applications. In this paper, we investigate a challenging yet practical scenario of BDD, Unsupervised Building Damage Detection (U-BDD), where only unlabelled pre- and post-disaster satellite image pairs are provided. As a pilot study, we have first proposed an advanced U-BDD baseline that leverages pre-trained vision-language foundation models to address the U-BDD task. However, the apparent domain gap between satellite and generic images causes low confidence in the foundation models used to identify buildings and their damages. In response, we further present a novel self-supervised framework, U-BDD++, which improves upon the U-BDD baseline by addressing domain-specific issues associated with satellite imagery. Extensive experiments on the widely used building damage assessment benchmark demonstrate the effectiveness of the proposed method for unsupervised building damage detection. The presented annotation-free and foundation model-based paradigm ensures an efficient learning phase. This study opens a new direction for real-world BDD and sets a strong baseline for future research.
Zijian Wang 0009, Yadan Luo, Xin Yu 0002, Zi Huang
ICDM2
2023 Exploring Active 3D Object Detection from a Generalization Perspective
Yadan Luo, Zhuoxiao Chen, Zijian Wang 0009, Xin Yu 0002, Zi Huang, Mahsa Baktash
ICLR3
2023 VLM-BCD: Unsupervised Building Change Detection
abstract
Building Change Detection (BCD) is one of the most important parts of remote sensing analysis. However, most of the existing BCD approaches require a large amount of pixel-level annotation, which limits their applicability due to intensive labour costs. To alleviate this issue, we propose a vision-language model-based framework, VLM-BCD, which performs BCD tasks without requiring any labels. Specifically, the proposed framework consists of two stages: 1) Bi-temporal building localisation by leveraging open-vocabulary DETR. 2) Unchanged mask suppressing by the Change Resolver module to detect the building change in bi-temporal satellite images. An application with an interactive dashboard is implemented to maximise the usability of the developed framework.
Zijian Wang 0009
MMAsia2
2023 Center-aware Adversarial Augmentation for Single Domain Generalization
abstract
Domain generalization (DG) aims to learn a model from multiple training (i.e., source) domains that can generalize well to the unseen test (i.e., target) data coming from a different distribution. Single domain generalization (Single-DG) has recently emerged to tackle a more challenging, yet realistic setting, where only one source domain is available at training time. The existing Single-DG approaches typically are based on data augmentation strategies and aim to expand the span of source data by augmenting out-of-domain samples. Generally speaking, they aim to generate hard examples to confuse the classifier. While this may make the classifier robust to small perturbation, the generated samples are typically not diverse enough to mimic a large domain shift, resulting in sub-optimal generalization performance. To alleviate this, we propose a center-aware adversarial augmentation technique that expands the source distribution by altering the source samples so as to push them away from the class centers via a novel angular center loss. We conduct extensive experiments to demonstrate the effectiveness of our approach on several benchmark datasets for Single-DG and show that our method outperforms the state-of-the-art in most cases.
Mahsa Baktash, Zijian Wang 0009, Mathieu Salzmann
WACV3
2023 FFM: Injecting Out-of-Domain Knowledge via Factorized Frequency Modification
abstract
This work investigates the Single Domain Generalization (SDG) problem and aims to generalize a model from a single source (i.e., training) domain to multiple target (i.e., test) domains coming from different distributions. Most of the existing SDG approaches focus on generating out-of-domain samples by either transforming the source images into different styles or optimizing adversarial noise perturbations applied on the source images. In this paper, we show that generating images with diverse styles can be complementary to creating hard samples when handling the SDG task, and propose our approach of Factorized Frequency Modification (FFM) to fulfill this requirement. Specifically, we design a unified framework consisting of a style transformation module, an adversarial perturbation module, and a dynamic frequency selection module. We seamlessly equip the framework with iterative adversarial training that facilitates learning discriminative features from hard and diverse augmented samples. Extensive experiments are performed on four image recognition benchmark datasets of Digits, CIFAR-10-C, CIFAR-100-C, and PACS, which demonstrates that our method outperforms existing state-of-the-art approaches.
Zijian Wang 0009, Yadan Luo, Zi Huang, Mahsa Baktash
WACV1
2023 Source-Free Progressive Graph Learning for Open-Set Domain Adaptation
abstract
Open-set domain adaptation (OSDA) aims to transfer knowledge from a label-rich source domain to a label-scarce target domain while addressing disturbances from irrelevant target classes not present in the source data. However, most OSDA approaches are limited due to the lack of essential theoretical analysis of generalization bound, reliance on the coexistence of source and target data during adaptation, and failure to accurately estimate model predictions' uncertainty. To address these limitations, the Progressive Graph Learning (PGL) framework is proposed. PGL decomposes the target hypothesis space into shared and unknown subspaces and progressively pseudo-labels the most confident known samples from the target domain for hypothesis adaptation. PGL guarantees a tight upper bound of the target error by integrating a graph neural network with episodic training and leveraging adversarial learning to close the gap between the source and target distributions. The proposed approach also tackles a more realistic source-free open-set domain adaptation (SF-OSDA) setting that makes no assumptions about the coexistence of source and target domains. In a two-stage framework, the SF-PGL model' uniformly selects the most confident target instances from each category at a fixed ratio, and the confidence thresholds in each class weigh the classification loss in the adaptation step. The proposed methods are evaluated on benchmark image classification and action recognition datasets, where they demonstrate superiority and flexibility in recognizing both shared and unknown categories. Additionally, balanced pseudo-labeling plays a significant role in improving calibration, making the trained model less prone to over- or under-confident predictions on the target data.
Yadan Luo, Zijian Wang 0009, Zhuoxiao Chen, Zi Huang, Mahsa Baktash
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Robust Domain Correction Latent Subspace Learning for Gas Sensor Drift Compensation
abstract
Subspace learning is a popular machine learning method that has been frequently applied for gas sensor calibration; however, there are the following limitations in the latent subspace learning process: 1) the existence of data distribution differences is not considered and 2) ignoring the inherent information of the original space, such as discriminative information and structure information. To overcome these issues, we design a novel domain correction latent subspace learning (DCLSL) algorithm for gas sensor drift compensation by integrating subspace learning and domain adaptation into a unified framework in this study. First, domain correction is applied to alleviate the distribution difference before and after sensor drift by exploiting the mean distribution discrepancy criterion. Second, for better-local representation and classification results, we consider both discriminative information of the source data and structural information of the target data, revealing the intrinsic geometric structure in the latent subspace and forming a compact intraclass and interclass separated discriminative data layout. In addition, to improve the consistency of features and labels, we use a projection term, a reconstruction term, and a regularization term to simultaneously implement joint learning of the label space and the latent subspace and impose a row-sparsity constraint to enhance the model robustness to noise. Inspired by non-negative matrix factorization, we skillfully adopt multiplication update rules and thus solve the proposed optimization problem. Finally, we conduct experiments for different drift compensation tasks on two common gas sensor datasets, and the results are encouraging.
Danhong Yi, Linxia Zhang, Zijian Wang 0009, Lidan Wang 0001, Shukai Duan 0001, Jia Yan 0002
IEEE Trans. Syst. Man Cybern. Syst.3
2022 ShowFace: Coordinated Face Inpainting with Memory-Disentangled Refinement Networks
Zhuojie Wu, Xingqun Qi, Zijian Wang 0009, Kun Yuan 0003, Muyi Sun, Zhenan Sun
BMVC3
2022 Self-supervised Correlation Mining Network for Person Image Generation
abstract
Person image generation aims to perform non-rigid deformation on source images, which generally requires unaligned data pairs for training. Recently, self-supervised methods express great prospects in this task by merging the disentangled representations for self-reconstruction. However, such methods fail to exploit the spatial correlation between the disentangled features. In this paper, we propose a Self-supervised Correlation Mining Network (SCM-Net) to rearrange the source images in the feature space, in which two collaborative modules are integrated, Decomposed Style Encoder (DSE) and Correlation Mining Module (CMM). Specifically, the DSE first creates unaligned pairs at the feature level. Then, the CMM establishes the spatial correlation field for feature rearrangement. Eventually, a translation module transforms the rearranged features to realistic results. Meanwhile, for improving the fidelity of cross-scale pose transformation, we propose a graph based Body Structure Retaining Loss (BSR Loss) to preserve reasonable body structures on half body to full body generation. Extensive experiments conducted on DeepFashion dataset demonstrate the superiority of our method compared with other supervised and unsupervised approaches. Furthermore, satisfactory results on face generation show the versatility of our method in other deformation tasks.
Zijian Wang 0009, Xingqun Qi, Kun Yuan 0003, Muyi Sun
CVPR1
2022 Contrastive Learning for Representation Degeneration Problem in Sequential Recommendation
abstract
Recent advancements of sequential deep learning models such as Transformer and BERT have significantly facilitated the sequential recommendation. However, according to our study, the distribution of item embeddings generated by these models tends to degenerate into an anisotropic shape, which may result in high semantic similarities among embeddings. In this paper, both empirical and theoretical investigations of this representation degeneration problem are first provided, based on which a novel recommender model DuoRec is proposed to improve the item embeddings distribution. Specifically, in light of the uniformity property of contrastive learning, a contrastive regularization is designed for DuoRec to reshape the distribution of sequence representations. Given the convention that the recommendation task is performed by measuring the similarity between sequence representations and item embeddings in the same space via dot product, the regularization can be implicitly applied to the item embedding distribution. Existing contrastive learning methods mainly rely on data level augmentation for user-item interaction sequences through item cropping, masking, or reordering and can hardly provide semantically consistent augmentation samples. In DuoRec, a model-level augmentation is proposed based on Dropout to enable better semantic preserving. Furthermore, a novel sampling strategy is developed, where sequences having the same target item are chosen hard positive samples. Extensive experiments conducted on five datasets demonstrate the superior performance of the proposed DuoRec model compared with baseline methods. Visualization results of the learned representations validate that DuoRec can largely alleviate the representation degeneration problem.
Ruihong Qiu, Zi Huang, Hongzhi Yin, Zijian Wang 0009
WSDM4
2021 PAENet: A Progressive Attention-Enhanced Network for 3D to 2D Retinal Vessel Segmentation
abstract
3D to 2D retinal vessel segmentation is a challenging problem in Optical Coherence Tomography Angiography (OCTA) images. Accurate retinal vessel segmentation is important for the diagnosis and prevention of ophthalmic diseases. However, making full use of the 3D data of OCTA volumes is a vital factor for obtaining satisfactory segmentation results. In this paper, we propose a Progressive Attention-Enhanced Network (PAENet) based on attention mechanisms to extract rich feature representation. Specifically, the framework consists of two main parts, the three-dimensional feature learning path and the two-dimensional segmentation path. In the three-dimensional feature learning path, we design a novel Adaptive Pooling Module (APM) and propose a new Quadruple Attention Module (QAM). The APM captures dependencies along the projection direction of volumes and learns a series of pooling coefficients for feature fusion, which efficiently reduces feature dimension. In addition, the QAM reweights the features by capturing four-group cross-dimension dependencies, which makes maximum use of 4D feature tensors. In the two-dimensional segmentation path, to acquire more detailed information, we propose a Feature Fusion Module (FFM) to inject 3D information into the 2D path. Meanwhile, we adopt the Polarized Self-Attention (PSA) block to model the semantic interdependencies in spatial and channel dimensions respectively. Experimentally, our extensive experiments on the OCTA-500 dataset show that our proposed algorithm achieves state-of-the-art performance compared with previous methods.
Zhuojie Wu, Zijian Wang 0009, Wenxuan Zou, Fan Ji, Hao Dang, Muyi Sun
BIBM2
2021 CoCo DistillNet: a Cross-layer Correlation Distillation Network for Pathological Gastric Cancer Segmentation
abstract
In recent years, deep convolutional neural networks have made significant advances in pathology image segmentation. However, pathology image segmentation encounters with a dilemma in which the higher-performance networks generally require more computational resources and storage. This phenomenon limits the employment of high-accuracy networks in real scenes due to the inherent high-resolution of pathological images. To tackle this problem, we propose CoCo DistillNet, a novel Cross-layer Correlation (CoCo) knowledge distillation network for pathological gastric cancer segmentation. Knowledge distillation, a general technique which aims at improving the performance of a compact network through knowledge transfer from a cumbersome network. Concretely, our CoCo DistillNet models the correlations of channel-mixed spatial similarity between different layers and then transfers this knowledge from a pre-trained cumbersome teacher network to a non-trained compact student network. In addition, we also utilize the adversarial learning strategy to further prompt the distilling procedure which is called Adversarial Distillation (AD). Furthermore, to stabilize our training procedure, we make the use of the unsupervised Paraphraser Module (PM) to boost the knowledge paraphrase in the teacher network. As a result, extensive experiments conducted on the Gastric Cancer Segmentation Dataset demonstrate the prominent ability of CoCo DistillNet which achieves state-of-the-art performance.
Wenxuan Zou, Xingqun Qi, Zhuojie Wu, Zijian Wang 0009, Muyi Sun, Caifeng Shan
BIBM4
2021 Learning to Diversify for Single Domain Generalization
abstract
Domain generalization (DG) aims to generalize a model trained on multiple source (i.e., training) domains to a distributionally different target (i.e., test) domain. In contrast to the conventional DG that strictly requires the availability of multiple source domains, this paper considers a more realistic yet challenging scenario, namely Single Domain Generalization (Single-DG), where only one source domain is available for training. In this scenario, the limited diversity may jeopardize the model generalization on unseen target domains. To tackle this problem, we propose a style-complement module to enhance the generalization power of the model by synthesizing images from diverse distributions that are complementary to the source ones. More specifically, we adopt a tractable upper bound of mutual information (MI) between the generated and source samples and perform a two-step optimization iteratively: (1) by minimizing the MI upper bound approximation for each sample pair, the generated images are forced to be diversified from the source samples; (2) subsequently, we maximize the MI between the samples from the same semantic category, which assists the network to learn discriminative features from diverse-styled images. Extensive experiments on three benchmark datasets demonstrate the superiority of our approach, which surpasses the state-of-the-art single-DG methods by up to 25.14%. The code will be publicly available at https://github.com/BUserName/Learning_to_diversify
Zijian Wang 0009, Yadan Luo, Ruihong Qiu, Zi Huang, Mahsa Baktash
ICCV1
2021 RoadAtlas: Intelligent Platform for Automated Road Defect Detection and Asset Management
abstract
With the rapid development of intelligent detection algorithms based on deep learning, much progress has been made in automatic road defect recognition and road marking parsing. This can effectively address the issue of an expensive and time-consuming process for professional inspectors to review the street manually. Towards this goal, we present RoadAtlas, a novel end-to-end integrated system that can support 1) road defect detection, 2) road marking parsing, 3) a web-based dashboard for presenting and inputting data by users, and 4) a backend containing a well-structured database and developed APIs.
Zhuoxiao Chen, Yadan Luo, Zijian Wang 0009, Jinjiang Zhong, Anthony Southon
MMAsia4
2021 Deep Collaborative Discrete Hashing With Semantic-Invariant Structure Construction
abstract
While deep hashing has made great progress in large-scale multimedia retrieval, most of the existing approaches under-explore the semantic correlations and neglect the effect of context-aware visual learning. In this paper, we propose a dual-stream learning framework, termed as Deep Collaborative Discrete Hashing (DCDH), which constructs a discriminative common discrete space by collaboratively incorporating the shared and individual semantics deduced from visual features and semantics. Specifically, DCDH generates context-aware representations by employing the outer product of visual embeddings and semantic encodings. To further preserve the original semantics and alleviate the class imbalance problem, we introduce the focal loss to take advantage of frequent and rare concepts. Furthermore, a common binary code space is constructed based on the joint learning of the visual representations, the context-aware representations, and the label distribution calibration. Three losses, i.e., the pairwise similarity loss, the quantization loss, and the balanced classification loss, are collaboratively optimized in the general learning framework of DCDH. Extensive experiments conducted on three large-scale benchmark datasets demonstrate the superiority of the proposed method, yielding the state-of-the-art image retrieval performance.
Zijian Wang 0009, Zheng Zhang 0006, Yadan Luo, Zi Huang, Heng Tao Shen
IEEE Trans. Multim.1
2020 Progressive Graph Learning for Open-Set Domain Adaptation
abstract
Domain shift is a fundamental problem in visual recognition which typically arises when the source and target data follow different distributions. The existing domain adaptation approaches which tackle this problem work in the "closed-set" setting with the assumption that the source and the target data share exactly the same classes of objects. In this paper, we tackle a more realistic problem of the "open-set" domain shift where the target data contains additional classes that were not present in the source data. More specifically, we introduce an end-to-end Progressive Graph Learning (PGL) framework where a graph neural network with episodic training is integrated to suppress underlying conditional shift and adversarial learning is adopted to close the gap between the source and target distributions. Compared to the existing open-set adaptation approaches, our approach guarantees to achieve a tighter upper bound of the target error. Extensive experiments on three standard open-set benchmarks evidence that our approach significantly outperforms the state-of-the-arts in open-set domain adaptation.
Yadan Luo, Zijian Wang 0009, Zi Huang, Mahsa Baktash
ICML2
2020 Adversarial Bipartite Graph Learning for Video Domain Adaptation
abstract
Domain adaptation techniques, which focus on adapting models between distributionally different domains, are rarely explored in the video recognition area due to the significant spatial and temporal shifts across the source (i.e. training) and target (i.e. test) domains. As such, recent works on visual domain adaptation which leverage adversarial learning to unify the source and target video representations and strengthen the feature transferability are not highly effective on the videos. To overcome this limitation, in this paper, we learn a domain-agnostic video classifier instead of learning domain-invariant representations, and propose an Adversarial Bipartite Graph (ABG) learning framework which directly models the source-target interactions with a network topology of the bipartite graph. Specifically, the source and target frames are sampled as heterogeneous vertexes while the edges connecting two types of nodes measure the affinity among them. Through message-passing, each vertex aggregates the features from its heterogeneous neighbors, forcing the features coming from the same class to be mixed evenly. Explicitly exposing the video classifier to such cross-domain representations at the training and test stages makes our model less biased to the labeled source data, which in-turn results in achieving a better generalization on the target domain. The proposed framework is agnostic to the choices of frame aggregation, and therefore, four different aggregation functions are investigated for capturing appearance and temporal dynamics. To further enhance the model capacity and testify the robustness of the proposed architecture on difficult transfer tasks, we extend our model to work in a semi-supervised setting using an additional video-level bipartite graph. Extensive experiments conducted on four benchmark datasets evidence the effectiveness of the proposed approach over the state-of-the-art methods on the task of video recognition.
Yadan Luo, Zi Huang, Zijian Wang 0009, Zheng Zhang 0006, Mahsa Baktash
ACM Multimedia3
2020 Prototype-Matching Graph Network for Heterogeneous Domain Adaptation
abstract
Even though the multimedia data is ubiquitous on the web, the scarcity of the annotated data and variety of data modalities hinder their usage by multimedia applications. Heterogeneous domain adaptation (HDA) has therefore arisen to address such limitations by facilitating the knowledge transfer between heterogeneous domains. Existing HDA methods only focus on aligning the cross-domain feature distributions and ignore the importance of maximizing the margin among different classes, which may lead to a sub-optimal classification performance. To tackle this problem, in this paper, we propose the Prototype-Matching Graph Network (PMGN), which gradually explores the domain-invariant class prototype representations. Specifically, we build an end-to-end Graph Prototypical Network, which computes the class prototypes through multiple layers of edge learning, node aggregation, and discrepancy minimization. Our framework utilizes the Swap training strategy to provide adequate supervision for training the edge learning component. Moreover, the proposed PMGN can be equipped with the clustering module that utilises the KL-divergence as a distance metric to reduce the distribution difference between the source and target data. Extensive experiments on three HDA tasks (i.e. object recognition, text-to-image classification, and text categorization) demonstrate the superiority of our approach over the state-of-the-art HDA methods.
Zijian Wang 0009, Yadan Luo, Zi Huang, Mahsa Baktash
ACM Multimedia1
2019 Deep Collaborative Discrete Hashing with Semantic-Invariant Structure
abstract
Existing deep hashing approaches fail to fully explore semantic correlations and neglect the effect of linguistic context on visual attention learning, leading to inferior performance. This paper proposes a dual-stream learning framework, dubbed Deep Collaborative Discrete Hashing (DCDH), which constructs a discriminative common discrete space by collaboratively incorporating the shared and individual semantics deduced from visual features and semantic labels. Specifically, the context-aware representations are generated by employing the outer product of visual embeddings and semantic encodings. Moreover, we reconstruct the labels and introduce the focal loss to take advantage of frequent and rare concepts. The common binary code space is built on the joint learning of the visual representations attended by language, the semantic-invariant structure construction and the label distribution correction. Extensive experiments demonstrate the superiority of our method.
Zijian Wang 0009, Zheng Zhang 0006, Yadan Luo, Zi Huang
SIGIR1