Jianyang Gu

dblp:241/7332 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-4060-7427ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 8 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Concord: Concept-Informed Diffusion for Dataset Distillation
abstract
Dataset distillation (DD) has witnessed significant progress in creating small datasets that encapsulate rich information from large original ones. Particularly, methods based on generative priors show promising performance while maintaining computational efficiency and cross-architecture generalization. However, the generation process lacks explicit controllability for each sample. Previous distillation methods primarily match the real distribution from the perspective of the entire dataset, whereas overlooking concept completeness at the instance level. The missing or incorrectly represented object details cannot be efficiently compensated due to the constrained sample amount typical in DD settings. To this end, we propose incorporating the concept understanding of large language models (LLMs) to perform Concept-Informed Diffusion (Concord) for dataset distillation. Specifically, distinguishable and fine-grained concepts are retrieved based on category labels to inform the denoising process and refine essential object details. These concepts can be applied to any diffusion-based DD framework to enhance both the controllability and interpretability of the distilled image generation, without relying on pre-trained classifiers. We demonstrate the efficacy of Concord by achieving state-of-the-art performance on ImageNet-1K and subsets. Code is released at Concord.
Jianyang Gu, Ruoxi Jia 0001, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001
WACV1
2026 Dataset distillation with pre-trained models: A contrastive approach
Yao Lu 0041, Xuguang Chen, Jianyang Gu, Qi Xuan 0001, Zhaowei Zhu
Neurocomputing3
2026 Stochastic style perturbation modelling for visible-Infrared person re-Identification with severely modality imbalance
Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009
Neural Networks3
2025 Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis
abstract
We present a simple approach to make pre-trained Vision Transformers (ViTs) interpretable for fine-grained analysis, aiming to identify and localize the traits that distinguish visually similar categories, such as bird species. Pretrained ViTs, such as DINO, have demonstrated remarkable capabilities in extracting localized, discriminative features. However, saliency maps like Grad-CAM often fail to identify these traits, producing blurred, coarse heatmaps that highlight entire objects instead. We propose a novel approach, Prompt Class Attention Map (Prompt-CAM), to address this limitation. Prompt-CAM learns class-specific prompts for a pre-trained ViT and uses the corresponding outputs for classification. To correctly classify an image, the true-class prompt must attend to unique image patches not present in other classes’ images (i.e., traits). As a result, the true class’s multi-head attention maps reveal traits and their locations. Implementation-wise, Prompt-CAM is almost a "free lunch," requiring only a modification to the prediction head of Visual Prompt Tuning (VPT). This makes Prompt-CAM easy to train and apply, in stark contrast to other interpretable methods that require designing specific models and training processes. Extensive empirical studies on a dozen datasets from various domains (e.g., birds, fishes, insects, fungi, flowers, food, and cars) validate the superior interpretation capability of Prompt-CAM. The source code and demo are available at https://github.com/Imageomics/Prompt_CAM.
Arpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu, Kazi Sajeed Mehrab, Elizabeth G. Campolongo, Daniel I. Rubenstein, Charles V. Stewart, Anuj Karpatne, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
CVPR4
2025 Finer-CAM: Spotting the Difference Reveals Finer Details for Visual Explanation
abstract
Class activation map (CAM) has been widely used to highlight image regions that contribute to class predictions. Despite its simplicity and computational efficiency, CAM often struggles to identify discriminative regions that distinguish visually similar fine-grained classes. Prior efforts address this limitation by introducing more sophisticated explanation processes, but at the cost of extra complexity. In this paper, we propose Finer-CAM, a method that retains CAM’s efficiency while achieving precise localization of discriminative regions. Our key insight is that the deficiency of CAM lies not in "how" it explains, but in "what" it explains. Specifically, previous methods attempt to identify all cues contributing to the target class’s logit value, which inadvertently also activates regions predictive of visually similar classes. By explicitly comparing the target class with similar classes and spotting their differences, Finer-CAM suppresses features shared with other classes and emphasizes the unique, discriminative details of the target class. Finer-CAM is easy to implement, compatible with various CAM methods, and can be extended to multi-modal models for accurate localization of specific concepts. Additionally, Finer-CAM allows adjustable comparison strength, enabling users to selectively highlight coarse object contours or fine discriminative details. Quantitatively, we show that masking out the top 5% of activated pixels by Finer-CAM results in a larger relative confidence drop compared to baselines. The source code and demo are available at https://github.com/Imageomics/Finer-CAM.
Jianyang Gu, Arpita Chowdhury, Zheda Mai, David Carlyn, Tanya Y. Berger-Wolf, Yu Su 0001, Wei-Lun Chao
CVPR2
2025 Group Distributionally Robust Dataset Distillation with Risk Minimization
abstract
Dataset distillation (DD) has emerged as a widely adopted technique for crafting a synthetic dataset that captures the essential information of a training dataset, facilitating the training of accurate neural models. Its applications span various domains, including transfer learning, federated learning, and neural architecture search. The most popular methods for constructing the synthetic data rely on matching the convergence properties of training the model with the synthetic dataset and the training dataset. However, using the empirical loss as the criterion must be thought of as auxiliary in the same sense that the training set is an approximate substitute for the population distribution, and the latter is the data of interest. Yet despite its popularity, an aspect that remains unexplored is the relationship of DD to its generalization, particularly across uncommon subgroups. That is, how can we ensure that a model trained on the synthetic dataset performs well when faced with samples from regions with low population density? Here, the representativeness and coverage of the dataset become salient over the guaranteed training error at inference. Drawing inspiration from distributionally robust optimization, we introduce an algorithm that combines clustering with the minimization of a risk measure on the loss to conduct DD. We provide a theoretical rationale for our approach and demonstrate its effective generalization and robustness across subgroups through numerical experiments.
Saeed Vahidian, Mingyu Wang 0004, Jianyang Gu, Vyacheslav Kungurtsev, Wei Jiang 0009, Yiran Chen 0001
ICLR3
2025 Taming Diffusion for Dataset Distillation with High Representativeness
abstract
Recent deep learning models demand larger datasets, driving the need for dataset distillation to create compact, cost-efficient datasets while maintaining performance. Due to the powerful image generation capability of diffusion, it has been introduced to this field for generating distilled images. In this paper, we systematically investigate issues present in current diffusion-based dataset distillation methods, including inaccurate distribution matching, distribution deviation with random noise, and separate sampling. Building on this, we propose D$^3$HR, a novel diffusion-based framework to generate distilled datasets with high representativeness. Specifically, we adopt DDIM inversion to map the latents of the full dataset from a low-normality latent domain to a high-normality Gaussian domain, preserving information and ensuring structural consistency to generate representative latents for the distilled dataset. Furthermore, we propose an efficient sampling scheme to better align the representative latents with the high-normality Gaussian distribution. Our comprehensive experiments demonstrate that D$^3$HR can achieve higher accuracy across different model architectures compared with state-of-the-art baselines in dataset distillation. Source code: https://github.com/lin-zhao-resoLve/D3HR.
Yushu Wu, Xinru Jiang, Jianyang Gu, Yanzhi Wang 0001, Xiaolin Xu 0001, Pu Zhao 0001, Xue Lin 0001
ICML4
2025 BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning
abstract
Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space.
Jianyang Gu, Samuel Stevens 0001, Elizabeth G. Campolongo, Matthew J. Thompson, Net Zhang, Jiaman Wu, Andrei Kopanev, Zheda Mai, Alexander E. White, James P. Balhoff, Wasila M. Dahdul, Daniel I. Rubenstein, Hilmar Lapp, Tanya Y. Berger-Wolf, Wei-Lun Chao, Yu Su 0001
NeurIPS1
2025 CoMix: Collaborative Mixed Learning via Style Fuzzy Normalization for Visible-Infrared Person Re-Identification
abstract
Visible–infrared person re-identification (VI-ReID) focuses on accurately matching individuals across different imaging modalities. Existing studies focus on generating modality-consistent images at the pixel level through the use of generative adversarial networks (GANs) to mitigate the impact of modality discrepancies. However, these methods face significant challenges in overcoming the limitation that synthesized samples from different modalities may suffer from semantic distortion. In this work, we propose an online one-stage style fuzzy normalization (SFN) method to generate modality-fuzzy features in the latent space while regularizing the model’s predictions. Specifically, SFN adaptively mixes the feature statistics of two random modality instances of the same identity in a single forward pass during training. In this process, to enhance the richness of modality interaction information, we design a novel causality balance loss, which enforces the generated fuzzy features to be independent of their initial modality while simultaneously encouraging them to align more closely with the other modality. Furthermore, we introduce an identity-aware consistency loss to regularize the predictions between the original and SFN-generated features to ensure semantic consistency. In contrast to prior work, SFN is a plug-and-play module that does not rely on any generative-based models, making it highly adaptable to various network architectures. Extensive experiments were performed on three public cross-modality datasets to ensure fair and reliable comparisons. The empirical results demonstrate the clear superiority of our method over previous state-of-the-art methods.
Jianyang Gu, Mingyu Wang 0004, Q. M. Jonathan Wu, Wei Jiang 0009
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Summarizing Stream Data for Memory-Constrained Online Continual Learning
abstract
Replay-based methods have proved their effectiveness on online continual learning by rehearsing past samples from an auxiliary memory. With many efforts made on improving training schemes based on the memory, however, the information carried by each sample in the memory remains under-investigated. Under circumstances with restricted storage space, the informativeness of the memory becomes critical for effective replay. Although some works design specific strategies to select representative samples, by only employing a small number of original images, the storage space is still not well utilized. To this end, we propose to Summarize the knowledge from the Stream Data (SSD) into more informative samples by distilling the training characteristics of real images. Through maintaining the consistency of training gradients and relationship to the past tasks, the summarized samples are more representative for the stream data compared to the original images. Extensive experiments are conducted on multiple online continual learning benchmarks to support that the proposed SSD method significantly enhances the replay effects. We demonstrate that with limited extra computational overhead, SSD provides more than 3% accuracy boost for sequential CIFAR-100 under extremely restricted memory buffer. Code in https://github.com/vimar-gu/SSD.
Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001
AAAI1
2024 Efficient Dataset Distillation via Minimax Diffusion
abstract
Dataset distillation reduces the storage and computational consumption of training a network by generating a small surrogate dataset that encapsulates rich information of the original large-scale one. However, previous distillation methods heavily rely on the sample-wise iterative optimization scheme. As the images-per-class (IPC) setting or image resolution grows larger, the necessary computation will demand overwhelming time and resources. In this work, we intend to incorporate generative diffusion techniques for computing the surrogate dataset. Observing that key factors for constructing an effective surrogate dataset are representativeness and diversity, we design additional minimax criteria in the generative training to enhance these facets for the generated images of diffusion models. We present a theoretical model of the process as hierarchical diffusion control demonstrating the flexibility of the diffusion process to target these criteria without jeopardizing the faithfulness of the sample to the desired distribution. The proposed method achieves state-of-the-art validation performance while demanding much less computational resources. Under the 100-IPC setting on Image Woof, our method requires less than one-twentieth the distillation time of previous methods, yet yields even better performance. Source code and generated data are available in https://github.com/vimar-gu/MinimaxDiffusion.
Jianyang Gu, Saeed Vahidian, Vyacheslav Kungurtsev, Wei Jiang 0009, Yang You 0001, Yiran Chen 0001
CVPR1
2024 InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data Pruning
abstract
Data pruning aims to obtain lossless performances with less overall cost. A common approach is to filter out samples that make less contribution to the training. This could lead to gradient expectation bias compared to the original data. To solve this problem, we propose InfoBatch, a novel framework aiming to achieve lossless training acceleration by unbiased dynamic data pruning. Specifically, InfoBatch randomly prunes a portion of less informative samples based on the loss distribution and rescales the gradients of the remaining samples to approximate the original gradient. As a plug-and-play and architecture-agnostic framework, InfoBatch consistently obtains lossless training results on classification, semantic segmentation, vision pertaining, and instruction fine-tuning tasks. On CIFAR10/100, ImageNet- 1K, and ADE20K, InfoBatch losslessly saves 40% overall cost. For pertaining MAE and diffusion model, InfoBatch can respectively save 24.8% and 27% cost. For LLaMA instruction fine-tuning, combining InfoBatch and the recent coreset selection method (DQ) can achieve 10 times acceleration. Our results encourage more exploration on the data efficiency aspect of large model training. Code is publicly available at NUS-HPC-AI-Lab/InfoBatch.
Ziheng Qin, Kai Wang 0036, Zangwei Zheng, Jianyang Gu, Zhaopan Xu, Daquan Zhou, Baigui Sun, Xuansong Xie, Yang You 0001
ICLR4
2024 Dynamic gradient reactivation for backward compatible person re-identification
Xiao Pan 0001, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li
Pattern Recognit.8
2023 MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReID
abstract
Neural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet.
Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006
CVPR1
2023 DREAM: Efficient Dataset Distillation by Representative Matching
abstract
Dataset distillation aims to synthesize small datasets with little information loss from original large-scale ones for reducing storage and training costs. Recent state-of-the-art methods mainly constrain the sample synthesis process by matching synthetic images and the original ones regarding gradients, embedding distributions, or training trajectories. Although there are various matching objectives, currently the strategy for selecting original images is limited to naive random sampling. We argue that random sampling overlooks the evenness of the selected sample distribution, which may result in noisy or biased matching targets. Besides, the sample diversity is also not constrained by random sampling. These factors together lead to optimization instability in the distilling process and degrade the training efficiency. Accordingly, we propose a novel matching strategy named as Dataset distillation by REpresentAtive Matching (DREAM), where only representative original images are selected for matching. DREAM is able to be easily plugged into popular dataset distillation frameworks and reduce the distilling iterations by more than 8 times without performance drop. Given sufficient training time, DREAM further provides significant improvements and achieves state-of-the-art performances.
Jianyang Gu, Kai Wang 0036, Wei Jiang 0009, Yang You 0001
ICCV2
2023 Dataset Quantization
abstract
State-of-the-art deep neural networks are trained with large amounts (millions or even billions) of data. The expensive computation and memory costs make it difficult to train them on limited hardware resources, especially for recent popular large language models (LLM) and computer vision models (CV). Recent popular dataset distillation methods are thus developed, aiming to reduce the number of training samples via synthesizing small-scale datasets via gradient matching. However, as the gradient calculation is coupled with the specific network architecture, the synthesized dataset is biased and performs poorly when used for training unseen architectures. To address these limitations, we present dataset quantization (DQ), a new framework to compress large-scale datasets into small subsets which can be used for training any neural network architectures. Extensive experiments demonstrate that DQ is able to generate condensed small datasets for training unseen network architectures with state-of-the-art compression ratios for lossless model training. To the best of our knowledge, DQ is the first method that can successfully distill large-scale datasets such as ImageNet-1k with a state-of-the-art compression ratio. Notably, with 60% data from ImageNet and 20% data from Alpaca’s instruction tuning data, the models can be trained with negligible or no performance drop for both vision tasks (including classification, semantic segmentation, and object detection) as well as language tasks (including instruction tuning tasks such as BBH and DROP).
Daquan Zhou, Kai Wang 0036, Jianyang Gu, Dongze Lian, Yifan Zhang 0004, Yang You 0001, Jiashi Feng
ICCV3
2023 3D-Guided Frontal Face Generation for Pose-Invariant Recognition
abstract
Although deep learning techniques have achieved extraordinary accuracy in recognizing human faces, the pose variances of images captured in real-world scenarios still hinder reliable model appliance. To mitigate this gap, we propose to recognize faces via generation frontal face images with a 3D -Guided Deep P ose- I nvariant Face Recognition M odel (3D-PIM) consisted of a simulator and a refiner module. The simulator employs a 3D Morphable Model (3D MM) to fit the shape and appearance features and recover primary frontal images with less training data. The refiner further enhances the image realism on both global facial structure and local details with adversarial training, while keeping the discriminative identity information consistent with original images. An Adaptive Weighting (AW) metric is then adopted to leverage the complimentary information from recovered frontal faces and original profile faces and to obtain credible similarity scores for recognition. Extended experiments verify the superiority of the proposed “recognition via generation” framework over state-of-the-art.
Hao Wu 0098, Jianyang Gu, Xiaojin Fan, He Li 0034, Lidong Xie, Jian Zhao 0006
ACM Trans. Intell. Syst. Technol.2
2023 Transformer-Based Domain-Specific Representation for Unsupervised Domain Adaptive Vehicle Re-Identification
abstract
Fully-supervised vehicle re-identification (re-ID) methods are faced with performance degradation when applied to new image domains. Therefore, developing unsupervised domain adaptation (UDA) to transfer the knowledge from learned source domain to new unlabeled target domain becomes an indispensable task. It is challenging because different domains have various image appearances, such as different backgrounds, illuminations and resolutions, especially when cameras have different viewpoints. To tackle this domain gap issue, a novel Transformer-based Domain-Specific Representation learning network (TDSR) is proposed to dynamically focus on corresponding detailed hints for each domain. Specifically, with the source and target domain being trained simultaneously, a domain encoding module is proposed to introduce domain information into the network. The original features of source and target domains are enriched with these domain encodings first, and then sequentially processed by a Transformer encoder to model contextual information and a decoder to summarize the encoded features into the final domain-specific feature representations. Moreover, we propose a Contrastive Clustering Loss (CCL) to directly optimize the distribution of features at cluster level. Instances are overall pulled closer to the prototype of the same identity, and pushed farther from the prototypes of different identities. It helps compact the clusters in the latent space and improve the discriminative capability of the network, leading to more accurate pseudo-label assignment in TDSR. Our method outperforms the state-of-the-art UDA methods on vehicle re-ID benchmark datasets VeRi and VehicleID on both real-world to real-world and synthetic to real-world settings.
Jianyang Gu, Shuting He, Wei Jiang 0009
IEEE Trans. Intell. Transp. Syst.2
2022 SFGN: Representing the sequence with one super frame for video person re-identification
Xiao Pan 0001, Hao Luo 0004, Wei Jiang 0009, Jianming Zhang 0005, Jianyang Gu, Peike Li
Knowl. Based Syst.5
2022 Multi-View Evolutionary Training for Unsupervised Domain Adaptive Person Re-Identification
abstract
Clustering-based approaches have been successfully applied to unsupervised domain adaptation (UDA) tasks for person re-identification (Re-ID), where no annotations are provided in target domain. However, the clustering process is sensitive to noises, leading to imperfect pseudo labels that could damage the training performance. In this work, we propose a Multi-view Evolutionary Training (MET) method to effectively reduce noises in clustering results from two dimensions. First, to improve the clustering accuracy at each time frame (i.e. snapshot quality), a Multi-view Diffusion (MvD) module is proposed. Through capturing data relationships from multiple viewpoints and aggregating their information, noises and bias from each individual viewpoint can be eliminated, and more reliable similarity matrix can be produced for clustering. Second, to improve the temporal consistency between clustering at different iterations, i.e. temporal consistency, we propose an Evolutionary Local Refinement (ELR) module, which utilizes the previous clustering results to guide and improve current results, and further make the training process more stable and robust. Extensive experiments demonstrate that our method can provide clustering results with high quality, and achieve state-of-the-art performance on UDA Re-ID.
Jianyang Gu, Hao Luo 0004, Fan Wang 0019, Hao Li 0030, Wei Jiang 0009
IEEE Trans. Inf. Forensics Secur.1
2021 An efficient global representation constrained by Angular Triplet loss for vehicle re-identification
Jianyang Gu, Wei Jiang 0009, Hao Luo 0004, Hongyan Yu
Pattern Anal. Appl.1
2020 A Strong Baseline and Batch Normalization Neck for Deep Person Re-Identification
abstract
This study proposes a simple but strong baseline for deep person re-identification (ReID). Deep person ReID has achieved great progress and high performance in recent years. However, many state-of-the-art methods design complex network structures and concatenate multi-branch features. In the literature, some effective training tricks briefly appear in several papers or source codes. The present study collects and evaluates these effective training tricks in person ReID. By combining these tricks, the model achieves 94.5% rank-1 and 85.9% mean average precision on Market1501 with only using the global features of ResNet50. The performance surpasses all existing global- and part-based baselines in person ReID. We propose a novel neck structure named as batch normalization neck (BNNeck). BNNeck adds a batch normalization layer after global pooling layer to separate metric and classification losses into two different feature spaces because we observe they are inconsistent in one embedding space. Extended experiments show that BNNeck can boost the baseline, and our baseline can improve the performance of existing state-of-the-art methods. Our codes and models are available at: https://github.com/michuanhaohao/reid-strong-baseline.
Hao Luo 0004, Wei Jiang 0009, Youzhi Gu, Fuxu Liu, Xingyu Liao, Shenqi Lai, Jianyang Gu
IEEE Trans. Multim.7
2018 RoboCup SSL 2018 Champion Team Paper
Zheyuan Huang, Yunkai Wang, Zexi Chen, Licheng Wen, Jianyang Gu, Rong Xiong
RoboCup7