Kuan Zhu

dblp:244/3040 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0002-1670-944XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Generalization in LLM Structured Pruning via Function-Aware Neuron Grouping
Tao Yu 0013, Yongqi An, Kuan Zhu, Guibo Zhu, Ming Tang 0001, Jinqiao Wang
AAAI3
2026 Continual Instruction Tuning for Large Multimodal Models
abstract
Instruction tuning has become a widely adopted approach for aligning large multimodal models (LMMs) with human intent. It enables multi-task joint training through unified data formats. However, as new vision-language tasks constantly emerge, exhaustive joint training of all tasks becomes impractical. Continual learning offers a more flexible and resource-efficient alternative, enabling incremental training of LMMs on emerging tasks. This study investigates two fundamental questions when applying continual learning to instruction tuning of LMMs: 1) Do LMMs suffer from catastrophic forgetting during continual instruction tuning? 2) Can existing continual learning methods be effectively applied to continual instruction tuning of LMMs? A comprehensive study was conducted to answer these questions. First, we establish the first benchmark for continual instruction tuning of LMMs and reveal the phenomenon of catastrophic forgetting in this setup. Second, we integrate and adapt traditional continual learning approaches to this setting, demonstrating the effectiveness of these strategies to varying degrees in different scenarios. Third, we explore task-similarity dynamics between pairs of vision-language tasks and propose task-similarity-informed regularization and model expansion methods. Experimental results show that our approach can consistently boost the model's performance.
Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Image Process.3
2025 Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
abstract
Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang, Tat-Seng Chua, Jinqiao Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jinghan He, Kuan Zhu, Haiyun Guo, Junfeng Fang, Zhenglin Hua, Yuheng Jia, Ming Tang 0001, Tat-Seng Chua, Jinqiao Wang
ACL (1)2
2025 FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
abstract
Pedestrian attribute recognition (PAR) is a fundamental perception task in intelligent transportation and security. To tackle this fine-grained task, most existing methods focus on extracting regional features to enrich attribute information. However, a regional feature is typically used to predict a fixed set of pre-defined attributes in these methods, which limits the performance and practicality in two aspects: 1) Regional features may compromise fine-grained patterns unique to certain attributes in favor of capturing common characteristics shared across attributes. 2) Regional features cannot generalize to predict unseen attributes in the test time. In this paper, we propose the Fine-grained Optimization with semantiC gUided underStanding (FOCUS) approach for PAR, which adaptively extracts fine-grained attribute-level features for each attribute individually, regardless of whether the attributes are seen or not during training. Specifically, we propose the Multi-Granularity Mix Tokens (MGMT) to capture latent features at varying levels of visual granularity, thereby enriching the diversity of the extracted information. Next, we introduce the Attribute-guided Visual Feature Extraction (AVFE) module, which leverages textual attributes as queries to retrieve their corresponding visual attribute features from the Mix Tokens using a cross-attention mechanism. To ensure that textual attributes focus on the appropriate Mix Tokens, we further incorporate a Region-Aware Contrastive Learning (RACL) method, encouraging attributes within the same region to share consistent attention maps. Extensive experiments on PA100K, PETA, and RAPv1 datasets demonstrate the effectiveness and strong generalization ability of our method.
Hongyan An, Kuan Zhu, Haiyun Guo, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
ICME2
2025 Semantic-aware Fine-grained Point Augmentation for 3D Multi-modal Object Detection
abstract
3D object detection aims to locate and recognize the object from the point cloud, which is a meaningful and foundation task in autonomous driving. However, the sparsity of the point cloud poses a significant challenge for this task, especially for distant and small objects. Existing methods employ depth estimation networks to generate pseudo points for improving the point density, but this introduces significant computational costs and noise, limiting performance gains. In this paper, we propose the Semantic-aware Fine-grained Point Augmentation (SFPA) approach for 3D object detection, which simultaneously enriches high-quality point clouds and filters noisy points, and incorporates multi-modal feature fusion to enhance detection performance. Specifically, we utilize a semantic segmentation model to generate object masks from RGB images and refine dense depth estimation maps, derived from sparse LiDAR points and RGB images, using these foreground object masks. Subsequently, high-quality pseudo point clouds, concentrated solely on foreground objects, are generated by projecting the refined dense depth maps back to 3D coordinates. Furthermore, we also employ the projection matrix as an alignment strategy to concatenate or add dense RGB features with point features, further improving detection performance for extremely sparse objects. Experimental results demonstrate that our method achieves state-of-the-art performance on KITTI 3D object detection leaderboard, i.e., 95.44%, 88.18%, 85.53% for the Car category at the easy, medium, and hard levels, respectively.
Wei Li 0315, Kuan Zhu, Haiyun Guo, Honghui Dong, Jinqiao Wang
ICME2
2025 Referring Expression Instance Retrieval and A Strong End-to-End Baseline
abstract
Text-Image Retrieval (TIR) retrieves a target image from a gallery based on an image-level description, while Referring Expression Comprehension (REC) localizes a target object within a given image using an instance-level description. However, real-world applications often present more complex demands. Users typically query an instance-level description across a large gallery and expect to receive both relevant image and the corresponding instance location. In such scenarios, TIR struggles with fine-grained descriptions and object-level localization, while REC is limited in its ability to efficiently search large galleries and lacks an effective ranking mechanism. In this paper, we introduce a new task called Referring Expression Instance Retrieval (REIR), which supports both instance-level retrieval and localization based on fine-grained referring expressions. First, we propose a large-scale benchmark for REIR, named REIRCOCO, constructed by prompting advanced vision-language models to generate high quality referring expressions for instances in the MSCOCO and RefCOCO datasets. Second, we present a baseline method, Contrastive Language Instance Alignment with Relation Experts (CLARE), which employs a dual-stream architecture to address REIR in an end-to-end manner. Given a referring expression, the textual branch encodes it into a query embedding, enhanced by a Mix of Relation Experts (MORE) module designed to better capture inter-instance relationships. The visual branch detects candidate objects and extracts their instance-level visual features. The most similar candidate to the query is selected for bounding box prediction. CLARE is first trained on object detection and REC datasets to establish initial grounding capabilities, then optimized via Contrastive Language Instance Alignment (CLIA) for improved retrieval across images. Experimental results demonstrate that CLARE outperforms existing methods on the REIR benchmark and generalizes well to both TIR and REC tasks, showcasing its effectiveness and versatility.
Xiangzhao Hao, Kuan Zhu, Haiyun Guo, Ming Tang 0001, Jinqiao Wang
ACM Multimedia2
2024 SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
abstract
Continual learning (CL) is crucial for language models to dynamically adapt to the evolving real-world demands.To mitigate the catastrophic forgetting problem in CL, data replay has been proven a simple and effective strategy, and the subsequent data-replay-based distillation can further enhance the performance.However, existing methods fail to fully exploit the knowledge embedded in models from previous tasks, resulting in the need for a relatively large number of replay samples to achieve good results.In this work, we first explore and emphasize the importance of attention weights in knowledge retention, and then propose a SElective attEntion-guided Knowledge Retention method (SEEKR) for data-efficient replay-based continual learning of large language models (LLMs).Specifically, SEEKR performs attention distillation on the selected attention heads for finer-grained knowledge retention, where the proposed forgettabilitybased and task-sensitivity-based measures are used to identify the most valuable attention heads.Experimental results on two continual learning benchmarks for LLMs demonstrate the superiority of SEEKR over the existing methods on both performance and efficiency.Explicitly, SEEKR achieves comparable or even better performance with only 1/10 of the replayed data used by other methods, and reduces the proportion of replayed data to 1%.The code is available at https: //github.com/jinghan1he/SEEKR.
Jinghan He, Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang
EMNLP3
2024 AAformer: Auto-Aligned Transformer for Person Re-Identification
abstract
In person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)," which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods.
Kuan Zhu, Haiyun Guo, Shiliang Zhang, Yaowei Wang 0001, Jing Liu 0001, Jinqiao Wang, Ming Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Learning Semantics-Consistent Stripes With Self-Refinement for Person Re-Identification
abstract
Aligning human parts automatically is one of the most challenging problems for person re-identification (re-ID). Recently, the stripe-based methods, which equally partition the person images into the fixed stripes for aligned representation learning, have achieved great success. However, the stripes with fixed height and position cannot well handle the misalignment problems caused by inaccurate detection and occlusion and may introduce much background noise. In this article, we aim at learning adaptive stripes with foreground refinement to achieve pixel-level part alignment by only using person identity labels for person re-ID and make two contributions. 1) A semantics-consistent stripe learning method (SCS). Given an image, SCS partitions it into adaptive horizontal stripes and each stripe is corresponding to a specific semantic part. Specifically, SCS iterates between two processes: i) clustering the rows to human parts or background to generate the pseudo-part labels of rows and ii) learning a row classifier to partition a person image, which is supervised by the latest pseudo-labels. This iterative scheme guarantees the accuracy of the learned image partition. 2) A self-refinement method (SCS+) to remove the background noise in stripes. We employ the above row classifier to generate the probabilities of pixels belonging to human parts (foreground) or background, which is called the class activation map (CAM). Only the most confident areas from the CAM are assigned with foreground/background labels to guide the human part refinement. Finally, by intersecting the semantics-consistent stripes with the foreground areas, SCS+ locates the human parts at pixel-level, obtaining a more robust part-aligned representation. Extensive experiments validate that SCS+ sets the new state-of-the-art performance on three widely used datasets including Market-1501, DukeMTMC-reID, and CUHK03-NP.
Kuan Zhu, Haiyun Guo, Songyan Liu, Jinqiao Wang, Ming Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 PASS: Part-Aware Self-Supervised Pre-Training for Person Re-Identification
Kuan Zhu, Haiyun Guo, Tianyi Yan, Yousong Zhu, Jinqiao Wang, Ming Tang 0001
ECCV (14)1
2022 Multi-Granularity Mutual Learning Network for Object Re-Identification
abstract
Object re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks.
Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Intell. Transp. Syst.2
2022 Hybrid Modality Metric Learning for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (Re-ID) has received increasing research attention for its great practical value in night-time surveillance scenarios. Due to the large variations in person pose, viewpoint, and occlusion in the same modality, as well as the domain gap brought by heterogeneous modality, this hybrid modality person matching task is quite challenging. Different from the metric learning methods for visible person re-ID, which only pose similarity constraints on class level, an efficient metric learning approach for visible-infrared person Re-ID should take both the class-level and modality-level similarity constraints into full consideration to learn sufficiently discriminative and robust features. In this article, the hybrid modality is divided into two types, within modality and cross modality. We first fully explore the variations that hinder the ranking results of visible-infrared person re-ID and roughly summarize them into three types: within-modality variation, cross-modality modality-related variation, and cross-modality modality-unrelated variation. Then, we propose a comprehensive metric learning framework based on four kinds of paired-based similarity constraints to address all the variations within and cross modality. This framework focuses on both class-level and modality-level similarity relationships between person images. Furthermore, we demonstrate the compatibility of our framework with any paired-based loss functions by giving detailed implementation of combing it with triplet loss and contrastive loss separately. Finally, extensive experiments of our approach on SYSU-MM01 and RegDB demonstrate the effectiveness and superiority of our proposed metric learning framework for visible-infrared person Re-ID.
La Zhang, Haiyun Guo, Kuan Zhu, Honglin Qiao, Gaopan Huang, Huichen Zhang, Jian Sun 0003, Jinqiao Wang
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Unsupervised cycle-consistent person pose transfer
Songyan Liu, Haiyun Guo, Kuan Zhu, Jinqiao Wang, Ming Tang 0001
Neurocomputing3
2020 Identity-Guided Human Semantic Parsing for Person Re-identification
Kuan Zhu, Haiyun Guo, Zhiwei Liu 0004, Ming Tang 0001, Jinqiao Wang
ECCV (3)1
2019 Two-Level Attention Network With Multi-Grain Ranking Loss for Vehicle Re-Identification
abstract
Vehicle re-identification (re-ID) aims to identify the same vehicle across multiple non-overlapping cameras, which is rather a challenging task. On the one hand, subtle changes in viewpoint and illumination condition can make the same vehicle look much different. On the other hand, different vehicles, even different vehicle models, may look quite similar. In this paper, we propose a novel Two-level Attention network supervised by a Multi-grain Ranking loss (TAMR) to learn an efficient feature embedding for the vehicle re-ID task. The two-level attention network consisting of hard part-level attention and soft pixel-level attention can adaptively extract discriminative features from the visual appearance of vehicles. The former one is designed to localize the salient vehicle parts, such as windscreen and car head. The latter one gives an additional attention refinement at pixel level to focus on the distinctive characteristics within each part. In addition, we present a multi-grain ranking loss to further enhance the discriminative ability of learned features. We creatively take the multi-grain relationship between vehicles into consideration. Thus, not only the discrimination between different vehicles but also the distinction between different vehicle models is constrained. Finally, the proposed network can learn a feature space, where both intra-class compactness and inter-class discrimination are well guaranteed. Extensive experiments demonstrate the effectiveness of our approach and we achieve state-of-the-art results on two challenging datasets, including VehicleID and Vehicle-1M.
Haiyun Guo, Kuan Zhu, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Image Process.2