VLDB 2026 Research / reviewers in the wild / expert
Shizhou Zhang
dblp:151/0743
· DBLP profile ↗
62ranked-venue papers
18as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 13 first-author · 34 since 2021Artificial intelligence and machine learning · 30 · 9 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do Large Language Models Reason About Uncertainty Like Humans? A Benchmark on Hurricane Forecast Visualization ComprehensionabstractUncertainty visualizations, such as hurricane cones and ensemble tracks, are essential for risk communication but are often misinterpreted, leading to harmful decisions. As AI assistants like large language models (LLMs) increasingly support understanding of graphics and decision-making, they offer a promising pathway to enhance the interpretation of complex visualizations and a new opportunity to examine and improve the interpretation of uncertainty. We introduce UnReason, the first benchmark that systematically compares how humans and LLMs reason about hurricane forecast uncertainty visualizations. UnReason spans two escalating phases, seven representative visualization formats, six real hurricane cases, and three agent types (humans, LLMs with context, and LLMs without context), including 880 visualizations and 117,600 structured question–answer pairs under matched evaluation conditions. Phase 1 evaluates reasoning across implicit and explicit uncertainty encodings; Phase 2 examines reasoning under single- versus multi-dimensional uncertainty representations. We thoroughly assess damage estimation, reasoning strategies, and comprehension patterns, revealing that LLMs have a stronger semantic and conceptual understanding of uncertainty, and are less misled by visual variability, but still replicate key human biases during decision-making. Our findings offer insights into aligning LLM behavior with human cognition in uncertainty-rich visual reasoning tasks. Le Liu 0008, Bohan Shen, Wei Zeng 0004, Shizhou Zhang, Di Xu 0010, Peng Wang 0015 |
AAAI | 5 |
| 2026 | Attention Retention for Continual Learning with Vision TransformersabstractContinual learning (CL) empowers AI systems to progressively acquire knowledge from non-stationary data streams. However, catastrophic forgetting remains a critical challenge. In this work, we identify attention drift in Vision Transformers as a primary source of catastrophic forgetting, where the attention to previously learned visual concepts shifts significantly after learning new tasks. Inspired by neuroscientific insights into the selective attention in the human visual system, we propose a novel attention-retaining framework to mitigate forgetting in CL. Our method constrains attention drift by explicitly modifying gradients during backpropagation through a two-step process: 1) extracting attention maps of the previous task using a layer-wise rollout mechanism and generating instance-adaptive binary masks, and 2) when learning a new task, applying these masks to zero out gradients associated with previous attention regions, thereby preventing disruption of learned visual concepts. For compatibility with modern optimizers, the gradient masking process is further enhanced by scaling parameter updates proportionally to maintain their relative magnitudes. Experiments and visualizations demonstrate the effectiveness of our method in mitigating catastrophic forgetting and preserving visual concepts. It achieves state-of-the-art performance and exhibits robust generalizability across diverse CL scenarios. Yue Lu 0008, Xiangyu Zhou 0002, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Wencong Zhang |
AAAI | 3 |
| 2026 | Better Matching, Less Forgetting: A Quality-Guided Matcher for Transformer-based Incremental Object DetectionabstractIncremental Object Detection (IOD) aims to continuously learn new object classes without forgetting previously learned ones. A persistent challenge is catastrophic forgetting, primarily attributed to background shift in conventional detectors. While pseudo-labeling mitigates this in dense detectors, we identify a novel, distinct source of forgetting specific to DETR-like architectures: background foregrounding. This arises from the exhaustiveness constraint of the Hungarian matcher, which forcibly assigns every ground truth target to one prediction, even when predictions primarily cover background regions (i.e., low IoU). This erroneous supervision compels the model to misclassify background features as specific foreground classes, disrupting learned representations and accelerating forgetting. To address this, we propose a Quality-guided Min-Cost Max-Flow (Q-MCMF) matcher. To avoid forced assignments, Q-MCMF builds a flow graph and prunes implausible matches based on geometric quality. It then optimizes for the final matching that minimizes cost and maximizes valid assignments. This strategy eliminates harmful supervision from background foregrounding while maximizing foreground learning signals. Extensive experiments on the COCO dataset under various incremental settings demonstrate that our method consistently outperforms existing state-of-the-art approaches. Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Lingyan Ran, Dahu Shi, Peng Wang 0015 |
AAAI | 2 |
| 2026 | DuGI-MAE: Improving Infrared Mask Autoencoders via Dual-Domain GuidanceabstractInfrared imaging plays a critical role in low-light and adverse weather conditions. However, due to the distinct characteristics of infrared images, existing foundation models such as Masked Autoencoder (MAE) trained on visible data perform suboptimal in infrared image interpretation tasks. To bridge this gap, an infrared foundation model known as InfMAE was developed and pre-trained on large-scale infrared datasets. Despite its effectiveness, InfMAE still faces several limitations, including the omission of informative tokens, insufficient modeling of global associations, and neglect of non-uniform noise. In this paper, we propose a Dual-domain Guided Infrared foundation model based on MAE (DuGI-MAE). First, we design a deterministic masking strategy based on token entropy, preserving only high-entropy tokens for reconstruction to enhance informativeness. Next, we introduce a Dual-Domain Guidance (DDG) module, which simultaneously captures global token relationships and adaptively filters non-uniform background noise commonly present in infrared imagery. To facilitate large-scale pretraining, we construct Inf-590K, a comprehensive infrared image dataset encompassing diverse scenes, various target types, and multiple spatial resolutions. Pretrained on Inf-590K, DuGI-MAE demonstrates strong generalization capabilities across various downstream tasks, including infrared object detection, semantic segmentation, and small target detection. Experimental results validate the superiority of the proposed method over both supervised and self-supervised comparison methods. Yinghui Xing, Xiaoting Su, Shizhou Zhang, Donghao Chu, Di Xu 0010 |
AAAI | 3 |
| 2026 | YOLO-IOD: Towards Real Time Incremental Object DetectionabstractCurrent methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to catastrophic forgetting in YOLO-based incremental detectors: foreground-background confusion, parameter interference, and misaligned knowledge distillation. Subsequently, we introduce YOLO-IOD, a real-time Incremental Object Detection (IOD) framework that is constructed upon the pretrained YOLO-World model, facilitating incremental learning via a stage-wise parameter-efficient finetuning process. Specifically, YOLO-IOD encompasses three principal components: 1) Conflict-Aware Pseudo-Label Refinement (CPR), which mitigates the foreground-background confusion by leveraging the confidence levels of pseudo labels and identifying potential objects relevant to future tasks. 2) Importance-based Kernel Selection (IKS), which identifies and updates the pivotal convolution kernels pertinent to the current task during the current learning stage. 3)Cross-Stage Asymmetric Knowledge Distillation (CAKD), which addresses the misaligned knowledge distillation conflict by transmitting the features of the student target detector through the detection heads of both the previous and current teacher detectors, thereby facilitating asymmetric distillation between existing and newly introduced categories. We further introduce LoCo COCO, a more realistic benchmark that eliminates data leakage across stages. Experiments on both conventional and LoCo COCO benchmarks show that YOLO-IOD achieves superior performance with minimal forgetting. Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Chen Zhao 0009, Yanning Zhang 0001 |
AAAI | 1 |
| 2026 | SpikeGate-YOLO: Spiking Object Detection With Dynamic Gating and Multigranularity FusionabstractSpiking Neural Networks (SNNs) transmit information via discrete spike events, offering advantages in energy efficiency and computational cost. However, current SNN-based object detectors suffer from limited feature expression and inefficient fusion due to temporal sparsity and the asynchronous nature of spike features. To address these challenges, we propose SpikeGate-YOLO, a spiking object detection architecture optimized for spike-driven processing. Specifically, we introduce the Reparam-Spike Gating (RSG) block to enhance feature expressiveness while maintaining computational efficiency. We also design the Spike Multi-Granularity Difference-aware Feature Harmonizer (SpikeMDFH), which improves multi-scale feature fusion through dynamic attention and biologically inspired gating mechanisms, preserving spike sparsity. Experiments on both the COCO and Gen1 datasets show that SpikeGate-YOLO achieves state-of-the-art results, reaching 63.4% mAP@50 and 46.3% mAP@50:95 on COCO, and 69.3% mAP@50 and 42.9% mAP@50:95 on Gen1. These results confirm the effectiveness of our architecture in overcoming spike-specific limitations in feature representation and fusion for object detection. Qiang Niu, Chen Zhao 0009, Shizhou Zhang, Wu Gao |
IEEE Internet Things J. | 3 |
| 2026 | Isolating Interference Factors for Robust Cloth-Changing Person Re-IdentificationabstractCloth-Changing Person Re-Identification (CC-ReID) aims to recognize individuals across camera views despite clothing variations, a crucial task for surveillance and security systems. Existing methods typically frame it as a cross-modal alignment problem but often overlook explicit modeling of interference factors such as clothing, viewpoints, and pedestrian actions. This oversight can distort their impact, compromising the extraction of robust identity features. To address these challenges, we propose a novel framework that systematically disentangles interference factors from identity features while ensuring the robustness and discriminative power of identity representations. Our approach consists of two key components. First, a dual-stream identity feature learning framework leverages a raw image stream and a cloth-isolated stream, to extract identity representations independent of clothing textures. An adaptive cloth-irrelevant contrastive objective is introduced to mitigate identity feature variations caused by clothing differences. Second, we propose a Text-Driven Conditional Generative Adversarial Interference Disentanglement Network (T-CGAIDN), to further suppress interference factors beyond clothing textures, such as finer clothing patterns, viewpoint, background, and lighting conditions. This network incorporates a multi-granularity interference recognition branch to learn interference-related features, a conditional adversarial module for bidirectional transformation between identity and interference feature spaces, and an interference decoupling objective to eliminate interference dependencies in identity learning. Extensive experiments on public benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, highlighting its effectiveness in CC-ReID. De Cheng, Chaowei Fang, Shizhou Zhang, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | VPT-NSP2++: Importance-Aware Visual Prompt Tuning in Null Space for Continual LearningabstractContinual learning (CL) enables AI models to adapt to evolving environments while mitigating catastrophic forgetting, which is a critical capability for dynamic real-world applications. With the growing popularity of pre-trained Vision Transformer (ViT) models and visual prompt tuning (VPT) technique in CL, this work explores a CL method on top of the ViT-based foundation model, through VPT mechanism with theoretical guarantees. Inspired by the orthogonal projection method, we aim to leverage this approach for VPT to enhance CL performance, particularly in long-term scenarios. However, since the orthogonal projection is originally designed for linear operations in CNNs, applying it to ViTs poses challenges induced by the non-linear self-attention mechanism and the distribution drift within LayerNorm. To address these issues, we deduced two orthogonality conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of maintaining stability. Considering the strict orthogonal constraints can diminish model capacity and reduce plasticity, we further propose an importance-aware orthogonal regularization framework. By applying varying degrees of orthogonal constraints to different parameters based on their importance to old and new tasks, the framework adaptively enhances model capacity and thereby promotes long-sequence CL while improving the stability-plasticity trade-off. To implement the proposed approach, a null-space-based approximation solution is employed to efficiently achieve the prompt gradient orthogonal projection. Extensive experiments on various class-incremental learning benchmarks demonstrate that our method achieves state-of-the-art performance across diverse CL scenarios. Shizhou Zhang, Yue Lu 0008, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Nearest-neighbor class prototype prompt and simulated logits for continual learning
Yue Lu 0008, Shizhou Zhang, Yinghui Xing, Guoqiang Liang 0001, Yanning Zhang 0001 |
Pattern Recognit. | 3 |
| 2026 | Multi-Level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCILabstractOnline Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams can be seen only once, it is more suitable for real-world scenarios compared to offline learning. However, this constraint intensifies the challenge for OCIL in maintaining an appropriate balance between stability and plasticity. Moreover, under stricter memory buffer constraints in real world, current replay-based methods are less effective. While ensemble methods improve plasticity, they often struggle with stability. Inspired by the Global Workspace Theory (GWT), we propose a novel approach that enhances ensemble learning through a Global Workspace Model (GWM)-a shared, implicit memory that guides the learning of multiple student models. The GWM is formed by fusing the parameters of all students within each training batch, capturing the historical learning trajectory and serving as a dynamic anchor for knowledge consolidation. Like the broadcasting mechanism of GWT, the GWM is redistributed periodically to students, stabilizing learning and promoting cross-task consistency. In addition, we introduce a multi-level collaborative distillation mechanism. It enforces peer-to-peer consistency among students and preserves historical knowledge by aligning each student with the GWM. As a result, student models remain adaptable to new tasks while maintaining previously learned knowledge, striking a better balance between stability and plasticity. Extensive experiments on three standard OCIL benchmarks show that our method delivers significant performance improvement for several OCIL models across various memory budgets. The code is available at https://github.com/susususushi/GWM. Shibin Su, Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Lingyan Ran |
IEEE Trans. Image Process. | 4 |
| 2026 | Less Is More: Infrared and Visible Images Fusion via Semantic-Guided Mixture of Multi-Feature ExpertsabstractInfrared (IR) and visible image fusion (IVIF) has become prevalent in recent years. By leveraging the complementary characteristics of infrared and visible images, we can obtain visually-appealing fused images, which further facilitate subsequent scene understanding and object detection from day to night. Integrating complementary information while simultaneously eliminating redundancy is a crucial challenge in fusion. Most of available deep learning based methods, after being trained, execute static inference on all pairs of infrared and visible images. They struggle to effectively handle redundancy of modality across diverse scenarios, resulting in superfluous information such as thermal noise in infrared images and artifacts in visible images. In this paper, we propose an IVIF method based on a semantic-guided mixture of multi-feature experts, where multiple types of features are extracted, each assigned to a dedicated expert network specialized in processing a specific type of features. Through an expert routing mechanism, these experts are chosen dynamically, ensuring that the most significant features of each image modality are routed to a specific group of experts. In order to align fusion task with subsequent semantic segmentation task, we introduce a segmentation head to semantically guide the selection of the complementary features. Extensive experiments on five infrared and visible image fusion and segmentation benchmarks demonstrate the effectiveness of our method, both for image fusion and subsequent semantic segmentation tasks. The code will be available at https://github.com/ZhilongNiu/SD-MoMFE. Yinghui Xing, Zhilong Niu, Shizhou Zhang, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual LearningabstractVisual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generators to preserve stability and minimize forgetting. In contrast, we introduce a novel approach that trains a consistent prompt generator to ensure stability during CL. Consistency means that for any instance from an old task, its corresponding instance-ware prompt generated by the prompt generator remains consistent even as the generator continually updates in a new task. This ensures that the representation of a specific instance remains stable across tasks and thereby prevents forgetting. We employ a mixture of experts (MoE) as the prompt generator, which contains a router and multiple experts. By deriving conditions sufficient to achieve the consistency for the MoE prompt generator, we demonstrate that: during training in a new task, if the router and experts update in the directions orthogonal to the subspaces spanned by old input features and gating vectors, respectively, the consistency can be theoretically guaranteed. To implement this orthogonality, we project parameter gradients to those orthogonal directions using the orthogonal projection matrices computed via the null space method. Extensive experiments on four class-incremental learning benchmarks validate the effectiveness and superiority of our approach. Yue Lu 0008, Shizhou Zhang, De Cheng, Guoqiang Liang 0001, Yinghui Xing, Nannan Wang 0001, Yanning Zhang 0001 |
AAAI | 2 |
| 2025 | ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge EditingabstractLarge multimodal language models (MLLMs) have revolutionized natural language processing and visual understanding, but often contain outdated or inaccurate information. Current multimodal knowledge editing evaluations are limited in scope and potentially biased, focusing on narrow tasks and failing to assess the impact on in-domain samples. To address these issues, we introduce ComprehendEdit, a comprehensive benchmark comprising eight diverse tasks from multiple datasets. We propose two novel metrics: Knowledge Generalization Index (KGI) and Knowledge Preservation Index (KPI), which evaluate editing effects on in-domain samples without relying on AI-synthetic samples. Based on insights from our framework, we establish Hierarchical In-Context Editing (HICE), a baseline method employing a two-stage approach that balances performance across all metrics. This study provides a more comprehensive evaluation framework for multimodal knowledge editing, reveals unique challenges in this field, and offers a baseline method demonstrating improved performance. Our work opens new perspectives for future research and provides a foundation for developing more robust and effective editing techniques for MLLMs. Yaohui Ma, Xiaopeng Hong, Shizhou Zhang, Huiyun Li, Zhilin Zhu 0001, Wei Luo 0014, Zhiheng Ma |
AAAI | 3 |
| 2025 | Dual-Granularity Semantic Guided Sparse Routing Diffusion Model for General PansharpeningabstractPansharpening aims at integrating complementary information from panchromatic and multispectral images. Available deep-learning based pansharpening methods typically perform exceptionally with particular satellite datasets. At the same time, it has been observed that these models also exhibit scene dependence, for example, if the majority of the training samples come from the urban scenes, the model’s performance may decline in the river scene. To address the domain gap produced by varying satellite sensors and distinct scenes, we propose a dual-granularity semantic guided sparse routing diffusion model for general pansharpening. By utilizing the large Vision-Language Models (VLMs) in the field of geoscience, e.g, GeoChat, we introduce the dual granularity semantics to generate dynamic sparse routing scores for adaptation of different satellite sensors and scenes. This scene-level and region-level dual-granularity semantic information serves as guidance for dynamically activating specialized experts within the diffusion model. Extensive experiments on WorldView-3, QuickBird, and GaoFen-2 datasets show the effectiveness of our proposed method. Notably, the proposed method outperforms the comparison approaches in adapting to new satellite sensors and scenes. The codes are available at https://github.com/codgodtao/SGDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Di Xu 0010, Yingkun Yang, Yanning Zhang 0001 |
CVPR | 3 |
| 2025 | Revisiting Generative Replay for Class Incremental Object DetectionabstractGenerative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation that the forgetting of prior knowledge is predominantly present in the classification sub-task as opposed to the localization sub-task, we revisit the generative replay method for class incremental object detection. Our method utilize a standard Stable Diffusion model to generate image-level replay data for all old and new tasks. Accordingly, the old detector and a stage-wise detector are conducted on the synthetic images respectively to determine the bounding box positions through pseudo-labeling. Furthermore, we propose to use a Similarity-based Cross Sampling mechanism to select valuable confusing data between old and new tasks to more effectively mitigate catastrophic forgetting and reduce the false alarm rate for the new task. Finally, all synthetic and real data are integrated for current-stage detector training, where the images generated for previous tasks are highly beneficial in minimizing the forgetting of existing knowledge, while those synthesized for the new task can help bridge the domain gap between real and synthetic images. We conducted extensive experiments on PASCAL VOC 2007 and MS COCO benchmark datasets in multiple settings to showcase the efficacy of our proposed approach, which achieves state-of-the-art results. The code is available at https://github.com/qiangzailv/RGR-IOD. Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Yanning Zhang 0001 |
CVPR | 1 |
| 2025 | Gradient Decomposition and Alignment for Incremental Object Detection
Wenlong Luo, Shizhou Zhang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
ICCV | 2 |
| 2025 | Demystifying Catastrophic Forgetting in Two-Stage Incremental Object DetectorabstractCatastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key insight: Catastrophic forgetting is predominantly localized to the RoI Head classifier, while regressors retain robustness across incremental stages. This finding challenges conventional assumptions, motivating us to develop a framework termed NSGP-RePRE. Regional Prototype Replay (RePRE) mitigates classifier forgetting via replay of two types of prototypes: coarse prototypes represent class-wise semantic centers of RoI features, while fine-grained prototypes model intra-class variations. Null Space Gradient Projection (NSGP) is further introduced to eliminate prototype-feature misalignment by updating the feature extractor in directions orthogonal to subspace of old inputs via gradient projection, aligning RePRE with incremental learning dynamics. Our simple yet effective design allows NSGP-RePRE to achieve state-of-the-art performance on the Pascal VOC and MS COCO datasets under various settings. Our work not only advances IOD methodology but also provide pivotal insights for catastrophic forgetting mitigation in IOD. Code will be available soon. Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Di Xu 0010, Peng Wang 0015, Yanning Zhang 0001 |
ICML | 2 |
| 2025 | Boosting Multi-Modal Alignment: Geometric Feature Separation for Class Incremental LearningabstractClass Incremental Learning (CIL) aims to continually learn new classes from a stream of data without forgetting previously learned ones. Recent approaches have leveraged pre-trained models (PTMs) to improve performance, especially vision-language models, which offer better generalization than models trained solely on visual data. Many of these methods rely on simple language templates to generate class representations, which then serve as classifiers. However, due to differences between the pre-training data and downstream tasks, these textual features can become too similar for certain classes, leading to prediction errors. To address this issue, we propose a method that optimizes the geometric structure of both visual and textual features across different classes. Inspired by neural collapse theory, we introduce a multi-modal alignment strategy: for each class, a reference vector is chosen from a simplex Equiangular Tight Frame, and both the visual and textual features of the class are aligned with this vector. To better capture intra-class variations, we also construct multiple visual prototypes for each class. A multi-prototype supervised contrastive loss is then employed to pull an image feature toward the closest matching prototype of its true class and push it away from prototypes of other classes. We evaluate our approach on five widely used CIL benchmarks. The results show that our method achieves state-of-the-art performance, demonstrating its effectiveness in addressing the challenges of class incremental learning. Our code is available at https://github.com/qcNPU/NCSCMP. Guoqiang Liang 0001, De Cheng, Shizhou Zhang, Yanning Zhang 0001 |
ACM Multimedia | 4 |
| 2025 | Amplitude-aware Domain Style Replay for Lifelong Person Re-identificationabstractLifelong Person Re-identification (LReID) focuses on continuously adapting to new domains over time while preserving knowledge from previously seen domains, particularly under the domain incremental learning setting. The major challenge of LReID is catastrophic forgetting, typically caused by large domain shifts during training. To address this, we propose a novel Amplitude-aware Domain Style Replay (ADSR) framework, which introduces a Fourier-based Style Transfer (FST) mechanism to generate synthetic data that reflects the style of previously encountered domains. These proxy images help retain prior knowledge without the need to store actual past data. Our method transfers stylistic information-mainly encoded in the amplitude spectrum-from old domains to new ones, creating old-stylized images that preserve the content of new domain data while adopting the visual style of earlier domains. To further boost generalization, we design a Self-Stylization Normalization (SSN) module that adapts the current domain's style distribution, making the model more robust to stylistic variations. Additionally, we introduce a Multi-Granularity Transfer (MGT) module that uses K-Means clustering to extract multiple representative style features from each domain, enabling compact yet comprehensive storage and replay of domain-specific information. Extensive experiments on multiple LReID benchmarks show that ADSR achieves superior performance over existing approaches, effectively reducing forgetting and improving cross-domain generalization. Our code is available at https://github.com/cclong8/MM2025-ADSR. De Cheng, Shizhou Zhang, Yinghui Xing, Di Xu 0010, Yanning Zhang 0001 |
ACM Multimedia | 3 |
| 2025 | SUVIS: A Depth- and Motion-Encoded Stereoscopic System for Communicating Forecast UncertaintyabstractEffectively communicating uncertainty in ensemble hurricane forecasts poses a significant multimedia challenge, requiring the integration of spatial, temporal, and perceptual dimensions. We introduce SUVIS, a stereoscopic visualization system that encodes forecast ensembles into an immersive, layered media experience. SUVIS transforms multidimensional ensemble data into animated stereoscopic representations, mapping time to vertical depth, intensity to texture color, and forward speed to motion flow, while semi-transparent glyphs represent evolving impact areas. A progressive sampling strategy ensures spatial clarity across depth layers. Rendered on a glasses-free stereoscopic display, SUVIS frames uncertainty visualization as a media encoding problem, synthesizing motion, depth, and spatial abstraction to align with human perception. A user study with 51 participants demonstrates that SUVIS supports high accuracy in spatial tasks and enables interpretation of dynamic storm attributes. These results highlight the system's potential to advance perceptual uncertainty communication through multimedia representation and immersive visual encoding. Le Liu 0008, Shizhou Zhang, Di Xu 0010 |
ACM Multimedia | 2 |
| 2025 | ChatHSI: Reliable LLM-Powered Human-Swarm Interaction FrameworkabstractHuman-swarm interaction (HSI) is critical for scalable control of UAV swarm systems. Traditional interfaces struggle with generalization and user workload, especially in immersive environments. Hence, we present ChatHSI, a framework leveraging large language models (LLMs) for swarm task planning. ChatHSI integrates prompt engineering, action validation, and a human-in-the-loop mechanism to improve planning feasibility and executability. We implement ChatHSI in an immersive simulation to improve users’ spatial and situational awareness. Our method shows improved task efficiency, reduced workload, and higher usability in user studies. Ablation study proves the effectiveness of prompt context and action validation. The results show the feasibility of LLM-driven interaction for immersive swarm control and point toward adaptive, intuitive, and scalable HSI systems. Bohan Shen, Le Liu 0008, Shizhou Zhang, Peng Wang 0015, Lingyun Yu 0001, Di Xu 0010 |
VINCI | 4 |
| 2025 | A masking, linkage and guidance framework for online class incremental learning
Guoqiang Liang 0001, Zhaojie Chen, Shibin Su, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2025 | Joint Memory Optimization for Continual LearningabstractContinual learning, focusing on sequential knowledge acquisition and retention, necessitates efficient memory management. This paper introduces a holistic approach, diverging from traditional methods that separately optimize neural network and replay buffer memory. We aim to enhance overall memory efficiency, addressing neural network parameters and replay buffer concurrently within strict memory constraints. This is achieved by harnessing neural network parameter redundancies and employing compression techniques like pruning and quantization, allowing data replay storage without extra memory overhead. Balancing memory use across components is challenging due to the complex search space of combined tasks. We tackle this by conceptualizing it as a bi-level optimization problem, integrating all tasks under a single objective, thus optimizing memory use and managing the interplay between different components. We employ a synergy of optimization techniques to solve this challenging bi-level optimization problem. Our experimental findings affirm the superior performance of our proposed method, outperforming existing techniques such as prompt-based, feature-replay, exemplar-replay, and regularization-based methods under stringent memory constraints, consistently across various datasets and neural network architectures. Zhiheng Ma, Yaohui Ma, Xiaopeng Hong, Huiyun Li, Shizhou Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | AdaSemiCD: An Adaptive Semi-Supervised Change Detection Method Based on Pseudo-Label EvaluationabstractChange detection (CD) is an essential field in remote sensing, with a primary focus on identifying areas of change in bitemporal image pairs captured at varying intervals of the same region. The data annotation process for CD tasks is both time-consuming and labor-intensive. To better utilize the scarce labeled data and abundant unlabeled data, we introduce an adaptive semi-supervised learning (SSL) method, AdaSemiCD, to improve pseudo-label usage and optimize the training process. Initially, due to the extreme class imbalance inherent in CD, the model is more inclined to focus on the background class, and it is easy to confuse the boundary of the target object. Considering these two points, we develop a measurable evaluation metric for pseudo-labels that enhances the representation of information entropy by class rebalancing and amplification of ambiguous areas, assigning greater weights to prospective change objects. Subsequently, to enhance the reliability of sample wise pseudo-labels, we introduce the AdaFusion module, to dynamically identify the most uncertain region and substitute it with more trustworthy content. Lastly, to ensure better training stability, we introduce the AdaEMA module, which updates the teacher model using only batches of trusted samples. Experimental results on ten public CD datasets validate the efficacy and generalizability of our proposed adaptive training framework. Lingyan Ran, Wen Dongcheng, Tao Zhuo, Shizhou Zhang, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Enhancing Feature Learning With Hard Samples in Mutual Learning for Online Class Incremental LearningabstractOnline Class-Incremental Learning (OCIL) aims to solve the problem of incrementally learning new classes from a non-i.i.d. and single-pass data stream. Compared to the offline setting, OCIL is much closer to a live learning experience requiring higher model update frequency at less computational budget. Due to its one-epoch training constraint, the model is likely to learn non-essential features and encounter the under-fitting issue, which severely affects the model's stability. In this paper, we investigate how to use hard samples to improve data variability, eventually enhancing feature learning and addressing the under-fitting problem. Specifically, by introducing a scoring function assessing the sample value, we build an OCIL formulation that simultaneously generates high-value samples and optimizes the OCIL model, improving generalization ability within the constraint of single-epoch training. Empirically, we found that strong data augmentation is a simple but effective way to generate a higher proportion of high-score samples. To make the most of these augmented samples, we design an OCIL model based on mutual learning with two networks of identical structures. Moreover, a collaborative learning mechanism is developed by aligning the features and class probabilities from the two networks to promote their interaction. Extensive experiments on three widely used datasets for OCIL have demonstrated the effectiveness of our method, obtaining superior performance to state-of-the-art methods. The code is available at https://github.com/susususushi/SDA-MCL. Guoqiang Liang 0001, Shibin Su, De Cheng, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Prompt-Based Modality Alignment for Effective Multi-Modal Object Re-IdentificationabstractA critical challenge for multi-modal Object Re-Identification (ReID) is the effective aggregation of complementary information to mitigate illumination issues. State-of-the-art methods typically employ complex and highly-coupled architectures, which unavoidably result in heavy computational costs. Moreover, the significant distribution gap among different image spectra hinders the joint representation of multi-modal features. In this paper, we propose a framework named as PromptMA to establish effective communication channels between different modality paths, thereby aggregating modal complementary information and bridging the distribution gap. Specifically, we inject a series of learnable multi-modal prompts into the Image Encoder and introduce a prompt exchange mechanism to enable the prompts to alternately interact with different modal token embeddings, thus capturing and distributing multi-modal features effectively. Building on top of the multi-modal prompts, we further propose Prompt-based Token Selection (PBTS) and Prompt-based Modality Fusion (PBMF) modules to achieve effective multi-modal feature fusion while minimizing background interference. Additionally, due to the flexibility of our prompt exchange mechanism, our method is well-suited to handle scenarios with missing modalities. Extensive evaluations are conducted on four widely used benchmark datasets and the experimental results demonstrate that our method achieves state-of-the-art performances, surpassing the current benchmarks by over 15% on the challenging MSVR310 dataset and by 6% on the RGBNT201. The code is available at https://github.com/FHR-L/PromptMA. Shizhou Zhang, Wenlong Luo, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Frequency-Guided Spatial Adaptation for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to segment camouflaged objects which exhibit very similar patterns with the surrounding environment. Recent research works have shown that enhancing the feature representation via the frequency information can greatly alleviate the ambiguity problem between the foreground objects and the background. With the emergence of vision foundation models, like InternImage, Segment Anything Model etc, adapting the pretrained model on COD tasks with a lightweight adapter module shows a novel and promising research direction. Existing adapter modules mainly care about the feature adaptation in the spatial domain. In this paper, we propose a novel frequency-guided spatial adaptation method for COD task. Specifically, we transform the input features of the adapter into frequency domain. By grouping and interacting with frequency components located within non overlapping circles in the spectrogram, different frequency components are dynamically enhanced or weakened, making the intensity of image details and contour features adaptively adjusted. At the same time, the features that are conducive to distinguishing object and background are highlighted, indirectly implying the position and shape of camouflaged object. We conduct extensive experiments on four widely adopted benchmark datasets and the proposed method outperforms 26 state-of-the-art methods with large margins. Code will be released. Shizhou Zhang, Dexuan Kong, Yinghui Xing, Yue Lu 0008, Lingyan Ran, Guoqiang Liang 0001, Hexu Wang, Yanning Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Cross-Platform Video Person ReID: A New Benchmark Dataset and Adaptation Approach
Shizhou Zhang, Wenlong Luo, De Cheng, Qingchun Yang, Lingyan Ran, Yinghui Xing, Yanning Zhang 0001 |
ECCV (27) | 1 |
| 2024 | Dual Supervised Contrastive Learning Based on Perturbation Uncertainty for Online Class Incremental Learning
Shibin Su, Zhaojie Chen, Guoqiang Liang 0001, Shizhou Zhang, Yanning Zhang 0001 |
ICPR (9) | 4 |
| 2024 | Dual-Branch Task Residual Enhancement with Parameter-Free Attention for Zero-Shot Multi-label Image Recognition
Shizhou Zhang, Kairui Dang, De Cheng, Yinghui Xing, Qirui Wu, Dexuan Kong, Yanning Zhang 0001 |
ICPR (22) | 1 |
| 2024 | Bridging Fourier and Spatial-Spectral Domains for Hyperspectral Image DenoisingabstractRemarkable progresses have been made in hyperspectral image (HSI) denoising. However, the majority of existing methods are predominantly confined to the spatial-spectral domain, overlooking the untapped potential inherent in the Fourier domain. This paper presents a novel approach to address HSI denoising by bridging the information from the Fourier and spatial-spectral domains. Our method highlights key insights into the Fourier properties within spatial and spectral domains through the Fourier transform. Specifically, we note that the amplitude inherently embody noise and photon reflection characteristics, while the phase holds structural information. These insights unveil new perspectives on the physical properties of HSIs, motivating us to leverage complementary information exchange between Fourier and spatial-spectral domains. To this end, we introduce the Fourier-prior Integration Denoising Network (FIDNet), a potent yet straightforward approach that utilizes Fourier insights to synergistically interact with spatial-spectral domains for superior HSI denoising. In FIDNet, we independently extract spatial and Fourier features through dual branches and merge these representations to enhance spectral evolution modeling through the inherent structure consistency constraints and continuing reflection variation revealed in Fourier prior. Our proposed method demonstrates robust generalization across synthetic and real-world benchmark datasets, achieves comparable results with state-of-the-art methods in both quantitative quality and visual results. The code is available at https://github.com/MIV-XJTU/FIDNet. Jiahua Xiao, Yang Liu 0385, Shizhou Zhang, Xing Wei 0001 |
ACM Multimedia | 3 |
| 2024 | Visual Prompt Tuning in Null Space for Continual LearningabstractExisting prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace spanned by previous tasks' features, so as to ensure no interference on tasks that have been learned to overcome catastrophic forgetting in CL. However, different from the orthogonal projection in the traditional CNN architecture, the prompt gradient orthogonal projection in the ViT architecture shows completely different and greater challenges, i.e., 1) the high-order and non-linear self-attention operation; 2) the drift of prompt distribution brought by the LayerNorm in the transformer block. Theoretically, we have finally deduced two consistency conditions to achieve the prompt gradient orthogonal projection, which provide a theoretical guarantee of eliminating interference on previously learned knowledge via the self-attention mechanism in visual prompt tuning. In practice, an effective null-space-based approximation solution has been proposed to implement the prompt gradient orthogonal projection. Extensive experimental results demonstrate the effectiveness of anti-forgetting on four class-incremental benchmarks with diverse pre-trained baseline models, and our approach achieves superior performances to state-of-the-art methods. Our code is available at https://github.com/zugexiaodui/VPTinNSforCL Yue Lu 0008, Shizhou Zhang, De Cheng, Yinghui Xing, Nannan Wang 0001, Peng Wang 0015, Yanning Zhang 0001 |
NeurIPS | 2 |
| 2024 | Scale-aware local difference attention on pyramidal features for crowd counting
Qian Zhang 0046, Shizhou Zhang, Xinyao Liu, Yanning Zhang 0001 |
Multim. Tools Appl. | 2 |
| 2024 | Vehicle Re-Identification in Aerial Images and Videos: Dataset and ApproachabstractIn this work, we propose a large-scale dataset, VRAI, and an effective Orientation Adaptive and Salience Attentive (OASA) Network for vehicle re-identification (ReID) in aerial imagery. The VRAI dataset includes two subsets: VRAI-Image, which contains over 137,000 images of 13,000 vehicle instances, and VRAI-Video, which comprises more than 14,000 video trajectories of 7,000 identities. To our best knowledge, this is the largest dataset for UAV-based vehicle ReID, and the first dataset proposed for video-based ReID under UAV views. Based on the VRAI dataset, we design an OASA network to address two crucial challenges of vehicle ReID in aerial imagery. Firstly, the significant vehicle orientation variations in aerial images could cause great vehicle pattern deformations, making it difficult to identify vehicles across UAV views. To overcome this challenge, in our OASA, an orientation adaptive dynamic convolution module is designed, which constructs customized kernels for each vehicle instance to extract orientation-invariant features. Besides, the unique vertical view and long focal length of the UAV platform often render many salient vehicle attributes, such as logos and license plates, invisible, which brings a great challenge to ReID models to extract distinguishable vehicle features. To address this issue, in the OASA, we design a transformer-based salience attentive module (Trans-Attn) that guides the model to focus on subtle yet discriminative clues of vehicle instances in aerial imagery. Through extensive experiments, both of our designed modules are verified effective. Besides, our OASA model outperforms state-of-the-art algorithms both on our VRAI dataset and other surveillance-based datasets. Our VRAI dataset is available in https://github.com/JiaoBL1234/VRAI-Dataset. Bingliang Jiao, Lu Yang 0016, Liying Gao, Peng Wang 0015, Shizhou Zhang, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Empower Generalizability for Pansharpening Through Text-Modulated Diffusion ModelabstractPansharpening is crucial to remote sensing applications by fusing high-resolution (HR) panchromatic (PAN) images with low-resolution multispectral (LRMS) images to generate HR multispectral (HRMS) images. Recently, diffusion probabilistic models (DPMs) have provided high-quality results than regression-based methods when trained on specific pairwise data for their specific purpose. However, their performance degrades when applied to a new satellite dataset, which represents different imaging properties and spectral ranges, limiting the generalization ability of them. For better generalizability of pansharpening, in this article, we propose a text-modulated diffusion model (TMDiff) for unified pansharpening of different satellites. TMDiff takes a text-modulated 3-D UNet (TM3DU) as denoising network to gradually recover HRMS through iterative refinement over multiple time steps. By introducing satellite’s physical properties as text prompts, TM3DU is able to learn meta-knowledge across different satellites and thus can sharpen LRMS images with diverse spatial and spectral attributes. Extensive experiments on various satellite datasets demonstrate the state-of-the-art performance of our model in both qualitative and quantitative metrics. Furthermore, our model exhibits superior generalization ability to unseen datasets, highlighting its practical significance. Code is available athttps://github.com/codgodtao/TMDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Jiapeng Feng, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SCAFNet: Semantic-Guided Cascade Adaptive Fusion Network for Infrared Small Target DetectionabstractInfrared small target detection is a crucial component of infrared target tracking and search. It is challenging due to the complex backgrounds, low contrast between targets and backgrounds, and the small, dim nature of the targets. Therefore, effectively representing the targets and enhancing the distinction between targets and backgrounds is essential. Existing deep-learning (DL)-based methods struggle to capture the subtle details of weak targets, neglecting the complementary characteristics of multilevel features, which leads to inaccurate localization of targets. In this article, we propose a semantic-guided cascade adaptive fusion network (SCAFNet) to address these challenges. To improve the representation of small targets in the deeper layers, we introduce a multiresolution auxiliary enhancement (MAE) encoder to progressively enhance detailed information within the deep features. After extracting multiscale features, an adaptive fusion (AdaFus) decoder is proposed to fuse them. It has a semantic-guided cascade fusion (SGCF) module to integrate feature maps at three different resolutions. Specifically, SGCF first employs rich semantic features from the high-level feature map to guide the spatial distribution of the low-level feature maps, thereby improving the distinction between the target and the background. Then, AdaFus weights are generated to guide the fusion process, ensuring that the final feature map combines rich semantic information with precise spatial details. Furthermore, we perform long-distance modeling on the feature map to achieve detailed reconstruction, which aids in restoring the shape information of the target. The effectiveness of our method is validated through experiments on various public infrared small target detection datasets. Shizhou Zhang, Yinghui Xing, Liangkui Lin, Xiaoting Su, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | CrossDiff: Exploring Self-SupervisedRepresentation of Pansharpening via Cross-Predictive Diffusion ModelabstractFusion of a panchromatic (PAN) image and corresponding multispectral (MS) image is also known as pansharpening, which aims to combine abundant spatial details of PAN and spectral information of MS images. Due to the absence of high-resolution MS images, available deep-learning-based methods usually follow the paradigm of training at reduced resolution and testing at both reduced and full resolution. When taking original MS and PAN images as inputs, they always obtain sub-optimal results due to the scale variation. In this paper, we propose to explore the self-supervised representation for pansharpening by designing a cross-predictive diffusion model, named CrossDiff. It has two-stage training. In the first stage, we introduce a cross-predictive pretext task to pre-train the UNet structure based on conditional Denoising Diffusion Probabilistic Model (DDPM). While in the second stage, the encoders of the UNets are frozen to directly extract spatial and spectral features from PAN and MS images, and only the fusion head is trained to adapt for pansharpening task. Extensive experiments show the effectiveness and superiority of the proposed model compared with state-of-the-art supervised and unsupervised methods. Besides, the cross-sensor experiments also verify the generalization ability of proposed self-supervised representation learners for other satellite datasets. Code is available at https://github.com/codgodtao/CrossDiff. Yinghui Xing, Litao Qu, Shizhou Zhang, Kai Zhang 0010, Yanning Zhang 0001, Lorenzo Bruzzone |
IEEE Trans. Image Process. | 3 |
| 2024 | MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced OptimizationabstractMultispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR. Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Dual Modality Prompt Tuning for Vision-Language Pre-Trained ModelabstractWith the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method. Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Weakly Supervised Video Anomaly Detection Based on Cross-Batch Clustering GuidanceabstractWeakly supervised video anomaly detection (WSVAD) is a challenging task since only video-level labels are available for training. In previous studies, the discriminative power of the learned features is not strong enough, and the data imbalance resulting from the mini-batch training strategy is ignored. To address these two issues, we propose a novel WSVAD method based on cross-batch clustering guidance. To enhance the discriminative power of features, we propose a batch clustering based loss to encourage a clustering branch to generate distinct normal and abnormal clusters based on a batch of data. Meanwhile, we design a cross-batch learning strategy by introducing clustering results from previous minibatches to reduce the impact of data imbalance. In addition, we propose to generate more accurate segment-level anomaly scores based on batch clustering guidance to further improve the performance of WSVAD. Extensive experiments on two public datasets demonstrate the effectiveness of our approach. Congqi Cao, Xin Zhang 0168, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ICME | 3 |
| 2023 | Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReIDabstractUnsupervised visible-infrared person re-identification (USL-VI-ReID) aims to match pedestrian images of the same identity from different modalities without annotations. Existing works mainly focus on alleviating the modality gap by aligning instance-level features of the unlabeled samples. However, the relationships between cross-modality clusters are not well explored. To this end, we propose a novel bilateral cluster matching-based learning framework to reduce the modality gap by matching cross-modality clusters. Specifically, we design a Many-to-many Bilateral Cross-Modality Cluster Matching (MBCCM) algorithm through optimizing the maximum matching problem in a bipartite graph. Then, the matched pairwise clusters utilize shared visible and infrared pseudo-labels during the model training. Under such a supervisory signal, a Modality-Specific and Modality-Agnostic (MSMA) contrastive learning framework is proposed to align features jointly at a cluster-level. Meanwhile, the cross-modality Consistency Constraint (CC) is proposed to explicitly reduce the large modality discrepancy. Extensive experiments on the public SYSU-MM01 and RegDB datasets demonstrate the effectiveness of the proposed method, surpassing state-of-the-art approaches by a large margin of 8.76% mAP on average. De Cheng, Nannan Wang 0001, Shizhou Zhang, Zhen Wang 0037, Xinbo Gao 0001 |
ACM Multimedia | 4 |
| 2023 | Ground-to-Aerial Person Search: Benchmark Dataset and ApproachabstractIn this work, we construct a large-scale dataset for Ground-to-Aerial Person Search, named G2APS, which contains 31,770 images of 260,559 annotated bounding boxes for 2,644 identities appearing in both of the UAVs and ground surveillance cameras. To our knowledge, this is the first dataset for cross-platform intelligent surveillance applications, where the UAVs could work as a powerful complement for the ground surveillance cameras. To more realistically simulate the actual cross-platform Ground-to-Aerial surveillance scenarios, the surveillance cameras are fixed about 2 meters above the ground, while the UAVs capture videos of persons at different location, with a variety of view-angles, flight attitudes and flight modes. Therefore, the dataset has the following unique characteristics: 1) drastic view-angle changes between query and gallery person images from cross-platform cameras; 2) diverse resolutions, poses and views of the person images under 9 rich real-world scenarios. On basis of the G2APS benchmark dataset, we demonstrate detailed analysis about current two-step and end-to-end person search methods, and further propose a simple yet effective knowledge distillation scheme on the head of the ReID network, which achieves state-of-the-art performances on both of the G2APS and the previous two public person search datasets, i.e., PRW and CUHK-SYSU. The dataset and source code available on https://github.com/yqc123456/HKD_for_person_search. Shizhou Zhang, Qingchun Yang, De Cheng, Yinghui Xing, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001 |
ACM Multimedia | 1 |
| 2022 | Dynamically Transformed Instance Normalization Network for Generalizable Person Re-Identification
Bingliang Jiao, Lingqiao Liu, Liying Gao, Guosheng Lin, Lu Yang 0016, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
ECCV (14) | 6 |
| 2022 | Video summarization with a convolutional attentive adversarial network
Guoqiang Liang 0001, Yanbing Lv, Shucheng Li, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | Adaptive Graph Convolutional Networks for Weakly Supervised Anomaly Detection in VideosabstractFor weakly supervised anomaly detection, most existing work is limited to the problem of inadequate video representation due to the inability of modeling long-term contextual information. To solve this, we propose a novel weakly supervised adaptive graph convolutional network (WAGCN) to model the complex contextual relationship among video segments. By which, we fully consider the influence of other video segments on the current one when generating the anomaly probability score for each segment. Firstly, we combine the temporal consistency as well as feature similarity of video segments to construct a global graph, which makes full use of the association information among spatial-temporal features of anomalous events in videos. Secondly, we propose a graph learning layer in order to break the limitation of setting topology manually, which can extract graph adjacency matrix based on data adaptively and effectively. Extensive experiments on two public datasets (i.e., UCF-Crime dataset and ShanghaiTech dataset) demonstrate the effectiveness of our approach which achieves state-of-the-art performance. Congqi Cao, Xin Zhang 0168, Shizhou Zhang, Peng Wang 0015, Yanning Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Improving visible-thermal ReID with structural common space embedding and part models
Lingyan Ran, Yujun Hong, Shizhou Zhang, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2021 | Attend to the Difference: Cross-Modality Person Re-Identification via Contrastive CorrelationabstractThe problem of cross-modality person re-identification has been receiving increasing attention recently, due to its practical significance. Motivated by the fact that human usually attend to the difference when they compare two similar objects, we propose a dual-path cross-modality feature learning framework which preserves intrinsic spatial structures and attends to the difference of input cross-modality image pairs. Our framework is composed by two main components: a Dual-path Spatial-structure-preserving Common Space Network (DSCSN) and a Contrastive Correlation Network (CCN). The former embeds cross-modality images into a common 3D tensor space without losing spatial structures, while the latter extracts contrastive features by dynamically comparing input image pairs. Note that the representations generated for the input RGB and Infrared images are mutually dependant to each other. We conduct extensive experiments on two public available RGB-IR ReID datasets, SYSU-MM01 and RegDB, and our proposed method outperforms state-of-the-art algorithms by a large margin with both full and simplified evaluation modes. Shizhou Zhang, Peng Wang 0015, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Person Re-Identification in Aerial ImageryabstractNowadays, with the rapid development of consumer Unmanned Aerial Vehicles (UAVs), visual surveillance by utilizing the UAV platform has been very attractive. Most of the research works for UAV captured visual data are mainly focused on the tasks of object detection and tracking. However, limited attention has been paid to the task of person Re-identification (ReID) which has been widely studied in ordinary surveillance cameras with fixed emplacements. In this paper, to facilitate the research of person ReID in aerial imagery, we collect a large scale airborne person ReID dataset named as Person ReID in Aerial Imagery (PRAI-1581), which consists of 39,461 images of 1581 person identities. The images of the dataset are shot by two DJI consumer UAVs flying at an altitude ranging from 20 to 60 meters above the ground, which covers most of the real UAV surveillance scenarios. In addition, we propose to utilize subspace pooling of convolution feature maps to represent the input person images. Our method can learn a discriminative and compact feature representation for ReID in aerial imagery and can be trained in an end-to-end fashion efficiently. We conduct extensive experiments on the proposed dataset and the experimental results demonstrate that re-identifying persons in aerial imagery is a challenging problem, where our method performs favorably against state of the arts. Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Bingliang Jiao, Yanning Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Vehicle Re-Identification in Aerial Imagery: Dataset and ApproachabstractIn this work, we construct a large-scale dataset for vehicle re-identification (ReID), which contains 137k images of 13k vehicle instances captured by UAV-mounted cameras. To our knowledge, it is the largest UAV-based vehicle ReID dataset. To increase intra-class variation, each vehicle is captured by at least two UAVs at different locations, with diverse view-angles and flight-altitudes. We manually label a variety of vehicle attributes, including vehicle type, color, skylight, bumper, spare tire and luggage rack. Furthermore, for each vehicle image, the annotator is also required to mark the discriminative parts that helps them to distinguish this particular vehicle from others. Besides the dataset, we also design a specific vehicle ReID algorithm to make full use of the rich annotation information. It is capable of explicitly detecting discriminative parts for each specific vehicle and significantly outperforming the evaluated baselines and state-of-the-art vehicle ReID approaches. Peng Wang 0015, Bingliang Jiao, Lu Yang 0016, Shizhou Zhang, Wei Wei 0008, Yanning Zhang 0001 |
ICCV | 5 |
| 2019 | Person Re-identification with Neural Architecture Search
Shizhou Zhang, Xing Wei 0001, Peng Wang 0015, Yanning Zhang 0001 |
PRCV (1) | 1 |
| 2019 | EMS-Net: Ensemble of Multiscale Convolutional Neural Networks for Classification of Breast Cancer Histology Images
Zhanbo Yang, Lingyan Ran, Shizhou Zhang, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 3 |
| 2019 | Normalized Non-Negative Sparse Encoder for Fast Image RepresentationabstractImage representation based on sparse coding generalizes the bag of words model. Although it reduces the reconstruction error for local features to achieve the state-of-the-art image classification performance, the large computational cost hinders the application of sparse coding-based image features. In this paper, we propose approximating a sparse code using the output of a simple neural network. The resulting parameter learning model for the neural network automatically incorporates non-negative and shift-invariant constraints, leading to an efficient normalized non-negative sparse coding (N3SC) sparse encoder. Without the use of the traditional iterative process to solve the sparse coding objective, the sparse encoder directly “converts” each local feature into a sparse code. We also introduce a method for training the encoder based on the auto-encoder method. In addition, we formally propose the corresponding sparse coding scheme called N3SC, which enforces both the non-negative constraint and the shift-invariant constraint in addition to the traditional sparse coding criteria. As demonstrated by several experiments, the obtained N3SC encoder requires only 3%-10% of the processing time for image feature extraction compared with the standard sparse coding scheme. At the same time, the features extracted using the exact solutions of the N3SC coding scheme and the N3SC encoder offer superior image classification accuracy compared to the accuracy of many existing sparse coding-based representations. Shizhou Zhang, Jinjun Wang, Weiwei Shi 0003, Yihong Gong, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Pedestrian search in surveillance videos by learning discriminative deep features
Shizhou Zhang, De Cheng, Yihong Gong, Dahu Shi, Xi Qiu, Yong Xia 0001, Yanning Zhang 0001 |
Neurocomputing | 1 |
| 2018 | Person re-identification by the asymmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 4 |
| 2018 | Correction to: Person re-identification by the symmetric triplet and identification loss function
De Cheng, Yihong Gong, Weiwei Shi 0003, Shizhou Zhang |
Multim. Tools Appl. | 4 |
| 2017 | Combining local and global hypotheses in deep neural network for multi-label image classification
Qinghua Yu, Jinjun Wang, Shizhou Zhang, Yihong Gong, Jizhong Zhao |
Neurocomputing | 3 |
| 2017 | Constructing Deep Sparse Coding Network for image classification
Shizhou Zhang, Jinjun Wang, Yihong Gong, Nanning Zheng 0001 |
Pattern Recognit. | 1 |
| 2016 | Improving DCNN Performance with Sparse Category-Selective Objective Function
Shizhou Zhang, Yihong Gong, Jinjun Wang |
IJCAI | 1 |
| 2015 | Incorporating image degeneration modeling with multitask learning for image super-resolutionabstractLearning the non-linear image upscaling process has previously been considered as a simple regression process, where various models have been utilized to describe the correlations between high-resolution (HR) and low-resolution (LR) images/patches. In this paper, we present a multitask learning framework based on deep neural network for image super-resolution, where we jointly consider the image super-resolution process and the image degeneration process. By sharing parameters between the two highly relevant tasks, the proposed framework could effectively improve the obtained neural network based mapping model between HR and LR image patches. Experimental results have demonstrated clear visual improvement and high computational efficiency, especially with large magnification factors. Yudong Liang, Jinjun Wang, Shizhou Zhang, Yihong Gong |
ICIP | 3 |
| 2015 | Multi-cue Normalized Non-Negative Sparse Encoder for image classificationabstractRecently, the sparse coding based image representation has achieved state-of-the-art recognition results on many benchmarks. In this paper, we propose Multi-cue Normalized Non-Negative Sparse Encoder (MN3SE) which enforces both the non-negative constraint and the shift-invariant constraint on top of the traditional sparse coding criteria, and takes multi-cue to further boost the performance. The former constraint reduces information loose by the negative coefficients and improves the coding stability, and the latter allows the sparseness to be self-adaptive to the local feature. The proposed coding scheme is then approximated by an neural network based encoder for speed-up. More importantly, the multi-layer neural network architecture allows us to apply a multi-task learning strategy to fuse information from multi-cue. Specifically, we take one type of descriptor, such as SIFT as the input, and enforce the learned encoder to produce sparse code that can reconstruct not only SIFT but also other types of descriptors such as color moments. In this way, we could achieve not only 10 to 33 times speed up for sparse-coding, the multi-cue enforced learning strategy gives the image feature extracted by MN3SE superior image classification accuracy. Shizhou Zhang, Jinjun Wang, Yudong Liang, Yihong Gong, Nanning Zheng 0001 |
ICME | 1 |
| 2014 | Low Computation Face Verification Using Class Center AnalysisabstractDespite the existence of many state-of-the-art face verification systems, the use of complex features and/or high order recognition models in these systems limits their application in devices with low computation power or low latency requirement. In this paper, we approach the problem by performing verification using simple linear distance model. We introduce a novel probability-based distance metric learning algorithm called Class Center Analysis (CCA) to improve the matching performance in a transformed space. CCA generalizes the classic Neighborhood Components Analysis (NCA) from two aspects. First NCA often leads to distributed clusters, while CCA produces more concentrative clusters, And second, NCA sometimes gives over-fitted distance transformation model, while CCA has better generalization ability. With CCA, our system is able to directly project the difference between face image pair into a real-valued score as their similarity, using only simple matrix-vector operation, and thus consuming very low computation. Our comprehensive experimental evaluation show that, CCA outperforms several other benchmark algorithms in verification accuracy. We have also built the CCA algorithm into a mobile application that uses face image for user authentication. Xinzi Zhang, Jinjun Wang, Yihong Gong, Shizhou Zhang |
ICPR | 4 |
| 2014 | Image parsing by loopy dynamic programming
Shizhou Zhang, Jinjun Wang, Yihong Gong, Xinzi Zhang, Xuguang Lan |
Neurocomputing | 1 |