VLDB 2026 Research / reviewers in the wild / expert
Cairong Zhao
dblp:81/8614
· DBLP profile ↗
108ranked-venue papers
25as first author
81since 2021 · last 2026
0000-0001-6745-9674ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 13 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 10 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Long-Context Summarization with Multi-Granularity Retrieval OptimizationabstractRetrieval-Augmented Generation (RAG) is an effective solution to overcome the limitations of Large Language Models (LLMs) in terms of specific-domain knowledge and timely information updates. However, current RAG methods typically respond to queries based on isolated segments, lacking the ability to integrate information within the same document. This undermines performance in real-world tasks requiring coherent understanding across an entire document. Notably, the human brain naturally integrates and summarizes prior knowledge upon reading a given text, progressively formulating a comprehensive understanding. Motivated by this cognitive process, we propose the Hierarchical Two-Stage Summarization-based Information Retrieval (HTSIR) method, which preprocesses the corpus prior to retrieval, summarizes continuous texts to obtain integrated information, and constructs a retrieval tree with varying summary granularities. The retrieved information is then processed by a Reranker based on the current question to serve as a context for LLMs. Additionally, as single-step summarization is often imprecise in query-based summarization tasks, we further apply a Refinement module, allowing LLMs to reflect and revise their output to achieve the final result. By combining HTSIR with GPT-4o mini, we achieve state-of-the-art results on complex question tasks across four long-text datasets (NarrativeQA, QASPER, QuALITY, and QMSum), achieving an improvement of about 6 points on the Question Answering (QA) task in QuALITY-HRAD. Xueyu Chen, Kaitao Song, Zifan Song, Dongsheng Li 0002, Cairong Zhao |
AAAI | 5 |
| 2026 | Dual-Phase Visual-Language Pretraining and Adaptation for Long-Tailed Multi-Label RecognitionabstractLong-Tailed Multi-Label Recognition (LTML) is a critical yet challenging task due to two core issues: the severe scarcity of training samples for rare "tail" classes, and the complex co-occurrence patterns among labels that often lead to biased models. To address this, we propose DP-VLPA, a novel Dual-Phase Visual-Language Pretraining and Adaptation framework. In the first phase, our Structured Tail-Aware Generation (STAG) module employs a Large Language Model (LLM) to create detailed descriptions that explicitly emphasize tail classes and their contextual relationships, providing a strong and less-biased feature foundation. In the second adaptation phase, we ensure this knowledge is applied effectively. A Dynamic Query Reweighting (DQR) mechanism forces the model to attend to crucial tail-class evidence. Simultaneously, a Co-occurrence-Aware (COA) loss explicitly teaches the model the statistical dependencies between labels, correcting for co-occurrence biases. Extensive experiments on VOC-LT and COCO-LT datasets demonstrate state-of-the-art performance, achieving mAP scores of 90.72% and 74.42% respectively - surpassing previous best methods by 2.84% and 8.23%. Xuekuan Wang, Cairong Zhao |
AAAI | 4 |
| 2026 | Tuning Medical Foundation Models for Inner Ear Temporal CT Analysis with Plug-and-play Domain Knowledge Aggregator
Weixun Wan, Xinyang Jiang, Zilong Wang 0006, Cairong Zhao |
AAAI | 5 |
| 2026 | Person identity shift for privacy-preserving person re-identification
Shuguang Dou, Xinyang Jiang, Yansen Wang, Dongsheng Li 0002, Cairong Zhao |
Sci. China Inf. Sci. | 6 |
| 2026 | Frequency-aware and lifting-based efficient transformer for person search
Qilin Shu, Qixian Zhang, Duoqian Miao 0001, Qi Zhang 0020, Hongyun Zhang 0001, Cairong Zhao |
Expert Syst. Appl. | 6 |
| 2026 | DiffPano++: Scalable and Consistent Multi-View Panorama Generation with Spherical Epipolar-Aware Diffusion
Chenhao Ji, Weicai Ye, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Xuekuan Wang, Guofeng Zhang 0001, Song-Hai Zhang, Tong He 0001, Wanli Ouyang, Cairong Zhao |
Int. J. Comput. Vis. | 11 |
| 2026 | StyleShot: A Snapshot on Any StyleabstractImage Style Transfer aims to replicate the style of a reference image based on the content from a text description or another image. With the significant advancements in image generation through diffusion models, recent studies have attempted to either fine-tuning embeddings to learn the single style or utilizing the pre-trained CLIP image encoder to extract style representations. However, style-tuning requires substantial computational resources and the pre-trained CLIP image encoder is trained for semantic understanding rather than for style representation. To address these challenges, we introduce a style-aware encoder and a well-organized style dataset called StyleGallery to learn a good style representation that is crucial and sufficient for generalized style transfer without test-time tuning. With dedicated design for style learning, this style-aware encoder is trained to extract expressive style representation from multi-level patches with decoupling training strategy, and StyleGallery enables the generalization ability. Moreover, we employ a content extraction and content-fusion encoder to enhance image-driven style transfer. We highlight that, our approach, named StyleShot, is simple yet effective in mimicking various desired styles, i.e., 3D, flat, abstract or even fine-grained styles, without test-time tuning. Rigorous experiments validate that, StyleShot achieves superior performance across a wide range of styles compared to existing state-of-the-art text- and image-driven methods. Junyao Gao 0002, Yanan Sun 0005, Yinhao Tang, Yanhong Zeng, Ding Qi, Kai Chen 0026, Cairong Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Structural feature enhanced transformer for fine-grained image recognition
Cairong Zhao, Enhong Chen |
Pattern Recognit. | 3 |
| 2026 | Integrating Disparity Confidence Estimation Into Relative Depth Prior-Guided Unsupervised Stereo MatchingabstractUnsupervised stereo matching has garnered significant attention for its independence from costly disparity annotations. Typical unsupervised methods rely on the multi-view consistency assumption for training networks, which suffer considerably from stereo matching ambiguities, such as repetitive patterns and texture-less regions. A feasible solution lies in transferring 3D geometric knowledge from a relative depth map to the stereo matching networks. However, existing knowledge transfer methods learn depth ranking information from randomly built sparse correspondences, which makes inefficient utilization of 3D geometric knowledge and introduces noise from mistaken disparity estimates. This work proposes a novel unsupervised learning framework to address these challenges, which comprises a plug-and-play disparity confidence estimation algorithm and two depth prior-guided loss functions. Specifically, the local coherence consistency between neighboring disparities and their corresponding relative depths is first checked to obtain disparity confidence. Afterwards, quasi-dense correspondences are built using only confident disparity estimates to facilitate efficient depth ranking learning. Finally, a dual disparity smoothness loss is proposed to boost stereo matching performance at disparity discontinuities. Experimental results demonstrate that our method achieves state-of-the-art stereo matching accuracy on the KITTI Stereo benchmarks among all unsupervised stereo matching methods. Mingjian Sun, Cairong Zhao, Hanli Wang, Alexander V. Dvorkovich, Rui Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | UD-Gaussian: Uncertainty-Driven Gaussian Modeling for Occluded Person Re-IdentificationabstractOccluded person re-identification aims to address the identification challenges posed by pedestrians obscured by other individuals or objects. Existing methods often rely on incorporating pose or semantic information to improve model performance under occlusion. However, such information often depends on external models with inevitably cross-domain gaps, whose stability is limited in complex occlusion environments and prone to false results. In this paper, we propose a Transformer-based uncertainty-driven Gaussian model, termed as UD-Gaussian. Firstly, to enrich the detailed features of pedestrian images, a high-frequency enhancement module is introduced. The high-frequency components of the pedestrian image are extracted by Discrete Haar Wavelet Transform, and Top-K high-frequency patches are extracted to construct a graph Laplacian matrix to achieve high-frequency graph attention, which is fused with features learned from self-attention to enhance the high-frequency feature representation. Given the uncertainty in pedestrian feature learning induced by occlusion makes it challenging to obtain reliable and stable pedestrian features, we propose a probability distribution learning module. This module establishes a memory bank to build Gaussian distributions for each pedestrian identity and the entropy is introduced as a loss function to encourage the model to generate more deterministic and relatively independent probability distributions, thereby enhancing the discriminative ability of the model across different pedestrian identities. The high-frequency enhancement module provides a solid foundation for the probability distribution learning module, alleviating uncertainty caused by pedestrian images themselves. Experimental results on occluded and holistic person re-identification datasets demonstrate the superiority of the proposed method. Yizhang Liu, Hongyun Zhang 0001, Cairong Zhao, Zhihua Wei 0001, Duoqian Miao 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Active Dataset Distillation via Dual-Space Informative MatchingabstractDataset distillation improves neural network training efficiency by compressing large real datasets into compact synthetic datasets. Existing methods typically optimize matching objectives, such as aligning gradients, features, and trajectories between the synthetic and original datasets to ensure the distilled data retains essential properties for model training. However, many of these approaches rely on predefined distillation pools to streamline the process or treat all real data points equally, overlooking the dynamic nature of the synthetic dataset's training requirements during optimization. To address these limitations, we propose Active Dataset Distillation via Dual-Space Informative Matching (ACDD), an active learning-based algorithm that dynamically selects the most informative real data subset to align with the synthetic dataset's evolving needs. By adaptively refining the distillation pool, ACDD enhances training efficiency and generalization while ensuring the synthetic dataset effectively captures the original data's key characteristics. ACDD operates through two interconnected loops: the dual-space active loop (DAL) and the distillation loop. DAL plays a key role by dynamically selecting samples that balance diversity and uncertainty, adding them to the target distillation pool to meet the evolving informational needs of the current distillation loop. As a result, ACDD enables the synthetic dataset to achieve superior performance compared to SOTA methods across multiple benchmarks, including SVHN, CIFAR-10, CIFAR-100, TinyImageNet, and ImageNet subset. Moreover, ACDD reduces the required real dataset to just 20%-40% of the original, demonstrating its efficiency and effectiveness in data distillation. Ding Qi, Jian Li 0062, Shuguang Dou, Junyao Gao 0002, Yabiao Wang, Bo Zhao 0015, Cairong Zhao |
IEEE Trans. Image Process. | 7 |
| 2026 | Mask-Guided Asymmetric Contrastive and Semantic Alignment for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (ReID) aims to learn identity-discriminative representations without manual annotations, which is challenging due to noisy pseudo labels, background clutter, and large appearance variations. Recent studies have shown that exploiting fine-grained local cues is crucial for improving robustness in unsupervised ReID. In this context, random masking has emerged as a simple and annotation-free way to encourage the model to focus on informative regions. However, existing masking-based unsupervised ReID methods still suffer from two limitations: (1) Underused masked views: masked views are treated as degraded auxiliaries rather than exploited as fine-grained supervisory signals; (2) Weak cross-view alignment: feature alignment is restricted to mini-batch pairs, lacking explicit global alignment between masked and unmasked views across clusters. To address these issues, we propose the Mask-guided Asymmetric Contrastive and Semantic Alignment (ACSA) framework. Specifically, we introduce an Asymmetric Contrastive Learning (ACL) module with a dual-memory mechanism to separately encode masked and unmasked features, allowing masked views to serve as informative and discriminative supervision. In parallel, a Semantic Alignment Learning (SAL) module conducts multi-granularity distribution alignment by aligning both cluster-level prototypes and randomly sampled instance-level features, thereby preserving semantic consistency and intra-cluster diversity. Furthermore, to provide more reliable semantic anchors for SAL under noisy pseudo labels, we introduce a Progressive Refinement Module (PRM), which refines prototypes and features via exponential moving averaging for more stable semantic alignment. Extensive experiments validate the superiority of our method, even outperforming certain supervised counterparts. Code is available at https://github.com/Trangle12/ACSA. Ruijian Wei, Qixian Zhang, Ding Qi, Duoqian Miao 0001, Cairong Zhao |
IEEE Trans. Image Process. | 6 |
| 2026 | IRPP: Invariant Representation Learning With Progressive Prototype Refinement for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (USL-ReID) typically relies on clustering to generate pseudo-labels, but significant cross-view appearance variations often cause images of the same identity to be split into different clusters. Training on such noisy pseudo-labels severely degrades the learned representations. Therefore, learning robust view-invariant features is paramount. Data augmentation provides a direct way to enhance invariance, yet its trade-offs in USL-ReID remain under-explored: weak augmentations usually preserve identity semantics but lack diversity, whereas strong augmentations provide richer appearance diversity at the cost of partially corrupting identity-consistent semantic cues. To address this challenge, we propose Invariant Representation learning with Progressive Prototype Refinement (IRPP), a unified framework that learns invariant and discriminative features from noisy pseudo-labels. IRPP consists of three synergistic components. First, an Augmented Dual-Contrastive Learning (ADCL) module performs dataset-level prototype-guided invariant learning by contrasting weakly and strongly augmented views against cluster-derived prototypes. Second, an Alignment and Uniformity Learning (AUL) module regularizes the mini-batch-level weak-strong feature geometry, leading to more stable feature distributions under data augmentation. Third, a Progressive Prototype Refinement (PPR) mechanism progressively optimizes cluster centroids into cleaner prototypes, thereby mitigating the influence of noisy pseudo-labels and further strengthening invariant representation learning. This closed-loop design enables prototype-guided contrastive learning, weak-strong regularization, and prototype refinement to mutually reinforce each other. Extensive experiments on standard USL-ReID benchmarks demonstrate that IRPP achieves state-of-the-art performance with a simple and efficient training pipeline. Code is available at https://github.com/Trangle12/IRPP. Qixian Zhang, Ding Qi, Duoqian Miao 0001, Shiping Wang, Cairong Zhao |
IEEE Trans. Image Process. | 6 |
| 2026 | ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal GroundingabstractVideo temporal grounding, including moment retrieval and highlight detection, is an emerging topic aiming to identify specific clips within videos. In addition to pre-trained video models, contemporary methods utilize pre-trained vision-language models (VLMs) to capture detailed characteristics of diverse scenes and objects from video frames. However, as pre-trained on images, directly using pre-extracted VLM features neglects the domain gap between the pre-trained and temporal grounding datasets, thus inducing domain shifts due to the data-level distribution disparity. As a result, VLMs may struggle to distinguish action-sensitive patterns from static objects, making it necessary to adapt them to specific data domains for effective feature representation over temporal grounding. In this work, we address two primary challenges to achieve this goal. Specifically, to mitigate high adaptation costs, we propose an efficient preliminary in-domain fine-tuning paradigm for feature adaptation before standard downstream training, where downstream-adaptive features are learned through several well-designed pretext tasks that ensure improved performance. Furthermore, to integrate action-sensitive information into VLMs, we introduce Action-Cue-Injected Temporal Prompt Learning (ActPrompt), which injects action cues into the image encoder of VLMs to discover action-sensitive visual patterns better. This is followed by context-aware temporal prompt learning, which considers both action cues and temporal context to enhance the ability to recognize patterns associated with actions for downstream tasks. Extensive experiments demonstrate that ActPrompt is an off-the-shelf training framework that can be applied effectively to various SOTA methods, resulting in notable improvements. Xinyang Jiang, De Cheng, Dongsheng Li 0002, Cairong Zhao |
IEEE Trans. Image Process. | 5 |
| 2026 | ASDTracker: Adaptively Sparse Detection With Attention-Guided Refinement for Efficient Multi-Object TrackingabstractTracking-by-Detection paradigms shine in generic multi-object tracking (MOT), while their compact construction hinders the real-time applications. In this work, we attribute the substantial computational burden to two expensive components, i.e. detection and re-identification. Building upon the principle of adaptively maintaining acceptable inference efficiency, we present Adaptively Sparse Detection with attention-guided refinement (ASDTracker) for efficient tracking. In specific, our ASDTracker rapidly assess the short-term and long-term occlusion, dynamically determining the usage of the expensive detector. For non-key frames, we efficiently refine small-size crops out of Kalman Filter predictions and introduce the noisy shadow labels to robustly train this refinement network. Additionally, we substitute the lightweight appearance representation for the heavy ReID network, which efficiently extracts sufficient appearance cues in the coarsely quantized color spaces. Extensive experiments on four benchmarks demonstrate that ASDTracker achieves competitive performance in generalization and robustness under favorable inference speed. Moreover, the efficient tracking deployment is further implemented to an unmanned surface vehicle with high accuracy and low latency in real-world scenarios. Yueying Wang, Chenyang Yan, Cairong Zhao, Weidong Zhang 0004, Dan Zeng 0001 |
IEEE Trans. Image Process. | 3 |
| 2026 | Reliable Pseudo-Supervision for Unsupervised Domain Adaptive Person SearchabstractUnsupervised Domain Adaptation (UDA) person search aims to adapt models trained on labeled source data to unlabeled target domains. Existing approaches typically rely on clustering-based proxy learning, but their performance is often undermined by unreliable pseudo-supervision. This unreliability mainly stems from two challenges: (i) spectral shift bias, where low- and high-frequency components behave differently under domain shifts but are rarely considered, degrading feature stability; and (ii) static proxy updates, which make clustering proxies highly sensitive to noise and less adaptable to domain shifts. To address these challenges, we propose the Reliable Pseudo-supervision in UDA Person Search (RPPS) framework. At the feature level, a Dual-branch Wavelet Enhancement Module (DWEM) embedded in the backbone applies discrete wavelet transform (DWT) to decompose features into low- and high-frequency components, followed by differentiated enhancements that improve cross-domain robustness and discriminability. At the proxy level, a Dynamic Confidence-weighted Clustering Proxy (DCCP) employs confidence-guided initialization and a two-stage online-offline update strategy to stabilize proxy optimization and suppress proxy noise. Extensive experiments on the CUHK-SYSU and PRW benchmarks demonstrate that RPPS achieves state-of-the-art performance and strong robustness, underscoring the importance of enhancing pseudo-supervision reliability in UDA person search. Our code is accessible at https://github.com/zqx951102/RPPS. Qixian Zhang, Duoqian Miao 0001, Qi Zhang 0020, Hongyun Zhang 0001, Cairong Zhao |
IEEE Trans. Image Process. | 6 |
| 2026 | SC-DETR: A Text-Guided Small Object Detection via Scale Prompting and Centerpoint Localization
Mingzhu Li, Xuekuan Wang, Cairong Zhao |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Towards Universal Dataset Distillation via Task-Driven DiffusionabstractDataset distillation (DD) condenses key information from large-scale datasets into smaller synthetic datasets, reducing storage and computational costs for training networks. However, most recent research has primarily focused on image classification tasks, with limited exploration in detection and segmentation. Two key challenges remain: (i) Task Optimization Heterogeneity, where existing methods focus on class-level information but fail to address the diverse needs of detection and segmentation, and (ii) Inflexible Image Generation, where current generation methods rely on global updates for single-class targets and lack localized optimization for specific object regions. To address these challenges, we propose UniDD, a universal dataset distillation framework built on a task-driven diffusion model for diverse DD tasks, as shown in Fig. 1. Our approach operates in two stages: Universal Task Knowledge Mining, which captures task-relevant information through task-specific proxy model training, and Universal Task-Driven Diffusion, where these proxies guide the diffusion process to generate task-specific synthetic images. Extensive experiments across ImageNet-1K, Pascal VOC, and MS COCO demonstrate that UniDD consistently outperforms state-of-the-art methods. In particular, on ImageNet-1K with IPC-10, UniDD surpasses previous diffusion-based methods by 6.1%, while also reducing deployment costs. Ding Qi, Jian Li 0062, Junyao Gao 0002, Shuguang Dou, Ying Tai, Jianlong Hu, Bo Zhao 0015, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 10 |
| 2025 | One Object, Multiple Lies: A Benchmark for Cross-Task Adversarial Attack on Unified Vision-Language ModelsabstractUnified vision-language models(VLMs) have recently shown remarkable progress, enabling a single model to flexibly address diverse tasks through different instructions within a shared computational architecture. This instruction-based control mechanism creates unique security challenges, as adversarial inputs must remain effective across multiple task instructions that may be unpredictably applied to process the same malicious content. In this paper, we introduce CrossVLAD, a new benchmark dataset carefully curated from MSCOCO with GPT-4-assisted annotations for systematically evaluating cross-task adversarial attacks on unified VLMs. CrossVLAD centers on the object-change objective-consistently manipulating a target object's classification across four downstream tasks-and proposes a novel success rate metric that measures simultaneous misclassification across all tasks, providing a rigorous evaluation of adversarial transferability. To tackle this challenge, we present CRAFT (Cross-task Region-based Attack Framework with Token-alignment), an efficient region-centric attack method. Extensive experiments on Florence-2 and other popular unified VLMs demonstrate that our method outperforms existing approaches in both overall cross-task attack performance and targeted object-change success rates, highlighting its effectiveness in adversarially influencing unified VLMs across diverse tasks. Xinyang Jiang, Junyao Gao 0002, Yuhao Xue, Cairong Zhao |
ICCV | 5 |
| 2025 | Domain Generalizable Portrait Style Transfer
Xinyang Jiang, Junyao Gao 0002, Yuhao Xue, Cairong Zhao |
ICCV | 5 |
| 2025 | FaceShot: Bring Any Character into LifeabstractIn this paper, we present ***FaceShot***, a novel training-free portrait animation framework designed to bring any character into life from any driven video without fine-tuning or retraining.
We achieve this by offering precise and robust reposed landmark sequences from an appearance-guided landmark matching module and a coordinate-based landmark retargeting module.
Together, these components harness the robust semantic correspondences of latent diffusion models to produce facial motion sequence across a wide range of character types.
After that, we input the landmark sequences into a pre-trained landmark-driven animation model to generate animated video.
With this powerful generalization capability, FaceShot can significantly extend the application of portrait animation by breaking the limitation of realistic portrait landmark detection for any stylized character and driven video.
Also, FaceShot is compatible with any landmark-driven animation model, significantly improving overall performance.
Extensive experiments on our newly constructed character benchmark CharacBench confirm that FaceShot consistently surpasses state-of-the-art (SOTA) approaches across any character domain.
More results are available at our project website https://faceshot2024.github.io/faceshot/. Junyao Gao 0002, Yanan Sun 0005, Fei Shen 0004, Xin Jiang 0010, Zhening Xing, Kai Chen 0026, Cairong Zhao |
ICLR | 7 |
| 2025 | Uni2Det: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D DetectionabstractWe present Uni$^2$Det, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization to unseen domains. Due to substantial disparities in data distribution and variations in taxonomy across diverse domains, training such a detector by simply merging datasets poses a significant challenge. Motivated by this observation, we introduce multi-stage prompting modules for multi-dataset 3D detection, which leverages prompts based on the characteristics of corresponding datasets to mitigate existing differences. This elegant design facilitates seamless plug-and-play integration within various advanced 3D detection frameworks in a unified manner, while also allowing straightforward adaptation for universal applicability across datasets. Experiments are conducted across multiple dataset consolidation scenarios involving KITTI, Waymo, and nuScenes, demonstrating that our Uni$^2$Det outperforms existing methods by a large margin in multi-dataset training. Notably, results on zero-shot cross-dataset transfer validate the generalization capability of our proposed method. Our code is available at https://github.com/ThomasWangY/Uni2Det. Zhikang Zou, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Cairong Zhao |
ICLR | 6 |
| 2025 | A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual GroundingabstractOpen-vocabulary 3D visual grounding aims to localize target objects based on free-form language queries, which is crucial for embodied AI applications such as autonomous navigation, robotics, and augmented reality. Learning 3D language fields through neural representations enables accurate understanding of 3D scenes from limited viewpoints and facilitates the localization of target objects in complex environments. However, existing language field methods struggle to accurately localize instances using spatial relations in language queries, such as ''the book on the chair.'' This limitation mainly arises from inadequate reasoning about spatial relations in both language queries and 3D scenes. In this work, we propose SpatialReasoner, a novel neural representation-based framework with large language model (LLM)-driven spatial reasoning that constructs a visual properties-enhanced hierarchical feature field for open-vocabulary 3D visual grounding. To enable spatial reasoning in language queries, SpatialReasoner fine-tunes an LLM to capture spatial relations and explicitly infer instructions for the target, anchor, and spatial relation. To enable spatial reasoning in 3D scenes, SpatialReasoner incorporates visual properties (opacity and color) to construct a hierarchical feature field. This field represents language and instance features using distilled CLIP features and masks extracted via the Segment Anything Model (SAM). The field is then queried using the inferred instructions in a hierarchical manner to localize the target 3D instance based on the spatial relation in the language query. Notably, SpatialReasoner is not limited to a specific 3D neural representation; it serves as a framework adaptable to various representations, such as Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS). Extensive experiments show that our framework can be seamlessly integrated into different neural representations, outperforming baseline models in 3D visual grounding while empowering their spatial reasoning capability. Project Homepage:ZhenyangLiu.github.io/SpatialReasoner. Zhenyang Liu, Sixiao Zheng, Siyu Chen 0023, Cairong Zhao, Longfei Liang, Xiangyang Xue 0001, Yanwei Fu 0001 |
ACM Multimedia | 4 |
| 2025 | Boosting Adversarial Transferability via Commonality-Oriented Gradient Optimization
Yanting Gao, Qi Zhang 0020, Hongyun Zhang 0001, Duoqian Miao 0001, Cairong Zhao |
PRCV (2) | 7 |
| 2025 | Full-Lifecycle Data Governance for Embodied Intelligence
Chuanhou Liu, Ding Qi, Cairong Zhao |
PRCV (18) | 3 |
| 2025 | CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware DiffusionabstractRecently, camera-controlled video generation has seen rapid development, offering more precise control over video generation. However, existing methods predominantly focus on camera control in perspective projection video generation, while geometrically consistent panoramic video generation remains challenging. This limitation is primarily due to the inherent complexities in panoramic pose representation and spherical projection. To address this issue, we propose CamPVG, the first diffusion-based framework for panoramic video generation guided by precise camera poses. We achieve camera position encoding for panoramic images and cross-view feature aggregation based on spherical projection. Specifically, we propose a panoramic Plücker embedding that encodes camera extrinsic parameters through spherical coordinate transformation. This pose encoder effectively captures panoramic geometry, overcoming the limitations of traditional methods when applied to equirectangular projections. Additionally, we introduce a spherical epipolar module that enforces geometric constraints through adaptive attention masking along epipolar lines. This module enables fine-grained cross-view feature aggregation, substantially enhancing the quality and consistency of generated panoramic videos. Extensive experiments demonstrate that our method generates high-quality panoramic videos consistent with camera trajectories, far surpassing existing methods in panoramic video generation. Chenhao Ji, Chaohui Yu, Junyao Gao 0002, Fan Wang 0019, Cairong Zhao |
SIGGRAPH Asia | 5 |
| 2025 | Fusion4DAL: Offline Multi-modal 3D Object Detection for 4D Auto-labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jincheng Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao |
Int. J. Comput. Vis. | 8 |
| 2025 | Identity aware 3D face reconstruction from in-the-wild images
Ruigang Hu, Xuekuan Wang, Cairong Zhao |
Neurocomputing | 3 |
| 2025 | Multi-clues Adaptive Learning for Cloth-Changing Person Re-IdentificationabstractSolving long-term Cloth-Changing Person Re-identification (CC-ReID) requires extracting features insensitive to clothing such as face, silhouette, gait and pose estimation. Most current work focuses on modeling from a single feature, but we observe that CC-ReID problems in open environments are often difficult to solve solely based on a single feature, for instance, sometimes, contour features may be advantageous for recognition, while at other times gait features may be more valuable. In our paper, we suggest a novel multi-clues guided Adaptive Learning Transformer (ALT) which can adaptively select the most readily identifiable features based on different scenarios. The method comprises two parts: a Multi-clues Guiding Module (MGM) and a Feature Selection Module (FSM). We utilize clothes-irrelevant features from multi-modality information as clues, integrating multiple features to extract robust representations invariant to clothing changes for CC-ReID through cross-attention and Mixture of Experts (MoEs). We utilized contour sketch and gait as clues, conducting experiments on the CC-ReID dataset. The experimental results show that our recommended approach prevails over all other SOTA methods, particularly showing significant improvement compared to using contour sketch and gait alone. Xiang Zhou 0006, Junzhu Liu, Xinyang Jiang, Cairong Zhao |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2025 | DPL++: Advancing the Network Performance via Image and Label PerturbationsabstractRecent advances in supervised learning have predominantly focused on regularizations, optimizers, and architectures, yet the potential of simultaneously optimizing data distributions and supervisory signals for training samples remains underexplored. In this paper, we propose a novel paradigm that leverages the benefits of image perturbations for rectifying data distributions. Our method, called DPL (Deep Perturbation Learning), introduces new insights into utilizing image perturbations and focuses on improving generalizability on normal samples, rather than resisting adversarial attacks. DPL formulates a differentiable function w.r.t. image perturbations and implements an alternative optimization process that seamlessly integrates with downstream tasks. However, the limitations of DPL stem from the inefficiency in employing differentiable targets caused by the exclusive optimization of image perturbations, while neglecting the critical role of supervisory signals in training effectiveness. These lead to the excessive necessity of DPL iterations and yield inferior performance-cost trade-off. To track this, we extend DPL to DPL++ with synchronous optimization for image perturbations and label perturbations. In our DPL++ paradigm, the post-hoc application of perturbations to images and labels endows amendments toward both data distributions and supervisory signals, significantly furthering the generalizability of models over various benchmarks. Crucially, the proposed synchronous optimization process shares key differentiable objectives to reduce computational complexity, thereby achieving enhanced effectiveness within fewer optimization iterations. Theoretically, as a generic and flexible approach, DPL++ can be applied to a variety of backbone architectures (e.g., ResNet, DenseNet, and ViT) and downstream tasks (e.g., image classification and object detection). To validate the efficacy of DPL++, we conduct extensive performance experiments and in-depth analytical studies on 2 visual tasks over 5 mainstream benchmarks across 13 backbone networks. The comprehensive results verify the superiority of DPL++ over DPL and demonstrate its promising capabilities for advancing decision-making capacity, risk minimization, class distinguishability, and training convergence. Zifan Song, Guosheng Hu, Shuguang Dou, Cairong Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | EA-HAS-Bench and Language-Enhanced Shrinkage Search for Energy-Aware NASabstractThis paper takes a crucial step in the development of energy-aware (EA) NAS methods by offering a benchmark that enhances the reproducibility and accessibility of EA-NAS research. Specifically, we introduce EA-HAS-Bench, the first large-scale energy-aware benchmark designed to enable the study of AutoML methods in achieving improved trade-offs between performance and search energy consumption. EA-HAS-Bench offers a vast architecture/hyperparameter joint search space, encompassing diverse configurations relevant to energy consumption, and proposes a novel surrogate model based on Bézier curves for predicting learning curves with versatile shapes and lengths. On the other hand, recent studies have started integrating large language models (LLMs) into AutoML frameworks to enhance model search efficiency and configuration prediction, yet challenges remain in adapting these methods for energy-efficient searches across vast configuration spaces, as they often neglect energy consumption metrics. As a result, we introduce the Language-Enhanced Shrinkage Search (LESS), a plug-and-play method that utilizes the analytical capabilities of LLMs to enhance the energy efficiency of existing hyperparameter optimization techniques. Moreover, we adapt existing AutoML algorithms to construct baselines. Our experiments demonstrate that these modified energy-aware AutoML methods and LESS achieve an improved balance between energy consumption and model performance. Cairong Zhao, Shuguang Dou, Xinyang Jiang, Junyao Gao 0002, Yuge Zhang, Bo Li 0080, Dongsheng Li 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Explainability-based knowledge distillation
Tianli Sun, Haonan Chen 0003, Guosheng Hu, Cairong Zhao |
Pattern Recognit. | 4 |
| 2025 | Multi-definition Deepfake detection via semantics reduction and cross-domain training
Cairong Zhao, Chutian Wang, Zifan Song, Guosheng Hu, Duoqian Miao 0001 |
Pattern Recognit. | 1 |
| 2025 | Learning Label PerturbationsabstractSupervised learning typically uses hard labels for annotations, which may not fully capture the underlying distribution of the data. In the literature, label smoothing is a method that can reduce overconfidence and enhance the model generalization by using a weighted average of one-hot vectors and the uniform distribution, but it does not extract the intrinsic information from the data. Another approach is knowledge distillation, which uses the predicted probability distribution of a teacher network trained with hard labels as soft labels for a student network. However, this method lacks a theoretical explanation. In this work, we draw inspiration from the influence function and propose a post-hoc label perturbation learning method called Deep Soft Label Learning (DSLL). This method iteratively leverages the inherent information present in both the model and data to theoretically determine optimal labels for classification and regression problems. Our experiments demonstrate that DSLL consistently enhances model performance across various tasks, including image classification and object detection. Zifan Song, Guosheng Hu, Cairong Zhao |
IEEE Signal Process. Lett. | 4 |
| 2025 | Learning Discriminative Representations in Videos via Active Embedding Distance CorrelationabstractIn action recognition, models often suffer from representation bias, focusing too much on background context rather than the action itself, which limits their ability to generalize. Existing methods suggest that incorporating differential inputs and utilizing dual-path structural designs could separate spatial and temporal representations. However, these approaches still rely on spatial hints and struggle to capture fine-grained temporal features. We propose a novel regularization technique, called Active Embedding Distance Correlation (AEDC), which is integrated into dual-path networks. AEDC minimizes the distance correlation between temporal and spatial embeddings, enabling spatially and temporally independent modeling. Our experiments show AEDC improves performance by 0.6% on SSV2 and 2.4% on TA50 compared to existing dual-path baselines. Ablation studies confirm that AEDC reduces scene bias and boosts robustness against video input variations. Yi Wang 0033, Yinan He, Yu Qiao 0001, Cairong Zhao |
IEEE Signal Process. Lett. | 5 |
| 2025 | TGAvatar: Reconstructing 3D Gaussian Avatars With Transformer-Based Tri-PlaneabstractWe introduce TGAvatar, a novel framework for 3D head animation and reconstruction that revolutionizes the use of 3D Gaussian Splatting (3DGS). TGAvatar significantly advances rendering quality by leveraging the intricate properties of 3DGS to achieve detailed and realistic representations of human head geometries and textures. We use an innovative application of linear blending techniques to imitate 3D Morphable Model (3DMM) coefficients within 3DGS, thereby enabling precise and dynamic facial feature and expression modeling. Further enhancing TGAvatar’s capabilities, a transformer based tri-plane module is incorporated to accurately infer spherical harmonics and alpha parameters. This integration is pivotal for the method, as it allows allows us to efficiently and precisely represent the visual characteristics of gaussians, tailored specifically to the intricate details of the head’s components. Our exhaustive evaluations show that TGAvatar not only elevates the fidelity and realism of 3D head reconstructions but also sets a new standard by surpassing existing methods in rendering quality and computational efficiency. Please see our project page athttps://hrg0417.github.io/TGAvatar/ Ruigang Hu, Xuekuan Wang, Yichao Yan, Cairong Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Scene Text Image Super-Resolution Via Semantic Distillation and Text Perceptual LossabstractText Super-Resolution (SR) technology aims to recover lost information in low-resolution text images. With the proposal of TextZoom, which is the first dataset aiming at text super-resolution in real scenes, more and more scene text super-resolution models have been presented on the basis of it. Although these methods have achieved excellent performance, they do not consider how to make full and efficient use of semantic information. Out of this consideration, a Semantic-aware Trident Network (STNet) for Scene Text Image Super-Resolution is proposed. Specifically, pre-trained text recognition model ASTER (Attentional Scene Text Recognizer) is utilized to assist this process in two ways. Firstly, a novel basic block named Semantic-aware Trident Block (STB) is designed to build the STNet, which incorporates an added branch for semantic distillation to learn semantic information of pre-trained recognition model. Secondly, we expand our model in an adversarial training manner and propose new text perceptual loss based on ASTER to further enhance semantic information in SR images. Extensive experiments on TextZoom dataset show that compared with directly recognizing bicubic images, the proposed STNet boosts the recognition accuracy of ASTER, MORAN (Multi-Object Rectified Attention Network), and CRNN (Convolutional Recurrent Neural Network) by 17.4%, 18.2%, and 24.3%, respectively, which is higher than the performance of several existing state-of-the-art (SOTA) SR network models. Besides, experiments in real scenes (on ICDAR 2015 dataset) and in restricted scenarios (defense against adversarial attacks) validate that addition of semantic information enables the proposed method to achieve promising cross-dataset performance. Since the proposed method is trained on cropped images, when applied to real-world scenarios, locations of text in natural images are firstly localized through scene text detection methods, and then cropped text images are obtained based on detected text positions. Cairong Zhao, Shuyang Feng, Xuekuan Wang |
IEEE Trans. Multim. | 1 |
| 2024 | Diverse Person: Customize Your Own Dataset for Text-Based Person SearchabstractText-based person search is a challenging task aimed at locating specific target pedestrians through text descriptions. Recent advancements have been made in this field, but there remains a deficiency in datasets tailored for text-based person search. The creation of new, real-world datasets is hindered by concerns such as the risk of pedestrian privacy leakage and the substantial costs of annotation. In this paper, we introduce a framework, named Diverse Person (DP), to achieve efficient and high-quality text-based person search data generation without involving privacy concerns. Specifically, we propose to leverage available images of clothing and accessories as reference attribute images to edit the original dataset images through diffusion models. Additionally, we employ a Large Language Model (LLM) to produce annotations that are both high in quality and stylistically consistent with those found in real-world datasets. Extensive experimental results demonstrate that the baseline models trained with our DP can achieve new state-of-the-art results on three public datasets, with performance improvements up to 4.82%, 2.15%, and 2.28% on CUHK-PEDES, ICFG-PEDES, and RSTPReid in terms of Rank-1 accuracy, respectively. Zifan Song, Guosheng Hu, Cairong Zhao |
AAAI | 3 |
| 2024 | Self-Supervised Likelihood Estimation with Energy Guidance for Anomaly Segmentation in Urban ScenesabstractRobust autonomous driving requires agents to accurately identify unexpected areas (anomalies) in urban scenes. To this end, some critical issues remain open: how to design advisable metric to measure anomalies, and how to properly generate training samples of anomaly data? Classical effort in anomaly detection usually resorts to pixel-wise uncertainty or sample synthesis, which ignores the contextual information and sometimes requires auxiliary data with fine-grained annotations. On the contrary, in this paper, we exploit the strong context-dependent nature of segmentation task and design an energy-guided self-supervised frameworks for anomaly segmentation, which optimizes an anomaly head by maximizing likelihood of self-generated anomaly pixels. For this purpose, we design two estimators to model anomaly likelihood, one is a task-agnostic binary estimator and the other depicts the likelihood as residual of task-oriented joint energy. Based on proposed estimators, we devise an adaptive self-supervised training framework, which exploits the contextual reliance and estimated likelihood to refine mask annotations in anomaly areas. We conduct extensive experiments on challenging Fishyscapes and Road Anomaly benchmarks, demonstrating that without any auxiliary data or synthetic models, our method can still achieves comparable performance to supervised competitors. Code is available at https://github.com/yuanpengtu/SLEEG. Yuanpeng Tu, Yuxi Li 0009, Boshen Zhang, Liang Liu 0007, Jiangning Zhang, Yabiao Wang, Cairong Zhao |
AAAI | 7 |
| 2024 | Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language ModelsabstractPrompt learning has become a prevalent strategy for adapting vision-language foundation models to downstream tasks. As large language models (LLMs) have emerged, recent studies have explored the use of category-related descriptions as input to enhance prompt effectiveness. Nevertheless, conventional descriptions fall short of structured information that effectively represents the interconnections among entities or attributes linked to a particular category. To address this limitation and prioritize harnessing structured knowledge, this paper advocates for leveraging LLMs to build a graph for each description to model the entities and attributes describing the category, as well as their correlations. Preexisting prompt tuning methods exhibit inadequacies in managing this structured knowledge. Consequently, we propose a novel approach called Hierarchical Prompt Tuning (HPT), which enables simultaneous modeling of both structured and conventional linguistic knowledge. Specifically, we introduce a relationship-guided attention module to capture pair-wise associations among entities and attributes for low-level prompt learning. In addition, by incorporating high-level and global-level prompts modeling overall semantics, the proposed hierarchical structure forges cross-level interlinks and empowers the model to handle more complex and long-term relationships. Extensive experiments demonstrate that our HPT shows strong effectiveness and generalizes much better than existing SOTA methods. Our code is available at https://github.com/Vill-Lab/2024-AAAI-HPT. Xinyang Jiang, De Cheng, Dongsheng Li 0002, Cairong Zhao |
AAAI | 5 |
| 2024 | Online Video Quality Enhancement with Spatial-Temporal Look-Up Tables
Zefan Qu, Xinyang Jiang, Yifan Yang 0004, Dongsheng Li 0002, Cairong Zhao |
ECCV (72) | 5 |
| 2024 | Self-supervised Feature Adaptation for 3D Industrial Anomaly Detection
Yuanpeng Tu, Boshen Zhang, Liang Liu 0007, Yuxi Li 0009, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
ECCV (2) | 8 |
| 2024 | BF-UNet: Bi-level Routing Attention U-shaped Network Based on Explicit Visual Prompt
Yuanfei Xu, Zhihui Lai 0001, Shihuan He, Cairong Zhao, Heng Kong |
ICPR (5) | 5 |
| 2024 | Uni4DAL: A Unified Baseline for Multi-dataset 4D Auto-Labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jinchen Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao |
ICPR (30) | 9 |
| 2024 | Fetch and Forge: Efficient Dataset Condensation for Object DetectionabstractDataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements.
However, current research on DC mainly focuses on image classification, with less exploration of object detection.
This is primarily due to two challenges: (i) the multitasking nature of object detection complicates the condensation process, and (ii) Object detection datasets are characterized by large-scale and high-resolution data, which are difficult for existing DC methods to handle.
As a remedy, we propose DCOD, the first dataset condensation framework for object detection. It operates in two stages: Fetch and Forge, initially storing key localization and classification information into model parameters, and then reconstructing synthetic images via model inversion.
For the complex of multiple objects in an image, we propose Foreground Background Decoupling to centrally update the foreground of multiple instances and Incremental PatchExpand to further enhance the diversity of foregrounds.
Extensive experiments on various detection datasets demonstrate the superiority of DCOD. Even at an extremely low compression rate of 1\%, we achieve 46.4\% and 24.7\% $\text{AP}_{50}$ on the VOC and COCO, respectively, significantly reducing detector training duration. Ding Qi, Jian Li 0062, Jinlong Peng, Bo Zhao 0015, Shuguang Dou, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
NeurIPS | 10 |
| 2024 | AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataabstractOpen-source Large Language Models (LLMs) and their specialized variants, particularly Code LLMs, have recently delivered impressive performance. However, previous Code LLMs are typically fine-tuned on single-source data with limited quality and diversity, which may insufficiently elicit the potential of pre-trained Code LLMs. In this paper, we present AlchemistCoder, a series of Code LLMs with enhanced code generation and generalization capabilities fine-tuned on multi-source data. To achieve this, we pioneer to unveil inherent conflicts among the various styles and qualities in multi-source code corpora and introduce data-specific prompts with hindsight relabeling, termed AlchemistPrompts, to harmonize different data sources and instruction-response pairs. Additionally, we propose incorporating the data construction process into the fine-tuning data as code comprehension tasks, including instruction evolution, data filtering, and code review. Extensive experiments demonstrate that AlchemistCoder holds a clear lead among all models of the same size (6.7B/7B) and rivals or even surpasses larger models (15B/33B/70B), showcasing the efficacy of our method in refining instruction-following capabilities and advancing the boundaries of code intelligence. Source code and models are available at https://github.com/InternLM/AlchemistCoder. Zifan Song, Yudong Wang 0002, Kuikun Liu, Chengqi Lyu, Demin Song, Qipeng Guo, Hang Yan 0001, Dahua Lin, Kai Chen 0026, Cairong Zhao |
NeurIPS | 11 |
| 2024 | DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionabstractDiffusion-based methods have achieved remarkable achievements in 2D image or 3D object generation, however, the generation of 3D scenes and even $360^{\circ}$ images remains constrained, due to the limited number of scene datasets, the complexity of 3D scenes themselves, and the difficulty of generating consistent multi-view images. To address these issues, we first establish a large-scale panoramic video-text dataset containing millions of consecutive panoramic keyframes with corresponding panoramic depths, camera poses, and text descriptions. Then, we propose a novel text-driven panoramic generation framework, termed DiffPano, to achieve scalable, consistent, and diverse panoramic scene generation. Specifically, benefiting from the powerful generative capabilities of stable diffusion, we fine-tune a single-view text-to-panorama diffusion model with LoRA on the established panoramic video-text dataset. We further design a spherical epipolar-aware multi-view diffusion model to ensure the multi-view consistency of the generated panoramic images. Extensive experiments demonstrate that DiffPano can generate scalable, consistent, and diverse panoramic images with given unseen text descriptions and camera poses. Weicai Ye, Chenhao Ji, Zheng Chen 0016, Junyao Gao 0002, Xiaoshui Huang, Song-Hai Zhang, Wanli Ouyang, Tong He 0001, Cairong Zhao, Guofeng Zhang 0001 |
NeurIPS | 9 |
| 2024 | Does Video-Text Pretraining Help Open-Vocabulary Online Action Detection?abstractVideo understanding relies on accurate action detection for temporal analysis. However, existing mainstream methods have limitations in real-world applications due to their offline and closed-set evaluation approaches, as well as their dependence on manual annotations. To address these challenges and enable real-time action understanding in open-world scenarios, we propose OV-OAD, a zero-shot online action detector that leverages vision-language models and learns solely from text supervision. By introducing an object-centered decoder unit into a Transformer-based model, we aggregate frames with similar semantics using video-text correspondence. Extensive experiments on four action detection benchmarks demonstrate that OV-OAD outperforms other advanced zero-shot methods. Specifically, it achieves 37.5\% mean average precision on THUMOS’14 and 73.8\% calibrated average precision on TVSeries. This research establishes a robust baseline for zero-shot transfer in online action detection, enabling scalable solutions for open-world temporal understanding. The code will be available for download at \url{https://github.com/OpenGVLab/OV-OAD}. Yi Wang 0074, Jilan Xu, Yinan He, Zifan Song, Limin Wang 0002, Yu Qiao 0001, Cairong Zhao |
NeurIPS | 8 |
| 2024 | Re-ID-leak: Membership Inference Attacks Against Person Re-identification
Junyao Gao 0002, Xinyang Jiang, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cairong Zhao |
Int. J. Comput. Vis. | 6 |
| 2024 | Adaptive Discriminative Regularization for Visual Classification
Yi Wang 0033, Shuguang Dou, Cairong Zhao |
Int. J. Comput. Vis. | 6 |
| 2024 | Attentive multi-granularity perception network for person search
Qixian Zhang, Jun Wu 0006, Duoqian Miao 0001, Cairong Zhao, Qi Zhang 0020 |
Inf. Sci. | 4 |
| 2024 | Learning adaptive shift and task decoupling for discriminative one-step person search
Qixian Zhang, Duoqian Miao 0001, Qi Zhang 0020, Changwei Wang 0001, Hongyun Zhang 0001, Cairong Zhao |
Knowl. Based Syst. | 7 |
| 2024 | Hierarchically Recognizing Vector Graphics and A New Chart-Based Vector Graphics DatasetabstractThe conventional approach to image recognition has been based on raster graphics, which can suffer from aliasing and information loss when scaled up or down. In this paper, we propose a novel approach that leverages the benefits of vector graphics for object localization and classification. Our method, called YOLaT (You Only Look at Text), takes the textual document of vector graphics as input, rather than rendering it into pixels. YOLaT builds multi-graphs to model the structural and spatial information in vector graphics and utilizes a dual-stream graph neural network (GNN) to detect objects from the graph. However, for real-world vector graphics, YOLaT only models in flat GNN with vertexes as nodes ignore higher-level information of vector data. Therefore, we propose YOLaT++ to learn Multi-level Abstraction Feature Learning from a new perspective: Primitive Shapes to Curves and Points. On the other hand, given few public datasets focus on vector graphics, data-driven learning cannot exert its full power on this format. We provide a large-scale and challenging dataset for Chart-based Vector Graphics Detection and Chart Understanding, termed VG-DCU, with vector graphics, raster graphics, annotations, and raw data drawn for creating these vector charts. Experiments show that the YOLaT series outperforms both vector graphics and raster graphics-based object detection methods on both subsets of VG-DCU in terms of both accuracy and efficiency, showcasing the potential of vector graphics for image recognition tasks. Shuguang Dou, Xinyang Jiang, Lu Liu 0019, Lu Ying, Yifei Shen 0004, Xuanyi Dong, Yun Wang 0012, Dongsheng Li 0002, Cairong Zhao |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2024 | Multi-granularity Cross Transformer Network for person re-identification
Duoqian Miao 0001, Hongyun Zhang 0001, Jie Zhou 0009, Cairong Zhao |
Pattern Recognit. | 5 |
| 2024 | Unified Multi-Modality Video Object Segmentation Using Reinforcement LearningabstractThe main task we aim to tackle is the multi-modality video object segmentation (VOS), which can be divided into two sub-tasks: mask-referred and language-referred VOS, where the first-frame mask-level or language-level label is utilized to provide the target information, respectively. Due to the huge gap between different modalities, existing works never come up with a unified framework for these two sub-tasks. In this work, such a unified framework is designed, where the visual and linguistic inputs are first spilt into a number of image patches and words, and then mapped into same-size tokens, which are equally processed by a self-attention based segmentation model. Furthermore, to highlight the significant information and discard the non-target or ambiguous one, unified multi-modality filter networks are further designed, and reinforcement learning is adopted to optimize such networks. Experiments show that new state-of-the-art performances are achieved by the proposed method: 52.8% ofJ&Fon Ref-YoutubeVOS dataset and 83.2% ofJSon YoutubeVOS dataset, respectively. The code will be released. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Cairong Zhao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Invisible Backdoor Attack With Dynamic Triggers Against Person Re-IdentificationabstractIn recent years, person Re-IDentification (ReID) has rapidly progressed with wide real-world applications but is also susceptible to various forms of attack, including proven vulnerability to adversarial attacks. In this paper, we focus on the backdoor attack on deep ReID models. Existing backdoor attack methods follow an all-to-one or all-to-all attack scenario, where all the target classes in the test set have already been seen in the training set. However, ReID is a much more complex fine-grained open-set recognition problem, where the identities in the test set are not contained in the training set. Thus, previous backdoor attack methods for classification are not applicable to ReID. To ameliorate this issue, we propose a novel backdoor attack on deep ReID under a new all-to-unknown scenario, called Dynamic Triggers Invisible Backdoor Attack (DT-IBA). Instead of learning fixed triggers for the target classes from the training set, DT-IBA can dynamically generate new triggers for any unknown identities. Specifically, an identity hashing network is proposed to first extract target identity information from a reference image, which is then injected into the benign images by image steganography. We extensively validate the effectiveness and stealthiness of the proposed attack on benchmark datasets and evaluate the effectiveness of several defense methods against our attack. Wenli Sun, Xinyang Jiang, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cheng Deng 0002, Cairong Zhao |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2024 | Learning Scene-Pedestrian Graph for End-to-End Person SearchabstractPerson search aims to find specific persons from visual scenes, including two subtasks, pedestrian detection, and person reidentification. The dominant fashion in this area is end-to-end networks that focus on analyzing the foreground (i.e., pedestrian) while ignoring the background (i.e., scene) information. However, the scene information often offers useful clues for person search. For example, pedestrians normally appear on the road rather than the top of a tree, and pedestrians appearing at the same location are likely to have similar occlusions. The interplay between the pedestrians and scenes can potentially improve the performance. In this article, a novel scene-pedestrian graph (SPG) is proposed, which can explicitly model the interplay between the pedestrians and scenes. To polish the quality of pedestrian bounding boxes, we pioneer a strategy of using the high-quality pedestrian bounding box to guide the low-quality one in the same scene. In addition, we design a contextual and temporal graph matching algorithm to effectively utilize the contextual and temporal information present in the constructed SPG to improve the performance of pedestrian matching. Benefiting from the robustness on complex scenes, our model achieves promising performance over the state-of-the-art methods on two popular person search benchmarks, CUHK-SYSU and PRW. Zifan Song, Cairong Zhao, Guosheng Hu, Duoqian Miao 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Occlusion-Aware Transformer With Second-Order Attention for Person Re-IdentificationabstractPerson re-identification (ReID) typically encounters varying degrees of occlusion in real-world scenarios. While previous methods have addressed this using handcrafted partitions or external cues, they often compromise semantic information or increase network complexity. In this paper, we propose a new method from a novel perspective, termed as OAT. Specifically, we first use a Transformer backbone with multiple class tokens for diverse pedestrian feature learning. Given that the self-attention mechanism in the Transformer solely focuses on low-level feature correlations, neglecting higher-order relations among different body parts or regions. Thus, we propose the Second-Order Attention (SOA) module to capture more comprehensive features. To address computational efficiency, we further derive approximation formulations for implementing second-order attention. Observing that the importance of semantics associated with different class tokens varies due to the uncertainty of the location and size of occlusion, we propose the Entropy Guided Fusion (EGF) module for multiple class tokens. By conducting uncertainty analysis on each class token, higher weights are assigned to those with lower information entropy, while lower weights are assigned to class tokens with higher entropy. The dynamic weight adjustment can mitigate the impact of occlusion-induced uncertainty on feature learning, thereby facilitating the acquisition of discriminative class token representations. Extensive experiments have been conducted on occluded and holistic person re-identification datasets, which demonstrate the effectiveness of our proposed method. Yizhang Liu, Hongyun Zhang 0001, Cairong Zhao, Zhihua Wei 0001, Duoqian Miao 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Learning Domain Invariant Prompt for Vision-Language ModelsabstractPrompt learning stands out as one of the most efficient approaches for adapting powerful vision-language foundational models like CLIP to downstream datasets by tuning learnable prompt vectors with very few samples. However, despite its success in achieving remarkable performance on in-domain data, prompt learning still faces the significant challenge of effectively generalizing to novel classes and domains. Some existing methods address this concern by dynamically generating distinct prompts for different domains. Yet, they overlook the inherent potential of prompts to generalize across unseen domains. To address these limitations, our study introduces an innovative prompt learning paradigm, called MetaPrompt, aiming to directly learn domain invariant prompt in few-shot scenarios. To facilitate learning prompts for image and text inputs independently, we present a dual-modality prompt tuning network comprising two pairs of coupled encoders. Our study centers on an alternate episodic training algorithm to enrich the generalization capacity of the learned prompts. In contrast to traditional episodic training algorithms, our approach incorporates both in-domain updates and domain-split updates in a batch-wise manner. For in-domain updates, we introduce a novel asymmetric contrastive learning paradigm, where representations from the pre-trained encoder assume supervision to regularize prompts from the prompted encoder. To enhance performance on out-of-domain distribution, we propose a domain-split optimization on visual prompts for cross-domain tasks or textual prompts for cross-class tasks during domain-split updates. Extensive experiments across 11 datasets for base-to-new generalization and 4 datasets for domain generalization exhibit favorable performance. Compared with the state-of-the-art method, MetaPrompt achieves an absolute gain of 1.02% on the overall harmonic mean in base-to-new generalization and consistently demonstrates superiority over all benchmarks in domain generalization. Cairong Zhao, Xinyang Jiang, Yifei Shen 0004, Kaitao Song, Dongsheng Li 0002, Duoqian Miao 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Explainability of Speech Recognition Transformers via Gradient-Based Attention VisualizationabstractIn vision Transformers, attention visualization methods are used to generate heatmaps highlighting the class-corresponding areas in input images, which offers explanations on how the models make predictions. However, it is not so applicable for explaining automatic speech recognition (ASR) Transformers. An ASR Transformer makes a particular prediction for every input token to form a sentence, but a vision Transformer only makes an overall classification for the input data. Therefore, traditional attention visualization methods may fail in ASR Transformers. In this work, we propose a novel attention visualization method in ASR Transformers and try to explain which frames of the audio result in the output text. Inspired by the model explainability, we also explore ways of improving the effectiveness of the ASR model. Comparing with other Transformer attention visualization methods, our method is more efficient and intuitively understandable, which unravels the attention calculation from information flow of Transformer attention modules. In addition, we demonstrate the utilization of visualization result in three ways: (1) We visualize attention with respect to connectionist temporal classification (CTC) loss to train an ASR model with adversarial attention erasing regularization, which effectively decreases the word error rate (WER) of the model and improves its generalization capability. (2) We visualize the attention on some specific words, interpreting the model by effectively demonstrating the semantic and grammar relationships between these words. (3) Similarly, we analyze how the model manage to distinguish homophones, using contrastive explanation with respect to homophones. Tianli Sun, Haonan Chen 0003, Guosheng Hu, Lianghua He, Cairong Zhao |
IEEE Trans. Multim. | 5 |
| 2023 | Similarity Distribution Based Membership Inference Attack on Person Re-identificationabstractWhile person Re-identification (Re-ID) has progressed rapidly due to its wide real-world applications, it also causes severe risks of leaking personal information from training data. Thus, this paper focuses on quantifying this risk by membership inference (MI) attack. Most of the existing MI attack algorithms focus on classification models, while Re-ID follows a totally different training and inference paradigm. Re-ID is a fine-grained recognition task with complex feature embedding, and model outputs commonly used by existing MI like logits and losses are not accessible during inference. Since Re-ID focuses on modelling the relative relationship between image pairs instead of individual semantics, we conduct a formal and empirical analysis which validates that the distribution shift of the inter-sample similarity between training and test set is a critical criterion for Re-ID membership inference. As a result, we propose a novel membership inference attack method based on the inter-sample similarity distribution. Specifically, a set of anchor images are sampled to represent the similarity distribution conditioned on a target image, and a neural network with a novel anchor selection module is proposed to predict the membership of the target image. Our experiments validate the effectiveness of the proposed approach on both the Re-ID task and conventional classification task. Junyao Gao 0002, Xinyang Jiang, Huishuai Zhang, Yifan Yang 0004, Shuguang Dou, Dongsheng Li 0002, Duoqian Miao 0001, Cheng Deng 0002, Cairong Zhao |
AAAI | 9 |
| 2023 | Cross-Modal Distillation for Speaker RecognitionabstractSpeaker recognition achieved great progress recently, however, it is not easy or efficient to further improve its performance via traditional solutions: collecting more data and designing new neural networks. Aiming at the fundamental challenge of speech data, i.e. low information density, multimodal learning can mitigate this challenge by introducing richer and more discriminative information as input for identity recognition. Specifically, since the face image is more discriminative than the speech for identity recognition, we conduct multimodal learning by introducing a face recognition model (teacher) to transfer discriminative knowledge to a speaker recognition model (student) during training. However, this knowledge transfer via distillation is not trivial because the big domain gap between face and speech can easily lead to overfitting. In this work, we introduce a multimodal learning framework, VGSR (Vision-Guided Speaker Recognition). Specifically, we propose a MKD (Margin-based Knowledge Distillation) strategy for cross-modality distillation by introducing a loose constrain to align the teacher and student, greatly reducing overfitting. Our MKD strategy can easily adapt to various existing knowledge distillation methods. In addition, we propose a QAW (Quality-based Adaptive Weights) module to weight input samples via quantified data quality, leading to a robust model training. Experimental results on the VoxCeleb1 and CN-Celeb datasets show our proposed strategies can effectively improve the accuracy of speaker recognition by a margin of 10% ∼ 15%, and our methods are very robust to different noises. Yufeng Jin, Guosheng Hu, Haonan Chen 0003, Duoqian Miao 0001, Liang Hu 0001, Cairong Zhao |
AAAI | 6 |
| 2023 | Learning from Noisy Labels with Decoupled Meta Label PurifierabstractTraining deep neural networks (DNN) with noisy labels is challenging since DNN can easily memorize inaccurate labels, leading to poor generalization ability. Recently, the meta-learning based label correction strategy is widely adopted to tackle this problem via identifying and correcting potential noisy labels with the help of a small set of clean validation data. Although training with purified labels can effectively improve performance, solving the meta-learning problem inevitably involves a nested loop of bi-level optimization between model weights and hyper-parameters (i.e., label distribution). As compromise, previous methods resort to a coupled learning process with alternating update. In this paper, we empirically find such simultaneous optimization over both model weights and label distribution can not achieve an optimal routine, consequently limiting the representation ability of backbone and accuracy of corrected labels. From this observation, a novel multi-stage label purifier named DMLP is proposed. DMLP decouples the label correction process into label-free representation learning and a simple meta label purifier, In this way, DMLP can focus on extracting discriminative feature and label correction in two distinctive stages. DMLP is a plug-and-play label purifier, the purified labels can be directly reused in naive end-to-end network retraining or other robust learning methods, where state-of-the-art results are obtained on several synthetic and real-world noisy datasets, especially under high noise levels. Code is available at https://github.com/yuanpengtu/DMLP. Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 8 |
| 2023 | Learning with Noisy labels via Self-supervised Adversarial Noisy MaskingabstractCollecting large-scale datasets is crucial for training deep models, annotating the data, however, inevitably yields noisy labels, which poses challenges to deep learning algorithms. Previous efforts tend to mitigate this problem via identifying and removing noisy samples or correcting their labels according to the statistical properties (e.g., loss values) among training samples. In this paper, we aim to tackle this problem from a new perspective, delving into the deep feature maps, we empirically find that models trained with clean and mislabeled samples manifest distinguishable activation feature distributions. From this observation, a novel robust training approach termed adversarial noisy masking is proposed. The idea is to regularize deep features with a label quality guided masking scheme, which adaptively modulates the input data and label simultaneously, preventing the model to overfit noisy samples. Further, an auxiliary task is designed to reconstruct input data, it naturally provides noise-free self-supervised signals to rein-force the generalization ability of models. The proposed method is simple yet effective, it is tested on synthetic and real-world noisy datasets, where significant improvements are obtained over previous methods. Code is available at https://github.com/yuanpengtu/SANM. Yuanpeng Tu, Boshen Zhang, Yuxi Li 0009, Liang Liu 0007, Jian Li 0062, Jiangning Zhang, Yabiao Wang, Chengjie Wang 0001, Cairong Zhao |
CVPR | 9 |
| 2023 | PatSTEG: Modeling Formation Dynamics of Patent Citation Networks via The Semantic-Topological Evolutionary GraphabstractPatent documents in the patent database (PatDB) are crucial for research, development, and innovation as they contain valuable technical information. However, PatDB presents a multifaceted challenge in comparison to publicly available preprocessed databases due to the intricate nature of patent text and the inherent sparsity within the patent citation network. Although patent text analysis and citation analysis bring new opportunities to explore patent data mining, no existing work exploits the complementation of them. To this end, we propose a joint semantic-topological evolutionary graph learning approach (PatSTEG) to model the formation dynamics of patent citation networks. More specifically, we first create a real-world dataset of Chinese patents named CNPat, and leveraging its patent texts and citations to construct a patent citation network. Then, PatSTEG is modeled to study the evolutionary dynamics of patent citation formation by jointly considering the semantic and topological information. Extensive experiments are conducted on both CNPat and public datasets to prove the superiority of PatSTEG over other state-of-the-art methods. All the results provide valuable references for patent literature research and technical exploration. Ran Miao, Xueyu Chen, Liang Hu 0004, Minghua Wan, Qi Zhang 0020, Cairong Zhao |
ICDM | 7 |
| 2023 | EA-HAS-Bench: Energy-aware Hyperparameter and Architecture Search Benchmark
Shuguang Dou, Xinyang Jiang, Cairong Zhao, Dongsheng Li 0002 |
ICLR | 3 |
| 2023 | Deep Perturbation Learning: Enhancing the Network Performance via Image PerturbationsabstractImage perturbation technique is widely used to generate adversarial examples to attack networks, greatly decreasing the performance of networks. Unlike the existing works, in this paper, we introduce a novel framework Deep Perturbation Learning (DPL), the new insights into understanding image perturbations, to enhance the performance of networks rather than decrease the performance. Specifically, we learn image perturbations to amend the data distribution of training set to improve the performance of networks. This optimization w.r.t data distribution is non-trivial. To approach this, we tactfully construct a differentiable optimization target w.r.t. image perturbations via minimizing the empirical risk. Then we propose an alternating optimization of the network weights and perturbations. DPL can easily be adapted to a wide spectrum of downstream tasks and backbone networks. Extensive experiments demonstrate the effectiveness of our DPL on 6 datasets (CIFAR-10, CIFAR100, ImageNet, MS-COCO, PASCAL VOC, and SBD) over 3 popular vision tasks (image classification, object detection, and semantic segmentation) with different backbone architectures (e.g., ResNet, MobileNet, and ViT). Zifan Song, Guosheng Hu, Cairong Zhao |
ICML | 4 |
| 2023 | Text-Enhanced Scene Image Super-Resolution via Stroke Mask and Orthogonal AttentionabstractLow-resolution text images are very commonplace in real life and their information is hard to be extracted by using existing text recognition methods only. Although this problem can be solved by introducing super-resolution (SR) techniques, most existing SR methods fail to process stroke regions and background regions of input text images distinctively. In this paper, we propose a text-specific super-resolution network named Text Enhanced Attention Network (TEAN) to solve this problem. First of all, we compensate for disadvantages of traditional thresholding mask operation proposed in Text Super-Resolution Network (TSRN) by utilizing deep-learning based semantic segmentation method to get correct masks as prior semantic information and propose a Text-Segmented-Contextual-Attention (TSCA) branch on the basis of them. Besides, we design an Orthogonal Contextual Attention Module (OCAM) working with TSCA to implicitly enhance stroke regions of LR images. Secondly, to effectively fuse shallow features and deep features of SR model, we propose a convolutional structure named Weight Balanced Fusion Module (WBFM) to improve traditional feature fusion methods of SR network. Finally, extensive experiments on TextZoom dataset demonstrate that the proposed network can improve the recognition accuracy of text images on existing text recognition models. Using TEAN to process low-resolution text images improves the recognition accuracy by 25.4% on CRNN, by 17.4% on ASTER, by 17.3% on MORAN, by 20.7% on NRTR, by 17.3% on SAR and by 15.9% on MASTER compared with directly recognizing them, which attains competitive performances against state-of-the-art methods. Furthermore, cross-dataset experiments on IC15_2077 demonstrate that TEAN is helpful for scene text recognition task, especially for low-resolution images even with the cross-domain issue. Cairong Zhao, Shuyang Feng, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | ISTVT: Interpretable Spatial-Temporal Video Transformer for Deepfake DetectionabstractWith the rapid development of Deepfake synthesis technology, our information security and personal privacy have been severely threatened in recent years. To achieve a robust Deepfake detection, researchers attempt to exploit the joint spatial-temporal information in the videos, like using recurrent networks and 3D convolutional networks. However, these spatial-temporal models remain room to improve. Another general challenge for spatial-temporal models is that people do not clearly understand what these spatial-temporal models really learn. To address these two challenges, in this paper, we propose an Interpretable Spatial-Temporal Video Transformer (ISTVT), which consists of a novel decomposed spatial-temporal self-attention and a self-subtract mechanism to capture spatial artifacts and temporal inconsistency for robust Deepfake detection. Thanks to this decomposition, we propose to interpret ISTVT by visualizing the discriminative regions for both spatial and temporal dimensions via the relevance (the pixel-wise importance on the input) propagation algorithm. We conduct extensive experiments on large-scale datasets, including FaceForensics++, FaceShifter, DeeperForensics, Celeb-DF, and DFDC datasets. Our strong performance of intra-dataset and cross-dataset Deepfake detection demonstrates the effectiveness and robustness of our method, and our visualization-based interpretability offers people insights into our model. Cairong Zhao, Chutian Wang, Guosheng Hu, Haonan Chen 0003, Chun Liu 0003, Jinhui Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Human Co-Parsing Guided Alignment for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) is a challenging task due to more background noises and incomplete foreground information. Although existing human parsing-based ReID methods can tackle this problem with semantic alignment at the finest pixel level, their performance is heavily affected by the human parsing model. Most supervised methods propose to train an extra human parsing model aside from the ReID model with cross-domain human parts annotation, suffering from expensive annotation cost and domain gap; Unsupervised methods integrate a feature clustering-based human parsing process into the ReID model, but lacking supervision signals brings less satisfactory segmentation results. In this paper, we argue that the pre-existing information in the ReID training dataset can be directly used as supervision signals to train the human parsing model without any extra annotation. By integrating a weakly supervised human co-parsing network into the ReID network, we propose a novel framework that exploits shared information across different images of the same pedestrian, called the Human Co-parsing Guided Alignment (HCGA) framework. Specifically, the human co-parsing network is weakly supervised by three consistency criteria, namely global semantics, local space, and background. By feeding the semantic information and deep features from the person ReID network into the guided alignment module, features of the foreground and human parts can then be obtained for effective occluded person ReID. Experiment results on two occluded and two holistic datasets demonstrate the superiority of our method. Especially on Occluded-DukeMTMC, it achieves 70.2% Rank-1 accuracy and 57.5% mAP. Shuguang Dou, Cairong Zhao, Xinyang Jiang, Shanshan Zhang 0001, Wei-Shi Zheng 0001, Wangmeng Zuo |
IEEE Trans. Image Process. | 2 |
| 2023 | Content-Adaptive Auto-Occlusion Network for Occluded Person Re-IdentificationabstractThe occluded person re-identification (ReID) aims to match person images captured in severely occluded environments. Current occluded ReID works mostly rely on auxiliary models or employ a part-to-part matching strategy. However, these methods may be sub-optimal since the auxiliary models are constrained by occlusion scenes and the matching strategy will deteriorate when both query and gallery set contain occlusion. Some methods attempt to solve this problem by applying image occlusion augmentation (OA) and have shown great superiority in their effectiveness and lightness. But there are two defects that existed in the previous OA-based method: 1) The occlusion policy is fixed throughout the entire training and cannot be dynamically adjusted based on the current training status of the ReID network. 2) The position and area of the applied OA are completely random, without reference to the image content to choose the most suitable policy. To address these challenges, we propose a novel Content-Adaptive Auto-Occlusion Network (CAAO), that is able to dynamically select the proper occlusion region of an image based on its content and the current training status. Specifically, CAAO consists of two parts: the ReID network and the Auto-Occlusion Controller (AOC) module. AOC automatically generates the optimal OA policy based on the feature map extracted from the ReID network and applies occlusion on the images for ReID network training. An on-policy reinforcement learning based alternating training paradigm is proposed to iteratively update the ReID network and AOC module. Comprehensive experiments on occluded and holistic person ReID benchmarks demonstrate the superiority of CAAO. Cairong Zhao, Zefan Qu, Xinyang Jiang, Yuanpeng Tu, Xiang Bai |
IEEE Trans. Image Process. | 1 |
| 2022 | Multi-Definition Video Deepfake Detection via Semantics Reduction and Cross-Domain TrainingabstractThe recent development of Deepfake videos directly threatens our information security and personal privacy. Although lots of previous works have made much progress on the Deepfake detection, we empirically find that the existing approaches do not perform well on the low definition (LD) and crossdefinition (high and low) videos. To address this problem, in this paper, we follow two motivations: (1) high-level semantics reduction and (2) cross-domain training. For (1), we propose the Facial Structure Destruction and Adversarial Jigsaw Loss to reduce our model to learn high-level semantics and focus on learning low-level discriminative information; For (2), we propose a domain generalization method based on adversarial learning. We conduct extensive experiments on the FaceForensics++ dataset. Results show the great effectiveness of our method and we also achieve very competitive performance against state-of-the-art methods. Chutian Wang, Cairong Zhao, Guosheng Hu |
ICME | 2 |
| 2022 | Part-Based Multi-Scale Attention Network for Text-Based Person Search
Ding Qi, Cairong Zhao |
PRCV (1) | 3 |
| 2022 | A new weakly supervised discrete discriminant hashing for robust data representation
Minghua Wan, Xueyu Chen, Cairong Zhao, Tianming Zhan, Guowei Yang 0002 |
Inf. Sci. | 3 |
| 2022 | Improved Instance Discrimination and Feature Compactness for End-to-End Person SearchabstractPerson search aims to locate and retrieve specific pedestrians in scene images, including two subtasks, pedestrian detection and person re-identification. Recently, triplet loss has been widely used in person re-identification, which effectively improves the pedestrian features embedding and achieves superior performance. However, forming triplet in the person search is not an easy task. Most of the existing end-to-end person search methods are based on Faster R-CNN. The training process of person re-identification part is affected by the detector. It is difficult to form pedestrian triplets within a limited batch size. Also, there are many pedestrian identities in the person search dataset, but each pedestrian identity only has a few samples. It is difficult to learn a robust pedestrian feature representation for person search. To resolve the problem discussed above, a novel Feature Compactness (FC) Loss for the person search is designed, which efficiently improves the inter-class discrimination and intra-class compactness of pedestrian features embedding without the need for positive or negative pairs. Besides, we propose a pedestrian attention module (PAM) to help the network focuses more on pedestrian information and suppresses irrelevant background information. Our method achieves comparable performance on two benchmarks, CUHK-SYSU and PRW, and achieves 91.96% of mAP and 93.34% of rank1 accuracy on CUHK-SYSU. Shaowei Hou, Cairong Zhao, Jun Wu 0006, Zhihua Wei 0001, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Context-Aware Feature Learning for Noise Robust Person SearchabstractPerson search aims to localize and identify specific pedestrians from numerous surveillance scene images. In this work, we focus on the noise in person search. We categorize the noise into scene-inherent noise and human-introduced noise. Scene-inherent noise comes from congestion, occlusion, and illumination changes. Human-introduced noise originates from the labeling process. For scene-inherent noise, we propose a novel context contrastive loss to take advantage of the latent contextual information from scene images. Features from context regions are utilized to construct contrastive pairs to constrain the feature discrimination among pedestrians in scene images while maintaining the feature consistency of the same identity. The network can thus learn to distinguish congested and overlapped pedestrians and more robust features can be obtained. For human-introduced noise, we propose a noise-discovery and noise-suppression training process for mislabeling robust person search. After the first training pass, the relation between feature prototypes of different identities is analyzed and the mislabeled pedestrians are discovered. During the second training pass, the label noise is suppressed to reduce the negative influence of mislabeled data. Experiments show that the proposed context-aware noise-robust (CANR) person search can achieve competitive performance. Further ablation studies confirm the effectiveness of CANR. Cairong Zhao, Shuguang Dou, Zefan Qu, Jiawei Yao, Jun Wu 0006, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Detecting Overlapped Objects in X-Ray Security Imagery by a Label-Aware MechanismabstractOne of the key challenges to the X-ray security check is to detect the overlapped items in backpacks or suitcases in the X-ray images. Most existing methods improve the robustness of models to the object overlapping problem by enhancing the underlying visual information such as colors and edges. However, this strategy ignores the situations that the objects have similar visual clues as to the background, and objects overlapping each other. Since the two cases rarely appear in existing datasets, we contribute a novel dataset – Cutters and Liquid Containers X-ray Dataset (CLCXray) to complete the related research. Furthermore, we propose a novel Label-aware Mechanism (LA) to tackle the object overlapping problem. Particularly, LA establishes the associations between feature channels and different labels and adjusts the features according to the assigned labels (or pseudo labels) to help improve the prediction results. Extensive experiments demonstrate that the LA is accurate and robust to detect overlapped objects, and also validate the effectiveness and the good generalization of the LA for arbitrary state-of-the-art (SOTA) methods. Furthermore, experimental results show that the network constructed by the LA is superior to the SOTA models on OPIXray and CLCXray, especially solving the challenges of the subset of the highly overlapped objects. Cairong Zhao, Shuguang Dou, Weihong Deng, Liang Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2021 | Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkabstractOptical degradation blurs text shapes and edges, so existing scene text recognition methods have difficulties in achieving desirable results on low-resolution (LR) scene text images acquired in real-world environments. The above problem can be solved by efficiently extracting sequential information to reconstruct super-resolution (SR) text images, which remains a challenging task. In this paper, we propose a Parallelly Contextual Attention Network (PCAN), which effectively learns sequence-dependent features and focuses more on high-frequency information of the reconstruction in text images. Firstly, we explore the importance of sequence-dependent features in horizontal and vertical directions parallelly for text SR, and then design a parallelly contextual attention block to adaptively select the key information in the text sequence that contributes to image super-resolution. Secondly, we propose a hierarchically orthogonal texture-aware attention module and an edge guidance loss function, which can help to reconstruct high-frequency information in text images. Finally, we conduct extensive experiments on TextZoom dataset, and the results can be easily incorporated into mainstream text recognition algorithms to further improve their performance in LR image recognition. Besides, our approach exhibits great robustness in defending against adversarial attacks on seven mainstream scene text recognition datasets, which means it can also improve the security of the text recognition pipeline. Compared with directly recognizing LR images, our method can respectively improve the recognition accuracy of ASTER, MORAN, and CRNN by 14.9%, 14.0%, and 20.1%. Our method outperforms eleven state-of-the-art (SOTA) SR methods in terms of boosting text recognition performance. Most importantly, it outperforms the current optimal text-orient SR method TSRN by 3.2%, 3.7%, and 6.0% on the recognition accuracy of ASTER, MORAN, and CRNN respectively. Cairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding, Jun Wu 0006, Fumin Shen, Heng Tao Shen |
ACM Multimedia | 1 |
| 2021 | Incremental Generative Occlusion Adversarial Suppression Network for Person ReIDabstractPerson re-identification (re-id) suffers from the significant challenge of occlusion, where an image contains occlusions and less discriminative pedestrian information. However, certain work consistently attempts to design complex modules to capture implicit information (including human pose landmarks, mask maps, and spatial information). The network, consequently, focuses on discriminative features learning on human non-occluded body regions and realizes effective matching under spatial misalignment. Few studies have focused on data augmentation, given that existing single-based data augmentation methods bring limited performance improvement. To address the occlusion problem, we propose a novel Incremental Generative Occlusion Adversarial Suppression (IGOAS) network. It consists of 1) an incremental generative occlusion block, generating easy-to-hard occlusion data, that makes the network more robust to occlusion by gradually learning harder occlusion instead of hardest occlusion directly. And 2) a global-adversarial suppression (G&A) framework with a global branch and an adversarial suppression branch. The global branch extracts steady global features of the images. The adversarial suppression branch, embedded with two occlusion suppression module, minimizes the generated occlusion's response and strengthens attentive feature representation on human non-occluded body regions. Finally, we get a more discriminative pedestrian feature descriptor by concatenating two branches' features, which is robust to the occlusion problem. The experiments on the occluded dataset show the competitive performance of IGOAS. On Occluded-DukeMTMC, it achieves 60.1% Rank-1 accuracy and 49.4% mAP. Cairong Zhao, Xinbi Lv, Shuguang Dou, Shanshan Zhang 0001, Jun Wu 0006, Liang Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Salience-Guided Iterative Asymmetric Mutual Hashing for Fast Person Re-IdentificationabstractPerson Re-identification (ReID) aims to retrieve the pedestrian with the same identity across different views. Existing studies mainly focus on improving accuracy, while ignoring their efficiency. Recently, several hash based methods have been proposed. Despite their improvement in efficiency, there still exists an unacceptable gap in accuracy between these methods and real-valued ones. Besides, few attempts have been made to simultaneously explicitly reduce redundancy and improve discrimination of hash codes, especially for short ones. Integrating Mutual learning may be a possible solution to reach this goal. However, it fails to utilize the complementary effect of teacher and student models. Additionally, it will degrade the performance of teacher models by treating two models equally. To address these issues, we propose a salience-guided iterative asymmetric mutual hashing (SIAMH) to achieve high-quality hash code generation and fast feature extraction. Specifically, a salience-guided self-distillation branch (SSB) is proposed to enable SIAMH to generate hash codes based on salience regions, thus explicitly reducing the redundancy between codes. Moreover, a novel iterative asymmetric mutual training strategy (IAMT) is proposed to alleviate drawbacks of common mutual learning, which can continuously refine the discriminative regions for SSB and extract regularized dark knowledge for two models as well. Extensive experiment results on five widely used datasets demonstrate the superiority of the proposed method in efficiency and accuracy when compared with existing state-of-the-art hashing and real-valued approaches. The code is released at https://github.com/Vill-Lab/SIAMH. Cairong Zhao, Yuanpeng Tu, Zhihui Lai 0001, Fumin Shen, Heng Tao Shen, Duoqian Miao 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | FLAG: feature learning with additional guidance for person search
Xinbi Lv, Tianli Sun, Cairong Zhao |
Vis. Comput. | 4 |
| 2020 | Path Aggregation and Dual Supervision Network for Scene Text Detection
Shuyang Feng, Cairong Zhao |
PRCV (3) | 3 |
| 2020 | Vehicle and wheel detection: a novel SSD-based approach and associated large-scale benchmark dataset
Jiayue Fu, Cairong Zhao |
Multim. Tools Appl. | 2 |
| 2020 | Similarity learning with joint transfer constraints for person re-identification
Cairong Zhao, Xuekuan Wang, Wangmeng Zuo, Fumin Shen, Ling Shao 0001, Duoqian Miao 0001 |
Pattern Recognit. | 1 |
| 2020 | Deep Fusion Feature Representation Learning With Hard Mining Center-Triplet Loss for Person Re-IdentificationabstractPerson re-identification (Re-ID) is a challenging task in the field of computer vision and focuses on matching people across images from different cameras. The extraction of robust feature representations from pedestrian images through CNNs with a single deterministic pooling operation is problematic as the features in real pedestrian images are complex and diverse. To address this problem, we propose a novel center-triplet (CT) model that combines the learning of robust feature representation and the optimization of metric loss function. Firstly, we design a fusion feature learning network (FFLN) with a novel fusion strategy consisting of max pooling and average pooling. Instead of adopting a single deterministic pooling operation, the FFLN combines two pooling operations that can learn high response values, bright features, and low response values, discriminative features simultaneously. Our model obtains more discriminative fusion features by adaptively learning the weights of the features learned by the corresponding pooling operations. In addition, we design a hard mining center-triplet loss (HCTL), a novel improved triplet loss, which effectively optimizes the intra/inter-class distance and reduces the cost of computing and mining hard training samples simultaneously, thereby enhancing the learning of robust feature representation. Finally, we proved our method can learn robust and discriminative feature representations for complex pedestrian images in real scenes. The experimental results also illustrate that our method achieves an 81.8% mAP and a 93.8% rank-1 accuracy on Market1501, a 68.2% mAP and an 83.3% rank-1 accuracy on DukeMTMC-ReID, and a 43.6% mAP and a 74.3% rank-1 accuracy on MSMT17, outperforming most state-of-the-art methods and achieving better performance for person re-identification. Cairong Zhao, Xinbi Lv, Zhang Zhang 0001, Wangmeng Zuo, Jun Wu 0006, Duoqian Miao 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | A Novel Hard Mining Center-Triplet Loss for Person Re-identification
Xinbi Lv, Cairong Zhao |
PRCV (3) | 2 |
| 2019 | Uncertainty-optimized deep learning model for small-scale person re-identification
Cairong Zhao, Di Zang, Zhaoxiang Zhang 0001, Wangmeng Zuo, Duoqian Miao 0001 |
Sci. China Inf. Sci. | 1 |
| 2019 | A self-adaptive cascade ConvNets model based on label relation mining
Zhihua Wei 0001, Wen Shen 0002, Cairong Zhao, Duoqian Miao 0001 |
Neurocomputing | 3 |
| 2019 | QRKISS: A Two-Stage Metric Learning via QR-Decomposition and KISS for Person Re-Identification
Cairong Zhao, Yipeng Chen, Zhihua Wei 0001, Duoqian Miao 0001, Xinjian Gu |
Neural Process. Lett. | 1 |
| 2019 | Improved adaptive image retrieval with the use of shadowed sets
Hongyun Zhang 0001, Witold Pedrycz, Cairong Zhao, Duoqian Miao 0001 |
Pattern Recognit. | 4 |
| 2019 | Multilevel triplet deep learning model for person re-identification
Cairong Zhao, Zhihua Wei 0001, Yipeng Chen, Duoqian Miao 0001 |
Pattern Recognit. Lett. | 1 |
| 2018 | Kernelized random KISS metric learning for person re-identification
Cairong Zhao, Yipeng Chen, Xuekuan Wang, Wai Keung Wong, Duoqian Miao 0001, Jingsheng Lei |
Neurocomputing | 1 |
| 2018 | Maximum decision entropy-based attribute reduction in decision-theoretic rough set model
Can Gao, Zhihui Lai 0001, Jie Zhou 0009, Cairong Zhao, Duoqian Miao 0001 |
Knowl. Based Syst. | 4 |
| 2018 | Maximal granularity structure and generalized multi-view discriminant analysis for person re-identification
Cairong Zhao, Xuekuan Wang, Duoqian Miao 0001, Hanli Wang, Wei-Shi Zheng 0001, Yong Xu 0001, David Zhang 0001 |
Pattern Recognit. | 1 |
| 2017 | Multiple metric learning based on bar-shape descriptor for person re-identification
Cairong Zhao, Xuekuan Wang, Wai Keung Wong, Wei-Shi Zheng 0001, Jian Yang 0003, Duoqian Miao 0001 |
Pattern Recognit. | 1 |
| 2016 | Mutli-channel micro-structure difference descriptor for image retrievalabstractThis paper presents a novel image feature representation method, called multi-channel micro-structure difference descriptor (MCMSDD) for image retrieval. With the local feature extraction from a micro-structure and MAX operator, MCMSDD integrates the advantages of multi-channel local binary encoding and color difference histogram , which are the fusion of color, texture and spatial distribution information. Although it extracts feature from full color image, the dimension of the feature vector is relatively low without learning and segmentation. To improve the performance of retrieval, a simple re-ranking algorithm is employed. Finally, the proposed MCMSDD is extensively tested on Corel-2K and Washington datasets, and the experimental results show that the proposed MCMSDD is more effective than the state-of-the-art. Xuekuan Wang, Cairong Zhao, Duoqian Miao 0001, Cuijun Liu, Yipeng Chen, Zhihui Lai 0001 |
ICPR | 2 |
| 2016 | Fusion of multiple channel features for person re-identification
Xuekuan Wang, Cairong Zhao, Duoqian Miao 0001, Zhihua Wei 0001, Renxian Zhang, Tingfei Ye |
Neurocomputing | 2 |
| 2015 | Feature extraction using adaptive slow feature discriminant analysis
Xingjian Gu, Chuancai Liu, Cairong Zhao |
Neurocomputing | 4 |
| 2015 | Uncorrelated slow feature discriminant analysis using globality preserving projections for feature extraction
Xingjian Gu, Chuancai Liu, Cairong Zhao, Songsong Wu |
Neurocomputing | 4 |
| 2014 | Graph embedding discriminant analysis for face recognition
Cairong Zhao, Zhihui Lai 0001, Duoqian Miao 0001, Zhihua Wei 0001, Caihui Liu |
Neural Comput. Appl. | 1 |
| 2014 | Sparse Alignment for Robust Tensor LearningabstractMultilinear/tensor extensions of manifold learning based algorithms have been widely used in computer vision and pattern recognition. This paper first provides a systematic analysis of the multilinear extensions for the most popular methods by using alignment techniques, thereby obtaining a general tensor alignment framework. From this framework, it is easy to show that the manifold learning based tensor learning methods are intrinsically different from the alignment techniques. Based on the alignment framework, a robust tensor learning method called sparse tensor alignment (STA) is then proposed for unsupervised tensor feature extraction. Different from the existing tensor learning methods, L1- and L2-norms are introduced to enhance the robustness in the alignment step of the STA. The advantage of the proposed technique is that the difficulty in selecting the size of the local neighborhood can be avoided in the manifold learning based tensor feature extraction algorithms. Although STA is an unsupervised learning method, the sparsity encodes the discriminative information in the alignment step and provides the robustness of STA. Extensive experiments on the well-known image databases as well as action and hand gesture databases by encoding object images as tensors demonstrate that the proposed STA algorithm gives the most competitive performance when compared with the tensor-based unsupervised learning methods. Zhihui Lai 0001, Wai Keung Wong, Yong Xu 0001, Cairong Zhao, Mingming Sun 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | Two-dimensional color uncorrelated discriminant analysis for face recognition
Cairong Zhao, Duoqian Miao 0001, Zhihui Lai 0001, Can Gao, Chuancai Liu, Jing-Yu Yang 0001 |
Neurocomputing | 1 |
| 2012 | Fisher Difference Discriminant Analysis: Determining the Effective Discriminant Subspace Dimensions for Face Recognition
Zhihui Lai 0001, Cairong Zhao, Minghua Wan |
Neural Process. Lett. | 2 |
| 2012 | Fuzzy local maximal marginal embedding for feature extraction
Cairong Zhao, Chuancai Liu, Xingjian Gu, Jianjun Qian |
Soft Comput. | 1 |
| 2011 | Multi-scale gist feature manifold for building recognition
Cairong Zhao, Chuancai Liu |
Neurocomputing | 1 |
| 2011 | Maximal local interclass embedding with application to face recognition
Cairong Zhao, Zhong Jin |
Mach. Vis. Appl. | 2 |
| 2010 | Fuzzy maximal marginal embedding and its applicationabstractIn this paper, we develops a new approach, called fuzzy maximal marginal embedding (FMME), combining LMME (local maximal marginal embedding) with fuzzy set theory, in which the fuzzy k-nearest neighbor (FKNN) is implemented to achieve the nature distribution information of original samples, and this information is utilized to redefine the affinity weights of neighborhood graph (intraclass and interclass ) instead of the weights of the binary pattern. We can reduce sensitivity of the method to substantial variations between samples caused by varying illumination and shape, viewing conditions. That makes FMME more powerful and robust than other method. The proposed algorithm is examined using Yale and ORL face image databases. The experimental results show FMME outperforms PCA, LDA, LPP and LMME. Cairong Zhao, Yue Sui, Chuancai Liu, Zhong Jin |
ICIP | 1 |
| 2010 | Sparse Embedding Visual Attention Systems Combined with Edge InformationabstractThe general computational models of visual attention are to obtain multi-scale feature maps in terms of visual properties like intensity, color and orientation, and then combine them to get one saliency map. But due to the lack of object edge information and reasonable feature combination strategy, the visual saliency map of the image is a blur map. Being aware of these, we propose a new scheme for saliency extraction. In this paper, we firstly put forward a sparse embedding feature combination strategy, inspired by sparse representation. The strategy is used to combine the salient regions from the individual feature maps based on a novel feature sparse indicator that measures the contribution of each map to saliency. Then we combine traditional visual attention with edge information. Results on different scene images show that our method outperforms other traditional feature combination strategies. Cairong Zhao, Chuancai Liu, Zhihui Lai 0001, Jing-Yu Yang 0001 |
ICPR | 1 |