Zhicheng Zhao 0001

dblp:55/7547-1 · also Zhi-Cheng Zhao 0001 · DBLP profile ↗
← Back
110ranked-venue papers
5as first author
59since 2021 · last 2026
0000-0001-6506-7298ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 2 first-author · 32 since 2021Artificial intelligence and machine learning · 40 · 3 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Modality and Task Adaptation for Enhanced Zero-shot Composed Image Retrieval
abstract
As a challenging vision-language task, Zero-Shot Composed Image Retrieval (ZS-CIR) is designed to retrieve target images using bi-modal (image+text) queries. Typical ZS-CIR methods employ an inversion network to generate pseudo-word tokens that effectively represent the input semantics. However, the inversion-based methods suffer from two inherent issues: First, the task discrepancy exists because inversion training and CIR inference involve different objectives. Second, the modality discrepancy arises from the input feature distribution mismatch between training and inference. To this end, we propose a lightweight post-hoc framework, consisting of two components: (1) A new text-anchored triplet construction pipeline leverages a large language model (LLM) to transform a standard image-text dataset into a triplet dataset, where a textual description serves as the target of each triplet. (2) The MoTa-Adapter, a novel parameter-efficient fine-tuning method, adapts the dual encoder to the CIR task using our constructed triplet data. Specifically, on the text side, multiple sets of learnable task prompts are integrated via a Mixture-of-Experts (MoE) layer to capture task-specific priors and handle different types of modifications. On the image side, MoTa-Adapter modulates the inversion network's input to better match the downstream text encoder. In addition, an entropy-based optimization strategy is proposed to assign greater weight to challenging samples, thus improving adaptation efficiency. Experiments show that, with the incorporation of our proposed components, inversion-based methods achieve significant improvements, reaching state-of-the-art performance across four widely-used benchmarks.
Haiwen Li, Delong Liu, Zhaohui Hou, Zeliang Ma, Zhicheng Zhao 0001
AAAI6
2026 RAA: Achieving Interactive Remove/Add Anything via Fully Synthetic Data
abstract
Precise and controllable image editing, especially object removal and insertion, represents one of the most common demands in image manipulation. However, existing methods suffer from severe limitations. Mask-based inpainting often introduces visual artifacts and semantic inconsistencies, while instruction-based approaches lack accurate spatial control and tend to unintentionally modify background regions. To address these issues, we propose two key contributions. First, we develop a fully automated and self-improving pipeline for synthetic data generation. This pipeline utilizes a Large Language Model (LLM) to generate diverse prompts, a Diffusion Transformer (DiT) fine-tuned evolutionarily to synthesize high-quality images, and a Multimodal LLM (MLLM) combined with open-set object detector for automated quality control and annotation. This process produces the Remove/Add Dataset (RAD), consisting of over 514,510 high-quality image pairs, each richly annotated with bounding boxes, segmentation masks, and a variety of editing instructions. Second, based on RAD, we introduce Remove/Add Anything (RAA), a novel editing framework with precise spatial control. Built upon a diffusion-based inpainting model, RAA achieves high editing accuracy by conditioning on both textual instructions and an explicitly defined region of interest (ROI), enabling efficient fine-tuning while maintaining global visual coherence. Extensive experiments demonstrate that RAA significantly outperforms existing open-source methods on both addition and removal tasks, and even slightly surpasses costly proprietary models.
Delong Liu, Haotian Hou, Zhaohui Hou, Shihao Han, Mingjie Zhan, Zhicheng Zhao 0001
AAAI8
2026 Mind2Word: Towards generalized visual neural representations for high-quality video reconstruction
Haiwen Li, Zhu Meng, Zhicheng Zhao 0001
Expert Syst. Appl.5
2026 Now and future of artificial intelligence-based signet ring cell diagnosis: A survey
Zhu Meng, Junhao Dong 0002, Limei Guo, Guangxi Wang, Zhicheng Zhao 0001
Expert Syst. Appl.7
2026 MA3C: Multi-attribute aesthetic assessment and captioning
Le Ruan, Minghua Luo, Zhicheng Zhao 0001
Expert Syst. Appl.3
2026 Data-efficient generalization for zero-shot composed image retrieval
Zining Chen, Zhicheng Zhao 0001, Shijian Lu
Pattern Recognit.2
2025 Filter or Compensate: Towards Invariant Representation from Distribution Shift for Anomaly Detection
abstract
Recent Anomaly Detection (AD) methods have achieved great success with In-Distribution (ID) data. However, real-world data often exhibits distribution shift, causing huge performance decay on traditional AD methods. From this perspective, few previous work has explored AD with distribution shift, and the distribution-invariant normality learning has been proposed based on the Reverse Distillation (RD) framework. However, we observe the misalignment issue between the teacher and the student network that causes detection failure, thereby propose FiCo, Filter or Compensate, to address the distribution shift issue in AD. FiCo firstly compensates the distribution-specific information to reduce the misalignment between the teacher and student network via the Distribution-Specific Compensation (DiSCo) module, and secondly filters all abnormal information to capture distribution-invariant normality with the Distribution-Invariant Filter (DiIFi) module. Extensive experiments on three different AD benchmarks demonstrate the effectiveness of FiCo, which outperforms all existing state-of-the-art (SOTA) methods, and even achieves better results on the ID scenario compared with RD-based methods.
Zining Chen, Xingshuang Luo, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men
AAAI4
2025 UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
abstract
Recently, diffusion-based video generation models have achieved significant success. However, existing models often suffer from issues like weak consistency and declining image quality over time. To overcome these challenges, inspired by aesthetic principles, we propose a non-invasive plug-in called Uniform Frame Organizer (UFO), which is compatible with any diffusion-based video generation model. The UFO comprises a series of adaptive adapters with adjustable intensities, which can significantly enhance the consistency between the foreground and background of videos and improve image quality without altering the original model parameters when integrated. The training for UFO is simple, efficient, requires minimal resources, and supports stylized training. Its modular design allows for the combination of multiple UFOs, enabling the customization of personalized video generation models. Furthermore, the UFO also supports direct transferability across different models of the same specification without the need for specific retraining. The experimental results indicate that UFO effectively enhances video generation quality and demonstrates its superiority in public video generation benchmarks.
Delong Liu, Zhaohui Hou, Mingjie Zhan, Shihao Han, Zhicheng Zhao 0001
AAAI5
2025 Object-Centric Discriminative Learning for Text-Based Person Retrieval
abstract
Text-based person retrieval (TBPR) is a vision-language task that aims to find specific pedestrians in a large image gallery using the textual description. However, due to the heterogeneity between modalities and the redundancy in visual representations, it remains a challenging task. Existing methods do not explicitly reduce the influence of the background regions in images, inevitably decreasing representation ability and reducing the image-text matching performance. In this paper, we propose a novel framework for text-based person retrieval, termed Object-Centric Discriminative Learning (OCDL), which incorporates person masks to indicate attentive regions, thereby enhancing the model’s focus on the pedestrians in images while suppressing the background noise. Additionally, a novel crossmodal matching loss, namely Soft Angular Distribution Matching (SADM), is introduced to learn discriminative visual and textual representations. Extensive experiments on three widely-used TBPR datasets demonstrate the effectiveness of our approach. The code is available at https://github.com/JThuge/OCDL.
Haiwen Li, Delong Liu, Zhicheng Zhao 0001
ICASSP4
2025 Slot Inversion for Asymmetric Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) is a challenging vision-language (VL) task that retrieves target images using multi-modal (image+text) queries. Although significant progress has been made in existing CIR approaches, their deployment in resource-constrained scenarios remains problematic. To address this issue, we propose a novel framework, named Slot Inversion for Asymmetric Composed Image Retrieval (Slot4ACir), where an asymmetric retrieval scheme is adopted: lightweight models are employed on the query side, while large-scale VL models operate on the gallery side. Specifically, we introduce a lightweight inversion module based on slot attention, which maps an image into multiple textual tokens with distinct semantics. Additionally, the LLM sampler is proposed to facilitate richer semantic interactions, and the distillation alignment (DTA) loss enables the extraction of more informative representations. Extensive experiments on two popular benchmarks demonstrate the effectiveness of our approach. The code is now available at https://github.com/JThuge/Slot4ACir.
Haiwen Li, Zining Chen, Zhicheng Zhao 0001
ICME5
2025 CE-LoRA: Consistent Person Synthesis by Exploring the Model's Spatial Consistency
abstract
In image generation, a large number of studies focus on advanced network architectures so as to adapt to various consistency requirements. However, the training of these models usually relies on labor-intensive real-world data collection. Additionally, they overlook the model’s inherent ability to generate consistent outputs, such as naturally producing identical objects within a single image. In contrast, we propose a scalable self-distillation-based consistency data generation pipeline, which progressively enhances spatial consistency and ultimately enables the automatic generation of high-quality consistency images. Taking person consistency as a case study, we first generate a high-quality open-source dataset named Per-400K. Secondly, based on this dataset, the Consistency-Enhanced Low-Rank Adaptation (CE-LoRA) module is presented to learn the spatial consistency by incorporating the guidance image into the same spatial dimension as the target person image, achieving high-fidelity generation of person images. Experimental results demonstrate that CE-LoRA achieves state-of-the-art performance across multiple consistency generation metrics.
Delong Liu, Zhicheng Zhao 0001
ICME4
2025 Think Twice: Empowering Action Recognition Models with Human-Like Deep Reasoning
abstract
When engaged in complex visual cognition, humans tend to rely on their experience and make decisions after thinking again and again. Inspired by this, we pour similar capability into action recognition and propose a new Think Twice framework, that is, think twice about similar categories that are easy to confuse, thus obtaining performance improvement. Firstly, based on visual similarity, a large language model is applied to cluster all categories of a given dataset into disjoint cliques. Accordingly, a textual prompt for each clique will be generated. Secondly, through the first inference, pseudo-labels are obtained, and then the prompt corresponding to its clique is assigned to each sample. Thirdly, a prompt learning method is integrated to enable the framework to simulate human-like iterative thinking, yielding a final decision. Our proposed framework requires minimal parameters while achieving state-of-the-art parameter-efficient fine-tuning(PEFT) performance across four datasets. Our code is available at https://github.com/KangRuan6/ThinkTwice.
Xiangning Ruan, Baoxing Xie, Zhaohui Hou, Qixiang Yin, Zhicheng Zhao 0001
ICME6
2025 Automatic Synthetic Data and Fine-grained Adaptive Feature Alignment for Composed Person Retrieval
abstract
Person retrieval has attracted rising attention. Existing methods are mainly divided into two retrieval modes, namely image-only and text-only. However, they are unable to make full use of the available information and are difficult to meet diverse application requirements. To address the above limitations, we propose a new Composed Person Retrieval (CPR) task, which combines visual and textual queries to identify individuals of interest from large-scale person image databases. Nevertheless, the foremost difficulty of the CPR task is the lack of available annotated datasets. Therefore, we first introduce a scalable automatic data synthesis pipeline, which decomposes complex multimodal data generation into the creation of textual quadruples followed by identity-consistent image synthesis using fine-tuned generative models. Meanwhile, a multimodal filtering method is designed to ensure the resulting SynCPR dataset retains 1.15 million high-quality and fully synthetic triplets. Additionally, to improve the representation of composed person queries, we propose a novel Fine-grained Adaptive Feature Alignment (FAFA) framework through fine-grained dynamic alignment and masked feature reasoning. Moreover, for objective evaluation, we manually annotate the Image-Text Composed Person Retrieval (ITCPR) test set. The extensive experiments demonstrate the effectiveness of the SynCPR dataset and the superiority of the proposed FAFA framework when compared with the state-of-the-art methods. All code and data will be provided at https://github.com/Delong-liu-bupt/Composed_Person_Retrieval.
Delong Liu, Haiwen Li, Zhaohui Hou, Zhicheng Zhao 0001
NeurIPS4
2025 Automated text annotation: a new paradigm for generalizable text-to-image person retrieval
Delong Liu, Zhicheng Zhao 0001
Appl. Intell.3
2025 Token Embeddings Augmentation benefits Parameter-Efficient Fine-Tuning under long-tailed distribution
Weiqiu Wang, Zining Chen, Zhicheng Zhao 0001
Neurocomputing3
2025 PDFL: Progressive Discriminative Feature Learning for long-tailed recognition
Weiqiu Wang, Zining Chen, Zhicheng Zhao 0001
Neurocomputing3
2025 OpenDriver: An open-road driver state detection benchmark
Delong Liu, Zhu Meng, Zhicheng Zhao 0001
J. Netw. Comput. Appl.6
2025 MindShot: A few-shot brain decoding framework via transferring cross-subject prior and distilling frequency domain knowledge
Zhu Meng, Haiwen Li, Delong Liu, Zhicheng Zhao 0001
Knowl. Based Syst.6
2025 Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval
Delong Liu, Haiwen Li, Zhicheng Zhao 0001
Neural Networks3
2025 GLD: Global-Local Dynamic Frame Selection for Action Recognition
abstract
Video data exhibits significant redundancy. Existing frame sampling techniques, relying solely on global or local strategies, suffer from stability issue and struggle to balance efficiency with accuracy. To address these limitations, we propose a Global-local dynamic (GLD) frame selection method, which adaptively preserves continuous salient action clips while capturing distributed activities through cross-clip sampling. Specifically, we innovatively construct a global-local subsequence (GLS) by splicing the most salient clip and a fixed number of discrete frames randomly sampled from other segments. Then, the GLS is further condensed into the optimal subsequences using beam-pruned dynamic programming. Finally, a lightweight mapping function, comprising a Transformer layer and an MLP, is trained to align the full video features with those of the selected frames for inference. Our method dynamically balances global and local information, improving action recognition accuracy with lower inference costs. Extensive experiments on four datasets demonstrate that GLD outperforms existing sampling strategies.
Xiangning Ruan, Baoxing Xie, Qixiang Yin, Zhicheng Zhao 0001
IEEE Signal Process. Lett.5
2025 EVA: Enabling Video Attributes With Hierarchical Prompt Tuning for Action Recognition
abstract
The pretraining and fine-tuning paradigm has excelled in action recognition. However, full fine-tuning is computationally and storage costly, while parameter-efficient fine-tuning (PEFT) always sacrifices accuracy and stability. To address these challenges, we propose a novel method, Enabling Video Attributes with Hierarchical Prompt Tuning (EVA), to guide action recognition. Firstly, instead of focusing solely on temporal features, EVA sparsely extracts six types of video attributes across two modalities, capturing the relatively gradual attribute changes in actions. Secondly, a hierarchical prompt tuning architecture with multiscale attribute prompts is introduced to learn the differences in actions. Finally, by adjusting only a small number of additional parameters, EVA outperforms all PEFT and most full fine-tuning methods across four widely used datasets (Something-Something V2, ActivityNet, HMDB51, and UCF101), demonstrating its effectiveness.
Xiangning Ruan, Qixiang Yin, Zhicheng Zhao 0001
IEEE Signal Process. Lett.4
2025 S2FCNet: Semantic and Spatial Feature Compensation Network for Tiny-Object Detection
abstract
Tiny object detection (TOD) in remote sensing images remains an extremely challenging task, primarily due to the severely limited feature availability and susceptibility to interference from complex background. Recently, the multi-scale feature based methods have demonstrated effectiveness in tiny object detection. However, they often neglect that the low-level features struggle to activate the discriminative local semantics and lack global semantic information due to the limited local receptive fields. To address these issues, this paper proposes a Semantic and Spatial Feature Compensation Network (S2FCNet) for tiny object detection. To mitigate the gradual degradation of semantic information from high-level to low-level features in multi-scale representations, we propose a Local Semantic Reactivation Module (LSRM), which reactivates low-level local semantic features through top-down guidance from high-level semantic features. To enhance spatial perception capabilities, we develop a Foreground Spatial Sense Module (FSSM) that captures precise spatial location information, effectively suppresses the background noise and enhances the foreground features. Meanwhile, we introduce a Spatial Guidance Mechanism (SGM) to compensate for the loss of spatial awareness caused by downsampling operations. Additionally, we synergistically combine semantic and spatial features through a Multi-level Fusion Mechanism (MLFM), enabling more accurate detection of tiny objects. Extensive experiments on three challenging datasets demonstrate the effectiveness and superiority of the S2FCNet in comparison with the state-of-the-art methods. The code will be released at https://github.com/DetectionTiny/S2FCNet.
Yuhui Zhang 0005, Zhicheng Zhao 0001, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 DRFormer: A Discriminable and Reliable Feature Transformer for Person Re-Identification
abstract
As person image variations are likely to cause a part misalignment problem, most previous person Re-Identification (ReID) works may adopt local feature partition or additional landmark annotations to acquire aligned person features and boost ReID performance. However, such approaches either only achieve coarse-grained part alignments without considering detailed image variations within each part, or require extra annotated landmarks to train an available pose estimation model. In this work, we propose an effective Discriminable and Reliable Transformer (DRFormer) framework to learn part-aligned person representations with only person identity labels. Specifically, the DRFormer framework consists of Discriminable Feature Transformer (DFT) and Reliable Feature Transformer (RFT) modules, which generate discriminable and reliable high-order features, respectively. For reducing the dimension of high-order features, the DFT module utilizes a Self-Attentive Kronecker Product (SAKP) algorithm to promote the representational capabilities of compressed features via a self-attention strategy. For eliminating the background noise, the RFT module mines the foreground regions to adaptively aggregate foreground features via a Gumbel-Softmax strategy. Moreover, the proposed framework derives from an interpretable motivation and elegantly solves part misalignments without using feature partition or pose estimation. This paper theoretically and experimentally demonstrates the superiority of the proposed DRFormer framework, achieving state-of-the-art performance on various person ReID datasets.
Pingyu Wang, Xingjian Zheng, Linbo Qing, Bonan Li, Zhicheng Zhao 0001, Honggang Chen
IEEE Trans. Inf. Forensics Secur.6
2025 Bring Adaptive Binding Prototypes to Generalized Referring Expression Segmentation
abstract
Referring Expression Segmentation (RES), which aims to identify and segment objects based on natural language expressions is garnering increased research attention. While substantial progress has been made in RES, the emergence of Generalized Referring Expression Segmentation (GRES) introduces new challenges by allowing the expressions to describe multiple objects or lack specific object references. Existing RES methods usually rely on sophisticated encoder-decoder and feature fusion modules, and have difficulty generating class prototypes that match each instance individually when confronted with the complex referent and binary labels of GRES. In this paper, reevaluating the differences between RES and GRES, we propose a novel Model with Adaptive Binding Prototypes (MABP) that adaptively binds queries to object features in the corresponding region. It enables different query vectors to match instances of different categories, or different parts of the same instance, significantly expanding the decoder's flexibility, dispersing global pressure across all the queries, and easing the demands on the encoder. The experimental results demonstrate that MABP significantly outperforms the state-of-the-art methods in all three splits on the gRefCOCO dataset. Moreover, MABP outperforms the state-of-the-art methods on the RefCOCO+ and G-Ref datasets, and achieves very competitive results on RefCOCO. The code is available athttps://github.com/buptLwz/MABP.
Zhicheng Zhao 0001, Haochen Bai
IEEE Trans. Multim.2
2025 GAReID: Grouped and Attentive High-Order Representation Learning for Person Re-Identification
abstract
As person parts are frequently misaligned between detected human boxes, an image representation that can handle this part misalignment is required. In this work, we propose an effective grouped attentive re-identification (GAReID) framework to learn part-aligned and background robust representations for person re-identification (ReID). Specifically, the GAReID framework consists of grouped high-order pooling (GHOP) and attentive high-order pooling (AHOP) layers, which generate high-order image and foreground features, respectively. In addition, a novel grouped Kronecker product (GKP) is proposed to use both channel group and shuffle strategies for high-order feature compression, while promoting the representational capabilities of compressed high-order features. We show that our method derives from an interpretable motivation and elegantly reduces part misalignments without using landmark detection or feature partition. This article theoretically and experimentally demonstrates the superiority of the GAReID framework, achieving state-of-the-art performance on various person ReID datasets.
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Nikolaos V. Boulgouris
IEEE Trans. Neural Networks Learn. Syst.3
2024 Structural Information Guided Multimodal Pre-training for Vehicle-Centric Perception
abstract
Understanding vehicles in images is important for various applications such as intelligent transportation and self-driving system. Existing vehicle-centric works typically pre-train models on large-scale classification datasets and then fine-tune them for specific downstream tasks. However, they neglect the specific characteristics of vehicle perception in different tasks and might thus lead to sub-optimal performance. To address this issue, we propose a novel vehicle-centric pre-training framework called VehicleMAE, which incorporates the structural information including the spatial structure from vehicle profile information and the semantic structure from informative high-level natural language descriptions for effective masked vehicle appearance reconstruction. To be specific, we explicitly extract the sketch lines of vehicles as a form of the spatial structure to guide vehicle reconstruction. The more comprehensive knowledge distilled from the CLIP big model based on the similarity between the paired/unpaired vehicle image-text sample is further taken into consideration to help achieve a better understanding of vehicles. A large-scale dataset is built to pre-train our model, termed Autobot1M, which contains about 1M vehicle images and 12693 text information. Extensive experiments on four vehicle-based downstream tasks fully validated the effectiveness of our VehicleMAE. The source code and pre-trained models will be released at https://github.com/Event-AHU/VehicleMAE.
Xiao Wang 0014, Chenglong Li 0002, Zhicheng Zhao 0001, Zhe Chen 0013, Yukai Shi, Jin Tang 0001
AAAI4
2024 Prompting vision-language fusion for Zero-Shot Composed Image Retrieval
Zining Chen, Zhicheng Zhao 0001
ACML3
2024 PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization
abstract
Domain Generalization (DG) aims to resolve distribution shifts between source and target domains, and current DG methods are default to the setting that data from source and target domains share identical categories. Nevertheless, there exists unseen classes from target domains in practical scenarios. To address this issue, Open Set Domain Generalization (OSDG) has emerged and several methods have been exclusively proposed. However, most existing methods adopt complex architectures with slight improvement compared with DG methods. Recently, vision-language models (VLMs) have been introduced in DG following the fine-tuning paradigm, but consume huge training overhead with large vision models. Therefore, in this paper, we innovate to transfer knowledge from VLMs to lightweight vision models and improve the robustness by introducing Perturbation Distillation (PD) from three perspectives, including Score, Class and Instance (SCI), named SCI-PD. Moreover, previous methods are oriented by the benchmarks with identical and fixed splits, ignoring the divergence between source domains. These methods are revealed to suffer from sharp performance decay with our proposed new benchmark Hybrid Domain Generalization (HDG) and a novel metric H2-CV, which construct various splits to comprehensively assess the robustness of algorithms. Extensive experiments demonstrate that our method outperforms state-of-the-art algorithms on multiple datasets, especially improving the robustness when confronting data scarcity.
Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men, Hongying Meng
CVPR3
2024 iKUN: Speak to Trackers Without Retraining
abstract
Referring multi-object tracking (RMOT) aims to track multiple objects based on input textual descriptions. Previous works realize it by simply integrating an extra textual module into the multi-object tracker. However, they typically need to retrain the entire framework and have difficulties in optimization. In this work, we propose an insertable Knowledge Unification Network, termed iKUN, to enable communication with off-the-shelf trackers in a plug-and-play manner. Concretely, a knowledge unification module (KUM) is designed to adaptively extract visual features based on textual guidance. Meanwhile, to improve the localization accuracy, we present a neural version of Kalman filter (NKF) to dynamically adjust process noise and observation noise based on the current motion status. More-over, to address the problem of open-set long-tail distribution of textual descriptions, a test-time similarity calibration method is proposed to refine the confidence score with pseudo frequency. Extensive experiments on Refer-KITTI dataset verify the effectiveness of our framework. Finally, to speed up the development of RMOT, we also contribute a more challenging dataset, Refer-Dance, byex-tending public DanceTrack dataset with motion and dressing descriptions. The codes and dataset are available at https://github.com/dyhBUPT/iKUN.
Yunhao Du, Zhicheng Zhao 0001
CVPR3
2024 Dynamic Clustering and Cluster Contrastive Learning for Unsupervised Person Re-Id With Feature Distribution Alignment
abstract
Unsupervised Re-ID methods aim at learning robust and discriminative features from unlabeled data. However, existing methods often ignore the noise from distribution discrepancy during network training, which may lead to feature misalignment and hinder the model performance. To address this problem, we propose a Dynamic Clustering and Cluster Contrastive Learning (DCCC) method. Specifically, we first design a Dynamic Clustering Parameters Scheduler (DCPS) which adjust the clustering algorithm to fit the variation of feature distances to alleviate the distribution noise caused by unreasonable hyper-parameter settings in a global aspect. Then, a Dynamic Cluster Contrastive Learning (DyCL) method is proposed to tackle the distribution discrepancy in batch training with re-weighting allocation in a local aspect. We also introduce a Label Smoothing Soft Contrastive Loss (Lss) to combine the DyCL loss and self-supervised loss with low consumption and high efficiency on computing. Experiments on several public datasets validate the effectiveness of our proposed DCCC which outperforms previous state-of-the-art methods by achieving the best performance. Code is available at https://github.com/theziqi/DCCC.
Ziqi He, Mengjia Xue, Yunhao Du, Zhicheng Zhao 0001
ICASSP4
2024 Selective Cross-Correlation Consistency Loss for Out-of-Distribution Generalization
abstract
Deep learning methods usually succeed in independent and identically distributed (IID) data distribution, but suffer from sharp performance decay in real-world out-of-distribution (OOD) data. OOD generalization emerges to alleviate the large distribution shift between source and target domains. Recently, domain-invariant learning has boosted the research on OOD generalization, but most methods indulge complex architectures and training strategies. Hence, we propose a simple yet effective Selective Cross-Correlation Consistency (SC3) loss to align the cross-correlation matrix of features from identical categories. Specifically, we design the Semantic-Oriented Selection (SOS) algorithm in SC3loss to eliminate negative effects on spurious channels. Extensive experiments demonstrate that the SC3loss achieves superior performance on multiple OOD scenarios, including domain generalization (DG) and single domain gener-alization (SDG) tasks. Also, our loss consumes negligible computational resource which conforms to real-world applications. Source code is available at https://github.com/znchen666/SC3.
Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men
ICME3
2024 The Root Element of Human Poses is Radian: MCPRL is All You Need
abstract
3D Human Pose Estimation aims to determine the spatial coordinates of key anatomical landmarks on human body. Common benchmarks for this task include Human3.6M and MPI-INF-3DHP, while substantial redundancy exists. Additionally, both datasets are confined indoors due to equipment limitations, compromising diversity and generalization. To address the above issues, we first quantify dataset redundancy by introducing Generalized Radian Pruning (GRP), which employs a novel radians-based method to categorize and optimize human poses. Secondly, based on merging H36M and MPII datasets, we construct a new Manifold Cadre Poses with Radian List-systematically (MCPRL) dataset where missing outdoor scenarios, particularly dynamic collision actions are supplemented. Five strong baselines are compared and the experimental results demonstrate the effectiveness of the GRP method and the superiority of the MCPRL dataset. Our released dataset reduces data redundancy by 35%, surpasses H36M and MPII by 114% in diversity, with an average 11.19mm reduction in the MPJPE indicator. MCPRL dataset can be found at https://github.com/Rxn666/MCPRL.
Ziming Cheng, Xiangning Ruan, Qixiang Yin, Zhicheng Zhao 0001
ICME4
2024 MMFENet:Multi-Modal Feature Enhancement Network with Transformer for Human-Object Interaction Detection
abstract
Transformer model has been successfully applied to human-object interaction (HOI) detection in a one-stage mode. However, this mode has not fully leveraged rich and valuable clues such as scenes and linguistics etc, thereby weakening the Transformer’s representation power. In this paper, Transformer is introduced to a novel two-stage HOI detection framework, named MMFENet, where an interaction subnet and a feature enhancement subnet are constructed to jointly detect HOIs. Specifically, the interaction subnet is firstly constructed to effectively fuse visual (or semantic) features and spatial information, thus enhancing the relationships of all human-object pairs. In addition, the feature enhancement subnet is built to improve the contextual comprehension of interaction features by incorporating a diverse range of scene information encoded in both visual and linguistic modalities. Extensive experimental results show that MMFENet achieves very competitive results on the two public HOI detection benchmarks (HICO-DET and V-COCO).
Zhicheng Zhao 0001
IJCNN3
2024 Boundary-refined prototype generation: A general end-to-end paradigm for semi-supervised semantic segmentation
Junhao Dong 0002, Zhu Meng, Delong Liu, Zhicheng Zhao 0001
Eng. Appl. Artif. Intell.5
2024 Text-guided Fourier Augmentation for long-tailed recognition
Weiqiu Wang, Zining Chen, Zhicheng Zhao 0001
Pattern Recognit. Lett.4
2024 Instance Paradigm Contrastive Learning for Domain Generalization
abstract
Domain Generalization (DG) aims to develop models that can learn from data in source domains and generalize to unseen target domains. Recently, some domain generalization algorithms have emerged, but most of them were designed with complex modules. Among all the prior methods under DG settings, contrastive learning has become a promising solution for simplicity and efficiency. However, existing contrastive learning neglects distribution shifts that causes severe domain confusions. In this paper, we propose an instance paradigm contrastive learning framework, introducing contrast between original features and novel paradigms to alleviate domain-specific distractions. And then we explore hard-pair information, an essential factor in contrastive learning, based on domain label and feature similarity. Moreover, to produce domain-invariant instance paradigms, we generate multiple views of the original images and design a novel channel-wise attention mechanism to dynamically combine features from all the views. Furthermore, a test-time feature integration module is designed to mimic the paradigms during the training process to improve generalization ability. Extensive experiments show that our method achieves state-of-the-art performance. The proposed algorithm can also serve as a plug-and-play module which improves performance of existing methods with a relatively large margin.
Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men
IEEE Trans. Circuits Syst. Video Technol.3
2024 Video-Based Visible-Infrared Person Re-Identification With Auxiliary Samples
abstract
Visible-infrared person re-identification (VI-ReID) aims to match persons captured by visible and infrared cameras, allowing person retrieval and tracking in 24-hour surveillance systems. Previous methods focus on learning from cross-modality person images in different cameras. However, temporal information and single-camera samples tend to be neglected. To crack this nut, in this paper, we first contribute a large-scale VI-ReID dataset named BUPTCampus. Different from most existing VI-ReID datasets, it 1) collects tracklets instead of images to introduce rich temporal information, 2) contains pixel-aligned cross-modality sample pairs for better modality-invariant learning, 3) provides one auxiliary set to help enhance the optimization, in which each identity only appears in a single camera. Based on our constructed dataset, we present a two-stream framework as baseline and apply Generative Adversarial Network (GAN) to narrow the gap between the two modalities. To exploit the advantages introduced by the auxiliary set, we propose a curriculum learning based strategy to jointly learn from both primary and auxiliary sets. Moreover, we design a novel temporal k-reciprocal re-ranking method to refine the ranking list with fine-grained temporal correlation cues. Experimental results demonstrate the effectiveness of the proposed methods. We also reproduce 9 state-of-the-art image-based and video-based VI-ReID methods on BUPTCampus and our methods show substantial superiority to them. The codes and dataset are available at:https://github.com/dyhBUPT/BUPTCampus.
Yunhao Du, Zhicheng Zhao 0001
IEEE Trans. Inf. Forensics Secur.3
2024 NuSEA: Nuclei Segmentation With Ellipse Annotations
abstract
OBJECTIVE: Nuclei segmentation is a crucial pre-task for pathological microenvironment quantification. However, the acquisition of manually precise nuclei annotations for improving the performance of deep learning models is time-consuming and expensive. METHODS: In this paper, an efficient nuclear annotation tool called NuSEA is proposed to achieve accurate nucleus segmentation, where a simple but effective ellipse annotation is applied. Specifically, the core network U-Light of NuSEA is lightweight with only 0.86 M parameters, which is suitable for real-time nuclei segmentation. In addition, an Elliptical Field Loss and a Texture Loss are proposed to enhance the edge segmentation and constrain the smoothness simultaneously. RESULTS: Extensive experiments on three public datasets (MoNuSeg, CPM-17, and CoNSeP) demonstrate that NuSEA is superior to the state-of-the-art (SOTA) methods and better than existing algorithms based on point, rectangle, and text annotations. CONCLUSIONS: With the assistance of NuSEA, a new dataset called NuSEA-dataset v1.0, encompassing 118,857 annotated nuclei from the whole-slide images of 12 organs is released. SIGNIFICANCE: NuSEA provides a rapid and effective annotation tool for nuclei in histopathological images, benefiting future explorations in deep learning algorithms.
Zhu Meng, Junhao Dong 0002, Binyu Zhang, Ruixiao Wu, Guangxi Wang, Limei Guo, Zhicheng Zhao 0001
IEEE J. Biomed. Health Informatics9
2024 MTCSNet: One-Stage Learning and Two-Point Labeling are Sufficient for Cell Segmentation
abstract
Deep convolution neural networks have been widely used in medical image analysis, such as lesion identification in whole-slide images, cancer detection, and cell segmentation, etc. However, it is often inevitable that researchers try their best to refine annotations so as to enhance the model performance, especially for cell segmentation task. Weakly supervised learning can greatly reduce the workload of annotations, while there is still a huge performance gap between the weakly and fully supervised learning approaches. In this work, we propose a weakly-supervised cell segmentation method, namely Multi-Task Cell Segmentation Network (MTCSNet), for multi-modal medical images, including pathological, brightfield, fluorescent, phase-contrast and differential interference contrast images. MTCSNet is learnt in a single-stage training manner, where only two annotated points for each cell provide supervision information, and the first one is the centroid, the second one is its boundary. Additionally, five auxiliary tasks are elaborately designed to train the network, including two pixel-level classifications, a pixel-level regression, a local temperature scaling and an instance-level distance regression task, which is proposed to regress the distances between the cell centroid and its boundaries in eight orientations. The experimental results indicate that our method outperforms all state-of-the-art weakly-supervised cell segmentation approaches on public multi-modal medical image datasets. The promising performance also shows that a single-stage learning with two-point labeling approach are sufficient for cell segmentation, instead of fine contour delineation. The codes are available at: https://github.com/binging512/MTCSNet.
Binyu Zhang, Zhu Meng, Hongyuan Li, Zhicheng Zhao 0001
IEEE Trans. Medical Imaging4
2024 Cluster-Instance Normalization: A Statistical Relation-Aware Normalization for Generalizable Person Re-Identification
abstract
Person re-identification (ReID) has achieved great improvement under supervised settings, but suffers from considerable degradation when large distribution shifts between training and testing sets exist. Domain generalization (DG ReID) emerges to promote the generalization ability of models, overcoming the distribution shifts issue between source domains and unseen target domains. Among most prior methods in DG ReID, instance normalization (IN) serves as a promising solution for removing domain-specific information, however, it damages the discriminative ability simultaneously. In this article, we propose a new normalization method called Cluster-Instance Normalization (CINorm) to extract information from clusters for information compensation. The relations between samples in a batch can be mined to establish evolving clusters with aggregated samples during the forward training process. In this way, high intra-cluster congregation can eliminate the impacts of outliers to avoid overfitting, and high inter-cluster variances can synthesize diverse novel statistics to compensate discriminative information. Therefore, a Relation-Aware Normalization (RANorm) with a Dynamic ReCalibration (DRC) module is designed to integrate normalized features between evolving clusters and instances efficiently. Furthermore, a novel Group-based Triplet (G-Triplet) loss is proposed to divide a batch into multiple groups with greater compactness for hard-pair mining. Extensive experiments show that our method outperforms state-of-the-art algorithms on multiple DG benchmarks by a large margin. The proposed method can also achieve superior performance on image classification tasks under DG settings without using domain labels.
Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men
IEEE Trans. Multim.3
2023 MotionMLP: End-to-End Action Recognition with Motion Aware Vision MLP
abstract
Action recognition aims to interpret complex spatiotemporal patterns in the video. Current methods utilize CNN or Transformer structures, requiring extensive pre-training methods and optical flow to capture motion information. Such approaches are computationally expensive, necessitate significant storage, cannot be trained end-to-end, and typically neglect joint learning of temporal and spatial streams. In this paper, we propose MotionMLP, a novel MLP architecture that extracts motion information from videos, then dynamically adjusts the connection between tokens and static weights within the MLP structure. The MotionMLP solely relies on video frames as input and is independent of any pre-training method or optical flow computation. The experimental results indicate that MotionMLP outperforms the previous SOTA real-time end-to-end methods on UCF101 and HMDB51, and relies on one-tenth of the parameters compared with typical two-stream CNN approaches, while operating ten times faster.
Xiangning Ruan, Zhicheng Zhao 0001
VCIP2
2023 SwinDAE: Electrocardiogram Quality Assessment Using 1D Swin Transformer and Denoising AutoEncoder
abstract
OBJECTIVE: Electrocardiogram (ECG) signals have wide-ranging applications in various fields, and thus it is crucial to identify clean ECG signals under different sensors and collection scenarios. Despite the availability of a variety of deep learning algorithms for ECG quality assessment, these methods still lack generalization across different datasets, hindering their widespread use. METHODS: In this paper, an effective model named Swin Denoising AutoEncoder (SwinDAE) is proposed. Specifically, SwinDAE uses a DAE as the basic architecture, and incorporates a 1D Swin Transformer during the feature learning stage of the encoder and decoder. SwinDAE was first pre-trained on the public PTB-XL dataset after data augmentation, with the supervision of signal reconstruction loss and quality assessment loss. Specially, the waveform component localization loss is proposed in this paper and used for joint supervision, guiding the model to learn key information of signals. The model was then fine-tuned on the finely annotated BUT QDB dataset for quality assessment. RESULTS: SwinDAE achieved 0.02-0.13 mean F1 score improvement on the BUT QDB dataset compared to multiple deep learning methods, and demonstrated applicability on two other datasets. CONCLUSION: The proposed SwinDAE shows strong generalization ability on different datasets, and surpasses other state-of-the-art deep learning methods on multiple evaluation metrics. In addition, the statistical analysis for SwinDAE prove the significance of the performance and the rationality of the prediction. SIGNIFICANCE: SwinDAE can learn the commonality between high-quality ECG signals, exhibiting excellent performance in the application of cross-sensors and cross-collection scenarios.
Baoxing Xie, Zhicheng Zhao 0001, Zhu Meng, Yadong Huang
IEEE J. Biomed. Health Informatics4
2023 StrongSORT: Make DeepSORT Great Again
abstract
Recently, Multi-Object Tracking (MOT) has attracted rising attention, and accordingly, remarkable progresses have been achieved. However, the existing methods tend to use various basic models (e.g, detector and embedding model), and different training or inference tricks, etc. As a result, the construction of a good baseline for a fair comparison is essential. In this paper, a classic tracker, i.e., DeepSORT, is first revisited, and then is significantly improved from multiple perspectives such as object detection, feature embedding, and trajectory association. The proposed tracker, named StrongSORT, contributes a strong and fair baseline for the MOT community. Moreover, two lightweight and plug-and-play algorithms are proposed to address two inherent “missing” problems of MOT: missing association and missing detection. Specifically, unlike most methods, which associate short tracklets into complete trajectories at high computation complexity, we propose an appearance-free link model (AFLink) to perform global association without appearance information, and achieve a good balance between speed and accuracy. Furthermore, we propose a Gaussian-smoothed interpolation (GSI) based on Gaussian process regression to relieve the missing detection. AFLink and GSI can be easily plugged into various trackers with a negligible extra computational cost (1.7 ms and 7.1 ms per image, respectively, on MOT17). Finally, by fusing StrongSORT with AFLink and GSI, the final tracker (StrongSORT++) achieves state-of-the-art results on multiple public benchmarks, i.e., MOT17, MOT20, DanceTrack and KITTI. Codes are available athttps://github.com/dyhBUPT/StrongSORTandhttps://github.com/open-mmlab/mmtracking.
Yunhao Du, Zhicheng Zhao 0001, Yang Song 0036, Yanyun Zhao, Hongying Meng
IEEE Trans. Multim.2
2023 LTReID: Factorizable Feature Generation With Independent Components for Long-Tailed Person Re-Identification
abstract
With the rapid increase of large-scale and real-world person datasets, it is crucial to address the problem of long-tailed data distributions,i.e., head classes have large number of images while tail classes occupy extremely few samples. We observe that the imbalanced data distribution is likely to distort the overall feature space and impair the generalization capability of trained models. Nevertheless, this long-tailed problem has been rarely investigated in previous person Re-Identification (ReID) works. In this paper, we propose a novelLong-Tailed Re-Identification(LTReID) framework to simultaneously alleviate class-imbalance and hard-imbalance problems. Specifically, each real feature is decomposed into multiple independent components with two decorrelation losses. Then these components are randomly aggregated to generate more fake features for tail classes than head ones, resulting in the class-balance between head and tail classes. For the hard-balance between easy and hard samples, we utilize adversarial learning to generate more hard features than easy ones. The proposed framework can be trained in an end-to-end manner and avoids increasing the space and time complexity of inference models. Moreover, comprehensive experiments are conducted on the four ReID datasets so as to validate the effectiveness of the overall framework and the advantage of each module. Our results show that when trained with either balanced or imbalanced datasets, the LTReID achieves superior performance over the state-of-the-art methods.
Pingyu Wang, Zhicheng Zhao 0001, Hongying Meng
IEEE Trans. Multim.2
2022 Change Detection Converter: Using Semantic Segmantation Models to Tackle Change Detection Task
abstract
In recent years, change detection (CD) in remote sensing images has achieved a huge success by using deep learning, and it is essentially a subtask of semantic segmentation. Both of them are tightly related. However, the complementarity of them has not been fully explored. In this paper, we propose a new change detection converter (CDC), which can be easily inserted into the existing semantic segmentation networks and be transformed into new networks suitable for change detection task. In addition, a new task-specific data augmentation method named Crossover is proposed, which can enhance the feature representation ability in sequence-form features without any additional cost. Extensive experiments demonstrate that the proposed method obtains significantly better performance and more efficiency than previous methods on two public datasets, and can achieve continuous performance improvement than other complex-designed solutions.
Bowei Ye, Zhicheng Zhao 0001, Feihong Wang, Weida Xu, Wenjun Yin
ICME3
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR32
2022 iCGPN: Interaction-centric graph parsing network for human-object interaction detection
Zhicheng Zhao 0001, Hongying Meng
Neurocomputing3
2022 Attentive Feature Augmentation for Long-Tailed Visual Recognition
abstract
Deep neural networks have achieved great success on many visual recognition tasks. However, training data with a long-tailed distribution dramatically degenerates the performance of recognition models. In order to relieve this imbalance problem, an effective Long-Tailed Visual Recognition (LTVR) framework is proposed based on learned balance and robust features under long-tailed distribution circumstances. In this framework, a plug-and-play Attentive Feature Augmentation (AFA) module is designed to mine class-related and variation-related features of original samples via a novel hierarchical channel attention mechanism. Then, those features are aggregated to synthesize fake features to cope with the imbalance of the original dataset. Moreover, a Lay-Back Learning Schedule (LBLS) is developed to ensure a good initialization of feature embedding. Extensive experiments are conducted with a two-stage training method to verify the effectiveness of the proposed framework on both feature learning and classifier rebalancing in the long-tailed image recognition task. Experimental results show that, when trained with imbalanced datasets, the proposed framework achieves superior performance over the state-of-the-art methods.
Weiqiu Wang, Zhicheng Zhao 0001, Pingyu Wang, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.2
2021 Pruned-YOLO: Learning Efficient Object Detector Using Model Pruning
Pingyu Wang, Zhicheng Zhao 0001
ICANN (4)3
2021 Diagnosing Covid-19 from CT Images Based on an Ensemble Learning Framework
abstract
Research on automated diagnosis of Coronavirus Disease 2019 (COVID-19) has increased in recent months. SPGC COVID19 aims at classifying the grouped images of the same patient into COVID, Community Acquired Pneumonia(CAP) or normal. In this paper, we propose a novel ensemble learning framework to solve this problem. Moreover, adaptive boosting and dataset clustering algorithms are introduced to improve the classification performance. In our experiments, we demonstrate that our framework is superior to existing networks in terms of both accuracy and sensitivity.
Yinan Song, Zhicheng Zhao 0001, Zhu Meng
ICASSP4
2021 SANet++: Enhanced Scale Aggregation with Densely Connected Feature Fusion for Crowd Counting
abstract
Crowd counting has gained considerable attention recently but remains challenging mainly due to large scale variations. In this paper, we present SANet++ with a novel architecture to generate high-quality density maps and further perform accurate counting. SANet++ obtains enhanced multi-scale representation with densely connected feature fusion between branches. Our approach avoids information redundancy while exploits complementary features at different scales. In addition, we introduce a novel Bulk loss which incorporates the spatial correlation within a whole patch. This global structural supervision enforces the network to learn the interactions between pixels without limitations on region size. Our SANet++ outperforms state-of-the-art crowd counting approaches according to extensive experiments conducted on three major datasets.
Siyang Pan, Yanyun Zhao, Zhicheng Zhao 0001
ICASSP4
2021 GSLD: A Global Scanner with Local Discriminator Network for Fast Detection of Sparse Plasma Cell in Immunohistochemistry
abstract
Compared with abundant application of deep learning on hematoxylin and eosin (H&E) images, the study on immunohistochemical (IHC) images is almost blank, while the diagnosis of chronic endometritis mainly relies on the detection of plasma cells in IHC images. In this paper, a novel framework named Global Scanner with Local Discriminator (GSLD) is proposed to detect plasma cells with highly sparse distribution in IHC whole slide images (WSI) effectively and efficiently. Firstly, input an IHC image, the Global Scanner subnetwork (GSNet) predicts a distribution map, where the candidate plasma cells are localized quickly. Secondly, based on the distribution map, the Local Discriminator subnetwork (LDNet)discriminates true plasma cells by adopting only local information, which greatly speeds up the detection. Moreover, a novel grid-oversampling strategy for WSI preprocessing is proposed to relieve sample imbalance problem. Experimentas show that the proposed framework outperforms the representative object detection networks in both speed and accuracy.
Zhu Meng, Zhicheng Zhao 0001
ICIP3
2021 A Method of Stable Long-Term Single Object Tracking
abstract
We propose a stable long-term tracking method to deal with visual tracking in multiple complex scenarios to solve the problem of frequent disappearance and reappearance of targets in long-term tracking. In our method, we do not blindly start the global tracker once the target disappears, but use it only when necessary and in reasonable scope with the assistance of localization module. In addition, we designed an FP-verifier based on feature pools to reevaluate the candidate bounding boxes given by our local tracker and global tracker to ensure that online learning local tracker can be updated stably. Our method outperforms the state-of-the-art results on the VOT2020LT challenge. In addition, the control experiments show that our FP-verifier is more effective than the RT-MDNet verifier used by the top three winners of VOT2020LT challenge.
Zitong Yi, Zhihang Tong, Yanyun Zhao, Zhicheng Zhao 0001
ICME4
2021 Rethinking Anchor-Object Matching and Encoding in Rotating Object Detection
abstract
Rotating object detection is more challenging than horizontal object detection because of the multi-orientation of the objects involved. In the recent anchor-based rotating object detector, the IoU-based matching mechanism has some mismatching and wrong-matching problems. Moreover, the encoding mechanism does not correctly reflect the location relationships between anchors and objects. In this paper, RBox-Diff-based matching (RDM) mechanism and angle-first encoding (AE) method are proposed to solve these problems. RDM optimizes the anchor-object matching by replacing IoU (Intersection-over-Union) with a new concept called RBox-Diff, while AE optimizes the encoding mechanism to make the encoding results consistent with the relative position between objects and anchors more. The proposed methods can be easily applied to most of the anchor-based rotating object detectors without introducing extra parameters. The extensive experiments on DOTA-v1.0 dataset show the effectiveness of the proposed methods over other advanced methods.
Zhaohui Hou, Pingyu Wang, Zhicheng Zhao 0001
VCIP5
2021 Dynamic proposal sampling for weakly supervised object detection
Wenhui Jiang 0001, Zhicheng Zhao 0001, Yuming Fang 0001
Neurocomputing2
2021 HOReID: Deep High-Order Mapping Enhances Pose Alignment for Person Re-Identification
abstract
Despite the remarkable progress in recent years, person Re-Identification (ReID) approaches frequently fail in cases where the semantic body parts are misaligned between the detected human boxes. To mitigate such cases, we propose a novel High-Order ReID (HOReID) framework that enables semantic pose alignment by aggregating the fine-grained part details of multilevel feature maps. The HOReID adopts a high-order mapping of multilevel feature similarities in order to emphasize the differences of the similarities between aligned and misaligned part pairs in two person images. Since the similarities of misaligned part pairs are reduced, the HOReID enhances pose-robustness within the learned features. We show that our method derives from an intuitive and interpretable motivation and elegantly reduces the misalignment problem without using any prior knowledge from human pose annotations or pose estimation networks. This paper theoretically and experimentally demonstrates the effectiveness of the proposed HOReID, achieving superior performance over the state-of-the-art methods on the four large-scale person ReID datasets.
Pingyu Wang, Zhicheng Zhao 0001, Xingyu Zu, Nikolaos V. Boulgouris
IEEE Trans. Image Process.2
2021 Triple Up-Sampling Segmentation Network With Distribution Consistency Loss for Pathological Diagnosis of Cervical Precancerous Lesions
abstract
OBJECTIVE: Cervical cancer, as one of the most frequently diagnosed cancers in women, is curable when detected early. However, automated algorithms for cervical pathology precancerous diagnosis are limited. METHODS: In this paper, instead of popular patch-wise classification, an end-to-end patch-wise segmentation algorithm is proposed to focus on the spatial structure changes of pathological tissues. Specifically, a triple up-sampling segmentation network (TriUpSegNet) is constructed to aggregate spatial information. Second, a distribution consistency loss (DC-loss) is designed to constrain the model to fit the inter-class relationship of the cervix. Third, the Gauss-like weighted post-processing is employed to reduce patch stitching deviation and noise. RESULTS: The algorithm is evaluated on three challenging and public datasets: 1) MTCHI for cervical precancerous diagnosis, 2) DigestPath for colon cancer, and 3) PAIP for liver cancer. The Dice coefficient is 0.7413 on the MTCHI dataset, which is significantly higher than the published state-of-the-art results. CONCLUSION: Experiments on the public dataset MTCHI indicate the superiority of the proposed algorithm on cervical pathology precancerous diagnosis. In addition, the experiments on two other pathological datasets, i.e., DigestPath and PAIP, demonstrate the effectiveness and generalization ability of the TriUpSegNet and weighted post-processing on colon and liver cancers. SIGNIFICANCE: The end-to-end TriUpSegNet with DC-loss and weighted post-processing leads to improved segmentation in pathology of various cancers.
Zhu Meng, Zhicheng Zhao 0001, Limei Guo, Haiying Wang 0005
IEEE J. Biomed. Health Informatics2
2021 A Cervical Histopathology Dataset for Computer Aided Diagnosis of Precancerous Lesions
abstract
Cervical cancer, as one of the most frequently diagnosed cancers worldwide, is curable when detected early. Histopathology images play an important role in precision medicine of the cervical lesions. However, few computer aided algorithms have been explored on cervical histopathology images due to the lack of public datasets. In this article, we release a new cervical histopathology image dataset for automated precancerous diagnosis. Specifically, 100 slides from 71 patients are annotated by three independent pathologists. To show the difficulty of the task, benchmarks are obtained through both fully and weakly supervised learning. Extensive experiments based on typical classification and semantic segmentation networks are carried out to provide strong baselines. In particular, a strategy of assembling classification, segmentation, and pseudo-labeling is proposed to further improve the performance. The Dice coefficient reaches 0.7833, indicating the feasibility of computer aided diagnosis and the effectiveness of our weakly supervised ensemble algorithm. The dataset and evaluation codes are publicly available. To the best of our knowledge, it is the first public cervical histopathology dataset for automated precancerous segmentation. We believe that this work will attract researchers to explore novel algorithms on cervical automated diagnosis, thereby assisting doctors and patients clinically.
Zhu Meng, Zhicheng Zhao 0001, Limei Guo
IEEE Trans. Medical Imaging2
2021 Deep Multi-Patch Matching Network for Visible Thermal Person Re-Identification
abstract
Visible Thermal Person Re-Identification(VTReID) is a cross-modality retrieval problem in computer vision. Accurate VTReID is very challenging due to large modality discrepancies. In this work, we design a novelMulti-Patch Matching Network(MPMN) framework to simultaneously mitigate the heterogeneity of coarse-grained and fine-grained visual semantics. In view of cross-modality matching, we verify that aligning modality distributions of the original features is likely to suffer from the selective alignment behavior, i.e., only focuses on easiest dimensions or subspaces. Inspired by adversarial learning, we propose a newMulti-Patch Modality Alignment(MPMA) loss to jointly balance and reduce the modality discrepancies of multi-patch features by mining hard subspaces and abandoning easy subspaces. Since multi-patch features are potentially complementary to each other, the semantic correlations between different patches should be exploited during training. Motivated by knowledge distillation, we put forward a newCross-Patch Correlation Distillation(CPCD) loss to transfer the semantic knowledges across different patches. To balance multi-patch tasks, an effectivePatch-Aware Priority Attention(PAPA) method is further introduced to dynamically prioritize hard patch tasks during training. This paper experimentally demonstrates the effectiveness of the proposed methods, achieving superior performance over the state-of-the-art methods on RegDB and SYSU-MM01 datasets.
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Haiying Wang 0005, Lei Yang 0063
IEEE Trans. Multim.2
2020 Adaptive Elastic Loss Based on Progressive Inter-Class Association for Cervical Histology Image Segmentation
abstract
Cervical cancer is one of the most commonly diagnosed cancer types worldwide, while is curable if detected early. However, few computer-aided algorithms have been explored on cervical histology image, which is vital for abnormality assessment. In this paper, an end-to-end deep segmentation network for complex cervical histology images is proposed, and a benchmark evaluation is contributed. Specifically, we observe that four-category cervical histology images possess a progressive inter-class association. To model the relationship, inspired by the elasticity, an adaptive elastic loss is proposed to reduce the deviation between difficult samples and their true categories. Moreover, five evaluation metrics are designed to measure the segmentation performance, and the Window Precision is particularly valuable for the evaluation of semi-supervised algorithms due to its robustness to the mislabeling. Finally, on a cervical histology dataset, benchmark experiments based on deep networks are conducted, and the results demonstrate the superiority of our new loss.
Zhu Meng, Zhicheng Zhao 0001, Weibao Wang
ICASSP2
2020 Human-Centric Parsing Network for Human-Object Interaction Detection
abstract
Human-object interactions detection is an essential task of image inference, but current methods can't efficiently make use of global knowledge in the image. To tackle this challenge, in this paper, we propose a Human-Centric Parsing Network (HCPN), which integrates global structural knowledge to infer human-object interactions. In HCPN, a semantic parse graph is first constructed by binding human-object relationships, edge features and node features, where the detected human box in image is regarded as the center node and other detected boxes are linked to it. Second, based on the message passing mechanism, edge features and node features with the relation graph are updated and finally, HCPN predicts human-object interactions and associated locations by a readout function. We evaluate our model on V-COCO dataset, and a great improvement is achieved compared with state-of-the-art methods.
Zhicheng Zhao 0001
ICPR3
2020 Triplet-path Dilated Network for Detection and Segmentation of General Pathological Images
abstract
Deep learning has been widely applied in the field of medical image processing. However, compared with flourishing visual tasks in natural images, the progress achieved in pathological images is not remarkable, and detection and segmentation, which are among basic tasks of computer vision, are regarded as two independent tasks. In this paper, we make full use of existing datasets and construct a triplet-path network using dilated convolutions to cooperatively accomplish one-stage object detection and nuclei segmentation for general pathological images. First, in order to meet the requirement of detection and segmentation, a novel structure called triplet feature generation (TFG) is designed to extract high-resolution and multiscale features, where features from different layers can be properly integrated. Second, considering that pathological datasets are usually small, a location-aware and partially truncated loss function is proposed to improve the classification accuracy of datasets with few images and widely varying targets. We compare the performance of both object detection and instance segmentation with state-of-the-art methods. Experimental results demonstrate the effectiveness and efficiency of the proposed network on two datasets collected from multiple organs.
Jiaqi Luo, Zhicheng Zhao 0001, Limei Guo
ICPR2
2020 Efficient-Receptive Field Block with Group Spatial Attention Mechanism for Object Detection
abstract
Object detection has been paid rising attention in computer vision field. Convolutional Neural Networks (CNNs) extract high-level semantic features of images, which directly determine the performance of object detection. As a common solution, embedding integration modules into CNNs can enrich extracted features and thereby improve the performance. However, the instability and inconsistency of internal multiple branches exist in these modules. To address this problem, we propose a novel multibranch module called Efficient-Receptive Field Block (E-RFB), in which multiple levels of features are combined for network optimization. Specifically, by downsampling and increasing depth, the E-RFB provides sufficient RF. Second, in order to eliminate the inconsistency across different branches, a novel spatial attention mechanism, namely, Group Spatial Attention Module (GSAM) is proposed. The GSAM gradually narrows a feature map by channel grouping; thus it encodes the information between spatial and channel dimensions into the final attention heat map. Third, the proposed module can be easily joined in various CNNs to enhance feature representation as a plug-and-play component. With SSD-style detectors, our method halves the parameters of the original detection head and achieves high accuracy on the PASCAL VOC and MS COCO datasets. Moreover, the proposed method achieves superior performance compared with state-of-the-art methods based on similar framework.
Zhicheng Zhao 0001
ICPR2
2020 Deep hard modality alignment for visible thermal person re-identification
Pingyu Wang, Zhicheng Zhao 0001, Yanyun Zhao, Lei Yang 0063
Pattern Recognit. Lett.3
2019 Multiple Saliency and Channel Sensitivity Network for Aggregated Convolutional Feature
abstract
In this paper, aiming at two key problems of instance-level image retrieval, i.e., the distinctiveness of image representation and the generalization ability of the model, we propose a novel deep architecture - Multiple Saliency and Channel Sensitivity Network(MSCNet). Specifically, to obtain distinctive global descriptors, an attention-based multiple saliency learning is first presented to highlight important details of the image, and then a simple but effective channel sensitivity module based on Gram matrix is designed to boost the channel discrimination and suppress redundant information. Additionally, in contrast to most existing feature aggregation methods, employing pre-trained deep networks, MSCNet can be trained in two modes: the first one is an unsupervised manner with an instance loss, and another is a supervised manner, which combines classification and ranking loss and only relies on very limited training data. Experimental results on several public benchmark datasets, i.e., Oxford buildings, Paris buildings and Holidays, indicate that the proposed MSCNet outperforms the state-of-the-art unsupervised and supervised methods.
Xuanlu Xiang, Zhipeng Wang 0008, Zhicheng Zhao 0001
AAAI3
2019 Jointly Predicting Future Sequence and Steering Angles for Dynamic Driving Scenes
abstract
Generative Adversarial Network (GAN) has attracted rising attention for video future sequence prediction in driving scenes. However, the images generated by GAN often miss the target for lack of any constraints for its generated target. In this paper, an encoder-decoder based multi-task video prediction network - SegVAE is proposed by simultaneously accomplishing the predictions (generations) of both future sequence and steering angles for egocentric driving videos at pixel-level. Specifically, the encoder is constructed based on Varitional Auto-Encoder (VAE) to learn the complex latent distribution of real driving scenes. The decoder is exploited with a multi-task manner to jointly predict the future sequence and steering angles of dynamic driving scenes, where an enhanced generation mechanism is also proposed. Varitional Auto-Encoder (VAE) and Long Short Term Memory Networks (LSTM) are introduced to optimize the learning of SegVAE. The experimental results on public KITTI and NVIDIA driving datasets indicate that the proposed Seg-VAE can effectively mimic humans prediction mechanism, and outperform standard VAE and CNN-based generative adversarial network.
Zhicheng Zhao 0001, Leiquan Wang, Chaohong An
ICASSP2
2019 Multi-classification of Breast Cancer Histology Images by Using Gravitation Loss
abstract
The scarcity of professional doctors stimulates the progress of breast cancer classification. However, there are still numerous challenges such as varied appearances (color, texture etc.) of microscopy images and the ambiguous category boundaries. In this paper, we propose an efficient and effective method to achieve multi-classification for H&E stained breast cancer images. Firstly, to restrain color noises in the staining stage, data augmentation in HSV color space is used to increase the diversity of color distribution. In addition, inspired by the principle of gravitation, a Gravitation Loss (G-loss) is proposed to maximize inter-class difference and minimize intra-class variance. The experimental results on public BACH 2018 dataset indicate that the proposed algorithm achieves the state-of-the-art performance, which demonstrates its effectiveness.
Zhu Meng, Zhicheng Zhao 0001
ICASSP2
2019 Pedestrian Attribute Recognition Based on Mtcnn with Online Batch Weighted Loss
abstract
Due to the large number and huge diversity of attributes, pedestrian attribute recognition in video surveillance scenarios is a challenging task in the field of computer vision. Different from most previous works which only focus on extremely imbalanced attribute distribution problem, a new grouping way of attributes based multi-task convolutional neural network (MTCNN) is put forward, which exploits the spatial correlations among attributes and guarantees some independence of each attribute as well. Meanwhile, we propose a novel online batch weighted loss to narrow the performance differences among attributes and boost the model to gain a higher average recognition accuracy. The whole network can be trained end to end, and experimental results on PETA and RAP datasets show that our method achieves significant performance, comparing with those state-of-the-art methods.
Xingting He, Qiuyue Shi, Zhicheng Zhao 0001, Bojin Zhuang
ICIP4
2019 An End-to-End Future Frame Prediction Method for Vehicle-Centric Driving Videos
abstract
In the field of autonomous driving, training an agent to watch and think similar to human drivers is an efficient way to solve self-driving problems. Inspired by NVIDIA's frame-level command generation task [2] and the discovery of humans memory capacity [13], we propose a future frame prediction method for vehicle-centric driving videos. An end-to-end deep learning architecture called future frame prediction (FFPRE) network is proposed, which can generate a future frame following the input video sequence. In particular, we develop a general memory preserving module to extract meaningful history information from input data. This module consists of two parts, namely, memory recall and memory refine. We train this module to generate the short-term spatiotemporal information of a given video batch, which is a concatenation of history appearance and temporal clues. Thereafter, the two history clues will be transformed into future representations by a long-term prediction module. Thus, humans' driving prediction progress is mimicked in a completely modular manner. Given the FFPRE network's effective long-short spatiotemporal feature learning ability, the proposed network can construct an internal representation (content and dynamic) of vehicle-centric driving videos without tracking the trajectory of every pixel. Experimental results on publicly released datasets of NVIDIA and DR(eye)VE indicate that our proposed method is efficient.
Kaikun Ji, Zhicheng Zhao 0001, Bojin Zhuang
VCIP3
2019 GRNet: An Efficient Group Ranking Network for Face Attribute Recognition
abstract
Face attribute recognition methods have been far from real-world applications despite remarkable progress in recent years. Most of these methods fail to mine attribute relationships with traditional cross entropy loss and hold considerable computational complexity. In this work, we propose a group ranking network (GRNet) to investigate substantial relationships among face attributes. First, an attribute grouping manner is designed to capture interactions among spatially related attributes and enhance computational efficiency. Second, a new supervision signal is presented to model attribute ranking relationships. Under the supervision of ranking loss, GRNet learns intra-attribute and inter-attribute ranking features to intensify the representational capability of attribute models. We evaluate our approach on aligned and unaligned CelebA datasets. Results show that the performance of the proposed approach is superior to other start-of-the-art methods on attribute recognition.
Dingchang Hu, Pingyu Wang, Zhicheng Zhao 0001
VCIP5
2019 Vehicle Re-Identification: Logistic Triplet Embedding Regularized by Label Smoothing
abstract
The explosive increasing of vehicles cause amount of traffic problems. Although vehicle re-identification (Re-ID) can help to acquire and manage vehicles, some intrinsic difficulties hinder the application of vehicle Re-ID. For example, vehicles have little inter-instance discrepancy due to their rigid structures and finite models. To address this problem, in this paper, a logistic triplet loss is proposed to fuse a label-smoothing cross entropy to extract fine-grained feature embeddings. Via exploring deeper into the inter-instance variances, the novel loss combines advantages of classification and metric learning, and reveals more stable performance than popular triplet loss. The experimental results on public datasets demonstrate the effectiveness of the proposed loss compared with state-of-the-art approaches.
Chenggang Li, Yinhao Wang, Zhicheng Zhao 0001
VCIP3
2019 A Spatio-temporal Hybrid Network for Action Recognition
abstract
Convolutional Neural Networks (CNNs) are powerful in learning spatial information for static images, while they appear to lose their abilities for action recognition in videos because of the neglecting of long-term motion information. Traditional 3D convolution has high computation complexity and the used Global Average Pooling (GAP) on the bottom of network can also lead to unwanted content loss or distortion. To address above problems, we propose a novel action recognition algorithm by effectively fusing 2D and Pseudo-3D CNN to learn spatio-temporal features of video. First, we use Pseudo-3D CNN with proposed Multi-level pooling module to learn spatio-temporal features. Second, the features output by multi-level pooling module are passed through our proposed processing module to make full use of the rich features. Third, a 2D CNN fed with motion vectors is designed to extract motion patterns, which can be regarded as a supplement of Pseudo-3D CNN to make up for the information lost by RGB images. Fourth, a dependency-based fusion method is proposed to fuse the multi-stream features. Finally, the effectiveness of our proposed action recognition algorithm is demonstrated on public UCF101 and HMDB51 datasets.
Zhicheng Zhao 0001
VCIP2
2019 Saliency-based Deep Multi-level Semantic Feature Fusion for Person Re-identification
abstract
Person re-identification (Re-ID) is a challenging problem due to external environmental disturbances and significant intra-class appearance variations. Discriminative person features should cover global and partial representations, and accordingly can describe high-level and middle-level semantic information of persons. To achieve this goal, a saliency-based deep multi-level semantic feature representation and fusion algorithm is proposed. Firstly, a feature fusion scheme at middle layer of our deep network is presented to effectively fuse global CNN features. Secondly, a part-based feature extraction method is designed to extract high-level semantic features. In addition, a parameter-free multi-scale saliency-based enhancement algorithm is proposed to compute patch-level saliency scores to enhance the distinctiveness of part-based features. Experimental results on three public datasets, namely, Market-1501, DukeMTMC-reID, and CHUK03, demonstrate the effectiveness of the proposed method compared with state-of-the-art Re-ID approaches.
Yinhao Wang, Chenggang Li, Qingwen Hu, Zhicheng Zhao 0001
VCIP4
2019 Deep class-skewed learning for face recognition
Pingyu Wang, Zhicheng Zhao 0001, Yandong Guo, Yanyun Zhao, Bojin Zhuang
Neurocomputing3
2019 Two-stage deep learning for supervised cross-modal retrieval
Jie Shao 0014, Zhicheng Zhao 0001
Multim. Tools Appl.2
2019 A comprehensive solution for detecting events in complex surveillance videos
Yandong Zhu, Kaihui Zhou, Menglai Wang, Yanyun Zhao, Zhicheng Zhao 0001
Multim. Tools Appl.5
2018 Deep Image Retrieval: Indicator and Gram Matrix Weighting for Aggregated Convolutional Features
abstract
Convolutional Neural Network (CNN) has been proven to be an effective feature extractor for multiple computer vision tasks such as image classification and object detection etc. However, image retrieval in realistic scenarios, usually faces large-scale unlabeled datasets, thus the learning of a good model is often infeasible. In this paper, we propose a novel and interpretable image representation via spatial-channel weighting for aggregated deep convolutional features. Specifically, we first determine discriminative regions of an image by computing the Indicator matrix, and then, the distinctive features are extracted from salient areas by calculating the Gram matrix, in which high-order features are learnt. Finally, a compact image representation is generated by fusing spatial saliency and channel sensitivity of CNN features. The experimental results on several benchmark datasets, i.e., Oxford buildings, Paris buildings and Holidays, indicate that the proposed approach outperforms state-of-the-art methods based on pre-trained deep networks.
Zhipeng Wang 0008, Xuanlu Xiang, Zhicheng Zhao 0001
ICME3
2018 A unified framework with a benchmark dataset for surveillance event detection
Zhicheng Zhao 0001, Xuanchong Li, Xingzhong Du, Qi Chen 0014, Yanyun Zhao, Xiaojun Chang, Alex Hauptmann 0001
Neurocomputing1
2018 Weakly supervised detection with decoupled attention-based deep representation
Wenhui Jiang 0001, Zhicheng Zhao 0001
Multim. Tools Appl.2
2018 Complex event detection via attention-based video representation and classification
Zhicheng Zhao 0001, Rui Xiang
Multim. Tools Appl.1
2018 Person re-identification via integrating patch-based metric learning and local salience learning
Zhicheng Zhao 0001, Binlin Zhao
Pattern Recognit.1
2017 Joint multi-feature fusion and attribute relationships for facial attribute prediction
abstract
Predicting facial attributes from wild images is very challenging due to complex face variations. The key to this problem is to construct rich facial representations and take advantage of attribute relationships. In this paper, we propose a novel multi-task convolutional neural network (MTCNN) and a supervision signal called Online Batch Relation Loss (OBRL) for face attribute prediction in the wild. In particular, MTCNN builds informative facial features by embedding identity, age and race features from IdentityNet, AgeNet and RaceNet respectively. In addition, OBRL can diminish distribution shift of attribute relationships by mining attribute correlation within each minibatch, while it penalizes the probability divergence between a pair of attributes. In order to learn discriminative attribute features, we feed AttributeNet with fused facial features and partition attributes into nine groups to share intra-group features and reduce redundant computation. Finally, AttributeNet is optimized with the joint supervision of Cross Entropy Loss and OBRL. Experiments on CelebA and LFWA show that the proposed method outperforms the state-of-the-art methods with a significant margin.
Pingyu Wang, Zhicheng Zhao 0001
VCIP3
2017 Gram matrix based representation for image retrieval
abstract
In the field of image retrieval, most of image representations based on convolutional neural network (CNN) are first-order forms, i.e., the pooling or encoding methods are adopted on feature maps directly to produce compact image representations, while the high-order representations, such as the dependencies between different channels in the same layer are often neglected. In this paper, a novel image representation and retrieval algorithm based on Gram matrix is proposed. Specifically, based on Gram matrix of convolutional layers, second-order features are firstly constructed by considering the relationships between different channels of feature maps. Afterwards, two weighted schemes, that is, equal channel weighting and sparsity-sensitive channel weighting are presented respectively to aggregate them into the final representation. The extensive experiments on four public image datasets are conducted, and the promising results demonstrate the effectiveness of the proposed algorithm.
Shanwei Zhao, Zhicheng Zhao 0001
VCIP2
2017 Modeling intra- and inter-pair correlation via heterogeneous high-order preserving for cross-modal retrieval
Leiquan Wang, Weichen Sun, Zhicheng Zhao 0001
Signal Process.3
2016 ALADDIN: A locality aligned deep model for instance search
abstract
Most instance search systems are based on modeling local features. It remains a challenge to apply deep learning techniques into this task because of the asymmetrical similarity between the query region and dataset images. In this paper, we propose ALADDIN, A Locality Aligned Deep moDel for INstance search. This model deals with the asymmetrical similarity by searching query instances at the scale of aligned target regions instead of the whole image. Towards discriminative region representations, we utilize a deep convolutional network which captures both intra-class and inter-class distinctions of the regions. In addition, we propose a semi-supervised method to collect appropriate data to train the network. Extensive experiments confirm that our method is more suitable for generic instance search than most conventional methods, and outperforms the best CNNs-based method in both accuracy and efficiency.
Wenhui Jiang 0001, Zhicheng Zhao 0001, Anni Cai
ICASSP2
2016 Homemade TS-Net for Automatic Face Recognition
abstract
Inspired by how human being accomplishes face recognition task, a new architecture, called transfer and specialized net (TS-Net) is proposed in this paper, which fuses the general and specialized knowledge by combining a Transfer FaceNet and a Specialized FaceNet. The former is obtained by fine-tuning the pre-trained GoogleNet to transfer object-recognition knowledge to face recognition, and the latter is trained on global and local face patches to provide the discriminative specialized knowledge for face recognition. The final face representation is formed by fusing the features from both FaceNets. The advantages of our proposed architecture come from that: (i) By explicitly assigning different learning rates to different layers we successfully transfer the well-trained GoogleNet from object recognition task to a distinctly different task - face recognition; (ii) We construct the Specialized FaceNet with 6 simple networks to imitate the capture of featured-based and configural information in human vision process; (iii) Both Transfer FaceNet and Specialized FaceNets can be trained with a relatively small amount of training data (about 0.4 million samples) and a low configuration hardware (for example, a Titan-Z GPU). Experimental results show that TS-Net achieves competitive performance on both LFW and CASIA-Webface datasets. Also, it is promising that only slight dropping is found on verification and identification accuracy when 300 dimensional binary face representations are applied with Cosine distance as measure, which is implemental to develop practical human face retrieval and recognition system.
Shilun Lin, Zhicheng Zhao 0001
ICMR2
2016 Adaptive Synopsis of Non-Human Primates' Surveillance Video Based on Behavior Classification
Zhicheng Zhao 0001
MMM (1)3
2016 Deep canonical correlation analysis with progressive and hypergraph learning for cross-modal retrieval
Jie Shao 0014, Leiquan Wang, Zhicheng Zhao 0001, Anni Cai
Neurocomputing3
2016 Efficient multi-modal hypergraph learning for social image classification with complex label correlations
Leiquan Wang, Zhicheng Zhao 0001
Neurocomputing2
2016 Specific video identification via joint learning of latent semantic concept, scene and temporal structure
Zhicheng Zhao 0001, Yi-Fan Song
Neurocomputing1
2016 Bayes pooling of visual phrases for object retrieval
Wenhui Jiang 0001, Zhicheng Zhao 0001
Multim. Tools Appl.2
2015 Large scale cross-media data retrieval based on Hadoop
Wenchen Cheng, Zhicheng Zhao 0001
QSHINE3
2015 Part-based deep network for pedestrian detection in surveillance videos
abstract
Accurate pedestrian detection in highly crowded surveillance videos is a challenging task, since the regions of pedestrians in the videos may be largely occluded by other pedestrians. In this paper, we propose an effective part-based deep network cascade (HsNet) to solve this problem. In this model, the part-based scheme effectively restrains the appearance variations of pedestrians caused by heavy occlusion. The deep network captures discriminative information of visible body parts. In addition, the cascade architecture enables very fast detection. We make experiments on one of the largest surveillance video dataset, namely TRECVid SED Pedestrian Dataset (SED-PD). It is shown that in highly crowded surveillance videos, our proposed method achieves very competitive performance compared with state-of-the-art methods. More importantly, our method is significantly faster.
Qi Chen 0014, Wenhui Jiang 0001, Yanyun Zhao, Zhicheng Zhao 0001
VCIP4
2015 Fast Uyghur text detection in videos based on learning of baseline feature
abstract
Text detection in image is always a significant part in image semantic understanding, and detection of Uyghur text is a special and extensible application. In this paper, we propose a Uyghur text detection on the basis of the learning of a baseline structure, which generated from texture feature of the text. Firstly, texture features of the image are extracted and texts are classed by a SVM classifier, and then the baseline of the text is structured and represented. Finally, another SVM classifier is trained for Uyghur text detection. The experimental results on user-built dataset including news, entertainment videos and movies show that the proposed algorithm is fast and effective, and better than several typical approaches.
Yi-Fan Song, Zhicheng Zhao 0001
VCIP3
2015 3View deep canonical correlation analysis for cross-modal retrieval
abstract
This paper investigates the problem of modeling Internet images and associated text for cross-modal retrieval tasks such as text-to-image search, and image-to-text search. Canonical correlation analysis (CCA), a classic two view approach for mapping text and image into a common latent space, does not make use of the semantic information of text and image pairs. We use CCA to map text, image and semantic information into a common latent space, in which the correlation of the three views is maximized. To improve the performance of CCA, in this paper, 3view-Deep Canonical Correlation Analysis (3view-DCCA), a nonlinear expansion of CCA is proposed to learn the complex nonlinear transformations between the three views. Like most deep learning methods, DCCA is easy to over-fitting. To overcome over-fitting, we add the reconstruct loss of each view into the loss function, which include the correlation loss of every two views and regularization of parameters. Inspired by PageRank, we propose a search-based similarity method to score relevance. The proposed model (3view-DCCA) is evaluated on three publicly available data sets from real scenes. We demonstrate that our deep model performs significantly better than traditional canonical correlation analysis based models and several other deep learning models on cross-modal retrieval tasks.
Jie Shao 0014, Zhicheng Zhao 0001, Ting Yue
VCIP2
2015 On semantic-instructed attention: From video eye-tracking dataset to memory-guided probabilistic saliency model
Yan Hua, Zhicheng Zhao 0001, Renlai Zhou, Anni Cai
Neurocomputing3
2015 Datum-Adaptive Local Metric Learning for Person Re-identification
abstract
Person re-identification (PRID) is a challenging problem in multi-camera surveillance systems. In this paper, we propose a novel Datum-Adaptive Local Metric learning method for PRID, which learns individual local feature projection for each image sample according to the current data distribution and projects all samples into a common discriminative space for similarity measure. We adopt an approximate strategy based on Local Coordinate Coding to learn local projections. Anchor points are first generated by clustering and the local projection of each sample is then approximated by the linear combination of a set of projection bases, which are associated with the anchor points. Experimental results demonstrate that the proposed approach obtains superior performance compared with state-of-the-art methods on public benchmarks.
Zhicheng Zhao 0001, Anni Cai
IEEE Signal Process. Lett.2
2014 Cross modal metric learning with multi-level semantic relevance
abstract
The Mahalanobis metric learning is an effective tool for constructing semantic consistent distance among data in single modal data analysis. However, distance metric learning is a more challenging issue for cross modal data, where less attention has been paid in previous studies. In this paper, we propose Cross mOdal Large mArgin metric leaRning (COLAR) with multi-level semantic relevance. With large margin principle, we model different levels of the semantic relations across modalities, e.g., the one-to-one correspondence and intra-class relation, while traditional correlation learning approaches (such as CCA and its variants) can only handle the one-to-one correspondence or treat them indiscriminatively. As a result, the distances of multi-level relevance among cross modal data are optimized based on a regularized learning framework. Promising performance is achieved on cross modal retrieval, i.e., image-to-text retrieval and text-to-image retrieval.
Yan Hua, Shuhui Wang, Zhicheng Zhao 0001, Qingming Huang, Anni Cai
ICIP3
2014 Bagging based metric learning for person re-identification
abstract
Person re-identification is a challenging problem in computer vision due to large variations of appearance among different cameras. Recently, metric learning is widely used to model the transformation between cameras. However, traditional metric learning based methods only learn one metric for the whole feature space, which cannot model different kinds of appearance variations well. In this paper, we introduce bagging into metric learning, and propose a bagging-based large margin nearest neighbor (LMNN) method for person re-identification. That is, multiple LMNN predictors are generated on sub-regions of the feature space and leveraged to obtain an aggregated predictor for performance improvement. Two bagging strategies, sample-bagging and feature-bagging, are proposed and compared. Extensive experiments on three benchmarks demonstrate the superiority of proposed approach over state-of-the-art methods.
Bohuai Yao, Zhicheng Zhao 0001, Anni Cai
ICME2
2014 Parametric Local Multi-modal Metric Learning for Person Re-identification
abstract
Person re-identification is a challenging problem in multi-camera surveillance systems. Most current methods always aim at learning a global distance metric to overcome the visual appearance changes between images from different cameras. However, the feature variations between images are not constant over the entire feature space, thus one global metric is not always applicable to all feature variation conditions. Moreover, few of the current methods take the modality discrepancy between different cameras into consideration during the metric learning process. To address these drawbacks, we propose Parametric Local MultiModal (PLMM) metric learning in this paper. We consider the feature difference value of one pair of images, which come from different cameras, characterizing one kind of feature variations. Accordingly, we learn a unique local metric for each image pair. Moreover, images from different cameras are regarded as lying on different modalities, and thus in each local metric, two different projection matrixes are learned for the cross-modality similarity measures. To balance the locality and computation efficiency, the local metrics are parameterized as weighted linear combinations of basis metrics, which correspond to a small set of anchor image pairs. The local metrics are capable of modeling the cross-modality feature variations among different cameras, making the distance of intra-class image pairs minimized and simultaneously those of inter-class image pairs maximized. Experimental results demonstrate that the proposed approach obtains competitive performance compared with state-of-the-art methods on three publicly available benchmarks.
Zhicheng Zhao 0001, Anni Cai
ICPR2
2014 An adaptive symmetry detection algorithm based on local features
abstract
Local feature-based symmetry detection algorithms can simultaneously consider symmetries over all locations, scales and orientations and achieve state-of-the-art performance. This paper demonstrates the limitations of these algorithms in case of dealing with background clutters, low contrast and smooth surfaces, and presents an adaptive feature point detection algorithm to overcome those limitations. Quantitative evaluations and subjective comparisons against the state-of-the-art reflection symmetry detection algorithm on the image dataset released by "Symmetry Detection from Real World Images Competition 2013" show a significant improvement in detection accuracy and computation efficiency. Furthermore, the proposed algorithm is also tested on the non-human primates' (NHPs') video surveillance data as a preprocessing step before NHPs' behaviors analysis, and a good performance is obtained as well.
Zhicheng Zhao 0001
VCIP4
2014 A novel distributed compressive video sensing based on hybrid sparse basis
abstract
Distributed compressive video sensing (DCVS) is a new emerging video codec that incorporates advantages of distributed video coding (DVC) and compressive sensing (CS). However, due to the absence of a good sparse basis, the DCVS does not achieve ideal compressing efficiency compared with the traditional video codec, such as MPEG-4, H.264, etc. This paper proposes a new hybrid sparse basis, which combines the image-block prediction and DCT basis. Adaptive block-based prediction is employed to learn block-prediction basis by exploiting temporal correlation among successive frames. Based on linear DCT basis and predicted basis, the hybrid sparse basis can achieve sparser representation with lower complexity. The experiment results indicate that the proposal outperforms the state-of-the-art DCVS schemes on both visual quality and average PSNR. In addition, an iterative fashion proposed in the decoder can enhance the sparsity of the hybrid sparse basis and improve the rate-distortion performance significantly.
Haifeng Dong, Bojin Zhuang, Zhicheng Zhao 0001
VCIP4
2014 Tag-based social image search with hyperedges correlation
abstract
In social image search, most existing hypergraph methods use the visual and textual features in isolation by treating each feature term as a hyperedge. Nevertheless, they neglect the correlations of visual and textual hyperedges, which are more robust to represent the high-order relationship among vertices. In this paper, we propose a hypergraph with correlated hyperedges (CHH), which introduces high-order relationship of hyperedges into hypergraph learning. Based on CHH, a pairwise visual-textual correlation hypergraph (VTCH) model is used for tag-based social image search. To overcome the large number of newly generated hybrid hyperedges, a bagging-based method is adopted to balance the accuracy and speed. Finally, adaptive hyperedges learning method is used to obtain the relevance score for social image search. The experiments conducted on MIR Flickr show the effectiveness of our proposed method.
Leiquan Wang, Zhicheng Zhao 0001
VCIP2
2013 Person re-identification using matrix completion
abstract
Person re-identification is a challenging problem in multicamera surveillance systems. In this paper, we formulate person re-identification as a cross-camera feature construction problem to overcome the feature variation between different camera spaces. The linear transformation of color information between probe and gallery camera spaces makes the stacked matrix, which concatenates features from these two camera spaces, rank deficient. From the feature observed in probe camera space we can construct its corresponding feature in gallery camera space by completing unknown entries on the relevant positions of the stacked matrix, and then match the constructed probe feature with features in gallery camera space. We also introduce additive noise term into the model to deal with the adverse effects caused by illumination variation with time. Experimental results demonstrate the proposed approach outperforms the metric learning methods as well as simple nearest neighbor search, and obtains a competitive performance compared with the state-of-the-art methods.
Zhicheng Zhao 0001, Anni Cai
ICIP3
2013 A probabilistic saliency model with memory-guided top-down cues for free-viewing
abstract
Attention models for free-viewing of images are commonly developed in a bottom-up (BU) manner. However, in this paper, we propose to include memory-oriented top-down spatial attention cues into the model. The proposed generative saliency model probabilistically combines a BU module and a top-down (TD) module and can be applied to both static and dynamic scenes. For static scenes, the experience of attention distribution to similar scenes in long-term memory is mimicked by linearly mapping a global feature of the scene to the long-term top-down (LTD) saliency. And for dynamic scenes, we add on the influence of short-term memory to form a HMM-like chain to guide the attention distribution. Our improved BU module utilizes low-level feature contrast, spatial distribution and location information to highlight saliency regions. It is tested on 1000 benchmark images and outperforms state-of-the-art BU methods. The complete model is examined on two video datasets, one with manually labeled saliency regions, and another with recorded eye-movement fixations. Experimental results show that our model achieves significant improvement in predicting human visual attention compared with existing saliency models.
Yan Hua, Zhicheng Zhao 0001, Anni Cai
ICME2
2013 Person Re-identification by Local Feature Based on Super Pixel
Zhicheng Zhao 0001
MMM (1)2
2013 Anchor-supported multi-modality hashing embedding for person re-identification
abstract
Person re-identification is a challenging problem in multi-camera surveillance systems. Most existing methods focus on metric learning which aims to match images from different cameras in a common metric space. Boosted hashing projection provides a new way of identifying instances based on pairwise similarity. However, both of these approaches ignore the underlying fact that images captured by two cameras should be seen as in different modalities. To address this drawback, we formulate person re-identification as an Anchor-supported Multi-Modality Hashing Embedding (AMMHE) problem, in which different projections are used to map data from different cameras into a common Hamming space. The data are projected to binary bits by using boosted hash projections, making the weighted Hamming distance of intra-class data pairs minimized and simultaneously those of inter-class data pairs maximized. We also introduce an anchor-supported dimension reduction method to avoid the computational burden of high feature dimensionality. Our approach obtains competitive performance compared with state-of-the-art methods on publicly available benchmarks.
Zhicheng Zhao 0001, Anni Cai
VCIP2
2013 Tree-based Shape Descriptor for scalable logo detection
abstract
Detecting logos in real-world images is a great challenging task due to a variety of viewpoint or light condition changes and real-time requirements in practice. Conventional object detection methods, e.g., part-based model, may suffer from expensively computational cost if it was directly applied to this task. A promising alternative, triangle structural descriptor associated with matching strategy, offers an efficient way of recognizing logos. However, the descriptor fails to the rotation of logo images that often occurs when viewpoint changes. To overcome this shortcoming, we propose a new Tree-based Shape Descriptor (TSD) in this paper, which is strictly invariant to affine transformation in real-world images. The core of proposed descriptor is to encode the shape of logos by depicting both appearance and spatial information of four local key-points. In the training stage, an efficient algorithm is introduced to mine a discriminate subset of four tuples from all possible key-point combinations. Moreover, a root indexing scheme is designed to enable to detect multiple logos simultaneously. Extensive experiments on three benchmarks demonstrate the superiority of proposed approach over state-of-the-art methods.
Chengde Wan, Zhicheng Zhao 0001, Anni Cai
VCIP2
2012 Find dominant bins of a histogram by sparse representation
Zhicheng Zhao 0001, Anni Cai
ICPR2
2010 A computable structure model for hollywood film
abstract
Automatic parsing of video structure is one of the key techniques for video understanding and retrieval. In this paper, we propose a computable film structure model: the nine-plot model, which parses a film into three hierarchical semantic levels: act, plot and scene. The model has been motivated by “Hollywood mode” and generic narrative structure of the story. In addition, a set of modeling methods for film-making rules based on scene classification and segmentation are proposed. The experimental results on seven full-length Hollywood movies demonstrate the effectiveness of the model.
Zhicheng Zhao 0001, Xiaojuan Ge
ICIP1