Guibo Zhu

dblp:125/2113 · DBLP profile ↗
← Back
44ranked-venue papers
11as first author
26since 2021 · last 2026
0000-0001-8293-3952ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 22 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
abstract
Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performance when deployed outside their designated contexts. To overcome this limitation, we propose AnomalyMoE, a novel and universal anomaly detection framework based on a Mixture-of-Experts (MoE) architecture. Our key insight is to decompose the complex anomaly detection problem into three distinct semantic hierarchies: local structural anomalies, component-level semantic anomalies, and global logical anomalies. AnomalyMoE correspondingly employs three dedicated expert networks at the patch, component, and global levels, and is specialized in reconstructing features and identifying deviations at its designated semantic level. This hierarchical design allows a single model to concurrently understand and detect a wide spectrum of anomalies. Furthermore, we introduce an Expert Information Repulsion (EIR) module to promote expert diversity and an Expert Selection Balancing (ESB) module to ensure the comprehensive utilization of all experts. Experiments on 8 challenging datasets spanning industrial imaging, 3D point clouds, medical imaging, video surveillance, and logical anomaly detection demonstrate that AnomalyMoE establishes new state-of-the-art performance, significantly outperforming specialized methods in their respective domains.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI3
2026 Improving Generalization in LLM Structured Pruning via Function-Aware Neuron Grouping
Tao Yu 0013, Yongqi An, Kuan Zhu, Guibo Zhu, Ming Tang 0001, Jinqiao Wang
AAAI4
2026 Adversarial discriminant attack on text-to-image diffusion models
Hanxiao Wu, Shengwu Xiong 0001, Dong Yi, Lingxiang Wu, Jianqing Zhu, Guibo Zhu, Jinqiao Wang
Neural Networks6
2026 FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
abstract
Anomaly detection methods typically require extensive normal samples from the target class for training, limiting their applicability in scenarios that require rapid adaptation, such as cold start. Zero-shot and few-shot anomaly detection do not require labeled samples from the target class in advance, making them a promising research direction. Existing zero-shot and few-shot approaches often leverage powerful multimodal models to detect and localize anomalies by comparing image-text similarity. However, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous regions of varying shapes and sizes. To address these issues, this paper proposes the FiLo++ method, which consists of two key components. The first component, Fused Fine-Grained Descriptions (FusDes), utilizes large language models to generate anomaly descriptions for each object category, combines both fixed and learnable prompt templates and applies a runtime prompt filtering method, producing more accurate and task-specific textual descriptions. The second component, Deformable Localization (DefLoc), integrates the vision foundation model Grounding DINO with position-enhanced text descriptions and a Multi-scale Deformable Cross-modal Interaction (MDCI) module, enabling accurate localization of anomalies with various shapes and sizes. In addition, we design a position-enhanced patch matching approach to improve few-shot anomaly detection performance. Experiments on multiple datasets demonstrate that FiLo++ achieves significant performance improvements compared with existing methods. Code will be available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.3
2025 See Through Their Minds: Learning Transferable Brain Decoding Models from Cross-Subject fMRI
abstract
Deciphering visual content from fMRI sheds light on the human vision system, but data scarcity and noise limit brain decoding model performance. Traditional approaches rely on subject-specific models, which are sensitive to training sample size. In this paper, we address data scarcity by proposing shallow subject-specific adapters to map cross-subject fMRI data into unified representations. A shared deep decoding model then decodes these features into the target feature space. We use both visual and textual supervision for multi-modal brain decoding and integrate high-level perception decoding with pixel-wise reconstruction guided by high-level perceptions. Our extensive experiments reveal several interesting insights: 1) Training with cross-subject fMRI benefits both high-level and low-level decoding models; 2) Merging high-level and low-level information improves reconstruction performance at both levels; 3) Transfer learning is effective for new subjects with limited training data by training new adapters; 4) Decoders trained on visually-elicited brain activity can generalize to decode imagery-induced activity, though with reduced performance.
Guibo Zhu, Haodong Jing, Nanning Zheng 0001
AAAI3
2025 UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
abstract
Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with specialized detection techniques and model architectures that are difficult to generalize across different domains. Moreover, even within the same domain, current VAD approaches often require large amounts of normal samples to train class-specific models, resulting in poor generalizability and hindering unified evaluation across domains. To address this issue, we propose a generalized few-shot VAD method, UniVAD, capable of detecting anomalies across various domains, with a training-free unified model. UniVAD only needs few normal samples as references during testing to detect anomalies in previously unseen objects, without training on the specific domain. Specifically, UniVAD employs a Contextual Component Clustering (C3) module based on clustering and vision foundation models to segment components within the image accurately, and leverages Component-Aware Patch Matching (CAPM) and Graph-Enhanced Component Modeling (GECM) modules to detect anomalies at different semantic levels, which are aggregated to produce the final detection result. We conduct experiments on nine datasets spanning industrial, logical, and medical fields, and the results demonstrate that UniVAD achieves state-of-the-art performance in few-shot anomaly detection tasks across multiple domains, outperforming domain-specific anomaly detection models. Code is available at https://github.com/FantasticGNU/UniVAD.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
CVPR3
2025 SACPlace: Multi-Agent Deep Reinforcement Learning for Symmetry-Aware Analog Circuit Placement
abstract
The placement of analog Integrated Circuits (ICs) plays a critical role in their physical design. The objective is to minimize the Half-Perimeter Wire Length (HPWL) while satisfying complex analog IC constraints, such as symmetry. Unlike digital ICs, analog ICs are highly sensitive to parasitic effects, making device symmetry crucial for optimal circuit performance. However, existing methods, including both machine learning-based and analytical approaches, struggle to meet strict symmetry constraints. In machine learning-based methods, training a general model is challenging due to the limited diversity of the training data. In analytical methods, the difficulty lies in formulating symmetry constraints as a convex function, which is necessary for gradient-based optimization of the placement. To address the issue, we formulate the placement process as a Markov decision process and propose SACPlace, a multi-agent deep reinforcement learning method for Symmetry-Aware analog Circuit Placement. SACPlace initially extracts layout information and various constraints as the input information for placement refinement and evaluation. Subsequently, SACPlace constructs multi-agent policy networks for symmetry-aware placement by refining placement guided by the evaluation of optimal symmetry quality. Following this, SACPlace constructs multilayer perceptron-based critic networks to embed placement information for evaluating symmetry quality. This evaluation reward will be used for guiding placement refinement. Experimental results from four public analog ICs instances demonstrate that our method achieves the lowest actual wirelength and area while fully satisfying symmetry and common constraints, outperforming state-of-the-art methods. Additionally, simulation results on real-world analog ICs show better performance than these methods and even manual designs.
Guojing Ge, Guibo Zhu, Jixin Zhang, Jinqiao Wang, Ning Xu 0006
DATE3
2025 Extracting Sparse Specialist Models from Generalist Models
abstract
Recently, several generalist models such as Contrastive Language Image Pre-training (CLIP) have demonstrated their capabilities of performing diverse downstream tasks through zero-shot or few-shot guidance. When these generalist models are used for the specific downstream task where only a fraction of features is relevant, they would suffer from a significant redundancy of parameters. While existing methods aim to achieve sparsity and specialization, they often require additional training and large datasets. In this paper, we propose a novel framework to extract a sparse specialist model from a generalist model using only few-shot samples, without any training. Our task-specific pruning framework defines task relevance metrics and employs weighted layer-wise pruning, preserving relevant features while removing redundancies. Experiments show that our method maintains nearly identical zero-shot accuracy compared to the original generalist models at 30% sparsity, with only minimal decline at 50%.
Tao Yu 0013, Xu Zhao 0003, Yongqi An, Guibo Zhu, Ming Tang 0001, Jinqiao Wang
ICASSP4
2025 Dual-Chain Reasoning: Enhancing Multimodal Document VQA Through Positive and Negative Reasoning Paths
Hanxiao Wu, Zhaopeng Gu, Dong Yi, Guibo Zhu, Jinqiao Wang
ICIG (3)5
2025 BrainCLIP: Brain Representation via CLIP for Generic Natural Visual Stimulus Decoding
abstract
Functional Magnetic Resonance Imaging (fMRI) presents challenges due to limited paired samples and low signal-to-noise ratios, particularly in tasks involving reconstructing natural images or decoding their semantic content. To address these challenges, we introduce BrainCLIP, an innovative fMRI-based brain decoding model. BrainCLIP leverages Contrastive Language-Image Pre-training's (CLIP) cross-modal generalization abilities to bridge brain activity, images, and text for the first time. Our experiments demonstrate CLIP's effectiveness in diverse brain decoding tasks, including zero-shot visual category decoding, fMRI-image/text alignment, and fMRI-to-image generation. The core objective of BrainCLIP is to train a mapping network that translates fMRI patterns into a unified CLIP embedding space, achieved through visual and textual supervision integration. Our experiments highlight that this approach significantly enhances performance in tasks such as fMRI-text alignment and fMRI-based image generation. Notably, BrainCLIP surpasses BraVL, a recent multi-modal method, in zero-shot visual category decoding. Moreover, BrainCLIP demonstrates strong capability in reconstructing visual stimuli with high semantic fidelity, competing favorably with state-of-the-art methods in capturing high-level semantic features during fMRI-based natural image reconstruction.
Liangjun Chen, Guibo Zhu, Badong Chen, Nanning Zheng 0001
IEEE Trans. Medical Imaging4
2024 AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific domain knowledge and have a weaker understanding of localized details within objects, which hinders their effectiveness in the Industrial Anomaly Detection (IAD) task. On the other hand, most existing IAD methods only provide anomaly scores and necessitate the manual setting of thresholds to distinguish between normal and abnormal samples, which restricts their practical implementation. In this paper, we explore the utilization of LVLM to address the IAD problem and propose AnomalyGPT, a novel IAD approach based on LVLM. We generate training data by simulating anomalous images and producing corresponding textual descriptions for each image. We also employ an image decoder to provide fine-grained semantic and design a prompt learner to fine-tune the LVLM using prompt embeddings. Our AnomalyGPT eliminates the need for manual threshold adjustments, thus directly assesses the presence and locations of anomalies. Additionally, AnomalyGPT supports multi-turn dialogues and exhibits impressive few-shot in-context learning capabilities. With only one normal shot, AnomalyGPT achieves the state-of-the-art performance with an accuracy of 86.1%, an image-level AUC of 94.1%, and a pixel-level AUC of 95.3% on the MVTec-AD dataset.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI3
2024 A Set of Effective Strategies for Optimized Road Damage Detection
abstract
In this paper, we propose an optimized method for road damage detection using a lightweight YOLO model as the baseline. Our approach incorporates a set of effective strategies, including lightweight attention mechanisms, data augmentation, dynamic sampling, weight averaging, and multi-step knowledge distillation. Our method significantly improves inference speed while maintaining high accuracy compared to previous mainstream ensemble-based methods. Our approach achieves notable success in the IEEE Big Data 2024 Optimized Road Damage Detection Challenge (ORDDC’2024), securing second place with an F1 score of 0.7013 and an inference time of 0.0328 s per image. These results demonstrate a strong balance between accuracy and efficiency. Extensive experiments confirm that our method boosts detection accuracy and greatly accelerates inference, making it highly suitable for real-world applications. The source code and trained model are available at https://github.com/YinglongDu/ShiYu_Kunchuan_ORDDC2024.
Yinglong Du, Xu Zhao 0003, Bailin He, Bingke Zhu, Shuaihua Zhao, Guibo Zhu, Jinqiao Wang
IEEE Big Data6
2024 BFRFormer: Transformer-Based Generator for Real-World Blind Face Restoration
abstract
Blind face restoration is a challenging task due to the unknown and complex degradation. Although face prior-based methods and reference-based methods have recently demonstrated high-quality results, the restored images tend to contain over-smoothed results and lose identity-preserved details when the degradation is severe. It is observed that this is attributed to short-range dependencies, the intrinsic limitation of convolutional neural networks. To model long-range dependencies, we propose a Transformer-based blind face restoration method, named BFRFormer, to reconstruct images with more identity-preserved details in an end-to-end manner. In BFRFormer, to remove blocking artifacts, the wavelet discriminator and aggregated attention module are developed, and spectral normalization and balanced consistency regulation are adaptively applied to address the training instability and over-fitting problem, respectively. Extensive experiments show that our method outperforms state-of-the-art methods on a synthetic dataset and four real-world datasets. The source code, Casia-Test dataset, and pre-trained models is released at https://github.com/s8Znk/BFRFormer.
Guojing Ge, Qi Song 0003, Guibo Zhu, Yuting Zhang 0007, Jinglu Chen, Miao Xin, Ming Tang 0001, Jinqiao Wang
ICASSP3
2024 Auto DragGAN: Editing the Generative Image Manifold in an Autoregressive Manner
abstract
Pixel-level fine-grained image editing remains an open challenge. Previous works fail to achieve an ideal trade-off between control granularity and inference speed. They either fail to achieve pixel-level fine-grained control, or their inference speed requires optimization. To address this, this paper for the first time employs a regression-based network to learn the variation patterns of StyleGAN latent codes during the image dragging process. This method enables pixel-level precision in dragging editing with little time cost. Users can specify handle points and their corresponding target points on any GAN-generated images, and our method will move each handle point to its corresponding target point. Through experimental analysis, we discover that a short movement distance from handle points to target points yields a high-fidelity edited image, as the model only needs to predict the movement of a small portion of pixels. To achieve this, we decompose the entire movement process into multiple sub-processes. Specifically, we develop a transformer encoder-decoder based network named 'Latent Predictor' to predict the latent code motion trajectories from handle points to target points in an autoregressive manner. Moreover, to enhance the prediction stability, we introduce a component named 'Latent Regularizer', aimed at constraining the latent code motion within the distribution of natural images. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) inference speed and image editing performance at the pixel-level granularity.
Pengxiang Cai, Zhiwei Liu 0004, Guibo Zhu, Yunfang Niu, Jinqiao Wang
ACM Multimedia3
2024 FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization
abstract
Zero-shot anomaly detection (ZSAD) methods detect anomalies without prior access to known normal or abnormal samples within target categories. Existing methods typically rely on pretrained multimodal models, computing similarities between manually crafted textual features representing ''normal'' or ''abnormal'' semantics and image patch features to detect anomalies. However, the generic descriptions of ''abnormal'' often fail to precisely match diverse types of anomalies across different object categories. Additionally, computing feature similarities for single patches struggles to pinpoint specific locations of anomalies with various sizes and scales. To address these issues, we propose a novel ZSAD method called FiLo, comprising two components: adaptively learned Fine-Grained Description (FG-Des) and position-enhanced High-Quality Localization (HQ-Loc). FG-Des introduces fine-grained anomaly descriptions for each category using Large Language Models (LLMs) and employs adaptively learned textual templates to enhance the accuracy and interpretability of anomaly detection. HQ-Loc, utilizing Grounding DINO for preliminary localization, position-enhanced text prompts, and Multi-scale Multi-shape Cross-modal Interaction (MMCI) module, facilitates more accurate localization of anomalies of different sizes and shapes. Experimental results on datasets like MVTec and VisA demonstrate that FiLo significantly improves the performance of ZSAD in both detection and localization, achieving state-of-the-art performance with an image-level AUC of 83.9% and a pixel-level AUC of 95.9% on the VisA dataset. Code is available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Hao Li 0115, Ming Tang 0001, Jinqiao Wang
ACM Multimedia3
2024 Multi-Model Style-Aware Diffusion Learning for Semantic Image Synthesis
abstract
Semantic image synthesis aims to generate images from given semantic layouts, which is a challenging task that requires training models to capture the relationship between layouts and images. Previous works are usually based on Generative Adversarial Networks (GAN) or autoregressive (AR) models. However, the GAN model's training process is unstable, and the AR model’s performance is seriously affected by the independent image encoder and the unidirectional generation bias. Due to the above limitations, these methods tend to synthesize unrealistic, poorly aligned images and only consider single-style image generation. In this paper, we propose a Multi-model Style-aware Diffusion Learning (MSDL) framework for semantic image synthesis, including a training module and a sampling module. In the training module, a layout-to-image model is introduced to transfer the learned knowledge from a model pretrained with massive weak correlated text-image pairs data, making the training process more efficient. In the sampling module, we designed a map-guidance technique and creatively designed a multi-model style-guidance strategy for creating images in multiple styles, e.g., oil painting, Disney Cartoon, and pixel style. We evaluate our method on Cityscapes, ADE20K, and COCO-Stuff, making visual comparisons and computing with multiple metrics such as FID, LPIPS, etc. Experimental results demonstrate that our model is highly competitive, especially in terms of fidelity and diversity.
Yunfang Niu, Lingxiang Wu, Yousong Zhu, Guibo Zhu, Jinqiao Wang
ACM Trans. Multim. Comput. Commun. Appl.5
2023 ShiftFormer: Spatial-Temporal Shift Operation in Video Transformer
abstract
Transformers have achieved great success in various tasks, especially that introducing pure Transformers into video understanding shows powerful performance. However, video Transformer suffers from the problem of memory explosion: it is difficult to be deployed on hardware due to the intensive computation. To address this issue, we propose ST-shift (spatial-temporal) operation with zero computation and zero parameter. We are only shifting a small portion of the channels along the temporal and spatial dimensions. Based on this operation, we build an attention-free ShiftFormer, where ST-shift blocks substitute the attention layers in video Transformer. ShiftFormer is accurate and efficient: it can reduce 56.34% of memory usage and achieve 3.41× faster training. When both using random initialization, our model performs even better than Video Swin Transformer for video recognition on Something-Something v2.
Beiying Yang, Guibo Zhu, Guojing Ge, Jinzhao Luo, Jinqiao Wang
ICME2
2023 Temporal-Channel Topology Enhanced Network for Skeleton-Based Action Recognition
Jinzhao Luo, Guibo Zhu, Guojing Ge, Beiying Yang, Jinqiao Wang
PRCV (1)3
2022 An Ensemble of One-Stage and Two-Stage Detectors Approach for Road Damage Detection
abstract
With the growth of the city and the increase in the number of cars, the maintenance and management of roads attract more attention. Road damage detection of road images is the basic step of road maintenance. To reduce the cost of labor, it is crucial to make the best use of road damage images from different geographical environments and capturing devices. This paper describes our 1-st place solution used in the Crowd sensing-based Road Damage Detection Challenge of the 2022 IEEE International Conference on Big Data. We use YOLO-series models and Faster RCNN as our one-stage and two-stage baseline models respectively. Our model only needs to be trained directly on the datasets of the overall six countries. Besides, with ensemble learning and test time augmentation, our ensemble model achieves the best results on the learderboard of each single country (India, Japan, United States, and Norway) without fine-tuning. Our ensemble model achieves the F1 scores of 0.7699 and 0.7160 on Overall and Average leaderboard, which significantly outperformed the 2-nd p lace F1 s cores of 0.7432 and 0.6744. The source code and trained model are available at https://github.com/berry-ding/ShiYu_SeaView_GRDDC2022.
Wenchao Ding 0004, Xu Zhao 0003, Bingke Zhu, Yinglong Du, Guibo Zhu, Tao Yu 0013, Jinqiao Wang
IEEE Big Data5
2022 TaiSu: A 166M Large-scale High-Quality Dataset for Chinese Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) has been shown to be an efficient method to improve the performance of models on different vision-and-language downstream tasks. Substantial studies have shown that neural networks may be able to learn some general rules about language and visual concepts from a large-scale weakly labeled image-text dataset. However, most of the public cross-modal datasets that contain more than 100M image-text pairs are in English; there is a lack of available large-scale and high-quality Chinese VLP datasets. In this work, we propose a new framework for automatic dataset acquisition and cleaning with which we construct a new large-scale and high-quality cross-modal dataset named as TaiSu, containing 166 million images and 219 million Chinese captions. Compared with the recently released Wukong dataset, our dataset is achieved with much stricter restrictions on the semantic correlation of image-text pairs. We also propose to combine texts collected from the web with texts generated by a pre-trained image-captioning model. To the best of our knowledge, TaiSu is currently the largest publicly accessible Chinese cross-modal dataset. Furthermore, we test our dataset on several vision-language downstream tasks. TaiSu outperforms BriVL by a large margin on the zero-shot image-text retrieval task and zero-shot image classification task. TaiSu also shows better performance than Wukong on the image-retrieval task without using image augmentation for training. Results demonstrate that TaiSu can serve as a promising VLP dataset, both for understanding and generative tasks. More information can be referred to https://github.com/ksOAn6g5/TaiSu.
Guibo Zhu, Qi Song 0003, Guojing Ge, Guanhui Qiao, Ru Peng, Lingxiang Wu, Jinqiao Wang
NeurIPS2
2022 Dynamic Orthogonal Projection Constrained Discriminative Tracking
abstract
Due to the end-to-end feature learning with convolutional neural networks (CNNs), modern discriminative trackers improve the state of the art significantly. To achieve a strong discrimination, the learned features are usually high-dimensional, resulting in a massive number of parameters contained in the discriminative model and the increase of risk of over-fitting in the online tracking. In this letter, we try to alleviate the risk of over-fitting by means of the adaptive dimensionality reduction (DR) through CNNs. Specifically, an orthogonality constrained ridge regression model is proposed to reduce the dimensionality of features, and a dynamic sub-network (DOPNet) is designed to learn to perform DR. After trained with an orthogonality loss and a regression one, DOPNet generates a set of orthogonal bases (i. e., weights in FC layers) dynamically to reduce the feature dimensionality for a discriminative model in the online tracking. Based on the novel discriminative model and DOPNet, an effective and efficient tracker, DOPTracker, is developed. DOPTracker achieves the state-of-the-art results on four benchmarks, OTB-2015, VOT-2018, NfS, and GOT-10 k while running at 30 FPS.
Ming Tang 0001, Guibo Zhu, Jinqiao Wang, Hanqing Lu
IEEE Signal Process. Lett.3
2022 Multi-Granularity Mutual Learning Network for Object Re-Identification
abstract
Object re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks.
Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Intell. Transp. Syst.6
2021 Improving Multiple Object Tracking With Single Object Tracking
abstract
Despite considerable similarities between multiple object tracking (MOT) and single object tracking (SOT) tasks, modern MOT methods have not benefited from the development of SOT ones to achieve satisfactory performance. The major reason for this situation is that it is inappropriate and inefficient to apply multiple SOT models directly to the MOT task, although advanced SOT methods are of the strong discriminative power and can run at fast speeds.In this paper, we propose a novel and end-to-end trainable MOT architecture that extends CenterNet by adding an SOT branch for tracking objects in parallel with the existing branch for object detection, allowing the MOT task to benefit from the strong discriminative power of SOT methods in an effective and efficient way. Unlike most existing SOT methods which learn to distinguish the target object from its local backgrounds, the added SOT branch trains a separate SOT model per target online to distinguish the target from its surrounding targets, assigning SOT models the novel discrimination. Moreover, similar to the detection branch, the SOT branch treats objects as points, making its online learning efficient even if multiple targets are processed simultaneously. Without tricks, the proposed tracker achieves MOTAs of 0.710 and 0.686, IDF1s of 0.719 and 0.714, on MOT17 and MOT20 benchmarks, respectively, while running at 16 FPS on MOT17.
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Guibo Zhu, Jinqiao Wang, Hanqing Lu
CVPR4
2021 High-Performance Discriminative Tracking with Transformers
abstract
End-to-end discriminative trackers improve the state of the art significantly, yet the improvement in robustness and efficiency is restricted by the conventional discriminative model, i.e., least-squares based regression. In this paper, we present DTT, a novel single-object discriminative tracker, based on an encoder-decoder Transformer architecture. By self- and encoder-decoder attention mechanisms, our approach is able to exploit the rich scene information in an end-to-end manner, effectively removing the need for hand-designed discriminative models. In online tracking, given a new test frame, dense prediction is performed at all spatial positions. Not only location, but also bounding box of the target object is obtained in a robust fashion, streamlining the discriminative tracking pipeline. DTT is conceptually simple and easy to implement. It yields state-of-the-art performance on four popular benchmarks including GOT-10k, LaSOT, NfS, and TrackingNet while running at over 50 FPS, confirming its effectiveness and efficiency. We hope DTT may provide a new perspective for single-object visual tracking.
Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Xuetao Feng, Hanqing Lu
ICCV4
2021 Multi-initialization Optimization Network for Accurate 3D Human Pose and Shape Estimation
abstract
3D human pose and shape recovery from a monocular RGB image is a challenging task. Existing learning based methods highly depend on weak supervision signals, e.g. 2D and 3D joint location, due to the lack of in-the-wild paired 3D supervision. However, considering the 2D-to-3D ambiguities existed in these weak supervision labels, the network is easy to get stuck in local optima when trained with such labels. In this paper, we reduce the ambituity by optimizing multiple initializations. Specifically, we propose a three-stage framework named Multi-Initialization Optimization Network (MION). In the first stage, we strategically select different coarse 3D reconstruction candidates which are compatible with the 2D keypoints of input sample. Each coarse reconstruction can be regarded as an initialization leads to one optimization branch. In the second stage, we design a mesh refinement transformer (MRT) to respectively refine each coarse reconstruction result via a self-attention mechanism. Finally, a Consistency Estimation Network (CEN) is proposed to find the best result from mutiple candidates by evaluating if the visual evidence in RGB image matches a given 3D reconstruction. Experiments demonstrate that our Multi-Initialization Optimization Network outperforms existing 3D mesh based methods on multiple public benchmarks.
Zhiwei Liu 0004, Xiangyu Zhu 0001, Lu Yang 0006, Ming Tang 0001, Zhen Lei 0001, Guibo Zhu, Xuetao Feng, Yan Wang 0068, Jinqiao Wang
ACM Multimedia7
2021 High-Performance Discriminative Tracking with Target-Aware Feature Embeddings
Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Hanqing Lu
PRCV (1)4
2019 Feature Distilled Tracking
abstract
Feature extraction and representation is one of the most important components for fast, accurate, and robust visual tracking. Very deep convolutional neural networks (CNNs) provide effective tools for feature extraction with good generalization ability. However, extracting features using very deep CNN models needs high performance hardware due to its large computation complexity, which prohibits its extensions in real-time applications. To alleviate this problem, we aim at obtaining small and fast-to-execute shallow models based on model compression for visual tracking. Specifically, we propose a small feature distilled network (FDN) for tracking by imitating the intermediate representations of a much deeper network. The FDN extracts rich visual features with higher speed than the original deeper network. To further speed-up, we introduce a shift-and-stitch method to reduce the arithmetic operations, while preserving the spatial resolution of the distilled feature maps unchanged. Finally, a scale adaptive discriminative correlation filter is learned on the distilled feature for visual tracking to handle scale variation of the target. Comprehensive experimental results on object tracking benchmark datasets show that the proposed approach achieves 5× speed-up with competitive performance to the state-of-the-art deep trackers.
Guibo Zhu, Jinqiao Wang, Peisong Wang 0001, Yi Wu 0001, Hanqing Lu
IEEE Trans. Cybern.1
2019 Dynamic Collaborative Tracking
abstract
Correlation filter has been demonstrated remarkable success for visual tracking recently. However, most existing methods often face model drift caused by several factors, such as unlimited boundary effect, heavy occlusion, fast motion, and distracter perturbation. To address the issue, this paper proposes a unified dynamic collaborative tracking framework that can perform more flexible and robust position prediction. Specifically, the framework learns the object appearance model by jointly training the objective function with three components: target regression submodule, distracter suppression submodule, and maximum margin relation submodule. The first submodule mainly takes advantage of the circulant structure of training samples to obtain the distinguishing ability between the target and its surrounding background. The second submodule optimizes the label response of the possible distracting region close to zero for reducing the peak value of the confidence map in the distracting region. Inspired by the structure output support vector machines, the third submodule is introduced to utilize the differences between target appearance representation and distracter appearance representation in the discriminative mapping space for alleviating the disturbance of the most possible hard negative samples. In addition, a CUR filter as an assistant detector is embedded to provide effective object candidates for alleviating the model drift problem. Comprehensive experimental results show that the proposed approach achieves the state-of-the-art performance in several public benchmark data sets.
Guibo Zhu, Zhaoxiang Zhang 0001, Jinqiao Wang, Yi Wu 0001, Hanqing Lu
IEEE Trans. Neural Networks Learn. Syst.1
2018 Appearance features in Encoding Color Space for visual surveillance
Lingxiang Wu, Min Xu 0001, Guibo Zhu, Jinqiao Wang, Tianrong Rao
Neurocomputing3
2018 Bundled Local Features for Image Representation
abstract
Local features have been widely used for image representation. Traditional methods often treat each local feature independently or simply model the correlations of local features with spatial partition. However, local features are correlated and should be jointly modeled. Besides, due to the variety of images, predefined partition rules will probably introduce noisy information. To solve these problems, in this paper we propose a novel bundled local features method for efficient image representation and apply it for classification. Specially, we first extract local features and bundle them together with over-complete spatial shapes by viewing each local feature as the central point. Then, the most discriminatively bundling features are selected by reconstruction error minimization. The encoding parameters are then used for image representations in a matrix form. Finally, we train bi-linear classifiers with quadratic hinge loss to predict the classes of images. The proposed method can combine local features appropriately and efficiently for discriminative representations. Experimental results on three image data sets show the effectiveness of the proposed method compared with other local features combination strategies.
Chunjie Zhang 0001, Jitao Sang 0001, Guibo Zhu, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 Image Class Prediction by Joint Object, Context, and Background Modeling
abstract
State-of-the-art image classification methods often use spatial pyramid matching or its variants to make use of the spatial layout of visual features. However, objects may appear at various places with different scales and orientations. Besides, traditionally object-centric-based methods only consider objects and the background without fully exploring the context information. To solve these problems, in this paper we propose a novel image classification method by jointly modeling the object, context, and background information (OCB). OCB consists of three components: 1) locate the positions of objects; 2) determine the context areas of objects; and 3) treat the other areas as the background. We use objectness proposal techniques to select candidate bounding boxes. Boxes with high confidence scores are combined to determine objects' positions. To select the context areas, we use candidate boxes that have relatively lower confidence scores compared with boxes for object location selection. The other areas are viewed as the background. We jointly combine the object, context, and background for image representation and classification. Experiments on six data sets well demonstrate the superiority of the proposed OCB method over other spatial partition methods.
Chunjie Zhang 0001, Guibo Zhu, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.2
2017 Diverse Neuron Type Selection for Convolutional Neural Networks
abstract
The activation function for neurons is a prominent element in the deep learning architecture for obtaining high performance. Inspired by neuroscience findings, we introduce and define two types of neurons with different activation functions for artificial neural networks: excitatory and inhibitory neurons, which can be adaptively selected by self-learning. Based on the definition of neurons, in the paper we not only unify the mainstream activation functions, but also discuss the complementariness among these types of neurons. In addition, through the cooperation of excitatory and inhibitory neurons, we present a compositional activation function that leads to new state-of-the-art performance comparing to rectifier linear units. Finally, we hope that our framework not only gives a basic unified framework of the existing activation neurons to provide guidance for future design, but also contributes neurobiological explanations which can be treated as a window to bridge the gap between biology and computer science.
Guibo Zhu, Zhaoxiang Zhang 0001, Xu-Yao Zhang, Cheng-Lin Liu 0001
IJCAI1
2017 Image classification by search with explicitly and implicitly semantic representations
Chunjie Zhang 0001, Guibo Zhu, Qingming Huang, Qi Tian 0001
Inf. Sci.2
2016 MC-HOG Correlation Tracking with Saliency Proposal
abstract
Designing effective feature and handling the model drift problem are two important aspects for online visual tracking. For feature representation, gradient and color features are most widely used, but how to effectively combine them for visual tracking is still an open problem. In this paper, we propose a rich feature descriptor, MC-HOG, by leveraging rich gradient information across multiple color channels or spaces. Then MC-HOG features are embedded into the correlation tracking framework to estimate the state of the target. For handling the model drift problem caused by occlusion or distracter, we propose saliency proposals as prior information to provide candidates and reduce background interference. In addition to saliency proposals, a ranking strategy is proposed to determine the importance of these proposals by exploiting the learnt appearance filter, historical preserved object samples and the distracting proposals. In this way, the proposed approach could effectively explore the color-gradient characteristics and alleviate the model drift problem. Extensive evaluations performed on the benchmark dataset show the superiority of the proposed method.
Guibo Zhu, Jinqiao Wang, Yi Wu 0001, Xiaoyu Zhang 0002, Hanqing Lu
AAAI1
2016 Person re-identification via rich color-gradient feature
abstract
Person re-identification refers to match the same pedestrian across disjoint views in non-overlapping camera networks. Lots of local and global features in the literature are put forward to solve the matching problem, where color feature is robust to viewpoint variance and gradient feature provides a rich representation robust to illumination change. However, how to effectively combine the color and gradient features is an open problem. In this paper, to effectively leverage the color-gradient property in multiple color spaces, we propose a novel Second Order Histogram feature (SOH) for person reidentification in large surveillance dataset. Firstly, we utilize discrete encoding to transform commonly used color space into Encoding Color Space (ECS), and calculate the statistical gradient features on each color channel. Then, a second order statistical distribution is calculated on each cell map with a spatial partition. In this way, the proposed SOH feature effectively leverages the statistical property of gradient and color as well as reduces the redundant information. Finally, a metric learned by KISSME [1] with Mahalanobis distance is used for person matching. Experimental results on three public datasets, VIPeR, CAVIAR and CUHK01, show the promise of the proposed approach.
Lingxiang Wu, Jinqiao Wang, Guibo Zhu, Min Xu 0001, Hanqing Lu
ICME3
2016 Learning weighted part models for object tracking
Chaoyang Zhao, Jinqiao Wang, Guibo Zhu, Yi Wu 0001, Hanqing Lu
Comput. Vis. Image Underst.3
2016 Clustering based ensemble correlation tracking
Guibo Zhu, Jinqiao Wang, Hanqing Lu
Comput. Vis. Image Underst.1
2015 Collaborative Correlation Tracking
abstract
Correlation filter based tracking has attracted many researchers’ attention in recent years for high efficiency and robustness. Most existing works focus on exploiting different characteristics with correlation filters for visual tracking, e.g. circulant structure, kernel trick, effective feature representation and context information. However, how to handle the scale variation and the model drift is still an open problem. In this paper, we propose a collaborative correlation tracker to deal with the above problems. Firstly, we extend the correlation tracking filter by embedding the scale factor into the kernelized matrix to handle the scale variation. Then a novel long-term CUR filter for detection is learnt efficiently with random sampling to alleviate model drift by detecting effective object candidates in the collaborative tracker. In this way, the proposed approach could estimate the object state accurately and handle the model drift problem effectively. Extensive experiments show the superiority of the proposed method.
Guibo Zhu, Jinqiao Wang, Yi Wu 0001, Hanqing Lu
BMVC1
2015 DualDS: A dual discriminative rating elicitation framework for cold start recommendation
Xi Zhang 0018, Jian Cheng 0001, Shuang Qiu 0002, Guibo Zhu, Hanqing Lu
Knowl. Based Syst.4
2015 Weighted Part Context Learning for Visual Tracking
abstract
Context information is widely used in computer vision for tracking arbitrary objects. Most of the existing studies focus on how to distinguish the object of interest from background or how to use keypoint-based supporters as their auxiliary information to assist them in tracking. However, in most cases, how to discover and represent both the intrinsic properties inside the object and the surrounding context is still an open problem. In this paper, we propose a unified context learning framework that can effectively capture spatiotemporal relations, prior knowledge, and motion consistency to enhance tracker's performance. The proposed weighted part context tracker (WPCT) consists of an appearance model, an internal relation model, and a context relation model. The appearance model represents the appearances of the object and the parts. The internal relation model utilizes the parts inside the object to directly describe the spatiotemporal structure property, while the context relation model takes advantage of the latent intersection between the object and background regions. Then, the three models are embedded in a max-margin structured learning framework. Furthermore, prior label distribution is added, which can effectively exploit the spatial prior knowledge for learning the classifier and inferring the object state in the tracking process. Meanwhile, we define online update functions to decide when to update WPCT, as well as how to reweight the parts. Extensive experiments and comparisons with the state of the arts demonstrate the effectiveness of the proposed method.
Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu
IEEE Trans. Image Process.1
2014 Clustering Ensemble Tracking
Guibo Zhu, Jinqiao Wang, Hanqing Lu
ACCV (5)1
2014 Part Context Learning for Visual Tracking
Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu
BMVC1
2014 Object tracking with part-based discriminative context models
abstract
Object tracking is a classic problem in computer vision. Part-based appearance model has been applied to object tracking and shown good performance. However, how to initialize the parts is still an open question. In this paper, we believe that the selection of discriminative parts and effectively modeling the structural context information could improve the tracking performance. Therefore, we tackle the tracking problem by discovering discriminative parts through exemplar-SVM in the initialization, and then exploit the structural relationship between discriminative context parts and the object in the process of tracking, which is consensual in the spatio-temporal domain. Experimental results demonstrate that our approach outperforms state-of-the-art trackers on benchmark videos.
Guibo Zhu, Jinqiao Wang, Hanqing Lu
ICIP1
2013 Collaborative Tracking: Dynamically Fusing Short-Term Trackers and Long-Term Detector
Guibo Zhu, Jinqiao Wang, Hanqing Lu
MMM (2)1