Chaoyang Zhao

dblp:08/9467 · DBLP profile ↗
← Back
48ranked-venue papers
4as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 14 since 2021Computer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 ViPSN 2.0: A Reconfigurable Battery-Free IoT Platform for Vibration Energy Harvesting
abstract
Vibration energy harvesting is a promising solution for powering battery-free IoT systems; however, the instability of ambient vibrations presents significant challenges, such as limited harvested energy, intermittent power supply, and poor adaptability to various applications. To address these challenges, this paper proposes ViPSN2.0, a modular and reconfigurable IoT platform that supports multiple vibration energy harvesters (piezoelectric, electromagnetic, and triboelectric) and accommodates sensing tasks with varying application requirements through standardized hot-swappable interfaces. ViPSN 2.0 incorporates an energyindication power management framework tailored to various application demands, including light-duty discrete sampling, heavyduty high-power sensing, and complex-duty streaming tasks, thereby effectively managing fluctuating energy availability. The platform’s versatility and robustness are validated through three representative applications: ViPSN-Beacon, using an ultra-lowcost structural PZT (ϕ35 mm, <0.002 $) to enable a BLE advertisement from a single transient fingertip press with 100 m line of sight; ViPSN-LoRa, supporting wireless communication powered by wave vibrations in actual marine environments (Bohai Bay) with per-uplink task energy compatible with kilometer-scale field links; and ViPSN-Cam, enabling intermittent image capture and wireless transfer, delivering one frame approximately every 15 s under typical conditions. Experimental results demonstrate that ViPSN 2.0 can reliably meet a wide range of requirements in practical battery-free IoT deployments under energy-constrained conditions.
Xin Li 0097, Mianxin Xiao, Jiaqing Chu, Weifeng Huang, Jiashun Li, Yaoyi Li, Mingjing Cai, Daxing Zhang, Congsi Wang, Bao Zhao, Qitao Lu, Minyi Xu, Shitong Fang, Xuanyu Huang, Chaoyang Zhao, Yaowen Yang, Guobiao Hu, Junrui Liang, Wei-Hsin Liao
IEEE Internet Things J.21
2026 DAS-Accelerometer Data Fusion With Semi-Supervised Graph Variational Autoencoder for In-Service Train Wheel Flat Detection
abstract
Wheel flats (WF) are a common defect in railway systems, posing risks to operational safety, passenger comfort, and the longevity of infrastructure. Existing detection methods face significant challenges, including sparse labeled data, high noise interference, and limited adaptability to complex operational conditions. To address these issues, this study introduces a semi-supervised learning workflow integrating multi-sensor data from Distributed Acoustic Sensing (DAS) and accelerometers, with a novel Graph Vector-Quantization Variational AutoEncoder (GVQVAE) as the core component. The model combines time-frequency analysis for feature extraction, a graph-based architecture for data fusion, and a vector quantization mechanism to effectively leverage both labeled and unlabeled data. Experimental results from an operational subway system demonstrate the model’s robustness and high accuracy, with an average detection accuracy of 97.08%. These findings highlight the potential of the proposed DAS-accelerometer fusion and GVQVAE model as an effective, scalable solution for enhancing WF detection in modern railway systems.
Yiqing Dong, Chengjia Han, Shuai Qu, Chaoyang Zhao, Aayush Madan, Yuguang Fu, Yaowen Yang
IEEE Trans. Intell. Transp. Syst.4
2025 PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
abstract
Understanding the environment and a robot’s physical reachability is crucial for task execution. While state-of-the-art vision-language models (VLMs) excel in environmental perception, they often generate inaccurate or impractical responses in embodied visual reasoning tasks due to a lack of understanding of robotic physical reachability. To address this issue, we propose a unified representation of physical reachability across diverse robots, i.e., Space-Physical Reachability Map (S-P Map), and PhysVLM, a vision-language model that integrates this reachability information into visual reasoning. Specifically, the S-P Map abstracts a robot’s physical reachability into a generalized spatial representation, independent of specific robot configurations, allowing the model to focus on reachability features rather than robot-specific parameters. Subsequently, PhysVLM extends traditional VLM architectures by incorporating an additional feature encoder to process the S-P Map, enabling the model to reason about physical reachability without compromising its general vision-language capabilities. To train and evaluate PhysVLM, we constructed a large-scale multi-robot dataset, Phys100K, and a challenging benchmark, EQA-phys, which includes tasks for six different robots in both simulated and real-world environments. Experimental results demonstrate that PhysVLM outperforms existing models, achieving a 14% improvement over GPT-4o on EQA-phys and surpassing advanced embodied VLMs such as RoboMamba and SpatialVLM on the RoboVQA-val and OpenEQA benchmarks. Additionally, the S-P Map shows strong compatibility with various VLMs, and its integration into GPT-4o-mini yields a 7.1% performance improvement.
Manli Tao, Chaoyang Zhao, Haiyun Guo, Honghui Dong, Ming Tang 0001, Jinqiao Wang
CVPR3
2025 FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
abstract
Pedestrian attribute recognition (PAR) is a fundamental perception task in intelligent transportation and security. To tackle this fine-grained task, most existing methods focus on extracting regional features to enrich attribute information. However, a regional feature is typically used to predict a fixed set of pre-defined attributes in these methods, which limits the performance and practicality in two aspects: 1) Regional features may compromise fine-grained patterns unique to certain attributes in favor of capturing common characteristics shared across attributes. 2) Regional features cannot generalize to predict unseen attributes in the test time. In this paper, we propose the Fine-grained Optimization with semantiC gUided underStanding (FOCUS) approach for PAR, which adaptively extracts fine-grained attribute-level features for each attribute individually, regardless of whether the attributes are seen or not during training. Specifically, we propose the Multi-Granularity Mix Tokens (MGMT) to capture latent features at varying levels of visual granularity, thereby enriching the diversity of the extracted information. Next, we introduce the Attribute-guided Visual Feature Extraction (AVFE) module, which leverages textual attributes as queries to retrieve their corresponding visual attribute features from the Mix Tokens using a cross-attention mechanism. To ensure that textual attributes focus on the appropriate Mix Tokens, we further incorporate a Region-Aware Contrastive Learning (RACL) method, encouraging attributes within the same region to share consistent attention maps. Extensive experiments on PA100K, PETA, and RAPv1 datasets demonstrate the effectiveness and strong generalization ability of our method.
Hongyan An, Kuan Zhu, Haiyun Guo, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
ICME5
2025 LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task Planning
abstract
In recent years, lightweight large language models (LLMs) have garnered significant attention in the robotics field due to their low computational resource requirements and suitability for edge deployment. However, in task planning—particularly for complex tasks that involve dynamic semantic logic reasoning—lightweight LLMs have underperformed. To address this limitation, we propose a novel task planner, LightPlanner, which enhances the performance of lightweight LLMs in complex task planning by fully leveraging their reasoning capabilities. Unlike conventional planners that use fixed skill templates, LightPlanner controls robot actions via parameterized function calls, dynamically generating parameter values. This approach allows for fine-grained skill control and improves task planning success rates in complex scenarios. Furthermore, we introduce hierarchical deep reasoning. Before generating each action decision step, LightPlanner thoroughly considers three levels: action execution (feedback verification), semantic parsing (goal consistency verification), and parameter generation (parameter validity verification). This ensures the correctness of subsequent action controls. Additionally, we incorporate a memory module to store historical actions, thereby reducing context length and enhancing planning efficiency for long-term tasks. We train the LightPlanner-1.5B model on our LightPlan-40k dataset, which comprises 40,000 action controls across tasks with 2 to 13 action steps. Experiments demonstrate that our model achieves the highest task success rate despite having the smallest number of parameters. In tasks involving spatial semantic reasoning, the success rate exceeds that of ReAct by 14.9%. Moreover, we demonstrate LightPlanner’s potential to operate on edge devices.
Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang 0001, Jinqiao Wang
IROS3
2025 PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
abstract
Visual reasoning in multimodal large language models (MLLMs) has primarily been studied in passive, static settings, limiting their effectiveness in real-world physical environments where an embodied agent must contend with incomplete information due to occlusion or a limited field of view. Humans, in contrast, leverage their embodiment to actively explore and interact with their environment—moving, examining, and manipulating objects—to gather information through a closed-loop process integrating perception, reasoning, and action. Inspired by this capability, we introduce the Active Visual Reasoning (AVR) task, extending visual reasoning to a paradigm of embodied interaction in partially observable environments. AVR necessitates embodied agents to: (1) actively acquire information via sequential physical actions, (2) integrate observations across multiple steps for coherent reasoning, and (3) dynamically adjust decisions based on evolving visual feedback. To rigorously evaluate AVR, we introduce CLEVR-AVR, a simulation benchmark featuring multi-round interactive environments designed to assess both reasoning correctness and information-gathering efficiency. We present AVR-152k, a large-scale dataset that offers rich Chain-of-Thought (CoT) annotations detailing iterative reasoning for uncertainty identification, action-conditioned information gain prediction, and information-maximizing action selection, crucial for training agents in a higher-order Markov Decision Process. Building on this, we develop PhysVLM-AVR, an embodied MLLM achieving state-of-the-art performance on CLEVR-AVR, embodied reasoning (OpenEQA, RoboVQA), and passive visual reasoning (GeoMath, Geometry30K). Our analysis also reveals that current embodied MLLMs, despite detecting information incompleteness, struggle to actively acquire and integrate new information through interaction, highlighting a fundamental gap in active reasoning capabilities.
Xuantang Xiong, Manli Tao, Chaoyang Zhao, Honghui Dong, Ming Tang 0001, Jinqiao Wang
NeurIPS5
2025 MaSA: Mamba-Based Global Feature Selective Aggregator for Efficient Lane Detection
La Zhang, Haiyun Guo, Chaoyang Zhao, Jinqiao Wang
PRCV (3)5
2025 Multi-Context enhanced Lane-Changing prediction using a heterogeneous Graph Neural Network
Yiqing Dong, Chengjia Han, Chaoyang Zhao, Aayush Madan, Lipi Mohanty, Yaowen Yang
Expert Syst. Appl.3
2025 A Semi-Supervised Diffusion-Based Paradigm for Vehicle-Track System Health Monitoring With Distributed Acoustic Sensing
abstract
Monitoring the health of vehicle-track system using deep learning and distributed fiber optic sensing presents a significant challenge due to the vast volume of real-time data and the difficulty of directly assessing the system’s condition. This often results in a severe imbalance in the distribution of extreme samples within the dataset, as large-scale signal collection typically lacks manual labeling. Consequently, supervised deep learning models face limitations due to insufficient labeled training data, while unsupervised deep learning models struggle with contamination from ambiguous samples whose health status remains unclear, hindering the development of robust and accurate models. To address this challenge, we propose SemAnoDiffusion, a semi-supervised model based on blur diffusion and an enhanced contrastive loss training approach. SemAnoDiffusion leverages a small set of labeled data alongside a large amount of unlabeled samples to accurately differentiate between anomalous data, normal data, and ambiguous samples that fall between these categories. In a case study of a metro system in Singapore, Distributed Acoustic Sensing and accelerometer arrays were used to collect track vibration responses as trains passed, with wheel flats occurring in a small subset of the trains. SemAnoDiffusion achieved 100% accuracy in classifying manually labeled normal and anomalous samples and effectively identified semi-damaged samples with unclear damage levels from the labeled data, successfully detecting all trains with wheel flats.
Chengjia Han, Yiqing Dong, Shuai Qu, Chaoyang Zhao, Aayush Madan, Yuguang Fu, Yaowen Yang
IEEE Trans. Intell. Transp. Syst.5
2024 Self-Supervised Representation Learning from Arbitrary Scenarios
abstract
Current self-supervised methods can primarily be categorized into contrastive learning and masked image modeling. Extensive studies have demonstrated that combining these two approaches can achieve state-of-the-art performance. However, these methods essentially reinforce the global consistency of contrastive learning without taking into account the conflicts between these two approaches, which hinders their generalizability to arbitrary scenarios. In this paper, we theoretically prove that MAE serves as a patch-level contrastive learning, where each patch within an image is considered as a distinct category. This presents a significant conflict with global-level contrastive learning, which treats all patches in an image as an identical category. To address this conflict, this work abandons the non-generalizable global-level constraints and proposes explicit patch-level contrastive learning as a solution. Specifically, this work employs the encoder of MAE to generate dual-branch features, which then perform patch-level learning through a decoder. In contrast to global-level data aug-mentation in contrastive learning, our approach leverages patch-level feature augmentation to mitigate interference from global-level learning. Consequently, our approach can learn heterogeneous representations from a single image while avoiding the conflicts encountered by previous methods. Massive experiments affirm the potential of our method for learning from arbitrary scenarios.
Zhaowen Li, Yousong Zhu, Zhiyang Chen 0002, Zongxin Gao, Rui Zhao 0001, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
CVPR6
2024 The Devil is in Details: Delving Into Lite FFN Design for Vision Transformers
abstract
Transformer has demonstrated exceptional performance on a variety of vision tasks. However, its high computational complexity can become problematic. In this paper, we conduct a systematic analysis of the complexity of each component in vision transformers, and identify an easily overlooked detail: that the Feed-Forward Network (FFN) is the primary computational bottleneck, even more so than the Multi-Head Self-Attention (MHSA) mechanism. Inspired by this, we further propose a lightweight FFN module, named SparseFFN, that can reduce dense computations in both channel and spatial dimension. Specifically, SparseFFN consists of two components: Channel-Sparse FFN (CS-FFN) and Spatial-Sparse FFN (SS-FFN), which can be seamlessly incorporated into various vision transformers and even pure MLP models with significantly fewer FLOPs. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method. For example, our approach can reduce model complexity by 23%-39% for most of vision transformers and MLP models while keeping comparable accuracy.
Zhiyang Chen 0002, Yousong Zhu, Zhaowen Li, Fan Yang 0089, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001
ICASSP5
2024 Intelligent detection of loose fasteners in railway tracks using distributed acoustic sensing and machine learning
Chengjia Han, Shun Wang 0002, Aayush Madan, Chaoyang Zhao, Lipi Mohanty, Yuguang Fu, Ruihua Liang, Ean Seong Huang, Tony Zheng, Phui Kai Ong, Alvin Zhang, Khai Jhin Woon, Kai Xin Wong, Yaowen Yang
Eng. Appl. Artif. Intell.4
2024 Efficient Masked Autoencoders With Self-Consistency
abstract
Inspired by the masked language modeling (MLM) in natural language processing tasks, the masked image modeling (MIM) has been recognized as a strong self-supervised pre-training method in computer vision. However, the high random mask ratio of MIM results in two serious problems: 1) the inadequate data utilization of images within each iteration brings prolonged pre-training, and 2) the high inconsistency of predictions results in unreliable generations, i.e., the prediction of the identical patch may be inconsistent in different mask rounds, leading to divergent semantics in the ultimately generated outcomes. To tackle these problems, we propose the efficient masked autoencoders with self-consistency (EMAE) to improve the pre-training efficiency and increase the consistency of MIM. In particular, we present a parallel mask strategy that divides the image into K non-overlapping parts, each of which is generated by a random mask with the same mask ratio. Then the MIM task is conducted parallelly on all parts in an iteration and the model minimizes the loss between the predictions and the masked patches. Besides, we design the self-consistency learning to further maintain the consistency of predictions of overlapping masked patches among parts. Overall, our method is able to exploit the data more efficiently and obtains reliable representations. Experiments on ImageNet show that EMAE achieves the best performance on ViT-Large with only 13% of MAE pre-training time using NVIDIA A100 GPUs. After pre-training on diverse datasets, EMAE consistently obtains state-of-the-art transfer ability on a variety of downstream tasks, such as image classification, object detection, and semantic segmentation.
Zhaowen Li, Yousong Zhu, Zhiyang Chen 0002, Wei Li 0314, Rui Zhao 0001, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Objformer: Boosting 3D object detection via instance-wise interaction
Manli Tao, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
Pattern Recognit.2
2024 ImFusion: Boosting Two-Stage 3D Object Detection via Image Candidates
abstract
Multi-modal fusion methods combine the advantages of both point clouds and RGB images to boost the performance of 3D object detection. Despite the significant progress, we find that existing two-stage multi-modal fusion methods suffer from the 3D proposal missing in the first stage and projected-style feature fusion mechanism. To solve these problems, we propose a two-stage multi-modal feature fusion network, which improves the recall rate of hard targets in the first stage of network with pseudo 3D proposals generated from image candidates. Then, considering the complementary information between similar image foreground features across multiple objects, we design a multi-modal cross-target fusion module to pay more attention to the foreground objects. It enables a 3D proposal can aggregate the semantic features of multiple image candidates belonging to the same category. Finally, these enhanced fused proposals are processed in the second stage to further boost the performance of 3D detector. Experimental results on SUN RGB-D and KITTI datasets show the effectiveness of our proposed method.
Manli Tao, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001
IEEE Signal Process. Lett.2
2023 ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground Selection
abstract
Background subtraction (BGS) aims to extract all moving objects in the video frames to obtain binary foreground segmentation masks. Deep learning has been widely used in this field. Compared with supervised-based BGS methods, unsupervised methods have better generalization. However, previous unsupervised deep learning BGS algorithms perform poorly in sophisticated scenarios such as shadows or night lights, and they cannot detect objects outside the pre-defined categories. In this work, we propose an unsuper-vised BGS algorithm based on zero-shot object detection called Zero-shot Background Subtraction (ZBS). The proposed method fully utilizes the advantages of zero-shot object detection to build the open-vocabulary instance-level background model. Based on it, the foreground can be effectively extracted by comparing the detection results of new frames with the background model. ZBS performs well for sophisticated scenarios, and it has rich and extensible categories. Furthermore, our method can easily generalize to other tasks, such as abandoned object detection in unseen environments. We experimentally show that ZBS surpasses state-of-the-art unsupervised BGS methods by 4.70% F-Measure on the CDnet 2014 dataset. The code is released at https://github.com/CASIA-IVA-Lab/ZBS.
Yongqi An, Xu Zhao 0003, Tao Yu 0013, Haiyun Gu, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
CVPR5
2022 UniVIP: A Unified Framework for Self-Supervised Visual Pre-training
abstract
Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential.
Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang
CVPR5
2022 C2AM Loss: Chasing a Better Decision Boundary for Long-Tail Object Detection
abstract
Long-tail object detection suffers from poor performance on tail categories. We reveal that the real culprit lies in the extremely imbalanced distribution of the classifier's weight norm. For conventional softmax cross-entropy loss, such imbalanced weight norm distribution yields ill conditioned decision boundary for categories which have small weight norms. To get rid of this situation, we choose to maxi-mize the cosine similarity between the learned feature and the weight vector of target category rather than the inner-product of them. The decision boundary between any two categories is the angular bisector of their weight vectors. Whereas, the absolutely equal decision boundary is sub-optimal because it reduces the model's sensitivity to vari-ous categories. Intuitively, categories with rich data diver-sity should occupy a larger area in the classification space while categories with limited data diversity should occupy a slightly small space. Hence, we devise a Category-Aware Angular Margin Loss (C2AM Loss) to introduce an adaptive angular margin between any two categories. Specif-ically, the margin between two categories is proportional to the ratio of their classifiers' weight norms. As a result, the decision boundary is slightly pushed towards the cat-egory which has a smaller weight norm. We conduct comprehensive experiments on LVIS dataset. C2AM Loss brings 4.9~5.2 AP improvements on different detectors and back-bones compared with baseline.
Tong Wang 0015, Yousong Zhu, Yingying Chen 0003, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001
CVPR4
2022 Transfering Low-Frequency Features for Domain Adaptation
abstract
Previous unsupervised domain adaptation methods did not handle the cross-domain problem from the perspective of frequency for computer vision. The images or feature maps of different domains can be decomposed into the low-frequency component and high-frequency component. This paper pro-poses the assumption that low-frequency information is more domain-invariant while the high-frequency information con-tains domain-related information. Hence, we introduce an approach, named low-frequency module (LFM), to extract domain-invariant feature representations. The LFM is constructed with the digital Gaussian low-pass filter. Our method is easy to implement and introduces no extra hyperparame-ter. We design two effective ways to utilize the LFM for domain adaptation, and our method is complementary to other existing methods and formulated as a plug-and-play unit that can be combined with these methods. Experimental results demonstrate that our LFM outperforms state-of-the-art meth-ods for various computer vision tasks, including image clas-sification and object detection.
Zhaowen Li, Xu Zhao 0003, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang
ICME3
2022 Obj2Seq: Formatting Objects as Sequences with Class Prompt for Visual Tasks
abstract
Visual tasks vary a lot in their output formats and concerned contents, therefore it is hard to process them with an identical structure. One main obstacle lies in the high-dimensional outputs in object-level visual tasks. In this paper, we propose an object-centric vision framework, Obj2Seq. Obj2Seq takes objects as basic units, and regards most object-level visual tasks as sequence generation problems of objects. Therefore, these visual tasks can be decoupled into two steps. First recognize objects of given categories, and then generate a sequence for each of these objects. The definition of the output sequences varies for different tasks, and the model is supervised by matching these sequences with ground-truth targets. Obj2Seq is able to flexibly determine input categories to satisfy customized requirements, and be easily extended to different visual tasks. When experimenting on MS COCO, Obj2Seq achieves 45.7% AP on object detection, 89.0% AP on multi-label classification and 65.0% AP on human pose estimation. These results demonstrate its potential to be generally applied to different visual tasks. Code has been made available at: https://github.com/CASIA-IVA-Lab/Obj2Seq.
Zhiyang Chen 0002, Yousong Zhu, Zhaowen Li, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Rui Zhao 0001, Jinqiao Wang, Ming Tang 0001
NeurIPS7
2022 Global Patch Cross-Attention for Point Cloud Analysis
Manli Tao, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001
PRCV (3)2
2022 Multi-Granularity Mutual Learning Network for Object Re-Identification
abstract
Object re-identification (re-ID), which is key and fundamental technology for intelligent transportation systems, is a challenging task including person re-ID and vehicle re-ID. It aims to retrieve a given target object from the gallery images captured by different cameras. In this task, it is necessary to extract fine-grained and discriminative features to deal with complex inter-class and intra-class variations caused by the changes of camera viewpoints and object poses. Existing methods focus on learning discriminative local features to improve the re-ID performance. Some state-of-the-art methods use key point detection model to locate local features, which also increases the additional computational cost as side effect. Another type of method focuses on how to learn features of different granularity from rigid stripes of different scales. However, there is little attention paid to how to effectively coalesce multi-granularity features without additional calculation cost. To tackle this issue, this paper proposes the Multi-granularity Mutual Learning Network (MMNet) and makes two contributions. 1) We introduce the multi-granularity jigsaw puzzle module into object re-ID to impel the network to learn local discriminative features from multiple visual granularities by breaking spatial correlation in original images. 2) We propose a parameter-free multi-scale feature reconstruction module to facilitate mutual learning of features at multiple grain levels, thereby both global features and local features have strong representation capabilities. Extensive experiments demonstrate the effectiveness of our proposed modules and the superiority of our method over various state-of-the-art methods on both person and vehicle re-ID benchmarks.
Mingfei Tu, Kuan Zhu, Haiyun Guo, Qinghai Miao, Chaoyang Zhao, Guibo Zhu, Honglin Qiao, Gaopan Huang, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Intell. Transp. Syst.5
2021 Adaptive Class Suppression Loss for Long-Tail Object Detection
abstract
To address the problem of long-tail distribution for the large vocabulary object detection task, existing methods usually divide the whole categories into several groups and treat each group with different strategies. These methods bring the following two problems. One is the training inconsistency between adjacent categories of similar sizes, and the other is that the learned model is lack of discrimination for tail categories which are semantically similar to some of the head categories. In this paper, we devise a novel Adaptive Class Suppression Loss (ACSL) to effectively tackle the above problems and improve the detection performance of tail categories. Specifically, we introduce a statistic-free perspective to analyze the long-tail distribution, breaking the limitation of manual grouping. According to this perspective, our ACSL adjusts the suppression gradients for each sample of each class adaptively, ensuring the training consistency and boosting the discrimination for rare categories. Extensive experiments on long-tail datasets LVIS and Open Images show that the our ACSL achieves 5.18% and 5.2% improvements with ResNet50-FPN, and sets a new state of the art. Code and models are available at https://github.com/CASIA-IVA-Lab/ACSL.
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001
CVPR3
2021 Attention-Guided Knowledge Distillation for Efficient Single-Stage Detector
abstract
Knowledge distillation has been successfully applied in image classification for model acceleration. There are also some works employing this technique to object detection, but they all treat different feature regions equally when performing feature mimic. In this paper, we propose an end-to-end attention-guided knowledge distillation method to train efficient single-stage detectors with much smaller backbones. More specifically, we introduce an attention mechanism to prioritize the transfer of important knowledge by focusing on a sparse set of hard samples, leading to a more thorough distillation process. In addition, the proposed distillation method also provides an easy way to train efficient detectors without tedious ImageNet pre-training procedure. Extensive experiments on PASCAL VOC and CityPersons datasets demonstrate the effectiveness of the proposed approach. We achieve 57.96% and 69.48% mAP on VOC07 with the backbone of 1/8 VGG16 and 1/4 VGG16, greatly outperforming their ImageNet pre-trained counterparts by 11.7% and 7.1% respectively.
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Xu Zhao 0003, Jinqiao Wang, Ming Tang 0001
ICME3
2021 DPT: Deformable Patch-based Transformer for Visual Recognition
abstract
Transformer has achieved great success in computer vision, while how to split patches in an image remains a problem. Existing methods usually use a fixed-size patch embedding which might destroy the semantics of objects. To address this problem, we propose a new Deformable Patch (DePatch) module which learns to adaptively split the images into patches with different positions and scales in a data-driven way rather than using predefined fixed patches. In this way, our method can well preserve the semantics in patches. The DePatch module can work as a plug-and-play module, which can easily be incorporated into different transformers to achieve an end-to-end training. We term this DePatch-embedded transformer as Deformable Patch-based Transformer (DPT) and conduct extensive evaluations of DPT on image classification and object detection. Results show DPT can achieve 81.8% top-1 accuracy on ImageNet classification, and 43.7% box AP with RetinaNet, 44.3% with Mask R-CNN on MSCOCO object detection. Code has been made available at: https://github.com/CASIA-IVA-Lab/DPT.
Zhiyang Chen 0002, Yousong Zhu, Chaoyang Zhao, Guosheng Hu, Wei Zeng 0006, Jinqiao Wang, Ming Tang 0001
ACM Multimedia3
2021 SiWa: see into walls via deep UWB radar
abstract
Being able to see into walls is crucial for diagnostics of building health; it enables inspections of wall structure without undermining the structural integrity. However, existing sensing devices do not seem to offer a full capability in mapping the in-wall structure while identifying their status (e.g., seepage and corrosion). In this paper, we design and implement SiWa as a low-cost and portable system for wall inspections. Built upon a customized IR-UWB radar, SiWa scans a wall as a user swipes its probe along the wall surface; it then analyzes the reflected signals to synthesize an image and also to identify the material status. Although conventional schemes exist to handle these problems individually, they require troublesome calibrations that largely prevent them from practical adoptions. To this end, we equip SiWa with a deep learning pipeline to parse the rich sensory data. With innovative construction and training, the deep learning modules perform structural imaging and the subsequent analysis on material status, without the need for repetitive parameter tuning and calibrations. We build SiWa as a prototype and evaluate its performance via extensive experiments and field studies; results evidently confirm that SiWa accurately maps in-wall structures, identifies their materials, and detects possible defects, suggesting a promising solution for diagnosing building health with minimal effort and cost.
Tianyue Zheng, Zhe Chen 0015, Jun Luo 0001, Lin Ke, Chaoyang Zhao, Yaowen Yang
MobiCom5
2021 MST: Masked Self-Supervised Transformer for Visual Representation
abstract
Transformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only consider the high-level feature and learning representation from a global perspective, which may fail to transfer to the downstream dense prediction tasks focusing on local features. In this paper, we present a novel Masked Self-supervised Transformer approach named MST, which can explicitly capture the local context of an image while preserving the global semantic information. Specifically, inspired by the Masked Language Modeling (MLM) in NLP, we propose a masked token strategy based on the multi-head self-attention map, which dynamically masks some tokens of local patches without damaging the crucial structure for self-supervised learning. More importantly, the masked tokens together with the remaining tokens are further recovered by a global image decoder, which preserves the spatial information of the image and is more friendly to the downstream dense prediction tasks. The experiments on multiple datasets demonstrate the effectiveness and generality of the proposed method. For instance, MST achieves Top-1 accuracy of 76.9% with DeiT-S only using 300-epoch pre-training by linear evaluation, which outperforms supervised methods with the same epoch by 0.4% and its comparable variant DINO by 1.0%. For dense prediction tasks, MST also achieves 42.7% mAP on MS COCO object detection and 74.04% mIoU on Cityscapes segmentation only with 100-epoch pre-training.
Zhaowen Li, Zhiyang Chen 0002, Fan Yang 0089, Wei Li 0314, Yousong Zhu, Chaoyang Zhao, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang
NeurIPS6
2020 Large Batch Optimization for Object Detection: Training COCO in 12 minutes
Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Wei Zeng 0006, Yaowei Wang 0001, Jinqiao Wang, Ming Tang 0001
ECCV (21)3
2020 Task Decoupled Knowledge Distillation For Lightweight Face Detectors
abstract
Face detection is a hot topic in computer vision. The face detection methods usually consist of two subtasks, i.e. the classification subtask and the regression subtask, which are trained with different samples. However, current face detection knowledge distillation methods usually couple the two subtasks, and use the same set of samples in the distillation task. In this paper, we propose a task decoupled knowledge distillation method, which decouples the detection distillation task into two subtasks and uses different samples in distilling the features of different subtasks. We firstly propose a feature decoupling method to decouple the classification features and the regression features, without introducing any extra calculations at inference time. Specifically, we generate the corresponding features by adding task-specific convolutions in the teacher network and adding adaption convolutions on the feature maps of the student network. Then we select different samples for different subtasks to imitate. Moreover, we also propose an effective probability distillation method to joint boost the accuracy of the student network. We apply our distillation method on a lightweight face detector, EagleEye. Experimental results show that the proposed method effectively improves the student detector's accuracy by 5.1%, 5.1%, and 2.8% AP in Easy, Medium, Hard subsets respectively.
Xiaoqing Liang, Xu Zhao 0003, Chaoyang Zhao, Nanfei Jiang, Ming Tang 0001, Jinqiao Wang
ACM Multimedia3
2020 A novel data augmentation scheme for pedestrian detection with attribute preserving GAN
Songyan Liu, Haiyun Guo, Jian-Guo Hu, Xu Zhao 0003, Chaoyang Zhao, Tong Wang 0015, Yousong Zhu, Jinqiao Wang, Ming Tang 0001
Neurocomputing5
2020 Food det: Detecting foods in refrigerator with supervised transformer network
Yousong Zhu, Xu Zhao 0003, Chaoyang Zhao, Jinqiao Wang, Hanqing Lu
Neurocomputing3
2019 Cascade Attention Network for Person Re-Identification
abstract
Person re-identification is a challenging task due to the viewpoint, illumination and pose variations. Recent works focus on extracting part-level features to offer beneficial fine-grained information. However, the part misalignment as well as the multi-stage training process limits their performance. Inspired by the human visual attention mechanism, this paper builds a cascade attention network(CAN) to learn the discriminative person features in a coarse-to-fine manner. Firstly, we employ the human semantic parsing module to generate coarse-grained part-level attention, which corresponds to the division of human body parts and can effectively filter the background noise. Then, to extract the local detailed features within each part, we introduce spatial-channel attention module to generate fine-grained pixel-level attention, which can further highlight the distinctive characteristics and repress the irrelevant ones. Finally, we can obtain an efficient person feature descriptor by combining both the global and local features. The whole learning process is conducted end-to-end. Experimental results show that the proposed method not only considerably outperforms its counter part but also achieves competitive performance on Market-1501 and DukeMTMC.
Haiyun Guo, Huiyao Wu, Chaoyang Zhao, Huichen Zhang, Jinqiao Wang, Hanqing Lu
ICIP3
2019 Mask Guided Knowledge Distillation for Single Shot Detector
abstract
In this paper, we explore the idea of distilling small networks for object detection task. More specifically, we propose a two-stage approach to learn more compact and efficient detectors under the single-shot object detection framework by leveraging knowledge distillation. During the 1st stage, we learn the feature maps of the student model for each of the prediction head from the teacher model. Instead of fitting the whole feature map directly, here we propose the mask guided structure including not only the entire feature map (i.e. global features) but also region features covered by the object (i.e. local features), which can significantly improve the performance of the student network. For the 2nd stage, the ground-truth is used to further refine the performance. Experimental results on PASCAL VOC and KITTI dataset demonstrate the effectiveness of our proposed approach. We achieve 56.88% mAP on VOC2007 at 143 FPS with the backbone of 1/8 VGG16.
Yousong Zhu, Chaoyang Zhao, Chenxia Han, Jinqiao Wang, Hanqing Lu
ICME2
2019 Adversarial image generation by combining content and style
abstract
Images can be considered as the combination of two parts: the content and the style. The authors’ approach can leverage this property by extracting a certain unique style from the reference images and combining it to generate images with new contents. With a well‐defined style feature extraction module, they propose a novel framework to generate images with various styles and the same content. To train the style specific image generation model efficiently, a double‐cycle training strategy is proposed: they input two natural‐content pairs simultaneously, extract their style features, and exchange them twice to obtain the reconstruction of the input natural images. What is more, they apply the triplet margin loss to the style feature extracted from the images before and after style exchange and an adversarial discriminator to force the style‐exchanged images to be real. They perform experiments on licence‐plate image, Chinese characters, and shoes or handbags images generating, obtain photo‐realistic results and remarkably improve the corresponding supervised recognition task.
Songyan Liu, Chaoyang Zhao, Yunze Gao, Jinqiao Wang, Ming Tang 0001
IET Image Process.2
2019 Elite Loss for scene text detection
Xu Zhao 0003, Chaoyang Zhao, Haiyun Guo, Yousong Zhu, Ming Tang 0001, Jinqiao Wang
Neurocomputing2
2019 Attention CoupleNet: Fully Convolutional Attention Coupling Network for Object Detection
abstract
The field of object detection has made great progress in recent years. Most of these improvements are derived from using a more sophisticated convolutional neural network. However, in the case of humans, the attention mechanism, global structure information, and local details of objects all play an important role for detecting an object. In this paper, we propose a novel fully convolutional network, named as Attention CoupleNet, to incorporate the attention-related information and global and local information of objects to improve the detection performance. Specifically, we first design a cascade attention structure to perceive the global scene of the image and generate class-agnostic attention maps. Then the attention maps are encoded into the network to acquire object-aware features. Next, we propose a unique fully convolutional coupling structure to couple global structure and local parts of the object to further formulate a discriminative feature representation. To fully explore the global and local properties, we also design different coupling strategies and normalization ways to make full use of the complementary advantages between the global and local information. Extensive experiments demonstrate the effectiveness of our approach. We achieve state-of-the-art results on all three challenging data sets, i.e., a mAP of 85.7% on VOC07, 84.3% on VOC12, and 35.4% on COCO. Codes are publicly available at https://github.com/tshizys/CoupleNet.
Yousong Zhu, Chaoyang Zhao, Haiyun Guo, Jinqiao Wang, Xu Zhao 0003, Hanqing Lu
IEEE Trans. Image Process.2
2018 Learning Coarse-to-Fine Structured Feature Embedding for Vehicle Re-Identification
abstract
Vehicle re-identification (re-ID) is to identify the same vehicle across different cameras. It’s a significant but challenging topic, which has received little attention due to the complex intra-class and inter-class variation of vehicle images and the lack of large-scale vehicle re-ID dataset. Previous methods focus on pulling images from different vehicles apart but neglect the discrimination between vehicles from different vehicle models, which is actually quite important to obtain a correct ranking order for vehicle re-ID. In this paper, we learn a structured feature embedding for vehicle re-ID with a novel coarse-to-fine ranking loss to pull images of the same vehicle as close as possible and achieve discrimination between images from different vehicles as well as vehicles from different vehicle models. In the learnt feature space, both intra-class compactness and inter-class distinction are well guaranteed and the Euclidean distance between features directly reflects the semantic similarity of vehicle images. Furthermore, we build so far the largest vehicle re-ID dataset "Vehicle-1M," which involves nearly 1 million images captured in various surveillance scenarios. Experimental results on "Vehicle-1M" and "VehicleID" demonstrate the superiority of our proposed approach.
Haiyun Guo, Chaoyang Zhao, Zhiwei Liu 0004, Jinqiao Wang, Hanqing Lu
AAAI2
2018 DeepSearch: A Fast Image Search Framework for Mobile Devices
abstract
Content-based image retrieval (CBIR) is one of the most important applications of computer vision. In recent years, there have been many important advances in the development of CBIR systems, especially Convolutional Neural Networks (CNNs) and other deep-learning techniques. On the other hand, current CNN-based CBIR systems suffer from high computational complexity of CNNs. This problem becomes more severe as mobile applications become more and more popular. The current practice is to deploy the entire CBIR systems on the server side while the client side only serves as an image provider. This architecture can increase the computational burden on the server side, which needs to process thousands of requests per second. Moreover, sending images have the potential of personal information leakage. As the need of mobile search expands, concerns about privacy are growing. In this article, we propose a fast image search framework, named DeepSearch, which makes complex image search based on CNNs feasible on mobile phones. To implement the huge computation of CNN models, we present a tensor Block Term Decomposition (BTD) approach as well as a nonlinear response reconstruction method to accelerate the CNNs involving in object detection and feature extraction. The extensive experiments on the ImageNet dataset and Alibaba Large-scale Image Search Challenge dataset show that the proposed accelerating approach BTD can significantly speed up the CNN models and further makes CNN-based image search practical on common smart phones.
Peisong Wang 0001, Qinghao Hu 0001, Zhiwei Fang, Chaoyang Zhao, Jian Cheng 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2017 CoupleNet: Coupling Global Structure with Local Parts for Object Detection
Yousong Zhu, Chaoyang Zhao, Jinqiao Wang, Xu Zhao 0003, Yi Wu 0001, Hanqing Lu
ICCV2
2017 Automatic group activity annotation for mobile videos
Chaoyang Zhao, Jinqiao Wang, Jianqiang Li 0002, Hanqing Lu
Multim. Syst.1
2017 Learning discriminative context models for concurrent collective activity recognition
Chaoyang Zhao, Jinqiao Wang, Hanqing Lu
Multim. Tools Appl.1
2016 Scale-Adaptive Deconvolutional Regression Network for Pedestrian Detection
Yousong Zhu, Jinqiao Wang, Chaoyang Zhao, Haiyun Guo, Hanqing Lu
ACCV (2)3
2016 Learning weighted part models for object tracking
Chaoyang Zhao, Jinqiao Wang, Guibo Zhu, Yi Wu 0001, Hanqing Lu
Comput. Vis. Image Underst.1
2015 Weighted Part Context Learning for Visual Tracking
abstract
Context information is widely used in computer vision for tracking arbitrary objects. Most of the existing studies focus on how to distinguish the object of interest from background or how to use keypoint-based supporters as their auxiliary information to assist them in tracking. However, in most cases, how to discover and represent both the intrinsic properties inside the object and the surrounding context is still an open problem. In this paper, we propose a unified context learning framework that can effectively capture spatiotemporal relations, prior knowledge, and motion consistency to enhance tracker's performance. The proposed weighted part context tracker (WPCT) consists of an appearance model, an internal relation model, and a context relation model. The appearance model represents the appearances of the object and the parts. The internal relation model utilizes the parts inside the object to directly describe the spatiotemporal structure property, while the context relation model takes advantage of the latent intersection between the object and background regions. Then, the three models are embedded in a max-margin structured learning framework. Furthermore, prior label distribution is added, which can effectively exploit the spatial prior knowledge for learning the classifier and inferring the object state in the tracking process. Meanwhile, we define online update functions to decide when to update WPCT, as well as how to reweight the parts. Extensive experiments and comparisons with the state of the arts demonstrate the effectiveness of the proposed method.
Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu
IEEE Trans. Image Process.3
2014 Part Context Learning for Visual Tracking
Guibo Zhu, Jinqiao Wang, Chaoyang Zhao, Hanqing Lu
BMVC3
2014 Discriminative Context Models for Collective Activity Recognition
abstract
Context information has been widely studied for recognizing collective activities. Most existing works assume that all individuals in a single image share the same activity label. However, in many cases, multiple activities can be coexisted and serve as the context for each other in real-world scenarios. Based on this observation, we propose a novel approach to model both the intra-class and inter-class behavior interactions among persons in the scenario. By introducing the intra-class and inter-class context descriptors, we propose a unified discriminative model to jointly capture the individual appearance information and the context patterns around the focal person in a max-margin framework. Finally, a greedy forward search method is utilized to optimally label the activities in the testing scene. Experimental results demonstrate the superiority of our approach in activity recognition.
Chaoyang Zhao, Jinqiao Wang, Xiao Bai 0001, Qingshan Liu 0001, Hanqing Lu
ICPR1
2013 Fusing multi-modal features for gesture recognition
abstract
This paper proposes a novel multi-modal gesture recognition framework and introduces its application to continuous sign language recognition. A Hidden Markov Model is used to construct the audio feature classifier. A skeleton feature classifier is trained to provided complementary information based on the Dynamic Time Warping model. The confidence scores generated by two classifiers are firstly normalized and then combined to produce a weighted sum for the final recognition. Experimental results have shown that the precision and recall scores for 20 classes of our multi-modal recognition framework can achieve 0.8829 and 0.8890 respectively, which proves that our method is able to correctly reject false detection caused by single classifier. Our approach scored 0.12756 in mean Levenshtein distance and was ranked 1st in the Multi-modal Gesture Recognition Challenge in 2013.
Jiaxiang Wu 0001, Jian Cheng 0001, Chaoyang Zhao, Hanqing Lu
ICMI3
2012 Object-centered narratives for video surveillance
abstract
Effective video presentation and summarization techniques are critical for fast browsing of video content. In this paper, we propose a novel presentation approach to vividly depict the moving process of a specific object in a surveillance video, which aims at effectively summarizing video content by a static image named narrative. Firstly, the object of interest is extracted and segmented from the video to form a spatio-temporal object tube. Then three criteria are proposed to select the most representative objects from this tube. We formulate the object selecting process as an energy minimization problem, in which each energy term measures a corresponding criterion cost. We maximally preserve the changes of appearance and behavior while remove other redundant content as much as possible. Finally, the selected representative objects are stitched to the background image by Poisson editing. Experimental results show the promise of the proposed approach.
Jinqiao Wang, Chaoyang Zhao, Hanqing Lu, Songde Ma
ICIP3