VLDB 2026 Research / reviewers in the wild / expert
Xu Zhao 0003
dblp:37/3580-3
· DBLP profile ↗
18ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-5888-8872ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Response-time analysis of a big data system with stochastic petri nets
Jian Sun 0003, Xu Zhao 0003, Yunyue Xie |
J. Supercomput. | 4 |
| 2025 | Extracting Sparse Specialist Models from Generalist ModelsabstractRecently, several generalist models such as Contrastive Language Image Pre-training (CLIP) have demonstrated their capabilities of performing diverse downstream tasks through zero-shot or few-shot guidance. When these generalist models are used for the specific downstream task where only a fraction of features is relevant, they would suffer from a significant redundancy of parameters. While existing methods aim to achieve sparsity and specialization, they often require additional training and large datasets. In this paper, we propose a novel framework to extract a sparse specialist model from a generalist model using only few-shot samples, without any training. Our task-specific pruning framework defines task relevance metrics and employs weighted layer-wise pruning, preserving relevant features while removing redundancies. Experiments show that our method maintains nearly identical zero-shot accuracy compared to the original generalist models at 30% sparsity, with only minimal decline at 50%. Tao Yu 0013, Xu Zhao 0003, Yongqi An, Guibo Zhu, Ming Tang 0001, Jinqiao Wang |
ICASSP | 2 |
| 2025 | Systematic Outliers in Large Language ModelsabstractOutliers have been widely observed in Large Language Models (LLMs), significantly impacting model performance and posing challenges for model compression. Understanding the functionality and formation mechanisms of these outliers is critically important. Existing works, however, largely focus on reducing the impact of outliers from an algorithmic perspective, lacking an in-depth investigation into their causes and roles. In this work, we provide a detailed analysis of the formation process, underlying causes, and functions of outliers in LLMs. We define and categorize three types of outliers—activation outliers, weight outliers, and attention outliers—and analyze their distributions across different dimensions, uncovering inherent connections between their occurrences and their ultimate influence on the attention mechanism. Based on these observations, we hypothesize and explore the mechanisms by which these outliers arise and function, demonstrating through theoretical derivations and experiments that they emerge due to the self-attention mechanism's softmax operation. These outliers act as implicit context-aware scaling factors within the attention mechanism. As these outliers stem from systematic influences, we term them systematic outliers. Our study not only enhances the understanding of Transformer-based LLMs but also shows that structurally eliminating outliers can accelerate convergence and improve model compression. The code is avilable at \url{https://github.com/an-yongqi/systematic-outliers}. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Ming Tang 0001, Jinqiao Wang |
ICLR | 2 |
| 2024 | Fluctuation-Based Adaptive Structured Pruning for Large Language ModelsabstractNetwork Pruning is a promising way to address the huge computing resource demands of the deployment and inference of Large Language Models (LLMs). Retraining-free is important for LLMs' pruning methods. However, almost all of the existing retraining-free pruning approaches for LLMs focus on unstructured pruning, which requires specific hardware support for acceleration. In this paper, we propose a novel retraining-free structured pruning framework for LLMs, named FLAP (FLuctuation-based Adaptive Structured Pruning). It is hardware-friendly by effectively reducing storage and enhancing inference speed. For effective structured pruning of LLMs, we highlight three critical elements that demand the utmost attention: formulating structured importance metrics, adaptively searching the global compressed model, and implementing compensation mechanisms to mitigate performance loss. First, FLAP determines whether the output feature map is easily recoverable when a column of weight is removed, based on the fluctuation pruning metric. Then it standardizes the importance scores to adaptively determine the global compressed model structure. At last, FLAP adds additional bias terms to recover the output feature maps using the baseline values. We thoroughly evaluate our approach on a variety of language benchmarks. Without any retraining, our method significantly outperforms the state-of-the-art methods, including LLM-Pruner and the extension of Wanda in structured pruning. The code is released at https://github.com/CASIA-IVA-Lab/FLAP. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Ming Tang 0001, Jinqiao Wang |
AAAI | 2 |
| 2024 | Knowledge Distillation Dealing with Sample-Wise Long-Tail Problem
Tao Yu 0013, Xu Zhao 0003, Yongqi An, Ming Tang 0001, Jinqiao Wang |
ACCV (10) | 2 |
| 2024 | A Set of Effective Strategies for Optimized Road Damage DetectionabstractIn this paper, we propose an optimized method for road damage detection using a lightweight YOLO model as the baseline. Our approach incorporates a set of effective strategies, including lightweight attention mechanisms, data augmentation, dynamic sampling, weight averaging, and multi-step knowledge distillation. Our method significantly improves inference speed while maintaining high accuracy compared to previous mainstream ensemble-based methods. Our approach achieves notable success in the IEEE Big Data 2024 Optimized Road Damage Detection Challenge (ORDDC’2024), securing second place with an F1 score of 0.7013 and an inference time of 0.0328 s per image. These results demonstrate a strong balance between accuracy and efficiency. Extensive experiments confirm that our method boosts detection accuracy and greatly accelerates inference, making it highly suitable for real-world applications. The source code and trained model are available at https://github.com/YinglongDu/ShiYu_Kunchuan_ORDDC2024. Yinglong Du, Xu Zhao 0003, Bailin He, Bingke Zhu, Shuaihua Zhao, Guibo Zhu, Jinqiao Wang |
IEEE Big Data | 2 |
| 2023 | ZBS: Zero-Shot Background Subtraction via Instance-Level Background Modeling and Foreground SelectionabstractBackground subtraction (BGS) aims to extract all moving objects in the video frames to obtain binary foreground segmentation masks. Deep learning has been widely used in this field. Compared with supervised-based BGS methods, unsupervised methods have better generalization. However, previous unsupervised deep learning BGS algorithms perform poorly in sophisticated scenarios such as shadows or night lights, and they cannot detect objects outside the pre-defined categories. In this work, we propose an unsuper-vised BGS algorithm based on zero-shot object detection called Zero-shot Background Subtraction (ZBS). The proposed method fully utilizes the advantages of zero-shot object detection to build the open-vocabulary instance-level background model. Based on it, the foreground can be effectively extracted by comparing the detection results of new frames with the background model. ZBS performs well for sophisticated scenarios, and it has rich and extensible categories. Furthermore, our method can easily generalize to other tasks, such as abandoned object detection in unseen environments. We experimentally show that ZBS surpasses state-of-the-art unsupervised BGS methods by 4.70% F-Measure on the CDnet 2014 dataset. The code is released at https://github.com/CASIA-IVA-Lab/ZBS. Yongqi An, Xu Zhao 0003, Tao Yu 0013, Haiyun Gu, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
CVPR | 2 |
| 2023 | FreConv: Frequency Branch-and-Integration Convolutional NetworksabstractRecent researches indicate that utilizing the frequency information of input data can enhance the performance of networks. However, the existing popular convolutional structure is not designed specifically for utilizing the frequency information contained in datasets. In this paper, we propose a novel and effective module, named FreConv (frequency branch-and-integration convolution), to replace the vanilla convolution. FreConv adopts a dual-branch architecture to extract and integrate high- and low-frequency information. In the high-frequency branch, a derivative-filter-like architecture is designed to extract the high-frequency information while a light extractor is employed in the low-frequency branch because the low-frequency information is usually redundant. FreConv is able to exploit the frequency information of input data in a more reasonable way to enhance feature representation ability and reduce the memory and computational cost significantly. Without any bells and whistles, experimental results on various tasks demonstrate that FreConv-equipped networks consistently outperform state-of-the-art baselines. Zhaowen Li, Xu Zhao 0003, Peigeng Ding, Zongxing Gao, Ming Tang 0001, Jinqiao Wang |
ICME | 2 |
| 2022 | An Ensemble of One-Stage and Two-Stage Detectors Approach for Road Damage DetectionabstractWith the growth of the city and the increase in the number of cars, the maintenance and management of roads attract more attention. Road damage detection of road images is the basic step of road maintenance. To reduce the cost of labor, it is crucial to make the best use of road damage images from different geographical environments and capturing devices. This paper describes our 1-st place solution used in the Crowd sensing-based Road Damage Detection Challenge of the 2022 IEEE International Conference on Big Data. We use YOLO-series models and Faster RCNN as our one-stage and two-stage baseline models respectively. Our model only needs to be trained directly on the datasets of the overall six countries. Besides, with ensemble learning and test time augmentation, our ensemble model achieves the best results on the learderboard of each single country (India, Japan, United States, and Norway) without fine-tuning. Our ensemble model achieves the F1 scores of 0.7699 and 0.7160 on Overall and Average leaderboard, which significantly outperformed the 2-nd p lace F1 s cores of 0.7432 and 0.6744. The source code and trained model are available at https://github.com/berry-ding/ShiYu_SeaView_GRDDC2022. Wenchao Ding 0004, Xu Zhao 0003, Bingke Zhu, Yinglong Du, Guibo Zhu, Tao Yu 0013, Jinqiao Wang |
IEEE Big Data | 2 |
| 2022 | Transfering Low-Frequency Features for Domain AdaptationabstractPrevious unsupervised domain adaptation methods did not handle the cross-domain problem from the perspective of frequency for computer vision. The images or feature maps of different domains can be decomposed into the low-frequency component and high-frequency component. This paper pro-poses the assumption that low-frequency information is more domain-invariant while the high-frequency information con-tains domain-related information. Hence, we introduce an approach, named low-frequency module (LFM), to extract domain-invariant feature representations. The LFM is constructed with the digital Gaussian low-pass filter. Our method is easy to implement and introduces no extra hyperparame-ter. We design two effective ways to utilize the LFM for domain adaptation, and our method is complementary to other existing methods and formulated as a plug-and-play unit that can be combined with these methods. Experimental results demonstrate that our LFM outperforms state-of-the-art meth-ods for various computer vision tasks, including image clas-sification and object detection. Zhaowen Li, Xu Zhao 0003, Chaoyang Zhao, Ming Tang 0001, Jinqiao Wang |
ICME | 2 |
| 2021 | Attention-Guided Knowledge Distillation for Efficient Single-Stage DetectorabstractKnowledge distillation has been successfully applied in image classification for model acceleration. There are also some works employing this technique to object detection, but they all treat different feature regions equally when performing feature mimic. In this paper, we propose an end-to-end attention-guided knowledge distillation method to train efficient single-stage detectors with much smaller backbones. More specifically, we introduce an attention mechanism to prioritize the transfer of important knowledge by focusing on a sparse set of hard samples, leading to a more thorough distillation process. In addition, the proposed distillation method also provides an easy way to train efficient detectors without tedious ImageNet pre-training procedure. Extensive experiments on PASCAL VOC and CityPersons datasets demonstrate the effectiveness of the proposed approach. We achieve 57.96% and 69.48% mAP on VOC07 with the backbone of 1/8 VGG16 and 1/4 VGG16, greatly outperforming their ImageNet pre-trained counterparts by 11.7% and 7.1% respectively. Tong Wang 0015, Yousong Zhu, Chaoyang Zhao, Xu Zhao 0003, Jinqiao Wang, Ming Tang 0001 |
ICME | 4 |
| 2020 | Task Decoupled Knowledge Distillation For Lightweight Face DetectorsabstractFace detection is a hot topic in computer vision. The face detection methods usually consist of two subtasks, i.e. the classification subtask and the regression subtask, which are trained with different samples. However, current face detection knowledge distillation methods usually couple the two subtasks, and use the same set of samples in the distillation task. In this paper, we propose a task decoupled knowledge distillation method, which decouples the detection distillation task into two subtasks and uses different samples in distilling the features of different subtasks. We firstly propose a feature decoupling method to decouple the classification features and the regression features, without introducing any extra calculations at inference time. Specifically, we generate the corresponding features by adding task-specific convolutions in the teacher network and adding adaption convolutions on the feature maps of the student network. Then we select different samples for different subtasks to imitate. Moreover, we also propose an effective probability distillation method to joint boost the accuracy of the student network. We apply our distillation method on a lightweight face detector, EagleEye. Experimental results show that the proposed method effectively improves the student detector's accuracy by 5.1%, 5.1%, and 2.8% AP in Easy, Medium, Hard subsets respectively. Xiaoqing Liang, Xu Zhao 0003, Chaoyang Zhao, Nanfei Jiang, Ming Tang 0001, Jinqiao Wang |
ACM Multimedia | 2 |
| 2020 | A novel data augmentation scheme for pedestrian detection with attribute preserving GAN
Songyan Liu, Haiyun Guo, Jian-Guo Hu, Xu Zhao 0003, Chaoyang Zhao, Tong Wang 0015, Yousong Zhu, Jinqiao Wang, Ming Tang 0001 |
Neurocomputing | 4 |
| 2020 | Food det: Detecting foods in refrigerator with supervised transformer network
Yousong Zhu, Xu Zhao 0003, Chaoyang Zhao, Jinqiao Wang, Hanqing Lu |
Neurocomputing | 2 |
| 2019 | Elite Loss for scene text detection
Xu Zhao 0003, Chaoyang Zhao, Haiyun Guo, Yousong Zhu, Ming Tang 0001, Jinqiao Wang |
Neurocomputing | 1 |
| 2019 | Attention CoupleNet: Fully Convolutional Attention Coupling Network for Object DetectionabstractThe field of object detection has made great progress in recent years. Most of these improvements are derived from using a more sophisticated convolutional neural network. However, in the case of humans, the attention mechanism, global structure information, and local details of objects all play an important role for detecting an object. In this paper, we propose a novel fully convolutional network, named as Attention CoupleNet, to incorporate the attention-related information and global and local information of objects to improve the detection performance. Specifically, we first design a cascade attention structure to perceive the global scene of the image and generate class-agnostic attention maps. Then the attention maps are encoded into the network to acquire object-aware features. Next, we propose a unique fully convolutional coupling structure to couple global structure and local parts of the object to further formulate a discriminative feature representation. To fully explore the global and local properties, we also design different coupling strategies and normalization ways to make full use of the complementary advantages between the global and local information. Extensive experiments demonstrate the effectiveness of our approach. We achieve state-of-the-art results on all three challenging data sets, i.e., a mAP of 85.7% on VOC07, 84.3% on VOC12, and 35.4% on COCO. Codes are publicly available at https://github.com/tshizys/CoupleNet. Yousong Zhu, Chaoyang Zhao, Haiyun Guo, Jinqiao Wang, Xu Zhao 0003, Hanqing Lu |
IEEE Trans. Image Process. | 5 |
| 2017 | CoupleNet: Coupling Global Structure with Local Parts for Object Detection
Yousong Zhu, Chaoyang Zhao, Jinqiao Wang, Xu Zhao 0003, Yi Wu 0001, Hanqing Lu |
ICCV | 4 |
| 2017 | Joint background reconstruction and foreground segmentation via a two-stage convolutional neural networkabstractForeground segmentation in video sequences is a classic topic in computer vision. Due to the lack of semantic and prior knowledge, it is difficult for existing methods to deal with sophisticated scenes well. Therefore, in this paper, we propose an end-to-end two-stage deep convolutional neural network (CNN) framework for foreground segmentation in video sequences. In the first stage, a convolutional encoder-decoder sub-network is employed to reconstruct the background images and encode rich prior knowledge of background scenes. In the second stage, the reconstructed background and current frame are input into a multi-channel fully-convolutional sub-network (MCFCN) for accurate foreground segmentation. In the two-stage CNN, the reconstruction loss and segmentation loss are jointly optimized. The background images and foreground objects are output simultaneously in an end-to-end way. Moreover, by incorporating the prior semantic knowledge of foreground and background in the pre-training process, our method could restrain the background noise and keep the integrity of foreground objects at the same time. Experiments on CDNet 2014 show that our method outperforms the state-of-the-art by 4.9%. Xu Zhao 0003, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang |
ICME | 1 |