EDBT 2026 Demo / reviewers in the wild / expert
Guanzhong Tian
dblp:240/6842
· DBLP profile ↗
21ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0002-7292-4056ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark
Jiangning Zhang, Chengjie Wang 0001, Xiangtai Li, Guanzhong Tian, Zhucun Xue, Yong Li 0008, Guansong Pang, Dacheng Tao |
Int. J. Comput. Vis. | 4 |
| 2026 | Learning Multi-View Anomaly Detection With Efficient Adaptive SelectionabstractThis study explores the recently proposed and challenging multi-view Anomaly Detection (AD) task. Single-view tasks will encounter blind spots from other perspectives, resulting in inaccuracies in sample-level prediction. Therefore, we introduce theMulti-ViewAnomalyDetection (MVAD) approach, which learns and integrates features from multi-views. Specifically, we propose aMulti-ViewAdaptiveSelection (MVAS) algorithm for feature learning and fusion across multiple views. The feature maps are divided into neighbourhood attention windows to calculate a semantic correlation matrix between single-view windows and all other views, which is an attention mechanism conducted for each single-view window and the top-$k$most correlated multi-view windows. Adjusting the window sizes and top-$k$can minimise the complexity to$O((hw)^\frac{4}{3})$. Extensive experiments on the Real-IAD dataset under the multi-class setting validate the effectiveness of our approach, achieving state-of-the-art performance with an average improvement of+2.5$\uparrow$across10 metricsat the sample/image/pixel levels, using only18Mparameters and requiring fewer FLOPs and training time. The codes are available athttps://github.com/lewandofskee/MVAD. Haoyang He, Jiangning Zhang, Guanzhong Tian, Chengjie Wang 0001, Lei Xie 0007 |
IEEE Trans. Multim. | 3 |
| 2025 | DISCOVERSE: Efficient Robot Simulation in Complex High-Fidelity EnvironmentsabstractWe present Discoverse, the first unified, modular, open-source 3DGS-based simulation framework for Real2Sim2Real robot learning. It features a holistic Real2Sim pipeline that synthesizes hyper-realistic geometry and appearance of complex real-world scenarios, paving the way for analyzing and bridging the Sim2Real gap. Powered by Gaussian Splatting and MuJoCo, Discoverse enables massively parallel simulation of multiple sensor modalities and accurate physics, with inclusive supports for existing 3D assets, robot models, and ROS plugins, empowering large-scale robot learning and complex robotic benchmarks. Through extensive experiments on imitation learning, Dis coverse demonstrates state-of-the-art zero-shot Sim2Real transfer performance compared to existing simulators. For code and demos: https://air-discoverse.github.io/. Yufei Jia, Junzhe Wu, Yupei Zeng, Haonan Lin, Haizhou Ge, Weibin Gu, Kairui Ding, Zike Yan, Yunjie Cheng, Chuxuan Li, Wei Sui, Guanzhong Tian, Ruqi Huang, Guyue Zhou |
IROS | 18 |
| 2025 | MCMC: Multi-Constrained Model Compression via One-Stage Envelope Reinforcement LearningabstractModel compression methods are being developed to bridge the gap between the massive scale of neural networks and the limited hardware resources on edge devices. Since most real-world applications deployed on resource-limited hardware platforms typically have multiple hardware constraints simultaneously, most existing model compression approaches that only consider optimizing one single hardware objective are ineffective. In this article, we propose an automated pruning method called multi-constrained model compression (MCMC) that allows for the optimization of multiple hardware targets, such as latency, floating point operations (FLOPs), and memory usage, while minimizing the impact on accuracy. Specifically, we propose an improved multi-objective reinforcement learning (MORL) algorithm, the one-stage envelope deep deterministic policy gradient (DDPG) algorithm, to determine the pruning strategy for neural networks. Our improved one-stage envelope DDPG algorithm reduces exploration time and offers greater flexibility in adjusting target priorities, enhancing its suitability for pruning tasks. For instance, on the visual geometry group (VGG)-16 network, our method achieved an 80% reduction in FLOPs, a reduction in memory usage, and a acceleration, with an accuracy improvement of 0.09% compared with the baseline. For larger datasets, such as ImageNet, we reduced FLOPs by 50% for MobileNet-V1, resulting in a faster speed and memory compression, while maintaining the same accuracy. When applied to edge devices, such as JETSON XAVIER NX, our method resulted in a 71% reduction in FLOPs for MobileNet-V1, leading to a faster speed, memory compression, and an accuracy improvement. Siqi Li 0009, Jun Chen 0023, Shanqi Liu, Chengrui Zhu, Guanzhong Tian, Yong Liu 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Adding Before Pruning: Sparse Filter Fusion for Deep Convolutional Neural Networks via Auxiliary AttentionabstractFilter pruning is a significant feature selection technique to shrink the existing feature fusion schemes (especially on convolution calculation and model size), which helps to develop more efficient feature fusion models while maintaining state-of-the-art performance. In addition, it reduces the storage and computation requirements of deep neural networks (DNNs) and accelerates the inference process dramatically. Existing methods mainly rely on manual constraints such as normalization to select the filters. A typical pipeline comprises two stages: first pruning the original neural network and then fine-tuning the pruned model. However, choosing a manual criterion can be somehow tricky and stochastic. Moreover, directly regularizing and modifying filters in the pipeline suffer from being sensitive to the choice of hyperparameters, thus making the pruning procedure less robust. To address these challenges, we propose to handle the filter pruning issue through one stage: using an attention-based architecture that adaptively fuses the filter selection with filter learning in a unified network. Specifically, we present a pruning method named adding before pruning (ABP) to make the model focus on the filters of higher significance by training instead of man-made criteria such as norm, rank, etc. First, we add an auxiliary attention layer into the original model and set the significance scores in this layer to be binary. Furthermore, to propagate the gradients in the auxiliary attention layer, we design a specific gradient estimator and prove its effectiveness for convergence in the graph flow through mathematical derivation. In the end, to relieve the dependence on the complicated prior knowledge for designing the thresholding criterion, we simultaneously prune and train the filters to automatically eliminate network redundancy with recoverability. Extensive experimental results on the two typical image classification benchmarks, CIFAR-10 and ILSVRC-2012, illustrate that the proposed approach performs favorably against previous state-of-the-art filter pruning algorithms. Guanzhong Tian, Yiran Sun, Yuang Liu, Xianfang Zeng, Mengmeng Wang 0005, Yong Liu 0007, Jiangning Zhang, Jun Chen 0023 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Locate N' Rotate: Two-Stage Openable Part Detection with Foundation Model Priors
Siqi Li 0009, Xiaoxue Chen, Haoyu Cheng, Guyue Zhou, Hao Zhao 0002, Guanzhong Tian |
ACCV (7) | 6 |
| 2024 | MambaAD: Exploring State Space Models for Multi-class Unsupervised Anomaly DetectionabstractRecent advancements in anomaly detection have seen the efficacy of CNN- and transformer-based approaches. However, CNNs struggle with long-range dependencies, while transformers are burdened by quadratic computational complexity. Mamba-based models, with their superior long-range modeling and linear efficiency, have garnered substantial attention. This study pioneers the application of Mamba to multi-class unsupervised anomaly detection, presenting MambaAD, which consists of a pre-trained encoder and a Mamba decoder featuring (Locality-Enhanced State Space) LSS modules at multi-scales. The proposed LSS module, integrating parallel cascaded (Hybrid State Space) HSS blocks and multi-kernel convolutions operations, effectively captures both long-range and local information. The HSS block, utilizing (Hybrid Scanning) HS encoders, encodes feature maps into five scanning methods and eight directions, thereby strengthening global connections through the (State Space Model) SSM. The use of Hilbert scanning and eight directions significantly improves feature sequence modeling. Comprehensive experiments on six diverse anomaly detection datasets and seven metrics demonstrate state-of-the-art performance, substantiating the method's effectiveness. The code and models are available at https://lewandofskee.github.io/projects/MambaAD. Haoyang He, Yuhu Bai, Jiangning Zhang, Qingdong He, Zhenye Gan, Chengjie Wang 0001, Xiangtai Li, Guanzhong Tian, Lei Xie 0007 |
NeurIPS | 9 |
| 2024 | Dual-path Frequency Discriminators for few-shot anomaly detection
Yuhu Bai, Jiangning Zhang, Zhaofeng Chen, Yunkang Cao, Guanzhong Tian |
Knowl. Based Syst. | 6 |
| 2024 | Learning Multi-Agent Cooperation via Considering Actions of TeammatesabstractRecently value-based centralized training with decentralized execution (CTDE) multi-agent reinforcement learning (MARL) methods have achieved excellent performance in cooperative tasks. However, the most representative method among these methods, Q-network MIXing (QMIX), restricts the joint action Q values to be a monotonic mixing of each agent's utilities. Furthermore, current methods cannot generalize to unseen environments or different agent configurations, which is known as ad hoc team play situation. In this work, we propose a novel Q values decomposition that considers both the return of an agent acting on its own and cooperating with other observable agents to address the nonmonotonic problem. Based on the decomposition, we propose a greedy action searching method that can improve exploration and is not affected by changes in observable agents or changes in the order of agents' actions. In this way, our method can adapt to ad hoc team play situation. Furthermore, we utilize an auxiliary loss related to environmental cognition consistency and a modified prioritized experience replay (PER) buffer to assist training. Our extensive experimental results show that our method achieves significant performance improvements in both challenging monotonic and nonmonotonic domains, and can handle the ad hoc team play situation perfectly. Shanqi Liu, Wenzhou Chen, Guanzhong Tian, Jun Chen 0023, Yong Liu 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | MixTeacher: Mining Promising Labels with Mixed Scale Teacher for Semi-Supervised Object DetectionabstractScale variation across object instances remains a key challenge in object detection task. Despite the remarkable progress made by modern detection models, this challenge is particularly evident in the semi-supervised case. While existing semi-supervised object detection methods rely on strict conditions to filter high-quality pseudo labels from network predictions, we observe that objects with extreme scale tend to have low confidence, resulting in a lack of positive supervision for these objects. In this paper, we propose a novel framework that addresses the scale variation problem by introducing a mixed scale teacher to improve pseudo label generation and scale-invariant learning. Additionally, we propose mining pseudo labels using score promotion of predictions across scales, which benefits from better predictions from mixed scale features. Our extensive experiments on MS COCO and PASCAL VOC benchmarks under various semi-supervised settings demonstrate that our method achieves new state-of-the-art performance. The code and models are available at https://github.com/lliuz/MixTeacher. Liang Liu 0007, Boshen Zhang, Jiangning Zhang, Wuhao Zhang, Zhenye Gan, Guanzhong Tian, Wenbing Zhu, Yabiao Wang, Chengjie Wang 0001 |
CVPR | 6 |
| 2023 | Lightweight Cryptography Implementation for Internet of Things Network on FPGAabstractWith the development of modern communication technology, traditional Internet of Things systems could not provide sufficient support for large data flow, hardware resource, power consumption, and security during transmission. Especially when it comes to the Industrial Internet of Things (IIoT), where the limitations of hardware resource and power consumption are more strict; the requirements of network security and hardware security are much higher. This paper implements a User Datagram Protocol (UDP) communication system, which is encrypted with one of the Lightweight Cryptography, Xoodyak. We aim to design a lightweight encrypted communication platform. Compared with a public key encryption system, our implementation costs much less hardware resource and power consumption, while providing excellent side-channel attack (SCA) protection. Besides, when there is a large burst length of data flow, two asynchronous FIFO can restore those data respectively, so our design can maintain throughput and data integrity in extreme cases. Based on the above properties, our design is relatively ideal in IIoT encryption scenarios. Guanzhong Tian, Longhua Ma, Zhishan Li, Shanqi Liu |
SMC | 2 |
| 2023 | Data-free quantization via mixed-precision compensation without fine-tuning
Jun Chen 0023, Shipeng Bai, Tianxin Huang, Mengmeng Wang 0005, Guanzhong Tian, Yong Liu 0007 |
Pattern Recognit. | 5 |
| 2023 | Towards accurate dense pedestrian detection via occlusion-prediction aware label assignment and hierarchical-NMS
Haoyang He, Zhishan Li, Guanzhong Tian, Lei Xie 0007, Shan Lu 0009 |
Pattern Recognit. Lett. | 3 |
| 2022 | CICC: Channel Pruning via the Concentration of Information and Contributions of Channels
Zhishan Li, Yingqing Yang, Lei Xie 0007, Yong Liu 0007, Longhua Ma, Shanqi Liu, Guanzhong Tian |
BMVC | 8 |
| 2022 | Designing One Unified Framework for High-Fidelity Face Reenactment and Swapping
Chao Xu 0023, Jiangning Zhang, Guanzhong Tian, Xianfang Zeng, Ying Tai, Yabiao Wang, Chengjie Wang 0001, Yong Liu 0007 |
ECCV (15) | 4 |
| 2022 | Multilevel Spatial-Temporal Feature Aggregation for Video Object DetectionabstractVideo object detection (VOD) focuses on detecting objects for each frame in a video, which is a challenging task due to appearance deterioration in certain video frames. Recent works usually distill crucial information from multiple support frames to improve the reference features, but they only perform at frame level or proposal level that cannot integrate spatial-temporal features sufficiently. To deal with this challenge, we treat VOD as a spatial-temporal hierarchical features interacting process and introduce a Multi-level Spatial-Temporal (MST) feature aggregation framework to fully exploit frame-level, proposal-level, and instance-level information in a unified framework. Specifically, MST first measures context similarity in pixel space to enhance all frame-level features rather than only update reference features. The proposal-level feature aggregation then models object relation to augment reference object proposals. Furthermore, to filter out irrelevant information from other classes and backgrounds, we introduce an instance ID constraint to boost instance-level features by leveraging support object proposal features that belong to the same object. Besides, we propose a Deformable Feature Alignment (DAlign) module before MST to achieve a more accurate pixel-level spatial alignment for better feature aggregation. Extensive experiments are conducted on ImageNet VID and UAVDT datasets that demonstrate the superiority of our method over state-of-the-art (SOTA) methods. Our method achieves 83.3% and 62.1% with ResNet-101 on two datasets, outperforming SOTA MEGA by 0.4% and 2.7%. Chao Xu 0023, Jiangning Zhang, Mengmeng Wang 0005, Guanzhong Tian, Yong Liu 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Delving Deeper Into Mask Utilization in Video Object SegmentationabstractThis paper focuses on the mask utilization of video object segmentation (VOS). The mask here mains the reference masks in the memory bank, i.e., several chosen high-quality predicted masks, which are usually used with the reference frames together. The reference masks depict the edge and contour features of the target object and indicate the boundary of the target against the background, while the reference frames contain the raw RGB information of the whole image. It is obvious that the reference masks could play a significant role in the VOS, but this is not well explored yet. To tackle this, we propose to investigate the mask advantages of both the encoder and the matcher. For the encoder, we provide a unified codebase to integrate and compare eight different mask-fused encoders. Half of them are inherited or summarized from existing methods, and the other half are devised by ourselves. We find the best configuration from our design and give valuable observations from the comparison. Then, we propose a new mask-enhanced matcher to reduce the background distraction and enhance the locality of the matching process. Combining the mask-fused encoder, mask-enhanced matcher and a standard decoder, we formulate a new architecture named MaskVOS, which sufficiently exploits the mask benefits for VOS. Qualitative and quantitative results demonstrate the effectiveness of our method. We hope our exploration could raise the attention of mask utilization in VOS. Mengmeng Wang 0005, Jianbiao Mei, Lina Liu 0010, Guanzhong Tian, Yong Liu 0007, Zaisheng Pan |
IEEE Trans. Image Process. | 4 |
| 2021 | Unpaired salient object translation via spatial attention prior
Xianfang Zeng, Yusu Pan, Mengmeng Wang 0005, Guanzhong Tian, Yong Liu 0007 |
Neurocomputing | 5 |
| 2021 | Pruning by Training: A Novel Deep Neural Network Compression Framework for Image ProcessingabstractFilter pruning for a pre-trained convolutional neural network is most normally performed through human-made constraints or criteria such as norms, ranks, etc. Typically, the pruning pipeline comprises two-stage: first learn a sparse structure from the original model, then optimize the weights in the new prune model. One disadvantage of using human-made criteria to prune filters is that the design and selection of threshold criteria depend on complicated prior knowledge. Besides, the pruning process is less robust due to the impact of directly regularizing on filters. To address the problems mentioned, we propose an effective one-stage pruning framework: introducing a trainable collaborative layer to jointly prune and learn neural networks in one go. In our framework, we first add a binary collaborative layer for each original filter. Then, a new type of gradient estimator - asymptotic gradient estimator is first introduced to pass the gradient in the binary collaborative layer. Finally, we simultaneously learn the sparse structure and optimize the weights from the original model in the training process. Our evaluation results on typical benchmarks, CIFAR and ImageNet, demonstrate very promising results against other state-of-the-art filter pruning methods. Guanzhong Tian, Jun Chen 0023, Xianfang Zeng, Yong Liu 0007 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Deep Superpixel Convolutional Network for Image RecognitionabstractDue to the high representational efficiency, superpixel largely reduces the number of image primitives for subsequent processing. However, superpixel is scarcely utilized in recent methods since its irregular shape is intractable for standard convolutional layer. In this paper, we propose an end-to-end trainable superpixel convolutional network, named SPNet, to learn high-level representation on image superpixel primitives. We start by treating irregular superpixel lattices as a 2D point cloud, where the low-level features inside one superpixel are aggregated to one feature vector. We replace the standard convolutional layer with the PointConv layer to handle the irregular and unordered point cloud. Besides, we propose grid based downsampling strategies to output uniform 2D sampling result. The resulting network largely utilizes the efficiency of superpixel and provides a novel view for image recognition task. Experiments on image recognition task show promising results compared with prominent image classification methods. The visualization of class activation mapping shows great accuracy at object localization and boundary segmentation. Xianfang Zeng, Guanzhong Tian, Fuxin Li, Yong Liu 0007 |
IEEE Signal Process. Lett. | 3 |
| 2019 | ObjectFusion: An object detection and segmentation framework with RGB-D SLAM and convolutional neural networks
Guanzhong Tian, Liang Liu 0007, JongHyok Ri, Yong Liu 0007, Yiran Sun |
Neurocomputing | 1 |