Qizhen Lan

dblp:312/5280 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0003-0496-5240ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
abstract
Tingyu Wu, Zhisheng Chen, Ziyan Weng, Shuhe Wang, Shuo Zhang, Sen Hu, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tingyu Wu, Zhisheng Chen 0003, Ziyan Weng, Shuhe Wang, Sen Hu 0005, Silin Wu, Qizhen Lan, Huacan Wang, Ronghao Chen
ACL (1)8
2026 Boundary- and Saliency-Aware Knowledge Distillation for Semantic Segmentation in Urban Driving Scenes
Xinyu Chu, Zhicheng Ding, Qizhen Lan, Qing Tian 0003
IV3
2026 Visual Detector Compression via Location-Aware Discriminant Analysis
abstract
Deep neural networks are powerful, yet their high complexity greatly limits their potential to be deployed on billions of resource-constrained edge devices. Pruning is a crucial network compression technique, yet most existing methods focus on classification models, with limited attention to detection. Even among those addressing detection, there is a lack of utilization of essential localization information. Also, many pruning methods passively rely on pre-trained models, in which useful and useless components are intertwined, making it difficult to remove the latter without harming the former at the neuron/filter level. To address the above issues, in this paper, we propose a proactive detection-discriminants-based network compression approach for deep visual detectors, which alternates between two steps: (1) maximizing and compressing detection-related discriminants and aligning them with a subset of neurons/filters immediately before the detection head, and (2) tracing the detection-related discriminating power across the layers and discarding features of lower importance. Object location information is exploited in both steps. Extensive experiments, employing four advanced detection models and four state-of-the-art competing methods on the KITTI and COCO datasets, highlight the superiority of our approach. Remarkably, our compressed models can even beat the original base models with a substantial reduction in complexity.
Qizhen Lan, Jung Im Choi, Qing Tian 0003
WACV1
2026 CLoCKDistill: Consistent Location and Context aware Knowledge Distillation for DETRs
abstract
Object detection has advanced significantly with Detection Transformers (DETRs). However, these models are computationally demanding, posing challenges for deployment in resource-constrained environments (e.g., self-driving cars). Knowledge distillation (KD) is an effective compression method widely applied to CNN detectors, but its application to DETR models has been limited. Most KD methods for DETRs fail to distill transformer-specific global context. Also, they blindly believe in the teacher model, which can sometimes be misleading. To bridge the gaps, this paper proposes Consistent Location-and-Context-aware Knowledge Distillation (CLoCKDistill) for DETR detectors, which includes both feature distillation and logit distillation components. For feature distillation, instead of distilling backbone features like many existing KD methods, we distill the transformer encoder output (i.e., memory) that contains valuable global context and long-range dependencies. Also, we enrich this memory with object location details during feature distillation so that the student model can prioritize relevant regions. To facilitate logit distillation, we create target-aware queries based on the ground truth, allowing both the student and teacher decoders to attend to consistent and accurate parts of encoder memory. Experiments on KITTI and COCO show our CLoCKDistill method’s efficacy across various DETRs. Our method boosts student detector performance by 2.2% to 6.4%. Our code is available at: https://github.com/lanqz7766/CLoCKDistill.
Qizhen Lan, Qing Tian 0003
WACV1
2026 Difficulty-Aware Agentic Orchestration for Query-Specific Multi-Agent Workflows
abstract
Large Language Model (LLM)-based agentic systems have shown strong capabilities across various tasks. However, existing multi-agent frameworks often rely on static or task-level workflows, which either over-process simple queries or underperform on complex ones, while also neglecting the efficiency-performance trade-offs across heterogeneous LLMs. To address these limitations, we propose Difficulty-Aware Agentic Orchestration (DAAO), which can dynamically generate query-specific multi-agent workflows guided by predicted query difficulty. DAAO comprises three interdependent modules: a variational autoencoder (VAE) for difficulty estimation, a modular operator allocator, and a cost- and performance-aware LLM router. A self-adjusting policy updates difficulty estimates based on workflow success, enabling simpler workflows for easy queries and more complex strategies for harder ones. Experiments on six benchmarks demonstrate that DAAO surpasses prior multi-agent systems in both accuracy and inference efficiency, validating its effectiveness for adaptive, difficulty-aware reasoning. Our code is open-sourced at https://github.com/AutoAgents-ai/DAAO
Jinwei Su, Qizhen Lan, Yinghui Xia, Lifan Sun, Weiyou Tian, Tianyu Shi 0003, Lewei He
WWW2
2026 Boosting deep detector efficiency and robustness through detection discriminant reorganization and compression
Jung Im Choi, Qizhen Lan, Qing Tian 0003
Neural Networks2
2025 ACAM-KD: Adaptive and Cooperative Attention Masking for Knowledge Distillation
Qizhen Lan, Qing Tian 0003
ICCV1
2025 Target-Driven and Student-Centered Knowledge Distillation for Traffic Object Tracking
abstract
Visual Object Tracking is crucial for autonomous driving, enabling real-time monitoring of dynamic environments. While Transformer-based trackers achieve state-of-the-art performance by modeling long-range dependencies, their high computational cost limits deployment in real-world autonomous systems. To address this, we propose Target-Driven and Student-Centered Knowledge Distillation (TDSC-KD), a novel framework designed to improve the efficiency of Transformer-based trackers while maintaining accuracy. Our framework consists of (1) target-driven distillation, which leverages a ground-truth query to guide knowledge transfer toward relevant and consistent regions, filtering out background noise, and (2) student-centered distillation, which employs a mask-and-reconstruct mechanism to encourage more active student learning and reduce over-reliance on the teacher. Experiments on the LaSOT-Traffic dataset demonstrate our TDSC-KD's efficacy, narrowing the gap between the strong performance of Transformer trackers and the strict efficiency constraints of real-world deployment.
Zhicheng Ding, Qizhen Lan, Qing Tian 0003
IV2
2025 ETT-CKGE: Efficient Task-Driven Tokens for Continual Knowledge Graph Embedding
Lijing Zhu, Qizhen Lan, Qing Tian 0003, Xi Xiao 0003, Tiehang Duan, Cui Tao, Shuteng Niu
ECML/PKDD (6)2
2025 Improving Deep Detector Robustness via Detection-Related Discriminant Maximization and Reorganization
abstract
Deep visual detectors are known to be vulnerable to adversarial attacks, raising concerns about their real-world applications (e.g., self-driving perception). We argue that this vulnerability arises from the spurious dependency of final detections on irrelevant/loophole latent dimensions. The greater the number of such dimensions, the higher the likelihood of the detector being compromised by adversarial attacks, making it more susceptible to input perturbations. To enhance detection robustness, we propose Detection-related Discriminant Maximization and Reorga-nization (DDMR), condensing the detection utility to a compressed number of relevant dimensions while deactivating the influence of irrelevant ones. This approach also alleviates the misalignment issue between the two task domains in visual detection and, consequently, their gradients. This enables the generation of more potent adversarial attacks and defenses for visual detectors within the adversarial training framework. Extensive experiments conducted with four cutting-edge visual detectors on the KITTI and COCO datasets showcase the efficacy of the proposed approach in improving the adversarial robustness of deep visual detectors against both white-box and black-box attacks. For example, on the KITTI dataset, our method demonstrates an increase in robustness of up to 12.4% and 28.0% without and with adversarial training, respectively.
Jung Im Choi, Qizhen Lan, Qing Tian 0003
WACV2
2024 Multi-dimension Transformer with Attention-based Filtering for Medical Image Segmentation
abstract
The accurate segmentation of medical images is crucial for diagnosing and treating diseases. Recent studies demonstrate that vision transformer-based methods have significantly improved performance in medical image segmentation, primarily due to their superior ability to establish global relationships among features and adaptability to various inputs. However, these methods struggle with the low signal-to-noise ratio inherent to medical images. Additionally, the effective utilization of channel and spatial information, which are essential for medical image segmentation, is limited by the representation capacity of self-attention. To address these challenges, we propose a Multi-dimension Transformer with Attention-based Filtering (MDT-AF), which redesigns the patch embedding and self-attention mechanism for medical image segmentation. MDT-AF incorporates an attention-based feature filtering mechanism into the patch embedding blocks and employs a coarse-to-fine process to mitigate the impact of a low signal-to-noise ratio. To better capture complex structures in medical images, MDT-AF extends self-attention and introduces an interaction mechanism to build and enhance feature relationships between dimensions, which can achieve richer feature representations across the spatial and channel dimensions. Experimental results on three public medical image segmentation benchmarks show that MDT-AF achieves state-of-the-art (SOTA) performance.
Xi Xiao 0003, Qizhen Lan, Xuanyao Huang, Qing Tian 0003, Swalpa Kumar Roy, Tianyang Wang 0004
ICTAI4
2024 Gradient-Guided Knowledge Distillation for Object Detectors
abstract
Deep learning models have demonstrated remarkable success in object detection, yet their complexity and computational intensity pose a barrier to deploying them in real-world applications (e.g., self-driving perception). Knowledge Distillation (KD) is an effective way to derive efficient models. However, only a small number of KD methods tackle object detection. Also, most of them focus on mimicking the plain features of the teacher model but rarely consider how the features contribute to the final detection. In this paper, we propose a novel approach for knowledge distillation in object detection, named Gradient-guided Knowledge Distillation (GKD). Our GKD uses gradient information to identify and assign more weights to features that significantly impact the detection loss, allowing the student to learn the most relevant features from the teacher. Furthermore, we present bounding-box-aware multi-grained feature imitation (BMFI) to further improve the KD performance. Experiments on the KITTI and COCO-Traffic datasets demonstrate our method’s efficacy in knowledge distillation for object detection. On one-stage and two-stage detectors, our GKD-BMFI leads to an average of 5.1% and 3.8% mAP improvement, respectively, beating various state-of-the-art KD methods. Our codes are available at: https://github.com/lanqz7766/GKD.
Qizhen Lan, Qing Tian 0003
WACV1
2022 Adaptive Instance Distillation for Object Detection in Autonomous Driving
abstract
In recent years, knowledge distillation (KD) has been widely used to derive efficient models. Through imitating a large teacher model, a lightweight student model can achieve comparable performance with more efficiency. However, most existing knowledge distillation methods are focused on classification tasks. Only a limited number of studies have applied knowledge distillation to object detection, especially in time-sensitive autonomous driving scenarios. In this paper, we propose Adaptive Instance Distillation (AID) to selectively impart teacher’s knowledge to the student to improve the performance of knowledge distillation. Unlike previous KD methods that treat all instances equally, our AID can attentively adjust the distillation weights of instances based on the teacher model’s prediction loss. We verified the effectiveness of our AID method through experiments on the KITTI and the COCO traffic datasets. The results show that our method improves the performance of state-of-the-art attention-guided and non-local distillation methods and achieves better distillation results on both single-stage and two-stage detectors. Compared to the baseline, our AID led to an average of 2.7% and 2.1% mAP increases for single-stage and two-stage detectors, respectively. Furthermore, our AID is also shown to be useful for self-distillation to improve the teacher model’s performance.
Qizhen Lan, Qing Tian 0003
ICPR1