Qi Li 0038

dblp:181/2688-38 · DBLP profile ↗
← Back
14ranked-venue papers
1as first author
14since 2021 · last 2026
0009-0000-7890-9220ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Seeing in Double: Dual-Granularity BEV Segmentation via Mamba-Driven Alignment and Polar-Decoupled Experts
abstract
Bird's Eye View (BEV) representation has become pivotal for autonomous driving, yet existing polar coordinate-based approaches face two critical limitations: (1) distant semantic misprojection caused by radial resolution decay, and (2) region-specific geometric distortions from non-uniform polar discretization. To address these issues, we propose a novel framework addressing these challenges through three key innovations. First, we present a bilateral heterogeneous network constructs multi-granularity BEV spaces, efficiently exploiting dual-resolution visual information for distant detail preservation. Second, we employ an align-fusion strategy for multi-granularity feature aggregation. Specifically, the Mamba-Based Cross-Resolution Alignment module establishes semantic consistency for perspective features through shared state-space optimization. In the later stage, the Adaptive BEV Space Selector dynamically aggregates multi-granularity BEV features. Third, we introduce a Mixture of Radial-Angular Decoupled Experts, which employs polar-aware expert routing to disentangle radial compression and angular shear distortions through specialized geometric refinement. Comprehensive experiments on nuScenes and Lyft L5 demonstrate the state-of-the-art performance of our model across various resolution settings, visibility filtering, and perception ranges.
Jingze Su, Qi Li 0038, Wenjie Yang 0005, Yuanlong Yu 0001, Wenxi Liu
AAAI4
2026 Adapting SAM to nuclei instance segmentation and classification via Cooperative Fine-Grained Refinement
Jingze Su, Tianle Zhu, Qi Li 0038, Tong Tong 0001, Wenxi Liu
Medical Image Anal.5
2025 Keep the Balance: A Parameter-Efficient Symmetrical Framework for RGB+X Semantic Segmentation
abstract
Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations of attention mechanism for modality fusion, which inadvertently introduces noise across modalities. To mitigate this noise, we propose a dynamic sparse cross-modality fusion module to facilitate effective and efficient cross-modality fusion. To further strengthen the above two modules, we propose a training strategy that leverages accurately predicted dual-modality results to self-teach the single-modality outcomes. In comprehensive experiments, we demonstrate that our method outperforms previous state-of-the-art approaches across six multimodal segmentation scenarios with minimal computation cost.
Jingze Su, Qi Li 0038, Wenjie Yang 0005, Tiesong Zhao, Shengfeng He, Wenxi Liu
CVPR3
2025 Regional crowd flow estimation from aerial view
Huibin Wei, Qi Li 0038, Xindai Lin, Shengfeng He, Antoni B. Chan, Wenxi Liu
Neural Networks2
2025 Category-Contrastive Fine-Grained Crowd Counting and Beyond
abstract
Crowd counting has drawn increasing attention across various fields. However, existing crowd counting tasks primarily focus on estimating the overall population, ignoring the behavioral and semantic information of different social groups within the crowd. In this paper, we aim to address a newly proposed research problem, namely fine-grained crowd counting, which involves identifying different categories of individuals and accurately counting them in static images. In order to fully leverage the categorical information in static crowd images, we propose a two-tier salient feature propagation module designed to sequentially extract semantic information from both the crowd and its surrounding environment. Additionally, we introduce a category difference loss to refine the feature representation by highlighting the differences between various crowd categories. Moreover, our proposed framework can adapt to a novel problem setup called few-example fine-grained crowd counting. This setup, unlike the original fine-grained crowd counting, requires only a few exemplar point annotations instead of dense annotations from predefined categories, making it applicable in a wider range of scenarios. The baseline model for this task can be established by substituting the loss function in our proposed model with a novel hybrid loss function that integrates point-oriented cross-entropy loss and category contrastive loss. Through comprehensive experiments, we present results in both the formulation and application of fine-grained crowd counting.
Meijing Zhang, Mengxue Chen, Qi Li 0038, Yanchen Chen, Xiaolian Li, Shengfeng He, Wenxi Liu
IEEE Trans. Multim.3
2024 Efficient Semantic Segmentation for Compressed Video
abstract
Robots, constrained by limited onboard computing resources, often encounter situations wherein high-resolution and high-bit-rate videos captured by their cameras necessitate compression before further analysis. In this paper, we propose a novel video semantic segmentation paradigm for compressed video. Specifically, our framework draws the inspiration from the principle of Wavelet Transform, and thus we design the network structure, WTDecomNet, approximating the decomposition of high-resolution image into its low-resolution counterpart and axial details. The aim is to well preserve the image content through decomposition and maintain model efficiency by obtaining semantics from low-resolution image. To facilitate this purpose, we propose an efficient axial subband approximation module for extracting axial details and a lightweight temporal alignment module for associating keyframes and non-keyframes of compressed video. Through comprehensive experiments, we show that our model can achieve the state-of-the-art performance on public benchmarks. Especially on CamVid, comparing to baseline, our proposed model reduces the computational overhead by ∼70% while improving mIoU by ∼4%.
Qi Li 0038, Jia Pan 0001, Wenxi Liu
ICRA2
2024 Ultra-High Resolution Image Segmentation via Locality-Aware Context Fusion and Alternating Local Enhancement
Wenxi Liu, Qi Li 0038, Xindai Lin, Weixiang Yang, Shengfeng He, Yuanlong Yu 0001
Int. J. Comput. Vis.2
2024 Pose-aware video action segmentation
Meijing Zhang, Chenyang Liao, Qi Li 0038, Wenxi Liu
Neural Comput. Appl.3
2024 Monocular BEV Perception of Road Scenes via Front-to-Top View Projection
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to expensive sensors and time-consuming computation. Camera-based methods usually need to perform road segmentation and view transformation separately, which often causes distortion and missing content. To push the limits of the technology, we present a novel framework that reconstructs a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. We propose a front-to-top view projection (FTVP) module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. In addition, we apply multi-scale FTVP modules to propagate the rich spatial information of low-level features to mitigate spatial deviation of the predicted object location. Experiments on public benchmarks show that our method achieves various tasks on road layout estimation, vehicle occupancy estimation, and multi-class semantic estimation, at a performance level comparable to the state-of-the-arts, while maintaining superior efficiency.
Wenxi Liu, Qi Li 0038, Weixiang Yang, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Prototype learning based generic multiple object tracking via point-to-box supervision
Wenxi Liu, Qi Li 0038, Yinhua She, Yuanlong Yu 0001, Jia Pan 0001, Jason Gu
Pattern Recognit.3
2024 Attentive and Contrastive Image Manipulation Localization With Boundary Guidance
abstract
In recent years, the rapid advancement of image generation techniques has resulted in the widespread abuse of manipulated images, leading to a crisis of trust and affecting social equity. Thus, the goal of our work is to detect and localize tampered regions in images. Many deep learning based approaches have been proposed to address this problem, but they can hardly handle the tampered regions that are manually fine-tuned to blend into image background. By observing that the boundaries of tempered regions are critical to separating tampered and non-tampered parts, we present a novel boundary-guided approach to image manipulation detection, which introduces an inherent bias towards exploiting the boundary information of tampered regions. Our model follows an encoder-decoder architecture, with multi-scale localization mask prediction, and is guided to utilize the prior boundary knowledge through an attention mechanism and contrastive learning. In particular, our model is unique in that 1) we propose a boundary-aware attention module in the network decoder, which predicts the boundary of tampered regions and thus uses it as crucial contextual cues to facilitate the localization; and 2) we propose a multi-scale contrastive learning scheme with a novel boundary-guided sampling strategy, leading to more discriminative localization features. Our state-of-art performance on several public benchmarks demonstrates the superiority of our model over prior works.
Wenxi Liu, Hao Zhang 0170, Xinyang Lin, Qi Li 0038, Xiaoxiang Liu, Ying Cao 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Learning Nighttime Semantic Segmentation the Hard Way
abstract
Nighttime semantic segmentation is an important but challenging research problem for autonomous driving. The major challenges lie in the small objects or regions from the under-/over-exposed areas or suffer from motion blur caused by the camera deployed on moving vehicles. To resolve this, we propose a novel hard-class-aware module that bridges the main network for full-class segmentation and the hard-class network for segmenting aforementioned hard-class objects. In specific, it exploits the shared focus of hard-class objects from the dual-stream network, enabling the contextual information flow to guide the model to concentrate on the pixels that are hard to classify. In the end, the estimated hard-class segmentation results will be utilized to infer the final results via an adaptive probabilistic fusion refinement scheme. Moreover, to overcome over-smoothing and noise caused by extreme exposures, our model is modulated by a carefully crafted pretext task of constructing an exposure-aware semantic gradient map, which guides the model to faithfully perceive the structural and semantic information of hard-class objects while mitigating the negative impact of noises and uneven exposures. In experiments, we demonstrate that our unique network design leads to superior segmentation performance over existing methods, featuring the strong ability of perceiving hard-class objects under adverse conditions.
Wenxi Liu, Qi Li 0038, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Projecting Your View Attentively: Monocular Road Scene Layout Estimation via Cross-View Transformation
abstract
HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird’s-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use of their correlation to strengthen the view transformation and scene understanding. Considering the relationship between vehicles and roads, we also design a context-aware discriminator to further refine the results. Experiments on public benchmarks show that our method achieves the state-of-the-art performance in the tasks of road layout estimation and vehicle occupancy estimation. Especially for the latter task, our model outperforms all competitors by a large margin. Furthermore, our model runs at 35 FPS on a single GPU, which is efficient and applicable for real-time panorama HD map reconstruction.
Weixiang Yang, Qi Li 0038, Wenxi Liu, Yuanlong Yu 0001, Yuexin Ma, Shengfeng He, Jia Pan 0001
CVPR2
2021 From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual Correlation
abstract
Ultra-high resolution image segmentation has raised increasing interests in recent years due to its realistic applications. In this paper, we innovate the widely used high-resolution image segmentation pipeline, in which an ultrahigh resolution image is partitioned into regular patches for local segmentation and then the local results are merged into a high-resolution semantic mask. In particular, we introduce a novel locality-aware contextual correlation based segmentation model to process local patches, where the relevance between local patch and its various contexts are jointly and complementarily utilized to handle the semantic regions with large variations. Additionally, we present a contextual semantics refinement network that associates the local segmentation result with its contextual semantics, and thus is endowed with the ability of reducing boundary artifacts and refining mask contours during the generation of final high-resolution mask. Furthermore, in comprehensive experiments, we demonstrate that our model outperforms other state-of-the-art methods in public benchmarks. Our released codes are available at https://github.com/liqiokkk/FCtL.
Qi Li 0038, Weixiang Yang, Wenxi Liu, Yuanlong Yu 0001, Shengfeng He
ICCV1