Quanmin Liang

dblp:359/4179 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-9935-5167ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Margin-Aware Prototype Debiasing for Generalized Category Discovery
abstract
Generalized Category Discovery (GCD) is a challenging task that aims to identify both seen and novel categories in unlabeled data. We argue that a clear margin between seen and novel class representations is essential for accurate recognition. However, existing methods often ignore this margin, mapping representations to prototypes without enforcing separation between seen and novel classes. This leads to a bias where seen samples are misclassified as novel. To address this issue, we propose DebiasGCD, a debiasing framework that enhances prototype separation through margin-aware learning. Unlike prior work that relies on static prototype learning and overlooks fine-grained representations, our method introduces Dynamic Prototype Debiasing (DPD) and Spatial-Aware Representation Distillation (SARD) to mitigate this bias. First, DPD dynamically enforces inter-prototype margins, improving class-specific feature learning and prototype discrimination. Meanwhile, SARD promotes local representation of spatial learning, supporting DPD to capture subtle details that further refine class-specific features. By synergizing these components, DebiasGCD significantly improves prototype discriminability, generating more reliable predictions for seen classes. Extensive experiments demonstrate that our approach effectively mitigates pseudo-labeling bias across datasets, especially on fine-grained ones, achieving +8.3% and +9.6% improvements on the ‘All’ classes in CUB and Stanford Cars, respectively.
Xinzi Cao, Feidiao Yang, Xiawu Zheng, Quanmin Liang, Yutong Lu, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal Imaging
abstract
Recent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modalityand task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique “modality signature” for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model.
Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Shahrooz Faghih Roohi, Kai Huang 0001, Nassir Navab, M. Ali Nasseri
BIBM5
2025 UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis Augmentation
abstract
Significant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus imagebased disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks.
Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri
BIBM5
2025 FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
abstract
Fully sparse 3D detectors have recently gained significant attention due to their efficiency in long-range detection. However, sparse 3D detectors extract features only from non-empty voxels, which impairs long-range interactions and causes the center feature missing. The former weakens the feature extraction capability, while the latter hinders network optimization. To address these challenges, we introduce the Fully Sparse Hybrid Network (FSHNet). FSHNet incorporates a proposed SlotFormer block to enhance the long-range feature extraction capability of existing sparse encoders. The SlotFormer divides sparse voxels using a slot partition approach, which, compared to traditional window partition, provides a larger receptive field. Additionally, we propose a dynamic sparse label assignment strategy to deeply optimize the network by providing more high-quality positive samples. To further enhance performance, we introduce a sparse upsampling module to refine downsampled voxels, preserving fine-grained details crucial for detecting small objects. Extensive experiments on the Waymo, nuScenes, and Argoverse2 benchmarks demonstrate the effectiveness of FSHNet. The code is available at https://github.com/Say2L/FSHNet.
Shuai Liu 0009, Mingyue Cui, Boyang Li 0009, Quanmin Liang, Tinghe Hong, Yunxiao Shan, Kai Huang 0001
CVPR4
2025 Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Shuai Liu 0009, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ICCV1
2025 ESOD: Event-Based Small Object Detection
abstract
Event-based object detection plays a crucial role in scenarios involving high-speed motion, extreme lighting conditions, and high-frequency detection. However, existing methods fail to address the challenges posed by small objects, including discriminative feature deficiency, the loss of critical information, and the inherent sparsity of event data. Moreover, the lack of benchmark datasets has significantly hindered progress in this field. To tackle these issues, we propose the Fully Deformable Detection Network (FDDNet), a lightweight framework that dynamically adapts to extract key features. First, we introduce a Long-Term Deformable Temporal Receptive Module (LDTR), which aligns critical features across consecutive event streams and leverages a State Space Model for long-range temporal modeling, enhancing the detection of high-speed small objects. Second, to address the sparsity of event data and the concentration of key features along object edges, we design a Sparse Feature Aggregation Block (SFAB) within the backbone and a coarse-to-fine deformable detection head, enabling hierarchical feature refinement from local to global, and improving the detection quality of sparse targets. Finally, to mitigate the lack of event-based small object datasets, we develop a high-quality, annotation-free data acquisition method and collect a real-world benchmark dataset for validation. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on event-based small object detection tasks, with a mAP of 37.4% (+2.4%) on our benchmark and runs at 88 FPS, showcasing both accuracy and real-time capability. Our code and Supplement are available at https://github.com/Lqm26/ESOD.
Quanmin Liang, Jinyi Lu, Shuai Liu 0009, Yinzheng Zhao, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ACM Multimedia1
2025 GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
abstract
Multi-sensor fusion is crucial for improving the performance and robustness of end-to-end autonomous driving systems. Existing methods predominantly adopt either attention-based flatten fusion or bird’s eye view fusion through geometric transformations. However, these approaches often suffer from limited interpretability or dense computational overhead. In this paper, we introduce GaussianFusion, a Gaussian-based multi-sensor fusion framework for end-to-end autonomous driving. Our method employs intuitive and compact Gaussian representations as intermediate carriers to aggregate information from diverse sensors. Specifically, we initialize a set of 2D Gaussians uniformly across the driving scene, where each Gaussian is parameterized by physical attributes and equipped with explicit and implicit features. These Gaussians are progressively refined by integrating multi-modal features. The explicit features capture rich semantic and spatial information about the traffic scene, while the implicit features provide complementary cues beneficial for trajectory planning. To fully exploit rich spatial and semantic information in Gaussians, we design a cascade planning head that iteratively refines trajectory predictions through interactions with Gaussians. Extensive experiments on the NAVSIM and Bench2Drive benchmarks demonstrate the effectiveness and robustness of the proposed GaussianFusion framework. The source code is included in the supplementary material and will be released publicly.
Shuai Liu 0009, Quanmin Liang, Zefeng Li, Boyang Li 0009, Kai Huang 0001
NeurIPS2
2024 Bilateral Event Mining and Complementary for Event Stream Super-Resolution
abstract
Event Stream Super-Resolution (ESR) aims to address the challenge of insufficient spatial resolution in event streams, which holds great significance for the application of event cameras in complex scenarios. Previous works for ESR often process positive and negative events in a mixed paradigm. This paradigm limits their ability to effectively model the unique characteristics of each event and mutually refine each other by considering their correlations. In this paper, we propose a bilateral event mining and complementary network (BMCNet) to fully leverage the potential of each event and capture the shared information to complement each other simultaneously. Specifically, we resort to a two-stream network to accomplish comprehensive mining of each type of events individually. To facilitate the exchange of information between two streams, we propose a bilateral information exchange (BIE) module. This module is layer-wisely embedded between two streams, enabling the effective propagation of hierarchical global information while alleviating the impact of invalid information brought by inherent characteristics of events. The experimental results demonstrate that our approach outperforms the previous state-of-the-art methods in ESR, achieving performance improvements of over 11% on both real and synthetic datasets. Moreover, our method significantly enhances the performance of event-based downstream tasks such as object recognition and video reconstruction. Our code is available at https://github.com/Lqm26/BMCNet-ESR.
Zhilin Huang, Quanmin Liang, Yijie Yu 0001, Chujun Qin, Xiawu Zheng, Kai Huang 0001, Zikun Zhou, Wenming Yang
CVPR2
2024 Efficient Event Stream Super-Resolution with Recursive Multi-Branch Fusion
Quanmin Liang, Zhilin Huang, Xiawu Zheng, Feidiao Yang, Jun Peng 0007, Kai Huang 0001, Yonghong Tian 0001
IJCAI1
2024 A Hybrid Routing Pattern in Human Brain Structural Network Revealed By Evolutionary Computation
abstract
The human brain functional connectivity network (FCN) is constrained and shaped by the communication processes in the structural connectivity network (SCN). The underlying communication mechanism thus becomes a critical issue for understanding the formation and organization of the FCN. A number of communication models supported by different routing strategies have been proposed, with shortest path (SP), random diffusion (DIF), and spatial navigation (NAV) as the most typical, respectively requiring network global knowledge, local knowledge, and both for path seeking. Yet these models all assumed every brain region to use one routing strategy uniformly, ignoring convergent evidence that supports the regional heterogeneity in both terms of biological substrates and functional roles. In this regard, the current study developed a hybrid communication model that allowed each brain region to choose a routing strategy from SP, DIF, and NAV independently. A genetic algorithm was designed to uncover the underlying region-wise hybrid routing strategy (namely HYB). The HYB was found to outperform the three typical routing strategies in predicting FCN and facilitating robust communication. Analyses on HYB further revealed that brain regions in lower-order functional modules inclined to route signals using global knowledge, while those in higher-order functional modules preferred DIF that requires only local knowledge. Compared to regions that used global knowledge for routing, regions using DIF had denser structural connections, participated in more functional modules, but played a less dominant role within modules. Together, our findings further evidenced that hybrid routing underpins efficient SCN communication and locally heterogeneous structure-function coupling.
Quanmin Liang, Junji Ma, Xi-tian Chen, Qixiang Lin, Ni Shu, Zhengjia Dai, Ying Lin 0001
IEEE Trans. Medical Imaging1
2023 Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion Models
abstract
Event cameras offer the advantages of low latency, high temporal resolution and HDR compared to conventional cameras. Due to the asynchronous and sparse nature of events, many existing algorithms cannot be directly applied, necessitating the reconstruction of intensity frames. However, existing reconstruction methods often result in artifacts and edge blurring due to noise and event accumulation. In this paper, we argue that the key to event-based image reconstruction is to enhance the edge information of objects and restore the artifacts in the reconstructed images. To explain, edge information is one of the most important features in the event stream, providing information on the shape and contour of objects. Considering the extraordinary capabilities of Denoising Diffusion Probabilistic Models (DDPMs) in image generation, reconstruction, and restoration, we propose a new framework which incorporate it into the reconstruction pipeline to obtain high-quality results which effectively remove artifacts and blur in reconstructed images. Specifically, we first extract edge information from the event stream using the proposed event-based denoising method. It employs the contrast maximization framework to remove noise from the event stream and extract clear object edge information. And then, the edge information is further adopted to our diffusion model, which is used to enhance the edges of objects in the reconstructed images, thus improving the restoration effect. Experimental results show that our method achieves significant improvements in the mean squared error (MSE), the structural similarity (SSIM), and the perceptual similarity (LPIPS) metrics, with average improvements of 40%, 15%, and 25%, respectively, compared to previous state-of-the-art models, and has good generalization performance.
Quanmin Liang, Xiawu Zheng, Kai Huang 0001, Yan Zhang 0109, Jie Chen 0001, Yonghong Tian 0001
ACM Multimedia1