Qiankun Li 0004

dblp:228/7339-4 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-5121-1682ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MPE: A Power-Efficient Edge-Device Mamba Processor with Multi-Dimensional Calculation-Compression Scheme
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Yongke Wang, Anil A. Bharath, Emm Mic Drakakis
ISCAS6
2026 GTPE: A 28nm 33.12 TFLOPS/W GNN Training Processor with Unstructured Multi Threshold Pruning, Hybrid Multi-mode Approximate Computing and QUIRE Number System Support
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Tian-Chun Ye 0001, Anil A. Bharath, Emm Mic Drakakis
ISCAS6
2026 GATPE: A High-Performance Edge-Device GAT Processor with Multi-Layer Data-Variation Mechanism
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Anil A. Bharath, Emm Mic Drakakis
ISCAS7
2026 GlareLane: A real-world benchmark and method for lane detection under challenging illumination
Xiaojie Yu, Qiankun Li 0004, Hao Li 0058, Ben-Guo He, Qiang He 0002, Junxin Chen 0001
Image Vis. Comput.2
2026 Advancing oral leukoplakia progression recognition: A benchmark with dataset, method, and application
Linfei Feng, Qiankun Li 0004, Hao Wang 0260, Xuanyu Li, Feng He 0008, Xin Ning 0001, Prayag Tiwari
Neural Networks2
2026 Multi-stage dual-domain progressive network with synergistic training for sparse-view CT reconstruction
Jingyuan Shao, Huabao Chen, Qiankun Li 0004, Jiong Shu, Lingling Liu, Shaobin Dou
Neural Networks3
2026 Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning
abstract
With the accumulation of resources in the era of Big Data and the rise of pre-trained models in deep learning, optimizing neural networks for various tasks often involves different strategies for fine-tuning pre-trained models versus training from scratch. However, existing optimizers primarily focus on reducing the loss function by updating model parameters, without fully addressing the unique demands of these two major paradigms. In this paper, we propose DualOpt, a novel approach that decouples optimization techniques specifically tailored for these distinct training scenarios. For training from scratch, we introduce real-time layer-wise weight decay, designed to enhance both convergence and generalization by aligning with the characteristics of weight updates and network architecture. For more importantly fine-tuning, we integrate weight rollback with the optimizer, incorporating a rollback term into each weight update step. This ensures consistency in the weight distribution between upstream and downstream models, effectively mitigating knowledge forgetting and improving fine-tuning performance. Additionally, we extend the layer-wise weight decay to dynamically adjust the rollback levels across layers, adapting to the varying demands of different downstream tasks. Extensive experiments across diverse tasks, including image classification, object detection, semantic segmentation, and instance segmentation, demonstrate the broad applicability and state-of-the-art performance of DualOpt.
Xin Ning 0001, Qiankun Li 0004, Xiaolong Huang 0001, Qiupu Chen, Feng He 0008, Weijun Li 0002, Prayag Tiwari, Xinwang Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Empowering 2D neural network for 3D medical image segmentation via neighborhood information fusion
Qiankun Li 0004, Xiaolong Huang 0001, Bo Fang 0005, Duo Hong, Junxin Chen 0001
Pattern Recognit.1
2025 Recognizing Diabetic Foot Progression via Multimodal Fusion of Infrared Thermography and Clinical Data
abstract
Diabetic foot (DF), a severe diabetes complication, remains a leading cause of lower-limb amputations, underscoring the urgent need for early-stage detection to enable timely intervention and improve outcomes. However, current approaches often fail to distinguish diabetic patients without foot complications (DM) from those with DF, as single-modality infrared thermography (IRT) struggles with subtle thermal cues and lacks localization of pathological regions. In this paper, we propose DFP-MMNet (Diabetic Foot Progression MultiModal Network), a novel two-stage multimodal framework that addresses this challenge. In the first stage, a deep registration model, guided by corresponding RGB images, accurately localizes lesion Regions of Interest (ROIs) on thermograms. In the second stage, a dual-branch network fuses deep visual features from the ROI (via RegNetY-16GF) with semantic embeddings derived from structured clinical data, encoded using CLIP text branch. Extensive experiments on the collected ITC dataset demonstrate that DFP-MMNet achieves an F1-score of 82.7%, outperforming unimodal baselines by over 15%, offering a robust and accurate solution for early DF recognition.
Huarui Liu, Qiankun Li 0004, Xianjun Yang, Yining Sun
BIBM3
2025 The Cyber-Physical System of Oral Health Monitoring: A Data-Driven Approach for Inference
Qiankun Li 0004, Wanxiu Xu, Dezhi Yuan, Baokang Wu, Qingqin Xu
ICIC (27)2
2025 Progressive Large-Scale Modeling via Temporal-Spatial Focus Connector for Micro-Action Recognition
abstract
Recent advances in video action recognition have achieved remarkable performance in coarse-grained macro-action classification by leveraging large-scale visual backbones and transformer architectures. However, extending these successes to fine-grained micro-action recognition remains a fundamental challenge due to the subtlety, brevity, and low motion intensity of micro-actions. In this paper, we propose a high-capacity framework for micro-action recognition, enhancing both representation learning and decision robustness. We scale to large-scale backbones using the VideoMAEv2 Giant model, enabling the extraction of finer spatial-temporal features. A Temporal-Spatial Connector (TSC) is introduced to dynamically highlight discriminative temporal frames and spatial regions, strengthening the model's focus on subtle motion cues critical for micro-action identification. To stabilize optimization and fully exploit the capacity of large models, we design a four-phase progressive training strategy, encompassing linear probing, full fine-tuning, connector-specific optimization, and classifier head refinement. Furthermore, we propose a novel ensemble decision mechanism that integrates Top-K predictions from diverse models via a Large Language Model (LLM), enhancing prediction consistency and robustness through multimodel consensus. Our method achieves an F1mean of 76.54% on the MA-52 dataset, ranking 3rd in the 2025 Micro-Action Analysis Grand Challenge and advancing the state of the art in fine-grained video understanding.
Qiankun Li 0004, Qiupu Chen, Huabao Chen, Feng He 0008, Depeng Li 0001, Zhigang Zeng
ACM Multimedia1
2025 From Pixels to Views: Learning Angular-Aware and Physics-Consistent Representations for Light Field Microscopy
abstract
Light field microscopy (LFM) has become an emerging tool in neuroscience for large-scale neural imaging in vivo, with XLFM (eXtended Light Field Microscopy) notable for its single-exposure volumetric imaging, broad field of view, and high temporal resolution. However, learning-based 3D reconstruction in XLFM remains underdeveloped due to two core challenges: the absence of standardized datasets and the lack of methods that can efficiently model its angular–spatial structure while remaining physically grounded. We address these challenges by introducing three key contributions. First, we construct the XLFM-Zebrafish benchmark, a large-scale dataset and evaluation suite for XLFM reconstruction. Second, we propose Masked View Modeling for Light Fields (MVM-LF), a self-supervised task that learns angular priors by predicting occluded views, improving data efficiency. Third, we formulate the Optical Rendering Consistency Loss (ORC Loss), a differentiable rendering constraint that enforces alignment between predicted volumes and their PSF-based forward projections. On the XLFM-Zebrafish benchmark, our method improves PSNR by 7.7\% over state-of-the-art baselines. Code and datasets are publicly available at: https://github.com/hefengcs/XLFM-Former.
Feng He 0008, Guodong Tan, Qiankun Li 0004
NeurIPS3
2025 Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific Domains
abstract
In the big data era, the computer vision field benefits from large-scale datasets such as LAION-2B, LAION-400M, and ImageNet-21K, Kinetics, on which popular models like the ViT and ConvNeXt series have been pre-trained, acquiring substantial knowledge. However, numerous downstream tasks in specialized and data-limited scientific domains continue to pose significant challenges. In this paper, we propose a novel Cluster Attention Adapter (CLAdapter), which refines and adapts the rich representations learned from large-scale data to various data-limited downstream tasks. Specifically, CLAdapter introduces attention mechanisms and cluster centers to personalize the enhancement of transformed features through distribution correlation and transformation matrices. This enables models fine-tuned with CLAdapter to learn distinct representations tailored to different feature sets, facilitating the models' adaptation from rich pre-trained features to various downstream scenarios effectively. In addition, CLAdapter's unified interface design allows for seamless integration with multiple model architectures, including CNNs and Transformers, in both 2D and 3D contexts. Through extensive experiments on 10 datasets spanning domains such as generic, multimedia, biological, medical, industrial, agricultural, environmental, geographical, materials science, out-of-distribution (OOD), and 3D analysis, CLAdapter achieves state-of-the-art performance across diverse data-limited scientific domains, demonstrating its effectiveness in unleashing the potential of foundation vision models via adaptive transfer. Code is available at https://github.com/qklee-lz/CLAdapter.
Qiankun Li 0004, Feng He 0008, Huabao Chen, Xin Ning 0001, Kun Wang 0056, Zengfu Wang
NeurIPS1
2025 Decoding text from electroencephalography signals: A novel Hierarchical Gated Recurrent Unit with Masked Residual Attention Mechanism
Qiupu Chen, Yimou Wang, Fenmei Wang, Duolin Sun, Qiankun Li 0004
Eng. Appl. Artif. Intell.5
2025 Fuzzy-ViT: A Deep Neuro-Fuzzy System for Cross-Domain Transfer Learning From Large-Scale General Data to Medical Image
abstract
The surge in visual general big data has notably advanced data-driven deep learning-based computer vision technologies. Transformer-based methods shine in this era of big data because of their attention mechanism architecture and demand for massive data. However, the difficulty of obtaining medical images has caused the field to continue facing the limited-data challenge. In this paper, we propose a novel deep neuro-fuzzy system named Fuzzy-ViT, which synergistically integrates fuzzy logic with the Vision Transformer (ViT) for cross-domain transfer learning from large-scale general data to medical image domain. Specifically, Fuzzy-ViT utilizes a ViT backbone pre-trained on extensive general datasets such as ImageNet-21K, LAION-400M, and LAION-2B to extract rich general features. Then, a Fuzzy Attention Cross-Domain Module (FACM) is presented to transfer general features to medical features, thereby enhancing the medical image analysis. Thanks to the Fuzzy System Transitioner (FST) in FACM, fuzzy and uninterpretable general domain features can be effectively converted into those needed in the medical domain. In addition, the Attention Mechanism Smoother (AMS) in FACM smoothes the conversion outcomes, ensuring a harmonious integration of the fuzzy system with the neural network architecture. Experimental results demonstrate that the proposed Fuzzy-ViT achieves state-of-the-art and satisfactory performance on popular medical image benchmarks (BreakHis and HCRF) with 93.37% and 97.22% F1 scores. Detailed ablation analysis demonstrates that the effectiveness of our method for bridging large general visual and medical images.
Qiankun Li 0004, Yimou Wang, Zhaoyu Zuo, Junxin Chen 0001, Wei Wang 0077
IEEE Trans. Fuzzy Syst.1
2024 One Step Learning, One Step Review
abstract
Visual fine-tuning has garnered significant attention with the rise of pre-trained vision models. The current prevailing method, full fine-tuning, suffers from the issue of knowledge forgetting as it focuses solely on fitting the downstream training set. In this paper, we propose a novel weight rollback-based fine-tuning method called OLOR (One step Learning, One step Review). OLOR combines fine-tuning with optimizers, incorporating a weight rollback term into the weight update term at each step. This ensures consistency in the weight range of upstream and downstream models, effectively mitigating knowledge forgetting and enhancing fine-tuning performance. In addition, a layer-wise penalty is presented to employ penalty decay and the diversified decay rate to adjust the weight rollback levels of layers for adapting varying downstream tasks. Through extensive experiments on various tasks such as image classification, object detection, semantic segmentation, and instance segmentation, we demonstrate the general applicability and state-of-the-art performance of our proposed OLOR. Code is available at https://github.com/rainbow-xiao/OLOR-AAAI-2024.
Xiaolong Huang 0001, Qiankun Li 0004, Xueran Li, Xuesong Gao
AAAI2
2024 ICAF-4: An Integrated Framework of Category-level Articulated Object Perception and Manipulation for Embodied Intelligence
Li Zhang 0104, Qiankun Li 0004, Qi Wu 0007, Lin Wu 0001, Liu Liu 0012
BMVC3
2024 Advancing Micro-Action Recognition with Multi-Auxiliary Heads and Hybrid Loss Optimization
abstract
Video action recognition has been a hot research direction in computer vision, with most existing technologies focusing on coarse-grained macro-action recognition. However, fine-grained action recognition remains challenging. Micro-actions, characterized by high fine-grained, low-intensity, and brief, are crucial for emotion recognition and psychological assessment applications. In this paper, we build on popular video action recognition frameworks as foundation models, introducing multi-auxiliary heads and hybrid loss optimization to advance micro-action recognition. Specifically, the Frame-Level pred and Coarse-Grained Body-Action auxiliary heads work collaboratively to enhance the model and Fine-Grained Micro-Action primary head for perceiving fine-grained and capturing keyframes. Incorporating F1 loss, ArcFace loss, and weighted multi-task loss improves training stability, convergence speed, and performance. Additionally, integrating the optical flow modality enriches the model's diversity, and ensemble learning across all foundational models. Finally, our method achieves a 75.37% F1-mean on the MA-52 dataset, ranking 1st in the Micro-Action Analysis Grand Challenge in conjunction with ACM MM'24. The code is available at https://github.com/qklee-lz/ACMMM2024-MAC.
Qiankun Li 0004, Xiaolong Huang 0001, Huabao Chen, Feng He 0008, Qiupu Chen, Zengfu Wang
ACM Multimedia1
2024 Enhancing Semi-Dense Feature Matching Through Probabilistic Modeling of Cascaded Supervision and Consistency
Hongchang Min, Yihong Tang, Qiankun Li 0004, Zengfu Wang
PRCV (15)3
2024 DeepSweep: Real-Time Multi-View 3D Pose Estimation Via Cross-View Deep Matching and Plane Sweeping
Wenrui Zhu, Qiankun Li 0004, Debin Liu, Zengfu Wang
PRCV (11)2
2024 Embracing Large Natural Data: Enhancing Medical Image Analysis via Cross-Domain Fine-Tuning
abstract
With the rapid advancements of Big Data and computer vision, many large-scale natural visual datasets are proposed, such as ImageNet-21K, LAION-400M, and LAION-2B. These large-scale datasets significantly improve the robustness and accuracy of models in the natural vision domain. However, the field of medical images continues to face limitations due to relatively small-scale datasets. In this article, we propose a novel method to enhance medical image analysis across domains by leveraging pre-trained models on large natural datasets. Specifically, a Cross-Domain Transfer Module (CDTM) is proposed to transfer natural vision domain features to the medical image domain, facilitating efficient fine-tuning of models pre-trained on large datasets. In addition, we design a Staged Fine-Tuning (SFT) strategy in conjunction with CDTM to further improve the model performance. Experimental results demonstrate that our method achieves state-of-the-art performance on multiple medical image datasets through efficient fine-tuning of models pre-trained on large natural datasets.
Qiankun Li 0004, Xiaolong Huang 0001, Bo Fang 0005, Huabao Chen, Siyuan Ding
IEEE J. Biomed. Health Informatics1
2024 PGA-Net: Polynomial Global Attention Network With Mean Curvature Loss for Lane Detection
abstract
Lane detection is an important task in the field of automatic driving. Since lane lines usually have complex topologies and exist in various complex scenes (e.g., damaged lanes, severe occlusion, etc.), lane detection remains challenging. In this work, we propose a Polynomial Global Attention Network (PGA-Net) for lane detection, which is an end-to-end model for mining global road information and predicting lanes shape parameter formulas simultaneously. We model lane shape with cubic polynomial function and use the transformer-based DETR model to introduce the context information of lanes and roads to better regress the lane parameters. For polynomial curve modeling, we propose Mean Curvature Loss (MCL) to constrain the curvature of the predicted lanes, thereby enhancing the quality of curve lanes prediction. In addition, we design an improved supervision strategy to eliminate information bias between our parametric prediction methods and the labeling methods of lane datasets. Our method achieves state-of-the-art performance on two popular benchmarks (TuSimple and LLAMAS) and a most challenging benchmark (CULane), while exhibiting accelerated speed (>140fps on 3090 GPU, 28.9% improvement in average) and lightweight model size (https://github.com/qklee-lz/PGA-Net.
Qiankun Li 0004, Xianwang Yu, Junxin Chen 0001, Ben-Guo He, Wei Wang 0077, Danda B. Rawat, Zhihan Lyu
IEEE Trans. Intell. Transp. Syst.1
2023 PnP-AE: A Plug-and-Play Module for Volumetric Medical Image Segmentation
abstract
In recent years, 3D volumetric medical images have been widely used in clinical diagnosis, however, the popular 2D networks were reported unsuitable for segmenting them. In this direction, we propose a plug-and-play (PnP-AE) module to improve the performance of using 2D network for 3D medical image segmentation. Our method takes advantage of the intrinsic correlation between adjacent slices, by multiple encoders and fusion components to decouple plane feature extraction and depth information integration. In addition, the proposed weight sharing and feature storage strategies make PnP-AE extremely efficient. Our method is able to conveniently incorporate with mainstream 2D networks to segment 3D volumetric medical images. Experimental results demonstrate the excellent performance of our method. The source code is available at https://github.com/qklee-lz/PnP-AE.
Qiankun Li 0004, Xiaolong Huang 0001, Bo Fang 0005, Yongyong Chen, Junxin Chen 0001
BIBM1
2023 LABANet: Lead-Assisting Backbone Attention Network for Oral Multi-Pathology Segmentation
abstract
This paper presents a Lead-Assisting Backbone Attention Network (LABANet), which is able to perform multi-pathology instance segmentation of dental panoramic X-rays. A Lead-Assisting Attention Backbone (LAAB), containing two Swin-Transformers, is first developed for feature extraction. The following Region Proposal Network (RPN) and RoIAlign modules further convert the extracted features to a fixed-size feature map. Finally, an improved attention head with a Squeeze-and-Excitation (SE) block is constructed for object classification, bounding-box regression, and mask segmentation. By taking advantage of the global attention mechanism, the LABANet can better achieve multiple pathology segmentation. Experiment results demonstrate its effectiveness and advantages over state-of-the-art methods.
Huabao Chen, Xiaolong Huang 0001, Qiankun Li 0004, Jianqing Wang, Bo Fang 0005, Junxin Chen 0001
ICASSP3
2023 Data-Efficient Masked Video Modeling for Self-supervised Action Recognition
abstract
Recently, self-supervised video representation learning based on Masked Video Modeling (MVM) has demonstrated promising results for action recognition. However, existing methods face two significant challenges: (1) video actions involve a crucial temporal dimension, yet current masking strategies adopt inefficient random approaches that undermine low-density dynamic motion clues in videos; (2) pre-training requires large-scale datasets and significant computing resources (including large batch sizes and enormous iterations). To address these issues, we propose a novel method named Data-Efficient Masked Video Modeling (DEMVM) for self-supervised action recognition. Specifically, a novel masking strategy named Flow-Guided Dense Masking (FGDM) is proposed to facilitate efficient learning by focusing more on the action-related temporal clues, which applies dense masking to dynamic regions based on optical flow priors, while sparse masking to background regions. Furthermore, DEMVM introduces a 3D video tokenizer to enhance the modeling of temporal clues. Finally, Progressive Masking Ratio (PMR) and 2D initialization strategies are presented to enable the model to adapt to the characteristics of the MVM paradigm during different training stages. Extensive experiments on multiple benchmarks, UCF101, HMDB51, and Mimetics, demonstrate that our method achieves state-of-the-art performance in the downstream action recognition task with both efficient data and low computational cost. More interestingly, the few-shot experiment on the Mimetics dataset shows that DEMVM can accurately recognize actions even in the presence of context bias.
Qiankun Li 0004, Xiaolong Huang 0001, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang 0071, Shiguang Shan, Zengfu Wang
ACM Multimedia1