Siqi Li 0009

dblp:34/180-9 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2026
0009-0000-4632-9010ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Automatic data-free pruning via channel similarity reconstruction
Siqi Li 0009, Jun Chen 0023, Jingyang Xiang, Chengrui Zhu, Jiandang Yang, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neurocomputing1
2026 SampleLLM: Prune LLMs via learnable structure sampling
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Longqi Wang, Chuang Guo, Jun Chen 0023, Baochang Zhu, Yong Liu 0007
Neurocomputing2
2026 WLR: Well-conditioned linear reconstruction for retraining-free pruning of LLMs
Siqi Li 0009, Jingyang Xiang, Jiateng Wei, Chengrui Zhu, Jiandang Yang, Jun Chen 0023, Jian Yang 0003, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neural Networks1
2026 OOPS: Outlier-aware and quadratic programming based structured pruning for large language models
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Jiandang Yang, Jun Chen 0023, Xiaobin Wei, Yunliang Jiang, Yong Liu 0007
Neural Networks2
2025 Chameleon: Fast-Slow Neuro-Symbolic Lane Topology Extraction
abstract
Lane topology extraction involves detecting lanes and traffic elements and determining their relationships, a key perception task for mapless autonomous driving. This task requires complex reasoning, such as determining whether it is possible to turn left into a specific lane. To address this challenge, we introduce neuro-symbolic methods powered by vision-language foundation models (VLMs). Existing approaches have notable limitations: (1) Dense visual prompting with VLMs can achieve strong performance but is costly in terms of both financial resources and carbon footprint, making it impractical for robotics applications. (2) Neuro-symbolic reasoning methods for 3D scene understanding fail to integrate visual inputs when synthesizing programs, making them ineffective in handling complex corner cases. To this end, we propose a fast-slow neuro-symbolic lane topology extraction algorithm, named Chameleon, which alternates between a fast system that directly reasons over detected instances using synthesized programs and a slow system that utilizes a VLM with a chain-of-thought design to handle corner cases. Chameleon leverages the strengths of both approaches, providing an affordable solution while maintaining high performance. We evaluate the method on the OpenLane-V2 dataset, showing consistent improvements across various baseline detectors. Our code, data, and models are publicly available at https://github.com/XR-Lee/neural-symbolic
Zongzheng Zhang, Xinrun Li, Sizhe Zou, Guoxuan Chi, Siqi Li 0009, Xuchong Qiu, Guoliang Wang 0002, Guantian Zheng, Leichen Wang, Hang Zhao 0021, Hao Zhao 0002
ICRA5
2025 MCMC: Multi-Constrained Model Compression via One-Stage Envelope Reinforcement Learning
abstract
Model compression methods are being developed to bridge the gap between the massive scale of neural networks and the limited hardware resources on edge devices. Since most real-world applications deployed on resource-limited hardware platforms typically have multiple hardware constraints simultaneously, most existing model compression approaches that only consider optimizing one single hardware objective are ineffective. In this article, we propose an automated pruning method called multi-constrained model compression (MCMC) that allows for the optimization of multiple hardware targets, such as latency, floating point operations (FLOPs), and memory usage, while minimizing the impact on accuracy. Specifically, we propose an improved multi-objective reinforcement learning (MORL) algorithm, the one-stage envelope deep deterministic policy gradient (DDPG) algorithm, to determine the pruning strategy for neural networks. Our improved one-stage envelope DDPG algorithm reduces exploration time and offers greater flexibility in adjusting target priorities, enhancing its suitability for pruning tasks. For instance, on the visual geometry group (VGG)-16 network, our method achieved an 80% reduction in FLOPs, a reduction in memory usage, and a acceleration, with an accuracy improvement of 0.09% compared with the baseline. For larger datasets, such as ImageNet, we reduced FLOPs by 50% for MobileNet-V1, resulting in a faster speed and memory compression, while maintaining the same accuracy. When applied to edge devices, such as JETSON XAVIER NX, our method resulted in a 71% reduction in FLOPs for MobileNet-V1, leading to a faster speed, memory compression, and an accuracy improvement.
Siqi Li 0009, Jun Chen 0023, Shanqi Liu, Chengrui Zhu, Guanzhong Tian, Yong Liu 0007
IEEE Trans. Neural Networks Learn. Syst.1
2024 Locate N' Rotate: Two-Stage Openable Part Detection with Foundation Model Priors
Siqi Li 0009, Xiaoxue Chen, Haoyu Cheng, Guyue Zhou, Hao Zhao 0002, Guanzhong Tian
ACCV (7)1
2024 MaxQ: Multi-Axis Query for N: m Sparsity Network
abstract
N:m sparsity has received increasing attention due to its remarkable performance and latency trade-off compared with structured and unstructured sparsity. How-ever, existing N:m sparsity methods do not differentiate the relative importance of weights among blocks and leave important weights underappreciated. Besides, they di-rectly apply N:m sparsity to the whole network, which will cause severe information loss. Thus, they are still sub-optimal. In this paper, we propose an efficient and effective Multi-Axis Query methodology, dubbed as MaxQ, to rectify these problems. During the training, MaxQ employs a dynamic approach to generate soft N:m masks, considering the weight importance across multiple axes. This method enhances the weights with more importance and ensures more effective updates. Meanwhile, a spar-sity strategy that gradually increases the percentage of N:m weight blocks is applied, which allows the network to heal from the pruning-induced damage progressively. During the runtime, the N:m soft masks can be precom-puted as constants and folded into weights without causing any distortion to the sparse pattern and incurring ad-ditional computational overhead. Comprehensive experi-ments demonstrate that MaxQ achieves consistent improve-ments across diverse CNN architectures in various com-puter vision tasks, including image classification, object detection and instance segmentation. For ResNet50 with 1:16 sparse pattern, MaxQ can achieve 74.6% top-1 ac-curacy on ImageNet and improve by over 2.8% over the state-of-the-art. Codes and checkpoints are available at https://github.com/JingyangXiang/MaxQ.
Jingyang Xiang, Siqi Li 0009, Zhuangzhi Chen, Tianxin Huang, Linpeng Peng, Yong Liu 0007
CVPR2
2024 OvSW: Overcoming Silent Weights for Accurate Binary Neural Networks
Jingyang Xiang, Zuohui Chen, Siqi Li 0009, Yong Liu 0007
ECCV (33)3
2024 Structured Optimal Brain Pruning for Large Language Models
abstract
The massive parameters and computational demands hinder the widespread application of Large Language Models (LLMs).Network pruning provides a practical solution to this problem.However, existing pruning works for LLMs mainly focus on unstructured pruning or necessitate post-pruning fine-tuning.The former relies on special hardware to accelerate computation, while the latter may need substantial computational resources.In this paper, we introduce a retraining-free structured pruning method called SoBP (Structured Optimal Brain Pruning).It leverages global first-order information to select pruning structures, then refines them with a local greedy approach, and finally adopts module-wise reconstruction to mitigate information loss.We assess the effectiveness of SoBP across 14 models from 3 LLM families on 8 distinct datasets.Experimental results demonstrate that SoBP outperforms current state-of-the-art methods.
Jiateng Wei, Siqi Li 0009, Jingyang Xiang, Jun Chen 0023, Yong Liu 0007
EMNLP4
2024 PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
abstract
Robotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categories and environments. To overcome these limitations, we introduce PreAfford, a novel pre-grasping planning framework incorporating a point-level affordance representation and a relay training approach. Our method significantly improves adaptability, allowing effective manipulation across a wide range of environments and object types. When evaluated on the ShapeNet-v2 dataset, PreAfford not only enhances grasping success rates by 69% but also demonstrates its practicality through successful real-world experiments. These improvements highlight PreAfford’s potential to redefine standards for robotic handling of complex manipulation tasks in diverse settings.
Kairui Ding, Boyuan Chen 0009, Ruihai Wu, Zongzheng Zhang, Huan-ang Gao, Siqi Li 0009, Guyue Zhou, Yixin Zhu 0001, Hao Dong 0003, Hao Zhao 0002
IROS7
2023 SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading Acceleration
Jingyang Xiang, Siqi Li 0009, Jun Chen 0023, Guang Dai, Shipeng Bai, Yukai Ma, Yong Liu 0007
NeurIPS2