EDBT 2026 Demo / reviewers in the wild / expert
Sifan Zhou
dblp:256/3342
· DBLP profile ↗
26ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0003-3602-7566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 12 since 2021Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object DetectionabstractCamera-based multi-view 3D detection is crucial for autonomous driving. PETR and its variants (PETRs) excel in benchmarks but face deployment challenges due to high computational cost and memory footprint. Quantization is an effective technique for compressing deep neural networks by reducing the bit width of weights and activations. However, directly applying existing quantization methods to PETRs leads to severe accuracy degradation. This issue primarily arises from two key challenges: (1) significant magnitude disparity between multi-modal features—specifically, image features and camera-ray positional embeddings (PE), and (2) the inefficiency and approximation error of quantizing non-linear operators, which commonly rely on hardware-unfriendly computations. In this paper, we propose FQ-PETR, a fully quantized framework for PETRs, featuring three key innovations: (1) Quantization-Friendly LiDAR-ray Position Embedding (QFPE): Replacing multi-point sampling with LiDAR-prior-guided single-point sampling and anchor-based embedding eliminates problematic non-linearities (e.g., inverse-sigmoid) and aligns PE scale with image features, preserving accuracy. (2) Dual-Lookup Table (DULUT): This algorithm approximates complex non-linear functions using two cascaded linear LUTs, achieving high fidelity with minimal entries and no specialized hardware. (3) Quantization After Numerical Stabilization (QANS): Performing quantization after softmax numerical stabilization mitigates attention distortion from large inputs. On PETRs (e.g., PETR, StreamPETR, PETRv2, MV2d), FQ-PETR under W8A8 achieves near-floating-point accuracy ( Jiangyong Yu, Changyong Shu, Sifan Zhou, Zichen Yu, Xing Hu 0010 |
AAAI | 3 |
| 2026 | CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Trackingabstract3D single object tracking (SOT) in LiDAR point clouds is a critical task in computer vision and autonomous driving. Despite great success having been achieved, the inherent sparsity of point clouds introduces a dual-redundancy challenge that limits existing trackers: (1) vast spatial redundancy from background noise impairs accuracy, and (2) informational redundancy within the foreground hinders efficiency. To tackle these issues, we propose CompTrack, a novel end-to-end framework that systematically eliminates both forms of redundancy in point clouds. First, CompTrack incorporates a Spatial Foreground Predictor (SFP) module to filter out irrelevant background noise based on information entropy, addressing spatial redundancy. Subsequently, its core is an Information Bottleneck-guided Dynamic Token Compression (IB-DTC) module that eliminates the informational redundancy within the foreground. Theoretically grounded in low-rank approximation, this module leverages an online SVD analysis to adaptively compress the redundant foreground into a compact and highly informative set of proxy tokens. Extensive experiments on KITTI, nuScenes and Waymo datasets demonstrate that CompTrack achieves top-performing tracking performance with superior efficiency, running at a real-time 90 FPS on a single RTX 3090 GPU. Sifan Zhou, Yichao Cao, Jiahao Nie 0001, Yuqian Fu, Xiaobo Lu, Shuo Wang 0030 |
AAAI | 1 |
| 2026 | FastPillars: A Deployment-Friendly Pillar-Based 3D DetectorabstractThe deployment of 3D detectors strikes one of the major challenges in real-world self-driving scenarios. Existing BEV-based (i.e., Bird Eye View) detectors favor sparse convolutions (known as SPConv) to speed up training and inference, which puts a hard barrier for deployment, especially for on-device applications. In this paper, in order to tackle the challenge of efficient 3D object detection from an industry perspective, we devise a deployment-friendly pillar-based 3D detector, termed FastPillars. Specifically, aiming to compensate the geometric information loss of pillar encoding. First, we design a novel lightweight Max-and-Attention Pillar Encoding (MAPE) module specially for enhancing small objects. Second, we propose a simple yet effective backbone design for pillar-based 3D detection, enhancing pillar representations. We construct FastPillars based on these designs, achieving high performance and low latency without SPConv. Extensive experiments on two large-scale datasets demonstrate the effectiveness and efficiency of FastPillars for on-device 3D detection regarding both performance and speed. Specifically, FastPillars delivers real-time state-of-the-art accuracy on Waymo Open Dataset with 1.8 × speed up and 3.8 mAPH/L2 improvement over CenterPoint (SPConv-based). Code will be opened soon in: https://github.com/StiphyJay/FastPillars. Sifan Zhou, Xinyu Zhang 0015, Xiangxiang Chu, Bo Zhang 0046, Xiaobo Lu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Tartan IMU: A Light Foundation Model for Inertial Positioning in RoboticsabstractDespite recent advances in deep learning, most existing learning IMU odometry methods are trained on specific datasets, lack generalization, and are prone to overfitting, which limits their real-world application. To address these challenges, we present Tartan IMU, a foundation model designed for generalizable, IMU-based state estimation across diverse robotic platforms. Our approach consists of three-stage: First, a pre-trained foundation model leverages over 100 hours of multi-platform data to establish general motion knowledge, achieving 36% improvement in ATE over specialized models. Second, to adapt to previously unseen tasks, we use Low-Rank Adaptation (LoRA), allowing positive transfer with only 1.1 M trainable parameters. Finally, to support robotics deployment, we introduce online test-time adaptation, which eliminates the boundary between training and testing, allowing the model to continuously "learn as it operates" at 200 FPS in real-time. Project page: https://superodometry.com/tartanimu. Shibo Zhao, Sifan Zhou, Raphael Blanchard, Yuheng Qiu, Sebastian A. Scherer |
CVPR | 2 |
| 2025 | PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware HistogramabstractReal-time and high-performance 3D object detection plays a critical role in autonomous driving and robotics. Recent pillar-based 3D object detectors have gained significant attention due to their compact representation and low computational overhead, making them suitable for onboard deployment and quantization. However, existing pillar-based detectors still suffer from information loss along height dimension and large numerical distribution difference during pillar feature encoding (PFE), which severely limits their performance and quantization potential. To address above issue, we first unveil the importance of different input information during PFE and identify the height dimension as a key factor in enhancing 3D detection performance. Motivated by this observation, we propose a heightaware pillar feature encoder, called PillarHist. Specifically, PillarHist statistics the discrete distribution of points at different heights within one pillar with the information entropy guidance. This simple yet effective design greatly preserves the information along the height dimension while significantly reducing the computation overhead of the PFE. Meanwhile, PillarHist also constrains the arithmetic distribution of PFE input to a stable range, making it quantization-friendly. Notably, PillarHist operates exclusively within the PFE stage to enhance performance, enabling seamless integration into existing pillar-based methods without introducing complex operations. Extensive experiments show the effectiveness of PillarHist in terms of both efficiency and performance. Sifan Zhou, Zhihang Yuan, Xing Hu 0010, Jian Qian |
CVPR | 1 |
| 2025 | OSTQuant: Refining Large Language Model Quantization with Orthogonal and Scaling Transformations for Better Distribution FittingabstractPost-training quantization (PTQ) has emerged as a widely adopted technique for compressing and accelerating Large Language Models (LLMs).
The major challenge in LLM quantization is that uneven and heavy-tailed data distributions can expand the quantization range, thereby reducing bit precision for most values.
Recent methods attempt to eliminate outliers and balance inter-channel differences by employing linear transformations; however, they remain heuristic and are often overlook optimizing the data distribution across the entire quantization space.
In this paper, we introduce Quantization Space Utilization Rate (QSUR), a novel metric that effectively assesses the quantizability of transformed data by measuring the space utilization of the data in the quantization space. We complement QSUR with mathematical derivations that examine the effects and limitations of various transformations, guiding our development of Orthogonal and Scaling Transformation-based Quantization (OSTQuant). OSTQuant employs a learnable equivalent transformation, consisting of an orthogonal transformation and a scaling transformation, to optimize the distributions of weights and activations across the entire quantization space. Futhermore, we propose the KL-Top loss function, designed to mitigate noise during optimization while retaining richer semantic information within the limited calibration data imposed by PTQ.
OSTQuant outperforms existing work on various LLMs and benchmarks. In the W4-only setting, it retains 99.5\% of the floating-point accuracy. In the more challenging W4A4KV4 configuration, OSTQuant reduces the performance gap by 32\% on the LLaMA-3-8B model compared to state-of-the-art methods. Code will be available. Xing Hu 0010, Zhixuan Chen, Zukang Xu, Jiangyong Yu, Zhihang Yuan, Zhe Jiang 0004, Sifan Zhou |
ICLR | 10 |
| 2025 | MambaQuant: Quantizing the Mamba Family with Variance Aligned Rotation MethodsabstractMamba is an efficient sequence model that rivals Transformers and demonstrates significant potential as a foundational architecture for various tasks. Quantization is commonly used in neural networks to reduce model size and computational latency. However, applying quantization to Mamba remains underexplored, and existing quantization methods, which have been effective for CNN and Transformer models, appear inadequate for Mamba models (e.g., Quarot suffers a 21% accuracy drop on Vim-T$\dagger$ even under W8A8). We have pioneered the exploration of this issue and identified several key challenges. First, significant outliers arepresent in gate projections, output projections, and matrix multiplications. Second, Mamba’s unique parallel scan further amplifies these outliers, leading to uneven and heavy-tailed data distributions. Third, even with the application of the Hadamard transform, the variance across channels in weights and activations still remains inconsistent. To these ends, we propose MambaQuant, a post-training quantization (PTQ) framework consisting of: 1) Karhunen-Lo`eve Transformation (KLT) enhanced rotation, rendering the rotation matrix adaptable to diverse channel distributions. 2) Smooth-Fused rotation, which equalizes channel variances and can merge additional parameters into model weights. Experiments show that MambaQuant can quantize both weights and activations into 8-bit with less than 1% accuracy loss for Mamba-based vision and language tasks. To our knowledge, MambaQuant is the first comprehensive PTQ design for the Mamba family, paving the way for further advancements in its application. Zukang Xu, Yuxuan Yue, Xing Hu 0010, Zhihang Yuan, Zixu Jiang, Zhixuan Chen, Jiangyong Yu, Sifan Zhou |
ICLR | 10 |
| 2025 | MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity GuidanceabstractMixture-of-Experts (MoE) large language models (LLMs), which leverage dynamic routing and sparse activation to enhance efficiency and scalability, have achieved higher performance while reducing computational costs. However, these models face significant memory overheads, limiting their practical deployment and broader adoption. Post-training quantization (PTQ), a widely used method for compressing LLMs, encounters severe accuracy degradation and diminished generalization performance when applied to MoE models. This paper investigates the impact of MoE’s sparse and dynamic characteristics on quantization and identifies two primary challenges: (1) Inter-expert imbalance, referring to the uneven distribution of samples across experts, which leads to insufficient and biased calibration for less frequently utilized experts; (2) Intra-expert imbalance, arising from MoE’s unique aggregation mechanism, which leads to varying degrees of correlation between different samples and their assigned experts. To address these challenges, we propose MoEQuant, a novel quantization framework tailored for MoE LLMs. MoEQuant includes two novel techniques: 1) Expert-Balanced Self-Sampling (EBSS) is an efficient sampling method that efficiently constructs a calibration set with balanced expert distributions by leveraging the cumulative probabilities of tokens and expert balance metrics as guiding factors. 2) Affinity-Guided Quantization (AGQ), which incorporates affinities between experts and samples into the quantization process, thereby accurately assessing the impact of individual samples on different experts within the MoE layer. Experiments demonstrate that MoEQuant achieves substantial performance gains (more than 10 points accuracy gain in the HumanEval for DeepSeekMoE-16B under 4-bit quantization) and boosts efficiency. Zhixuan Chen, Xing Hu 0010, Zukang Xu, Zhihang Yuan, Sifan Zhou, Jiangyong Yu |
ICML | 7 |
| 2025 | RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector QuantizationabstractRWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14$\times$ speed up. Yuxuan Yue, Zukang Xu, Xing Hu 0010, Jiangyong Yu, Zhixuan Chen, Sifan Zhou, Zhihang Yuan |
ICML | 7 |
| 2025 | MVCTrack: Boosting 3D Point Cloud Tracking via Multimodal-Guided Virtual Cues
Zhaofeng Hu, Sifan Zhou, Zhihang Yuan, Shibo Zhao, Ci-Jyun Liang |
ICRA | 2 |
| 2025 | PTQ4RIS: Post-Training Quantization for Referring Image SegmentationabstractReferring Image Segmentation (RIS), aims to segment the object referred by a given sentence in an image by understanding both visual and linguistic information. However, existing RIS methods tend to explore top-performance models, disregarding considerations for practical applications on resources-limited edge devices. This oversight poses a significant challenge for on-device RIS inference. To this end, we propose an effective and efficient post-training quantization framework termed PTQ4RIS. Specifically, we first conduct an in-depth analysis of the root causes of performance degradation in RIS model quantization and propose dual-region quantization (DRQ) and reorder-based outlier-retained quantization (RORQ) to address the quantization difficulties in visual and text encoders. Extensive experiments on three benchmarks with different bits settings (from 8 to 4 bits) demonstrates its superior performance. Importantly, we are the first PTQ method specifically designed for the RIS task, highlighting the feasibility of PTQ in RIS applications. The code is available at https://github.com/gugu511yy/PTQ4RIS. Kaiying Zhu, Xihe Qiu, Shibo Zhao, Sifan Zhou |
ICRA | 6 |
| 2025 | MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static Quantization
Jiangyong Yu, Sifan Zhou, Shuoyu Li, Shuo Wang 0001, Xing Hu 0010, Zukang Xu, Changyong Shu, Zhihang Yuan |
ACM Multimedia | 2 |
| 2025 | FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
Sifan Zhou, Jiahao Nie 0001, Yichao Cao, Xiaobo Lu |
ACM Multimedia | 1 |
| 2025 | Point4Bit: Post Training 4-bit Quantization for Point Cloud 3D DetectionabstractVoxel-based 3D object detectors have achieved remarkable performance in point cloud perception, yet their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Post-training quantization (PTQ) provides a practical means to compress models and accelerate inference; however, existing PTQ methods for point cloud detection are typically limited to INT8 and lack support for lower-bit formats such as INT4, which restricts their deployment potential. In this paper, we present Point4bit, the first general 4-bit PTQ framework tailored for voxel-based 3D object detectors. To tackle challenges in low-bit quantization, we propose two key techniques: (1) Foreground-aware Piecewise Activation Quantization (FA-PAQ), which leverages foreground structural cues to improve the quantization of sparse activations; and (2) Gradient-guided Key Weight Quantization (G-KWQ), which preserves task-critical weights through gradient-based analysis to reduce quantization-induced degradation. Extensive experiments demonstrate that Point4bit achieves INT4 quantization with minimal accuracy loss with less than 1.5\% accuracy drop. Moreover, we validate its generalization ability on point cloud classification and segmentation tasks, demonstrating broad applicability. Our method further advances the bit-width limitation of point cloud quantization to 4 bits, demonstrating strong potential for efficient deployment on resource-constrained edge devices. Jianyu Wang 0003, Yu Wang 0174, Shengjie Zhao 0001, Sifan Zhou |
NeurIPS | 4 |
| 2025 | A depth-aware geometric fusion based view synthesis method for sparse RGB-D input
Xiaobo Lu, Sifan Zhou, Patrick Chiang 0001 |
Expert Syst. Appl. | 3 |
| 2025 | P2P: Part-to-Part Motion Cues Guide a Strong Tracking Framework for LiDAR Point Clouds
Jiahao Nie 0001, Sifan Zhou, Xueyi Zhou, Dong-Kyu Chae, Zhiwei He 0001 |
Int. J. Comput. Vis. | 3 |
| 2025 | PINR: A physics-integrated neural representation for dynamic fluid scenes
Sifan Zhou, Xiaobo Lu, Jian Qian |
Neurocomputing | 2 |
| 2025 | Depth-free view synthesis from diffusion models for monocular 3D detector in autonomous driving
Yuguang Shi, Sifan Zhou, Xiaobo Lu |
Multim. Syst. | 2 |
| 2025 | Rethinking iterative stereo matching from a diffusion bridge model perspective
Yuguang Shi, Sifan Zhou, Xiaobo Lu |
Pattern Recognit. | 2 |
| 2025 | An Adaptive Beam-Steering dToF LiDAR System Using Addressable Multi-Channel VCSEL Transmitter, 128 × 80 SPAD Sensor, and ML-Based Edge-Computing Object DetectionabstractIn this work, a solid-state direct time-of-flight (dToF) and adaptive beam-steering Light Detection and Ranging (LiDAR) system is proposed for machine learning (ML) based object detection. To leverage the capabilities of software and hardware, a co-optimization design from a neural network based algorithm to the architecture of transmitter, receiver and optical components is realized. Firstly, an object detection neural network is proposed for the depth-only input algorithm, which indicates the Region of Interest (ROI) in the illuminating field and gives hints of opened scan channels in the next two frames to decrease the total cost of the laser driver and sensor array. Next, the proposed network utilizes the Cross-Stage-Patrial (CSP) block to replace the residual structure in the backbone to achieve a lightweight performance and is implemented on the NVIDIA-Jetson to verify the system-level adaptive beam steering feature. To realize the smart working mode, a customized multi-channel and addressable TX is designed for adaptive and optical control to save power consumption and extend the ranging distance. At the same time, a 128×80 resolution RX which consists of Single-Photon Avalanche Diodes (SPADs) and column-wise Time-to-Digital Converter (TDC) is incorporated to capture the returned photons for combining sub-regions into an entire depth map. Next, to customize the specific scanning mechanism, for the optical setup, a cylindrical lens array is designed to reshape the laser beam, which matches the pattern of the transmitter to illuminate different targeted objects. Both the laser driver chip and the sensor chip with a 128×80 SPAD array are fabricated in the 180-nm Bipolar-CMOS-DMOS (BCD) process. Finally, the laser driver chip realizes the power of 5 W with an adjustable pulse width of 1.5 ns and the SPAD array integrates the depth accuracy of 5 cm at 15 m. Due to that the neural network realizes an accuracy up to 0.8, a low-power solid-state LiDAR prototype with adaptive beam steering is demonstrated. Yifan Wu 0009, Sifan Zhou, Lei Wang 0187, Jier Wang, Yuan Li 0074, Rui Bai 0001, Xuefeng Chen 0004, Yuanjin Zheng, Patrick Chiang 0001, Shenglong Zhuo, Lei Qiu 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Sub-SA: Strengthen In-Context Learning via Submodular Selective AnnotationabstractIn-context learning (ICL) leverages in-context examples as prompts for the predictions of Large Language Models (LLMs). These prompts play a crucial role in achieving strong performance. However, the selection of suitable prompts from a large pool of labeled examples often entails significant annotation costs. To address this challenge, we propose Sub-SA (Submodular Selective Annotation), a submodule-based selective annotation method. The aim of Sub-SA is to reduce annotation costs while improving the quality of in-context examples and minimizing the time consumption of the selection process. In Sub-SA, we design a submodular function that facilitates effective subset selection for annotation and demonstrates the characteristics of monotonically and submodularity from the theoretical perspective. Specifically, we propose RPR (Reward and Penalty Regularization) to better balance the diversity and representativeness of the unlabeled dataset attributed to a reward term and a penalty term, respectively. Consequently, the selection for annotations can be effectively addressed with a simple yet effective greedy search algorithm based on the submodular function. Finally, we apply the similarity prompt retrieval to get the examples for ICL. Compared to existing selective annotation approaches, Sub-SA offers two main advantages. (1.) Sub-SA operates in an end-to-end, unsupervised manner, and significantly reduces the time consumption of the selection process (from hours-level to millisecond-level). (2.) Sub-SA enables a better balance between data diversity and representativeness and obtains state-of-the-art performance. Meanwhile, the theoretical support guarantees their reliability and scalability in practical scenarios. Extensive experiments conducted on diverse models and datasets demonstrate the superiority of Sub-SA over previous methods, achieving millisecond(ms)-level time selection and remarkable performance gains. The efficiency and effectiveness of Sub-SA make it highly suitable for real-world ICL scenarios. Our codes are available at https://github.com/JamesQian11/SubSA Jian Qian, Sifan Zhou, Ruizhi Hun, Patrick Chiang 0001 |
ECAI | 3 |
| 2024 | LK-UNet: Large Kernel Design for 3D Medical Image SegmentationabstractRecently, the medical image segmentation have made rapid progress. Specifically, the precision of medical image segmentation play a pivotal role in the realm of disease diagnosis and treatment. Therefore, it is vital to improve the segmentation performance. Generally, Transformer-based methods exhibit superior performance compared to CNN-based methods on 3D medical image segmentation tasks due to their inherent capability to capture global-aware context. However, the existing transformer-based models are still unsatisfactory in accuracy. In this paper, we propose a novel fully convolution architecture for medical image segmentation tasks, called LK-UNet. Specifically, the key of LK-UNet lies in its incorporation of a large kernel module, which can achieve comparable receptive fields to transformer module. Besides, we introduce Depth-wise Convolution Layer (DCL) and Point-wise Convolution Layer (PCL) to substitute the vanilla convolution layer to reduce the number of model parameters and enhance the feature representation. Extensive experiment shows that our method achieves state-of-the-art performance on the public BTCV dataset, which even outperforms hybrid transformer-based and CNN-based networks. Jiang Shang, Sifan Zhou |
ICASSP | 2 |
| 2024 | LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object DetectionabstractDue to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted in 2D vision tasks. However, applying it directly to 3D lidar-based tasks inevitably leads to performance degradation. As a remedy, we propose an effective PTQ method called LiDAR-PTQ, which is particularly curated for 3D lidar detection (both SPConv-based and SPConv-free). Our LiDAR-PTQ features three main components, (1) a sparsity-based calibration method to determine the initialization of quantization parameters, (2) an adaptive rounding-to-nearest operation to minimize the layerwise reconstruction error, (3) a Task-guided Global Positive Loss (TGPL) to reduce the disparity between the final predictions before and after quantization. Extensive experiments demonstrate that our LiDAR-PTQ can achieve state-of-the-art quantization performance when applied to CenterPoint (both Pillar-based and Voxel-based). To our knowledge, for the very first time in lidar-based 3D detection tasks, the PTQ INT8 model's accuracy is almost the same as the FP32 model while enjoying 3X inference speedup. Moreover, our LiDAR-PTQ is cost-effective being 6X faster than the quantization-aware training method. The code will be released. Sifan Zhou, Liang Li 0003, Xinyu Zhang 0015, Bo Zhang 0046, Shipeng Bai, Xiaobo Lu, Xiangxiang Chu |
ICLR | 1 |
| 2023 | Real-Time 3D Single Object Tracking With TransformerabstractLiDAR-based 3D single object tracking is a challenging issue in robotics and autonomous driving. Currently, existing approaches usually suffer from the problem that objects at long distance often have very sparse or partially-occluded point clouds, which makes the features extracted by the model ambiguous. Ambiguous features will make it hard to locate the target object and finally lead to bad tracking results. To solve this problem, we utilize the powerful Transformer architecture and propose aPoint-Track-Transformer (PTT)module for point cloud-based 3D single object tracking task. Specifically, PTT module generates fine-tuned attention features by computing attention weights, which guides the tracker focusing on the important features of the target and improves the tracking ability in complex scenarios. To evaluate our PTT module, we embed PTT into the dominant method and construct a novel 3D SOT tracker named PTT-Net. In PTT-Net, we embed PTT into the voting stage and proposal generation stage, respectively. PTT module in the voting stage could model the interactions among point patches, which learns context-dependent features. Meanwhile, PTT module in the proposal generation stage could capture the contextual information between object and background. We evaluate our PTT-Net on KITTI and NuScenes datasets. Experimental results demonstrate the effectiveness of PTT module and the superiority of PTT-Net, which surpasses the baseline by a noticeable margin,$\sim$10% in the Car category. Meanwhile, our method also has a significant performance improvement in sparse scenarios. In general, the combination of transformer and tracking pipeline enables our PTT-Net to achieve state-of-the-art performance on both two datasets. Additionally, PTT-Net could run in real-time at 40FPS on NVIDIA 1080Ti GPU. Our code is open-sourced for the research community athttps://github.com/shanjiayao/PTT. Jiayao Shan, Sifan Zhou, Yubo Cui, Zheng Fang 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | 3D Object Tracking with Transformer
Yubo Cui, Zheng Fang 0001, Jiayao Shan, Zuoxu Gu, Sifan Zhou |
BMVC | 5 |
| 2021 | PTT: Point-Track-Transformer Module for 3D Single Object Tracking in Point Cloudsabstract3D single object tracking is a key issue for robotics. In this paper, we propose a transformer module called Point-Track-Transformer (PTT) for point cloud-based 3D single object tracking. PTT module contains three blocks for feature embedding, position encoding, and self-attention feature computation. Feature embedding aims to place features closer in the embedding space if they have similar semantic information. Position encoding is used to encode coordinates of point clouds into high dimension distinguishable features. Self-attention generates refined attention features by computing attention weights. Besides, we embed the PTT module into the open-source state-of-the-art method P2B to construct PTT-Net. Experiments on the KITTI dataset reveal that our PTT-Net surpasses the state-of-the-art by a noticeable margin $\left( {\sim 10\% } \right)$. Additionally, PTT-Net could achieve real-time performance (~40FPS) on NVIDIA 1080Ti GPU. Our code is open-sourced for the robotics community at https://github.com/shanjiayao/PTT. Jiayao Shan, Sifan Zhou, Zheng Fang 0001, Yubo Cui |
IROS | 2 |