Zhiwei Dong

dblp:228/1648 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real-time multi-constraint control of autonomous flexible endoscope robots via finite-time neural optimization
Yisen Huang, Weibing Li, Jixiu Li, Zhiwei Dong, Weiping Ding, Philip W. Y. Chiu, Zheng Li 0012
Eng. Appl. Artif. Intell.5
2025 AtomNet: Designing Tiny Models from Operators Under Extreme MCU Constraints
abstract
Tiny machine learning (TinyML) has attracted heightened attention for its ability to provide low-cost and instantaneous performance on edge devices. Particularly, the commonly used microcontroller unit (MCU) imposes extreme constraints on peak memory (SRAM) and storage (Flash). Existing TinyML methods often rely on a customized and hard-to-obtain inference libraries, as well as necessitate a time-consuming search for a deployable architecture using advanced Neural Architecture Search (NAS) algorithms. To solve these problems, we fully exploit the resources on MCU and deduce hardware-oriented guidelines for designing models under extreme MCU constraints. In detail, we delve into thorough information about the atom operators by collecting the runtime data of Flash, SRAM, and latency to build a dataset named AtomDB. Based on AtomDB, several critical operator guidelines are established to fully utilize limited Flash and SRAM, while minimizing latency. By transferring the guidelines to analyze blocks, we propose a hybrid pattern that organizes appropriate blocks at different network stages to form the AtomNet, a more hardware-oriented architecture, to handle the former SRAM bottleneck and the latter Flash bottleneck. Extensive experiments demonstrate the effectiveness of the exploitation of the hardware characteristics. Remarkably, AtomNet pioneeringly achieve 3.5% accuracy enhancement and more than 15% latency reduction on 320KB MCU using readily available official inference libraries for ImageNet tasks, surpassing the current state-of-the-art method.
Zhiwei Dong, Mingzhu Shen, Shihao Bai, Xiuying Wei, Jinyang Guo 0002, Ruihao Gong, Song-Lu Chen, Xianglong Liu 0001, Xu-Cheng Yin
AAAI1
2025 AGTCNet: Hybrid Network Based on AGT and Curvature Information for Skin Lesion Detection
Zhiwei Dong, Genji Yuan, Jinjiang Li 0001
CVM (1)1
2025 Leveraging SD Map to Augment HD Map-based Trajectory Prediction
abstract
Trajectory prediction models in real-world autonomous driving often rely on online High-Definition (HD) maps to understand road environments, but online HD maps suffer from perception errors and feature redundancy, which hinder the performance of these models. To address this issue, we introduce a framework, termed SD map-Augmented Trajectory Prediction (SATP), which leverages Standard-Definition (SD) maps to enhance HD map-based trajectory prediction models. First, we propose an SD-HD fusion approach to leverage SD maps across the diverse range of HD map-based trajectory prediction models. Second, we design a novel AlignNet to align the SD map with the HD map, further improving the effectiveness of SD maps. Experiments on real-world autonomous driving benchmarks demonstrate that SATP not only improves the performance of HD map-based trajectory prediction up to 25% in real-world scenarios using online HD maps but also brings benefits in ideal scenarios with ground-truth HD maps.
Zhiwei Dong, Guobin Tang
CVPR1
2025 Tool Playgrounds: A Comprehensive and Analyzable Benchmark for LLM Tool Invocation
abstract
The rapid advancement of large language models (LLMs) has paved the way for their use in solving real-world problems, which in turn has significantly driven the development of tool-assisted LLMs. This progress necessitates thorough evaluation methods. However, existing benchmarks typically only provide end-to-end scores but lack in-depth analysis and often suffer from issues such as instability. To address this gap, we have meticulously designed the Tool Playgrounds framework, a comprehensive, analyzable, and extensible benchmark. This framework evaluates boundary dimensions such as parameter missing interaction, parameter correction, tool failover, and leveraging internal knowledge. Our findings indicate that even the most advanced commercial models frequently overlook these essential aspects and face challenges in managing complex tool usage. To foster further research and development, we have made our code, dataset, and leaderboard publicly available on https://github.com/zhiwei-dong/ToolPlaygrounds.
Zhiwei Dong, Ruihao Gong, Yang Yong, Yongqiang Yao, Song-Lu Chen, Xu-Cheng Yin
ICASSP1
2025 Data-Free Post-Training Quantization with Block-wise Enhanced Sample Generation
abstract
Data-free quantization is known for quantizing a pre-trained deep neural network without access to any training data, which applies to many real-world scenarios in that the training data is unavailable due to security, user privacy, or proprietary concerns. Most of the existing data-free quantization methods adopt a generator-quantization framework, which generator network to synthesize fake samples and Quantization-Aware Training (QAT) to quantize model. While the combination of the generator network and QAT can result in good accuracy for quantized models, the diversity of generated samples is lacking and quantizing a single model may take over 10 hours, which contrasts with Post-Training Quantization (PTQ)’s time-saving potential but poor accuracy. In order to address these issues, we have made improvements to the data generation and quantization process. In detail, 1) We propose Generator Exploration Enhancement (GEE) for utilizing the batch normalization statistics and adversarial sample exploration to enhance the quality and diversity of synthetic samples; 2) We introduce Block-wise Sample Generation (BSG) to collectively optimize individual blocks and the generator, leveraging PTQ as a foundation to boost workflow efficiency. Experiment results show that our proposed method improves both the performance and the efficiency of the data-free quantization compared to that of existing methods. Significantly, BSG achieves an 18% accuracy improvement and reduces quantization time by over 50% for 3-bit ResNet-18 in ImageNet tasks, surpassing the current state-of-the-art QAT method.
Ruiyao Zhang, Zhiwei Dong, Shutong Ti, Song-Lu Chen, Xu-Cheng Yin
ICASSP2
2025 BEV-MMC: Bird's-Eye-View-Based Multimodal Compression for Enhanced Visual Recognition
abstract
Today, visual data compression is essential not only for human perception but also for machine analysis. With the growing volumes of visual data, challenges in data storage and transmission demand advanced compression techniques for downstream tasks. Most current works focus on either single-modal compression or multimodal compression without unifying the modalities effectively. In this study, we introduce a bird’s-eye-view (BEV)-based multimodal compression framework that jointly compresses camera images and LiDAR point clouds to enhance the accuracy of downstream tasks. A core innovation of our approach is the inception-triggered soft element-wise mask, which leverages multi-scale feature extraction to effectively capture richer information and reduces spatial-channel redundancy. Compared to the existing state-of-the-art methods, our proposed framework has achieved BD-rate gains of 88.12% for 3D object detection and 98.71% for BEV map segmentation compared to the referenced BEV-based compression framework and 95.43% for 3D object detection compared to the referenced multimodal compression framework.
Zhiwei Dong
ICME1
2025 MASTP: More Accurate and Safer Trajectory Prediction Based on a Trajectory Planner
abstract
Accurately predicting future vehicle trajectories that comply with the traffic scenario is crucial for the safety of autonomous driving systems. In recent years, deep learning-based methods have become dominant in vehicle motion prediction by effectively capturing complex interactions within traffic scenarios. However, existing models often generate infeasible trajectory predictions in certain scenarios, such as driving in the wrong lane. In this paper, we propose MASTP, an innovative trajectory prediction method based on a trajectory planner, designed to generate more accurate and safer predictions. We introduce a method for extracting feasible lane centerlines and present a modal adaptive trajectory planner that incorporates speed estimation, guided by these centerlines. This planner explicitly integrates traffic scenario constraints, ego-vehicle kinematic constraints, and interaction constraints with other vehicles, providing a model-agnostic planner fusion approach. Our approach significantly reduces infeasible trajectory predictions, enhancing scenario compliance and improving prediction accuracy. Our method offers stronger interpretability, can be easily integrated into various existing prediction models, and generates more precise and safer multimodal vehicle trajectory predictions with reduced reliance on training data.
Yaode Wang, Zhiwei Dong, Qian Hou
IJCNN4
2025 Data-Bootstrapped, Physics-Informed Framework for Object Rearrangement
abstract
Object rearrangement, which involves arranging objects step-by-step to achieve tidy states, is critical in robotic applications. Progress in this area is often constrained by issues such as high-cost data collection and physically infeasible trajectory prediction. To address these challenges, we propose the Data-Bootstrapped, Physics-Informed Rearrangement (DPR) framework, which leverages a transformer for sequential decision making. Specifically, DPR integrates Enhanced Data Generation with a Physics Reward Feedback Transformer. Enhanced Data Generation consists of Random Trajectory Reverse for producing high-quality training data and Bootstrapped Trajectory Synthesis, which leverages the transformer’s sequence modeling to diversify training trajectories. To ensure the feasibility of the generated trajectories and to improve the transformer’s performance, we incorporate a Physical Reward Feedback mechanism into the transformer. Experiments on ball and room rearrangement tasks show that DPR significantly outperforms existing methods in terms of both efficiency and effectiveness. Code will be released soon.
Zhiwei Dong
IROS2
2025 SpectMamba: Remote sensing change detection network integrating frequency and visual state space model
Zhiwei Dong, Dapeng Cheng, Jinjiang Li 0001
Expert Syst. Appl.1
2025 Few-Shot Counting with Multi-Scale Vision Transformers and Attention Mechanisms
abstract
Object counting is a fundamental task in computer vision, with critical applications in areas such as crowd monitoring and ecological conservation. Traditional methods typically rely on large-scale annotated datasets, which are costly and time-consuming to obtain. Few-shot object counting has emerged as a promising solution, enabling accurate counting with minimal annotated samples. However, in real-world scenarios, objects often exhibit significant scale variations due to factors such as view distortion, varying shooting distances, and inherent size differences. Existing few-shot methods usually struggle to address this challenge effectively. To address these issues, we propose a Scale-Aware Vision Transformer (SAViT) framework. Specifically, we design a multi-scale dilated convolution module in SAViT, which can adaptively adjust convolution kernel sampling rates to handle objects of varying sizes. Additionally, we incorporate a global channel attention mechanism to strengthen the model’s ability to capture robust feature representations, thereby improving detection accuracy. For practical usability, we integrate the Segment Anything Model (SAM) to create an exemplar box selection module, simplifying the process by allowing users to generate precise exemplar boxes with a single line drawn on the target object. Extensive experiments on the FSC-147 dataset demonstrate the effectiveness of our approach, achieving a Mean Absolute Error (MAE) of 8.92 and a Root Mean Squared Error (RMSE) of 31.26. Compared to the state-of-the-art method, CACViT, our model reduces MAE by 0.21 (2.30% improvement) and RMSE by 17.7 (36.15% improvement). Our approach not only provides an effective solution for few-shot object counting but also provides a new practical paradigm for extending few-shot learning to complex vision tasks requiring multi-scale reasoning. The code of our paper is available at https://github.com/BlouseDong/SAViT .
Xiaopan Chen, Zhiwei Dong, Xiaoke Zhu, Fan Zhang 0028, Caihong Yuan
Int. J. Pattern Recognit. Artif. Intell.2
2025 Presegmentation of Point Cloud Based on Centroid Matrix of Streak Tube Imaging LiDAR
abstract
This letter presents a presegmentation method for the streak tube imaging LiDAR (STIL) before generating the point cloud. This method uses the spatial relationships between scan lines to construct a multilayer matrix data structure without more precise positional and posture data. The elements of this matrix not only correspond to the point cloud but also contain the proximity relationships between points. After the multilayer matrix has been segmented by the two-pass algorithm, each element will be assigned a distinct label, which will be incorporated into the 3-D reconstruction of the point cloud. In terms of segmentation results, this preprocessing is found to be consistent with the traditional point cloud analysis techniques.
Chaowei Dong, Zhiwei Dong, Rongwei Fan, Pengfei Hao, Deying Chen
IEEE Geosci. Remote. Sens. Lett.4
2024 QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models
abstract
Large Language Models (LLMs) have demonstrated unparalleled efficacy in natural language processing. However, their high computational demands and memory overheads hinder their broad deployment. To address this, two quantization strategies emerge, including Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ). For LLMs, the billions of parameters make the QAT impractical due to the prohibitive training cost and thus PTQ becomes more prevalent. In existing studies, activation outliers in particular channels are identified as the biggest challenge to PTQ accuracy. They propose to transform the magnitudes from activations to weights, which however offers limited alleviation or suffers from unstable gradients, resulting in a severe performance drop at low-bitwidth. In this paper, we propose QLLM, an accurate and efficient low-bitwidth PTQ method designed for LLMs. QLLM introduces an adaptive channel reassembly technique that reallocates the magnitude of outliers to other channels, thereby mitigating their impact on the quantization range. This is achieved by channel disassembly and channel assembly, which first breaks down the outlier channels into several sub-channels to ensure a more balanced distribution of activation magnitudes. Then similar channels are merged to maintain the original channel number for efficiency. Additionally, an adaptive strategy is designed to autonomously determine the optimal number of sub-channels for channel disassembly. To further compensate for the performance loss caused by quantization, we propose an efficient tuning method that only learns a small number of low-rank weights while freezing the pre-trained quantized model. After training, these low-rank parameters can be fused into the frozen weights without affecting inference. Extensive experiments on LLaMA-1 and LLaMA-2 show that QLLM is able to obtain accurate quantized models efficiently. For example, QLLM quantizes the 4-bit LLaMA-2-70B within 10 hours on a single A100-80G GPU, outperforming the previous state-of-the-art method by 7.89% on the average accuracy across five zero-shot tasks. Code is available at [ZIP Lab](https://github.com/ziplab/QLLM) and [ModelTC](https://github.com/ModelTC/QLLM).
Jing Liu 0048, Ruihao Gong, Xiuying Wei, Zhiwei Dong, Jianfei Cai 0001, Bohan Zhuang
ICLR4
2024 BézierFormer: A Unified Architecture for 2D and 3D Lane Detection
abstract
Lane detection have made significant progress in recent years, but there is not a unified architecture for its two sub-tasks: 2D lane detection and 3D lane detection. To fill this gap, we introduce BézierFormer, a unified 2D and 3D lane detection architecture based on Bézier curve lane representation. BézierFormer formulate queries as Bézier control points and incorporate a novel Bézier curve attention mechanism. This attention mechanism enables comprehensive and accurate feature extraction for slender lane curves via sampling and fusing multiple reference points on each curve. In addition, we propose a novel Chamfer IoU-based loss which is more suitable for the Bézier control points regression. The state-of-the-art performance of BézierFormer on widely-used 2D and 3D lane detection benchmarks verifies its effectiveness and suggests the worthiness of further exploration.
Zhiwei Dong, Xiya Cao, Caifa Zhou, Qiangbo Liu
ICME1
2024 HQOD: Harmonious Quantization for Object Detection
abstract
Task inharmony problem commonly occurs in modern object detectors, leading to inconsistent qualities between classification and regression tasks. The predicted boxes with high classification scores but poor localization positions or low classification scores but accurate localization positions will worsen the performance of detectors after Non-Maximum Suppression. Furthermore, when object detectors collaborate with Quantization- Aware Training (QAT), we observe that the task inharmony problem will be further exacerbated, which is considered one of the main causes of the performance degradation of quantized detectors. To tackle this issue, we propose the Harmonious Quantization for Object Detection (HQOD) framework, which consists of two components. Firstly, we propose a task-correlated loss to encourage detectors to focus on improving samples with lower task harmony quality during QAT. Secondly, a harmonious Intersection over Union (IoU) loss is incorporated to balance the optimization of the regression branch across different IoU levels. The proposed HQOD can be easily integrated into different QAT algorithms and detectors. Remarkably, on the MS COCO dataset, our 4-bit ATSS with ResNet-50 backbone achieves a state-of-the- art mAP of 39.6%, even surpassing the full-precision one. Codes are available at https://github.com/Menace-Dragon/VP-QOD.
Zhiwei Dong, Song-Lu Chen, Ruiyao Zhang, Shutong Ti, Feng Chen 0040, Xu-Cheng Yin
ICME2
2024 Efficient Long-Range Context Modeling for Motion Forecasting with State Space Models
Zhiwei Dong, Wei Li 0304
ICPR (26)1
2024 Transformer-based multi-attention hybrid networks for skin lesion segmentation
Zhiwei Dong, Jinjiang Li 0001, Zhen Hua
Expert Syst. Appl.1
2024 Diffusion model-based text-guided enhancement network for medical image segmentation
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
Expert Syst. Appl.1
2024 ConMamba: CNN and SSM High-Performance Hybrid Network for Remote Sensing Change Detection
abstract
Accurate remote sensing change detection (RSCD) tasks rely on comprehensively processing multiscale information from local details to effectively integrate global dependencies. Hybrid models based on convolutional neural networks (CNNs) and Transformers have become mainstream approaches in RSCD due to their complementary advantages in local feature extraction and long-term dependency modeling. However, the Transformer faces application bottlenecks due to the high secondary complexity of its attention mechanism. In recent years, state-space models (SSMs) with efficient hardware-aware design, represented by Mamba, have gained widespread attention for their excellent performance in long-series modeling and have demonstrated significant advantages in terms of improved accuracy, reduced memory consumption, and reduced computational cost. Based on the high match between the efficiency of SSM in long sequence data processing and the requirements of the RSCD task, this study explores the potential of its application in the RSCD task. However, relying on SSM alone is insufficient in recognizing fine-grained features in remote sensing images. To this end, we propose a novel hybrid architecture, ConMamba, which constructs a high-performance hybrid encoder (CS-Hybridizer) by realizing the deep integration of the CNN and SSM through the feature interaction module (FIM). In addition, we introduce the spatial integration module (SIM) in the feature reconstruction stage to further enhance the model’s ability to integrate complex contextual information. Extensive experimental results on three publicly available RSCD datasets show that ConMamba significantly outperforms existing techniques in several performance metrics, validating the effectiveness and foresight of the hybrid architecture based on the CNN and SSM in RSCD.
Zhiwei Dong, Genji Yuan, Zhen Hua, Jinjiang Li 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 Adaptive Rounding Compensation for Post-training Quantization
Jinhui Lin, Song-Lu Chen, Ruiyao Zhang, Zhiwei Dong, Feng Chen 0040, Xu-Cheng Yin
ICONIP (5)6
2020 CentripetalNet: Pursuing High-Quality Keypoint Pairs for Object Detection
abstract
Keypoint-based detectors have achieved pretty-well performance. However, incorrect keypoint matching is still widespread and greatly affects the performance of the detector. In this paper, we propose CentripetalNet which uses centripetal shift to pair corner keypoints from the same instance. CentripetalNet predicts the position and the centripetal shift of the corner points and matches corners whose shifted results are aligned. Combining position information, our approach matches corner points more accurately than the conventional embedding approaches do. Corner pooling extracts information inside the bounding boxes onto the border. To make this information more aware at the corners, we design a cross-star deformable convolution network to conduct feature adaption. Furthermore, we explore instance segmentation on anchor-free detectors by equipping our CentripetalNet with a mask prediction module. On COCO test-dev, our CentripetalNet not only outperforms all existing anchor-free detectors with an AP of 48.0% but also achieves comparable performance to the state-of-the-art instance segmentation approaches with a 40.2% Mask AP. Code is available at https: //github.com/KiveeDong/CentripetalNet.
Zhiwei Dong, Guoxuan Li, Yue Liao, Fei Wang 0032, Pengju Ren, Chen Qian 0006
CVPR1
2020 A Lightweight Multi-Label Segmentation Network for Mobile Iris Biometrics
abstract
This paper proposes a novel, lightweight deep convolutional neural network specifically designed for iris segmentation of noisy images acquired by mobile devices. Unlike previous studies, which only focused on improving the accuracy of segmentation mask using the popular CNN technology, our method is a complete end-to-end iris segmentation solution, i.e., segmentation mask and parameterized pupillary and limbic boundaries of the iris are obtained simultaneously, which further enables CNN-based iris segmentation to be applied in any regular iris recognition systems. By introducing an intermediate pictorial boundary representation, predictions of iris boundaries and segmentation mask have collectively formed a multi-label semantic segmentation problem, which could be well solved by a carefully adapted stacked hourglass network. Experimental results show that our method achieves competitive or state-of-the-art performance in both iris segmentation and localization on two challenging mobile iris databases.
Caiyong Wang, Yunlong Wang 0003, Boqiang Xu, Yong He 0009, Zhiwei Dong, Zhenan Sun
ICASSP5
2018 An Extracting Method of Symmetry Plane from Head CT images for Surgery Based on OBB and Image Mutual Information
Wenjun Tan, Ying Kang, Zhiwei Dong, Jinzhu Yang, Lisheng Xu, Dazhe Zhao
BIBM3
2018 Spiking Locality-Sensitive Hash: Spiking Computation with Phase Encoding Method
abstract
A novel similarity search method, named spiking locality sensitive hash (SLSH), a forward spiking neuron network(SNN) is proposed in this paper. The SLSH architecture is composed of successively connected encoding and fully connected layer. We optimize phase encoding to maximize the difference between corresponding pixels of any two different images. Then we test the performance of the encoding method and the SLSH model on graphic datasets. Experimental results prove that improved phase encoding method based on the difference exhibits the accuracy of 100%, 100% and 92%, which has superiority over previous phase encoding whose accuracies are 93%, 78% and 55% when the noise level is 5%, 20% and 40% respectively. Furthermore, experiments demonstrate that SLSH method is more capable than the traditional Locality-Sensitive Hash(LSH) and the FLY algorithm published in SCIENCE in similarity search. The mean average precision of SLSH is twice of FLY algorithm when the hash length is 5. In addition, the SLSH achieves a good recognition performance even under the influence of noise for MNIST, SVHN and SIFT datasets.
Ziru Wang, Zhiwei Dong, Nanning Zheng 0001, Pengju Ren
IJCNN3