Erkang Cheng

dblp:41/4416 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
23since 2021 · last 2026
0000-0001-7941-6911ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 10 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving
abstract
Unlike discriminative approaches in autonomous driving that predict a fixed set of candidate trajectories of the ego vehicle, generative methods, such as diffusion models, learn the underlying distribution of future motion, enabling more flexible trajectory prediction. However, since these methods typically rely on denoising human-craft trajectory anchors or random noise, there remains significant room for improvement. In this paper, we propose DiffRefiner, a novel two-stage trajectory prediction framework. The first stage employs a transformer-based Proposal Decoder to generate coarse trajectory predictions by regressing from sensor inputs using predefined trajectory anchors. The second stage applies a Diffusion Refiner that iteratively denoises and refines these initial predictions. In this way, we enhance the performance of diffusion-based planning by incorporating a discriminative trajectory proposal module, which provides strong guidance for the generative refinement process. Furthermore, we design a fine-grained denoising decoder to enhance scene compliance, enabling more accurate trajectory prediction through enhanced alignment with the surrounding environment. Experimental results demonstrate that DiffRefiner achieves state-of-the-art performance, attaining 87.4 EPDMS on NAVSIM v2, and 87.1 DS along with 71.4 SR on Bench2Drive, thereby setting new records on both public benchmarks. The effectiveness of each component is validated via ablation studies as well.
Liuhan Yin, Runkun Ju, Guodong Guo, Erkang Cheng
AAAI4
2026 Consistency-Guided Diffusion Model for Image Generation in Dual-Front-Camera System with Wide and Narrow Angles
Boya Zhou, Zhiyong Sun 0002, Erkang Cheng
IV5
2025 HiP-AD: Hierarchical and Multi-Granularity Planning with Deformable Attention for Autonomous Driving in a Single Decoder
abstract
Although end-to-end autonomous driving (E2E-AD) technologies have made significant progress in recent years, there remains an unsatisfactory performance on closed-loop evaluation. The potential of leveraging planning in query design and interaction has not yet been fully explored. In this paper, we introduce a multi-granularity planning query representation that integrates heterogeneous waypoints, including spatial, temporal, and driving-style waypoints across various sampling patterns. It provides additional supervision for trajectory prediction, enhancing precise closed-loop control for the ego vehicle. Additionally, we explicitly utilize the geometric properties of planning trajectories to effectively retrieve relevant image features based on physical locations using deformable attention. By combining these strategies, we propose a novel end-to-end autonomous driving framework, termed HiP-AD, which simultaneously performs perception, prediction, and planning within a unified decoder. HiP-AD enables comprehensive interaction by allowing planning queries to iteratively interact with perception queries in the BEV space while dynamically extracting image features from perspective views. Experiments demonstrate that HiP-AD outperforms all existing end-to-end autonomous driving methods on the closed-loop benchmark Bench2Drive and achieves competitive performance on the real-world dataset nuScenes.
Yingqi Tang, Zhaotie Meng, Erkang Cheng
ICCV4
2025 Asynchronous Rectification-Based Fast Local Imaging and Estimation Scheme for High-Speed Rotating States Observation of MM
abstract
Magnetic microrobots (MMs) have emerged as promising tools for targeted therapies, including non-invasive in vivo treatments and precise drug delivery, owing to their untethered controllability and biocompatibility. Current actuation strategies for MMs primarily rely on two magnetic field (MF) generation approaches: gradient-based and rotational methods. Unlike the gradient method, rotational actuation enables efficient manipulation of MMs under significantly weaker magnetic fields. To fully leverage the potential of rotationally driven MMs, a comprehensive understanding of their fundamental spin motility is essential. Achieving accurate characterization of these MMs necessitates the development of an MF generation system equipped with rapid motion-tracking and broad-range measurement capabilities. This study proposes a high-speed rotating states observation scheme by developing a tracking-based optimal local imaging and estimation scheme, simultaneously meeting the broad-range observation capability and the high imaging speed requirement. Specifically, the CSR-DCF tracking method is adopted to detect the MM’s location, and based on this, the observation system adjusts the imaging region optimally. An estimation scheme based on the asynchronous rectification method is derived to measure the MM rotating states consistently using measured MF data and local optical images of the target. Experimental studies are carried out to validate the effectiveness of the proposed scheme.
Zhiyong Sun 0002, Yu Cheng 0006, Gengliang Chen, Erkang Cheng
IROS4
2025 Edge-Aware Token Halting for Efficient and Accurate Medical Image Segmentation
Yuhao Guo, Erkang Cheng
MICCAI (4)4
2025 Cross Image Feature Perturbation with Pseudo Label Fusion for Semi-Supervised Medical Image Segmentation
Minxia Xu, Weida Hu, Jinshui Miao, Erkang Cheng
WACV6
2025 CurveFormer++: 3D Lane Detection by Curve Propagation With Temporal Curve Queries and Attention
abstract
In autonomous driving, accurate 3D lane detection using monocular cameras is important for downstream tasks. Recent CNN and Transformer approaches usually apply a two-stage model design. The first stage transforms the image feature from a front image into a bird’s-eye-view (BEV) representation. Subsequently, a sub-network processes the BEV feature to generate the 3D detection results. However, these approaches heavily rely on a challenging image feature transformation module from a perspective view to a BEV representation. In our work, we present CurveFormer++, a single-stage Transformer-based method that does not require the view transform module and directly infers 3D lane results from the perspective image features. Specifically, our approach models the 3D lane detection task as a curve propagation problem, where each lane is represented by a curve query with a dynamic and ordered anchor point set. By employing a Transformer decoder, the model can iteratively refine the 3D lane results. A curve cross-attention module is introduced to calculate similarities between image features and curve queries. To handle varying lane lengths, we employ context sampling and anchor point restriction techniques to compute more relevant image features. Furthermore, we apply a temporal fusion module that incorporates selected informative sparse curve queries and their corresponding anchor point sets to leverage historical information. In the experiments, we evaluate our approach on two publicly real-world datasets. The results demonstrate that our method provides outstanding performance compared with both CNN and Transformer based methods. We also conduct ablation studies to analyze the impact of each component.
Yifeng Bai, Zhirong Chen, Pengpeng Liang, Erkang Cheng
IEEE Trans. Intell. Transp. Syst.5
2025 An Error-Tolerant Design of Joint Planner for a Table Tennis Robot Using Variable-Sigmoid-Based Motion Template
abstract
Efficient motion planning with the error tolerance is crucial for dynamic robotic tasks, particularly robotic table tennis. This task demands simultaneous high efficiency and error tolerance. First, the incoming ball’s high speed allows only tens of milliseconds for motion planning. Second, two types of errors, namely, ball-paddle motion uncertainty and joint limit violation, must be tolerated to ensure a high success rate of planning (SRP) and striking. This article proposes an advanced joint planning framework designed for high efficiency and error tolerance. To tolerate the error of ball-paddle motion uncertainty, this work introduces a joint classification criterion according to the joint motion characteristics. To tolerate the error of joint limit violation, based on the classification criterion, this study also develops a robust reference trajectory generation scheme, named error-tolerant-variable-sigmoid-based motion template (ETVSMT), to fully consider the motion capabilities of different joints. The ETVSMT approach utilizes the variable-sigmoid-based motion template (VSMT) as the backbone and designs its basic and remedial portions to tolerate the error of hard joint limit violation. The implementation of the ETVSMT scheme results in an average SRP of 98% and a ball-striking success rate of up to 90%, with execution time on an Intel Xeon CPU as low as 15 ms. Furthermore, the proposed method ensures the outgoing ball lands on the opposite side of the table with an average landing error of 19.62 cm and a net-passing height error of 12.68 cm. The proposed method can also benefit other dynamic tasks like human–robot interaction to improve the planning efficiency.
Yuxin Wang 0007, Zhiyong Sun 0002, Chengeng Qu, Yongle Luo, Yu Liu 0101, Bin Cai 0006, Erkang Cheng
IEEE Trans. Syst. Man Cybern. Syst.10
2024 Enhancing 3D Object Detection with 2D Detection-Guided Query Anchors
abstract
Multi-camera-based 3D object detection has made no-table progress in the past several years. However, we observe that there are cases (e.g. faraway regions) in which popular 2D object detectors are more reliable than state-of-the-art 3D detectors. In this paper, to improve the performance of query-based 3D object detectors, we present a novel query generating approach termed QAF2D, which infers 3D query anchors from 2D detection results. A 2D bounding box of an object in an image is lifted to a set of 3D anchors by associating each sampled point within the box with depth, yaw angle, and size candidates. Then, the validity of each 3D anchor is verified by comparing its projection in the image with its corresponding 2D box, and only valid anchors are kept and used to construct queries. The class information of the 2D bounding box associated with each query is also utilized to match the predicted boxes with ground truth for the set-based loss. The image feature extraction backbone is shared between the 3D detector and 2D detector by adding a small number of prompt parameters. We integrate QAF2D into three popular query-based 3D object detectors and carry out comprehensive evaluations on the nuScenes dataset. The largest improvement that QAF2D can bring about on the nuScenes validation subset is 2.3% NDS and 2.7% mAP. Code is available at https://github.com/nullmax-vision/QAF2D.
Haoxuanye Ji, Pengpeng Liang, Erkang Cheng
CVPR3
2024 SimPB: A Single Model for 2D and 3D Object Detection from Multiple Cameras
Yingqi Tang, Zhaotie Meng, Erkang Cheng
ECCV (2)4
2024 Dyna-style Model-based reinforcement learning with Model-Free Policy Optimization
Yongle Luo, Yuxin Wang 0007, Yu Liu 0101, Chengeng Qu, Erkang Cheng, Zhiyong Sun 0002
Knowl. Based Syst.7
2024 Reinforcement Learning with Decoupled State Representation for Robot Manipulations
abstract
Abstract Deep reinforcement learning has significantly advanced robot manipulations by providing an alternative solution for designing control strategies using raw images as direct inputs. While images offer additional environmental information, the end-to-end policy training manner (from image to action) requires simultaneous representation and task learning by the agent. This often necessitates a substantial number of interaction samples to achieve satisfactory policy performance. Previous works has attempted to address this challenge by learning a visual representation model that encodes the entire image into a low-dimensional vector before the policy training. However, since this vector contains both robot and object information, it inevitably introduces coupling within the state, which can mislead the policy training process. In this study, a novel method called Reinforcement Learning with Decoupled State Representation is proposed to effectively decouple robot and object information within the state representation. Experimental results demonstrate that the proposed method exhibits faster learning speed and achieves superior performance compared to previous methods across various robot manipulation tasks. Moreover, with only 3096 offline images, the proposed method successfully applies to real-world robot pushing tasks, which demonstrates its high practicability.
Kun Wang 0036, Yongle Luo, Yuxin Wang 0007, Erkang Cheng, Zhiyong Sun 0002
Neural Process. Lett.6
2023 CurveFormer: 3D Lane Detection by Curve Propagation with Curve Queries and Attention
abstract
3D lane detection is an integral part of au-tonomous driving systems. Previous CNN and Transformer-based methods usually first generate a bird's-eye-view (BEV) feature map from the front view image, and then use a sub-network with BEV feature map as input to predict 3D lanes. Such approaches require an explicit view transformation between BEV and front view, which itself is still a challenging problem. In this paper, we propose CurveFormer, a single-stage Transformer-based method that directly calculates 3D lane pa-rameters and can circumvent the difficult view transformation step. Specifically, we formulate 3D lane detection as a curve propagation problem by using curve queries. A 3D lane query is represented by a dynamic and ordered anchor point set. In this way, queries with curve representation in Transformer decoder iteratively refine the 3D lane detection results. Moreover, a curve cross-attention module is introduced to compute the similarities between curve queries and image features. Additionally, a context sampling module that can capture more relative image features of a curve query is provided to further boost the 3D lane detection performance. We evaluate our method for 3D lane detection on both synthetic and real-world datasets, and the experimental results show that our method achieves promising performance compared with the state-of-the-art approaches. The effectiveness of each component is validated via ablation studies as well.
Yifeng Bai, Zhirong Chen, Zhangjie Fu 0002, Lang Peng, Pengpeng Liang, Erkang Cheng
ICRA6
2023 CircleFormer: Circular Nuclei Detection in Whole Slide Images with Circle Queries and Attention
Hengxu Zhang, Pengpeng Liang, Zhiyong Sun 0002, Erkang Cheng
MICCAI (8)5
2023 BEVSegFormer: Bird's Eye View Semantic Segmentation From Arbitrary Camera Rigs
abstract
Semantic segmentation in bird's eye view (BEV) is an important task for autonomous driving. Though this task has attracted a large amount of research efforts, it is still challenging to flexibly cope with arbitrary (single or multiple) camera sensors equipped on the autonomous vehicle. In this paper, we present BEVSegFormer, an effective transformer-based method for BEV semantic segmentation from arbitrary camera rigs. Specifically, our method first encodes image features from arbitrary cameras with a shared backbone. These image features are then enhanced by a deformable transformer-based encoder. Moreover, we introduce a BEV transformer decoder module to parse BEV semantic segmentation results. An efficient multi-camera deformable attention unit is designed to carry out the BEV-to-image view transformation. Finally, the queries are reshaped according to the layout of grids in the BEV, and upsampled to produce the semantic segmentation result in a supervised manner. We evaluate the proposed algorithm on the public nuScenes dataset and a self-collected dataset. Experimental results show that our method achieves promising performance on BEV semantic segmentation from arbitrary camera rigs. We also demonstrate the effectiveness of each component via ablation study.
Lang Peng, Zhirong Chen, Zhangjie Fu 0002, Pengpeng Liang, Erkang Cheng
WACV5
2023 Relay Hindsight Experience Replay: Self-guided continual reinforcement learning for sequential object manipulation tasks with sparse rewards
Yongle Luo, Yuxin Wang 0007, Erkang Cheng, Zhiyong Sun 0002
Neurocomputing5
2022 Traffic Context Aware Data Augmentation for Rare Object Detection in Autonomous Driving
abstract
Detection of rare objects (e.g., traffic cones, traffic barrels and traffic warning triangles) is an important perception task to improve the safety of autonomous driving. Training of such models typically requires a large number of annotated data which is expensive and time consuming to obtain. To address the above problem, an emerging approach is to apply data augmentation to automatically generate cost-free training samples. In this work, we propose a systematic study on simple Copy-Paste data augmentation for rare object detection in autonomous driving. Specifically, local adaptive instance-level image transformation is introduced to generate realistic rare object masks from source domain to the target domain. Moreover, traffic scene context is utilized to guide the placement of masks of rare objects. To this end, our data augmentation generates training data with high quality and realistic characteristics by leveraging both local and global consistency. In addition, we build a new dataset named NM10k consisting 10k training images, 4k validation images and the corresponding labels with a diverse range of scenarios in autonomous driving. Experiments on NM10k show that our method achieves promising results on rare object detection. We also present a thorough study to illustrate the effectiveness of our local-adaptive and global constraints based Copy-Paste data augmentation for rare object detection. The data, development kit and more information of NM10k dataset are available online at: https://nullmax-vision.github.io.
Naifan Li, Pengpeng Liang, Erkang Cheng
ICRA5
2022 An Indeterministic Vision-Based State Observer for Growing Magnetic Microrobot Motion Status Estimation
abstract
To date, untethered micro/nanorobots have attracted considerable attention in various aspects due to their unique potential for in-vivo applications such as the targeted therapy. One of the most promising types of micro/nanorobots is the class of ferromagnetic microrobots which can be efficiently actuated via gradient/rotational magnetic field generated by less costly electromagnetic coil systems. For performing successful operations, locomotion control of the magnetic microrobots is non-trivial. Modern controllers commonly require motion status-based feedback. To fully utilize those advanced approaches, motion state of one microrobot should be supplied, however it is still challenging in cases. It is noted that, during locomotion, one ferromagnetic microrobot can combine with others to form an unstructured larger one, namely growing magnetic microrobot (GMM), whose dynamic behavior keeps changing, and thus the model-based observers are never applicable. Besides, tracking and estimating states of those unstructured time-varying GMMs in complex surroundings are always challenging, especially for an uneven sampling scenario. In order to accurately estimate the GMM motion status in a complex environment via micro-vision, this study develops an indeterministic observer leveraging on the approach of discriminative correlation filter with channel/spatial reliability (CSR-DCF) and the variable-step finite-time sliding mode (FSM-V) state estimation theory. Experimental study verifies that the proposed observation scheme can effectively estimate motion states of one GMM moving in obstacle surroundings throughout.
Zhiyong Sun 0002, Yu Cheng 0006, Erkang Cheng, Gengliang Chen, Lixin Dong
ICRA4
2022 Pseudo Segmentation for Semantic Information-Aware Stereo Matching
abstract
Stereo matching plays an important role in computer vision and robotics. Though substantial progress has been made on deep learning-based algorithms, the inherent semantic information within the ground truth of the training data for stereo matching has not been well explored. In this letter, we propose to use a pseudo segmentation sub-network to extract additional semantic information. More specifically, we divide the disparity label into groups and let each group correspond to a class for pseudo segmentation. To assist stereo matching with the semantic information obtained from pseudo segmentation, we inject the feature maps at the end of the pseudo segmentation sub-network into the cost volume that is used to infer the pixel-level disparity. To validate the effectiveness of the proposed approach, we select PSMNet (Chang and Chen, 2018)and GwcNet (Guoet al., 2019) as baselines and enhance them with the pseudo segmentation sub-network. Comprehensive experiments are carried out on the Scene Flow, KITTI 2015, and KITTI 2012 datasets, and the results show that our proposed method can improve the performance notably.
Shengyou Hua, Zhiyong Sun 0002, Pengpeng Liang, Erkang Cheng
IEEE Signal Process. Lett.5
2021 3D Periodic Magnetic Servoing System for Microrobot Actuation Using Decoupled Asynchronous Repetitive Control Approach
abstract
To date, untethered microrobots have been receiving tremendous attention for playing implacable roles of maneuverable tools in fields such as microfabrication and biomanipulation. Typical actuation of such untethered tiny robots is the magnetic field-based approaches, including gradient and rotational methods. Compared to the gradient type method, the rotational approach requires much less magnetic field strength to generate efficient actuation for magnetic microrobots. To actuate microrobots desirably, a precise periodic magnetic field should be provided. To generate precise periodic magnetic field with enhanced strength, this paper develops a prototype of 3D magnetic servoing system based on integrated solenoids, performance of which are enhanced by employing iron cores and extended number of coils. Each solenoid is equipped with a Hall sensor to provide real-time feedback signal for performing precise magnetic field control. To precisely regulate this setup, a decoupled asynchronous repetitive control (DARC) scheme is established to generate a desirable 3D periodic magnetic field with noise-level tracking error under the situation of missing execution opportunity randomly. Experimental results demonstrate the effectiveness of the proposed magnetic servoing system, which is promising for dynamic properties characterization of magnetic microrobots.
Zhiyong Sun 0002, Yu Cheng 0006, Erkang Cheng, Gengliang Chen, Lixin Dong
ICRA4
2021 Coarse-to-fine Semantic Localization with HD Map for Autonomous Driving in Structural Scenes
abstract
Robust and accurate localization is an essential component for robotic navigation and autonomous driving. The use of cameras for localization with high definition map (HD Map) provides an affordable localization sensor set. Existing methods suffer from pose estimation failure due to error prone data association or initialization with accurate initial pose requirement. In this paper, we propose a cost-effective vehicle localization system with HD map for autonomous driving that uses cameras as primary sensors. To this end, we formulate vision-based localization as a data association problem that maps visual semantics to landmarks in HD map. Specifically, system initialization is finished in a coarse to fine manner by combining coarse GPS (Global Positioning System) measurement and fine pose searching. In tracking stage, vehicle pose is refined by implicitly aligning the semantic segmentation result between image and landmarks in HD maps with photometric consistency. Finally, vehicle pose is computed by pose graph optimization in a sliding window fashion. We evaluate our method on two datasets and demonstrate that the proposed approach yields promising localization results in different driving scenarios. Additionally, our approach is suitable for both monocular camera and multi-cameras that provides flexibility and improves robustness for the localization system.
Minjie Lin, Heyang Guo, Pengpeng Liang, Erkang Cheng
IROS5
2021 Joint Spinal Centerline Extraction and Curvature Estimation with Row-Wise Classification and Curve Graph Network
Long Huo, Bin Cai 0006, Pengpeng Liang, Zhiyong Sun 0002, Chi Xiong, Chaoshi Niu, Erkang Cheng
MICCAI (5)8
2021 Learning local descriptors with multi-level feature aggregation and spatial context pyramid
Pengpeng Liang, Haoxuanye Ji, Erkang Cheng, Yumei Chai, Haibin Ling
Neurocomputing3
2016 Structure-Aware Rank-1 Tensor Approximation for Curvilinear Structure Tracking Using Learned Hierarchical Features
Peng Chu, Erkang Cheng, Ying J. Zhu, Yefeng Zheng 0001, Haibin Ling
MICCAI (1)3
2014 Curvilinear Structure Tracking by Low Rank Tensor Approximation with Model Propagation
abstract
Robust tracking of deformable object like catheter or vascular structures in X-ray images is an important technique used in image guided medical interventions for effective motion compensation and dynamic multi-modality image fusion. Tracking of such anatomical structures and devices is very challenging due to large degrees of appearance changes, low visibility of X-ray images and the deformable nature of the underlying motion field as a result of complex 3D anatomical movements projected into 2D images. To address these issues, we propose a new deformable tracking method using the tensor-based algorithm with model propagation. Specifically, the deformable tracking is formulated as a multi-dimensional assignment problem which is solved by rank-1 l1tensor approximation. The model prior is propagated in the course of deformable tracking. Both the higher order information and the model prior provide powerful discriminative cues for reducing ambiguity arising from the complex background, and consequently improve the tracking robustness. To validate the proposed approach, we applied it to catheter and vascular structures tracking and tested on X-ray fluoroscopic sequences obtained from 17 clinical cases. The results show, both quantitatively and qualitatively, that our approach achieves a mean tracking error of 1.4 pixels for vascular structure and 1.3 pixels for catheter tracking.
Erkang Cheng, Ying J. Zhu, Jingyi Yu 0001, Haibin Ling
CVPR1
2014 Discriminative vessel segmentation in retinal images by fusing context-aware hybrid features
Erkang Cheng, Yi Wu 0001, Ying J. Zhu, Vasileios Megalooikonomou, Haibin Ling
Mach. Vis. Appl.1
2011 Blurred target tracking by Blur-driven Tracker
abstract
Visual tracking plays an important role in many computer vision tasks. A common assumption in previous methods is that the video frames are blur free. In reality, motion blurs are pervasive in the real videos. In this paper we present a novel BLUr-driven Tracker (BLUT) framework for tracking motion-blurred targets. BLUT actively uses the information from blurs without performing debluring. Specifically, we integrate the tracking problem with the motion-from-blur problem under a unified sparse approximation framework. We further use the motion information inferred by blurs to guide the sampling process in the particle filter based tracking. To evaluate our method, we have collected a large number of video sequences with significant motion blurs and compared BLUT with state-of-the-art trackers. Experimental results show that, while many previous methods are sensitive to motion blurs, BLUT can robustly and reliably track severely blurred targets.
Yi Wu 0001, Haibin Ling, Jingyi Yu 0001, Feng Li 0005, Xue Mei, Erkang Cheng
ICCV6
2008 An Improved Model of Producing Saliency Map for Visual Attention System
Jingang Huang, Erkang Cheng
ICIC (3)3