VLDB 2026 Research / reviewers in the wild / expert
Delong Zhu 0001
dblp:210/9651-1
· DBLP profile ↗
21ranked-venue papers
4as first author
14since 2021 · last 2023
0000-0002-1143-7860ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 4 first-author · 9 since 2021Systems, architecture and hardware · 14 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | CLC-Net: Contextual and local collaborative network for lesion segmentation in diabetic retinopathy images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Jun Zhang 0018, Jun Cheng 0003, Raymond Kai-Yu Tong, Xiao Han 0011 |
Neurocomputing | 4 |
| 2023 | ELWNet: An Extremely Lightweight Approach for Real-Time Salient Object DetectionabstractExisting lightweight salient object detection (SOD) methods aim to solve the problem of high computational costs that is prevalent with heavyweight methods. However, compared with heavyweight methods, the detection accuracy of lightweight methods is greatly reduced while real-time performance is not significantly improved. Therefore, we aim to establish a trade off between computational cost and detection performance by improving the network efficiency. We propose a fast and extremely lightweight end-to-end wavelet neural network (ELWNet) for real-time salient object detection. ELWNet can achieve salient object detection and segmentation at approximately 70FPS (GPU), 19FPS (CPU) with 76K parameters and 0.38G FLOPs. We introduce wavelet transform theory into a neural network, proposing a wavelet transform module (WTM), a wavelet transform fusion module (WTFM), a novel feature residual mechanism, and construct an efficient architecture. The wavelet transform theory is integrated into the neural network to realize the interaction between the features in the frequency and the time domain. Meanwhile, ELWNet does not rely on a pre-trained model, which significantly reduces redundant features. We validate the performance of ELWNet using five well-known datasets, and demonstrate state-of-the-art performance compared with 24 other SOD models in terms of being lightweight, detection accuracy and real-time capabilities. Our method maintains high detection performance while reducing the number of model parameters by approximately 99% compared with heavyweight methods. Zhenyu Wang 0010, Yunzhou Zhang, Yan Liu 0080, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Structure-Aware Feature Disentanglement With Knowledge Transfer for Appearance-Changing Place RecognitionabstractLong-term visual place recognition (VPR) is challenging as the environment is subject to drastic appearance changes across different temporal resolutions, such as time of the day, month, and season. A wide variety of existing methods address the problem by means of feature disentangling or image style transfer but ignore the structural information that often remains stable even under environmental condition changes. To overcome this limitation, this article presents a novel structure-aware feature disentanglement network (SFDNet) based on knowledge transfer and adversarial learning. Explicitly, probabilistic knowledge transfer (PKT) is employed to transfer knowledge obtained from the Canny edge detector to the structure encoder. An appearance teacher module is then designed to ensure that the learning of appearance encoder does not only rely on metric learning. The generated content features with structural information are used to measure the similarity of images. We finally evaluate the proposed approach and compare it to state-of-the-art place recognition methods using six datasets with extreme environmental changes. Experimental results demonstrate the effectiveness and improvements achieved using the proposed framework. Source code and some trained models will be available at http://www.tianshu.org.cn. Cao Qin, Yunzhou Zhang, Yingda Liu, Delong Zhu 0001, Sonya A. Coleman, Dermot Kerr |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | An Object SLAM Framework for Association, Mapping, and High-Level TasksabstractObject SLAM is considered increasingly significant for robot high-level perception and decision-making. Existing studies fall short in terms of data association, object representation, and semantic mapping and frequently rely on additional assumptions, limiting their performance. In this article, we present a comprehensive object SLAM framework that focuses on object-based perception and object-oriented robot tasks. First, we propose an ensemble data association approach for associating objects in complicated conditions by incorporating parametric and nonparametric statistic testing. In addition, we suggest an outlier-robust centroid and scale estimation algorithm for modeling objects based on the iForest and line alignment. Then a lightweight and object-oriented map is represented by estimated general object models. Taking into consideration the semantic invariance of objects, we convert the object map to a topological map to provide semantic descriptors to enable multimap matching. Finally, we suggest an object-driven active exploration strategy to achieve autonomous mapping in the grasping scenario. A range of public datasets and real-world results in mapping, augmented reality, scene matching, relocalization, and robotic manipulation have been used to evaluate the proposed object SLAM framework for its efficient performance. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Zhiqiang Deng, Wenkai Sun, Jian Zhang 0018 |
IEEE Trans. Robotics | 3 |
| 2022 | SemLoc: Accurate and Robust Visual Localization with Semantic and Structural Constraints from Prior MapsabstractSemantic information and geometrical structures of a prior map can be leveraged in visual localization to bound drift errors and improve accuracy. In this paper, we propose SemLoc, a pure visual localization system, for accurate localization in a prior semantic map. To tightly couple semantic and structure information from prior maps, a hybrid constraint is presented by using the Dirichlet distribution. Then, with the local landmarks and their semantic states tracked in the frontend, the camera poses and data associations are jointly optimized through Expectation-Maximization (EM) algorithm. We validate the effectiveness of our approach in both monocular and stereo modes on the public KITTI dataset. Experimental results demonstrate that our system can greatly reduce drift errors with an satisfying real-time performance. Compared with several state-of-the-art visual localization systems, the proposed framework achieves a competitive localization performance. Shiwen Liang, Yunzhou Zhang, Rui Tian 0002, Delong Zhu 0001, Linghao Yang, Zhenzhong Cao |
ICRA | 4 |
| 2022 | Online State-Time Trajectory Planning Using Timed-ESDF in Highly Dynamic EnvironmentsabstractOnline state-time trajectory planning in highly dynamic environments remains an unsolved problem due to the curse of dimensionality of the state-time space. Existing state-time planners are typically implemented based on randomized sampling approaches or path searching on discrete graphs. The smoothness, path clearance, or planning efficiency is sometimes not satisfying. In this work, we propose a gradient-based planner on the state-time space for online trajectory generation in highly dynamic environments. To enable the gradient-based optimization, we propose a Timed-ESDT that supports distance and gradient queries with state-time keys. Based on the Timed-ESDT, we also define a smooth prior and an obstacle likelihood function that are compatible with the state-time space. The trajectory planning is then formulated to a MAP problem and solved by an efficient numerical optimizer. Moreover, to improve the optimality of the planner, we also define a state-time graph and conduct path searching on it to find a better initialization for the optimizer. By integrating the graph searching, the planning quality is significantly improved. Experiments on simulated and benchmark datasets demonstrate the superior performance of our proposes method over conventional ones. Delong Zhu 0001, Tong Zhou 0005, Jiahui Lin, Yuqi Fang, Max Q.-H. Meng |
ICRA | 1 |
| 2021 | Object SLAM-Based Active Mapping and Robotic GraspingabstractThis paper presents the first active object mapping framework for complex robotic manipulation and autonomous perception tasks. The framework is built on an object SLAM system integrated with a simultaneous multi-object pose estimation process that is optimized for robotic grasping. Aiming to reduce the observation uncertainty on target objects and increase their pose estimation accuracy, we also design an object-driven exploration strategy to guide the object mapping process, enabling autonomous mapping and high-level perception. Combining the mapping module and the exploration strategy, an accurate object map that is compatible with robotic grasping can be generated. Additionally, quantitative evaluations also indicate that the proposed framework has a very high mapping accuracy. Experiments with manipulation (including object grasping and placement) and augmented reality significantly demonstrate the effectiveness and advantages of our proposed framework. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Sonya A. Coleman, Wenkai Sun, Xinggang Hu, Zhiqiang Deng |
3DV | 3 |
| 2021 | Search-Based Online Trajectory Planning for Car-like Robots in Highly Dynamic EnvironmentsabstractThis paper presents a search-based partial motion planner for generating feasible trajectories of car-like robots in highly dynamic environments. The planner searches for smooth, safe, and near-time-optimal trajectories by exploring a state graph built on motion primitives. To enable fast online planning, we propose an efficient path searching algorithm based on the aggregation and pruning of motion primitives. We then propose a fast collision checking algorithm that takes into account the motions of moving obstacles. The algorithm linearizes relative motions between the robot and obstacles, and then checks collisions by calculating a point-line distance. Benefiting from the fast searching and collision checking algorithms, the planner can effectively explore the state-time space to generate near-time-optimal solutions. Experiments show that the proposed method can generate feasible trajectories within milliseconds while maintaining a higher success rate than up-to-date methods, which significantly demonstrates its advantages. Jiahui Lin, Tong Zhou 0005, Delong Zhu 0001, Jianbang Liu 0002, Max Q.-H. Meng |
ICRA | 3 |
| 2021 | A Large-Scale Dataset for Benchmarking Elevator Button Segmentation and Character RecognitionabstractHuman activities are hugely restricted by COVID-19, recently. Robots that can conduct inter-floor navigation attract much public attention since they can substitute human workers to conduct the service work. However, current robots either depend on human assistance or elevator retrofitting, and fully autonomous inter-floor navigation is still not available. As the very first step of inter-floor navigation, elevator button segmentation and recognition hold an important position. Therefore, we release the first large-scale publicly available elevator panel dataset in this work, containing 3,718 panel images with 35,100 button labels, to facilitate more powerful algorithms on autonomous elevator operation. Together with the dataset, a number of deep learning based implementations for button segmentation and recognition are also released to benchmark future methods in the community. The dataset is available at https://github.com/zhudelong/elevator_button_recognition Jianbang Liu 0002, Yuqi Fang, Delong Zhu 0001, Nachuan Ma, Max Q.-H. Meng |
ICRA | 3 |
| 2021 | Generalized Point Set Registration with the Kent DistributionabstractPoint set registration (PSR) is an essential problem in communities of computer vision, medical robotics and biomedical engineering. This paper is motivated by considering the anisotropic characteristics of the error values in estimating both the positional and orientational vectors from the PSs to be registered. To do this, the multi-variate Gaussian and Kent distributions are utilized to model the positional and orientational uncertainties, respectively. Our contributions of this paper are three-folds: (i) the PSR problem using normal vectors is formulated as a maximum likelihood estimation (MLE) problem, where the anisotropic characteristics in both positional and normal vectors are considered; (ii) the matrix forms of the objective function and its associated gradients with respect to the desired parameters are provided, which can facilitate the computational process; (iii) two approaches of computing the normalizing constant in the Kent distribution are compared. We verify our proposed registration method on various PSs (representing pelvis and femur bones) in computer- assisted orthopedic surgery (CAOS). Extensive experimental results demonstrate that our method outperforms the state- of-the-art methods in terms of the registration accuracy and the robustness. Zhe Min, Delong Zhu 0001, Max Q.-H. Meng |
ICRA | 2 |
| 2021 | Accurate and Robust Scale Recovery for Monocular Visual Odometry Based on Plane GeometryabstractScale ambiguity is a fundamental problem in monocular visual odometry. Typical solutions include loop closure detection and environment information mining. For applications like self-driving cars, loop closure is not always available, hence mining prior knowledge from the environment becomes a more promising approach. In this paper, with the assumption of a constant height of the camera above the ground, we develop a light-weight scale recovery framework leveraging an accurate and robust estimation of the ground plane. The framework includes a ground point extraction algorithm for selecting high-quality points on the ground plane, and a ground point aggregation algorithm for joining the extracted ground points in a local sliding window. Based on the aggregated data, the scale is finally recovered by solving a least-squares problem using a RANSAC-based optimizer. Sufficient data and robust optimizer enable a highly accurate scale recovery. Experiments on the KITTI dataset show that the proposed framework can achieve state-of-the-art accuracy in terms of translation errors, while maintaining competitive performance on the rotation error. Due to the light-weight design, our framework also demonstrates a high frequency of 20 Hz on the dataset. Rui Tian 0002, Yunzhou Zhang, Delong Zhu 0001, Shiwen Liang, Sonya A. Coleman, Dermot Kerr |
ICRA | 3 |
| 2021 | PiPo-Net: A Semi-automatic and Polygon-based Annotation Method for Pathological ImagesabstractMetastatic involvement of lymph nodes is one of the most important prognostic variables for many cancers. Several deep learning based algorithms have been developed to segment metastatic regions in pathological images to help predict prognosis. However, the training of these methods requires a large amount of annotated data, and the labeling task is an extremely time-consuming process for human annotators. In order to reduce the annotation burden, we for the first time propose a semi-automatic annotation method (PiPo-Net) for the labeling of pathological images. The method is comprised of two subnetworks, a pixel-wise segmentation network (Pi-Net) and a polygon-based annotation network (Po-Net). The Pi-Net adopts an improved encoder-decoder architecture and can effectively aggregate multi-scale image features. The Po-Net is built on the Pi-Net and leverages a two-layer recurrent neural network to generate tight-bounded polygons for the metastatic regions. Corresponding to the proposed network architecture, a loss function called PiPo-loss is introduced to help optimize the whole network. The main advantage of our method is that it integrates human annotators into the prediction loop, allowing to iteratively refine the predictions according to the suggestions from human annotators. We evaluate our method on Camelyon16 database and achieve a Dice score of 91% in the initial annotation attempt. We also demonstrate the effectiveness of the human-network collaborative annotation, which achieves promising labeling results, verifying the advantages of our proposed method. Yuqi Fang, Delong Zhu 0001, Niyun Zhou, Li Liu 0017, Jianhua Yao 0001 |
IROS | 2 |
| 2021 | A hybrid network for automatic hepatocellular carcinoma segmentation in H&E-stained whole slide images
Yuqi Fang, Sen Yang 0006, Delong Zhu 0001, Jing Zhang 0051, Raymond Kai-Yu Tong, Xiao Han 0011 |
Medical Image Anal. | 4 |
| 2021 | Feature-Guided Nonrigid 3-D Point Set Registration Framework for Image-Guided Liver Surgery: From Isotropic Positional Noise to Anisotropic Positional NoiseabstractRegistration is an essential problem in image-guided surgery (IGS) since it brings different involved coordinate frames together. Nonrigid or deformable registration still faces many challenges, such as two point sets (PSs) are partially overlapped. To tackle the challenges in the nonrigid registration, we introduce a new two-step point-based registration pipeline that includes two steps. In the first step, the rigid transformation between the two spaces is recovered where the orientation vectors are adopted. In the second step, built on the nonrigid coherent point drift (CPD) approach, the anisotropic positional noise is also assumed. Registration results on the human liver verify the proposed approach' great improvements over the other methods. First, the rotation and translation are recovered with smaller error values than the existing methods. Second, our registration method's performance is much more robust to the partial overlapping between two PSs. Third, the two-step registration framework achieves the best performances in most test cases when there is a localization error in acquiring the intraoperative data. Note to Practitioners-A novel registration approach is presented for image-guided liver surgery (LGLS). Compared with existing nonrigid registration methods, two significant changes (or improvements) exist in the proposed registration framework: 1) the normal vectors are extracted and utilized in the rigid registration step and 2) the anisotropic positional uncertainties are considered. In both steps, the registration problems are formulated as a maximum likelihood (ML) problems and dealt with the expectation-maximization (EM) technique. In both steps, the matrix form of the updated positional covariance is provided and can speed up the computational process. The readers are reminded that with extra information and a more general positional error assumption, our approach demonstrates improved performances in the case of partial-to-full alignment. Zhe Min, Delong Zhu 0001, Hongliang Ren 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Robust and Accurate 3D Curve to Surface Registration with Tangent and Normal VectorsabstractThis paper presents a robust and accurate approach for the rigid registration of pre-operative and intraoperative point sets in image-guided surgery (IGS). Three challenges are identified in the pre-to-intraoperative registration: the intra-operative 3D data (usually forms a 3D curve in space) (1) is often contaminated with noise and outliers; (2) usually only covers a partial region of the whole pre-operative model; (3) is usually sparse. To tackle those challenges, we utilize the tangent vectors extracted from the sparse intraoperative data points and the normal vectors extracted from the pre-operative model points. Our first contribution is to formulate a novel probabilistic distribution of the error between a pair of corresponding tangent and normal vectors. The second contribution is, based on the novel distribution, we formulate the registration of two multi-dimensional (6D) point sets as a maximum likelihood (ML) problem and solve it under the expectation maximization (EM) framework. Our last contribution is, in order to facilitate the computation process, the derivatives of the objective function with respect to desired parameters are presented. We conduct extensive experiments to demonstrate that our approach outperforms the state-of-the-art methods. Importantly, in the context of anteriro cruciate ligament (ACL) reconstruction, our method can achieve as low as 0.6795 mm mean target registration error (TRE) value with considerable noises and very limited overlapping ratios. Zhe Min, Delong Zhu 0001, Max Q.-H. Meng |
ICRA | 2 |
| 2020 | HouseExpo: A Large-scale 2D Indoor Layout Dataset for Learning-based Algorithms on Mobile RobotsabstractAs one of the most promising areas, mobile robots draw much attention these years. Current work in this field is often evaluated in a few manually designed scenarios, due to the lack of a common experimental platform. Meanwhile, with the recent development of deep learning techniques, some researchers attempt to apply learning-based methods to mobile robot tasks, which requires a substantial amount of data. To satisfy the underlying demand, in this paper we build HouseExpo, a large-scale indoor layout dataset containing 35, 126 2D floor plans including 252, 550 rooms in total. Together we develop PseudoSLAM, a lightweight and efficient simulation platform to accelerate the data generation procedure, thereby speeding up the training process. In our experiments, we build models to tackle obstacle avoidance and autonomous exploration from a learning perspective in simulation as well as real-world experiments to verify the effectiveness of our simulator and dataset. All the data and codes are available online and we hope HouseExpo and PseudoSLAM can feed the need for data and benefit the whole community. Tingguang Li, Danny Ho, Delong Zhu 0001, Chaoqun Wang 0009, Max Q.-H. Meng |
IROS | 4 |
| 2020 | TartanAir: A Dataset to Push the Limits of Visual SLAMabstractWe present a challenging dataset, the TartanAir, for robot navigation tasks and more. The data is collected in photo-realistic simulation environments with the presence of moving objects, changing light and various weather conditions. By collecting data in simulations, we are able to obtain multi-modal sensor data and precise ground truth labels such as the stereo RGB image, depth image, segmentation, optical flow, camera poses, and LiDAR point cloud. We set up large numbers of environments with various styles and scenes, covering challenging viewpoints and diverse motion patterns that are difficult to achieve by using physical data collection platforms. In order to enable data collection at such a large scale, we develop an automatic pipeline, including mapping, trajectory sampling, data processing, and data verification. We evaluate the impact of various factors on visual SLAM algorithms using our data. The results of state-of-the-art algorithms reveal that the visual SLAM problem is far from solved. Methods that show good performance on established datasets such as KITTI do not perform well in more difficult scenarios. Although we use the simulation, our goal is to push the limits of Visual SLAM algorithms in the real world by providing a challenging benchmark for testing new methods, while also using a large diverse training data for learning-based methods. Our dataset is available at http://theairlab.org/tartanair-dataset. Delong Zhu 0001, Yaoyu Hu, Yuheng Qiu, Chen Wang 0033, Yafei Hu, Ashish Kapoor, Sebastian A. Scherer |
IROS | 2 |
| 2020 | EAO-SLAM: Monocular Semi-Dense Object SLAM Based on Ensemble Data AssociationabstractObject-level data association and pose estimation play a fundamental role in semantic SLAM, which remain unsolved due to the lack of robust and accurate algorithms. In this work, we propose an ensemble data associate strategy for integrating the parametric and nonparametric statistic tests. By exploiting the nature of different statistics, our method can effectively aggregate the information of different measurements, and thus significantly improve the robustness and accuracy of data association. We then present an accurate object pose estimation framework, in which an outliers-robust centroid and scale estimation algorithm and an object pose initialization algorithm are developed to help improve the optimality of pose estimation results. Furthermore, we build a SLAM system that can generate semi-dense or lightweight object-oriented maps with a monocular camera. Extensive experiments are conducted on three publicly available datasets and a real scenario. The results show that our approach significantly outperforms state-of-the-art techniques in accuracy and robustness. The source code is available on https://github.com/yanmin-wu/EAO-SLAM. Yanmin Wu, Yunzhou Zhang, Delong Zhu 0001, Yonghui Feng, Sonya A. Coleman, Dermot Kerr |
IROS | 3 |
| 2018 | Deep Reinforcement Learning Supervised Autonomous Exploration in Office EnvironmentsabstractExploration region selection is an essential decision making process in autonomous robot exploration task. While a majority of greedy methods are proposed to deal with this problem, few efforts are made to investigate the importance of predicting long-term planning. In this paper, we present an algorithm that utilizes deep reinforcement learning (DRL) to learn exploration knowledge over office blueprints, which enables the agent to predict a long-term visiting order for unexplored subregions. On the basis of this algorithm, we propose an exploration architecture that integrates a DRL model, a next-best-view (NBV) selection approach and a structural integrity measurement to further improve the exploration performance. At the end of this paper, we evaluate the proposed architecture against other methods on several new office maps, showing that the agent can efficiently explore uncertain regions with a shorter path and smarter behaviors. Delong Zhu 0001, Tingguang Li, Danny Ho, Chaoqun Wang 0009, Max Q.-H. Meng |
ICRA | 1 |
| 2018 | A Novel OCR-RCNN for Elevator Button RecognitionabstractAutonomous elevator operation is considered an intelligent solution in handling the inter-floor navigation problem of service robots. As one of the most fundamental steps, elevator button recognition starts to receive more and more attention. However, due to the challenging image conditions and severe class imbalance problem, the performance of existing results is unsatisfying. In this paper, we propose to combine an optical character recognition (OCR) network and the Faster RCNN architecture into a single neural network, called OCR-RCNN to facilitate an end-to-end training and elevator button recognition procedure. To verify our method, we collect a large dataset of elevator panels and carry out extensive comparative experiments. The experiment results show that our method can greatly outperform the traditional recognition pipelines, yielding an accurate and robust performance on recognizing untrained elevator buttons. Delong Zhu 0001, Tingguang Li, Danny Ho, Tong Zhou 0005, Max Q.-H. Meng |
IROS | 1 |
| 2017 | Hawkeye: Open source framework for field surveillanceabstractThis paper introduces a generic framework for field surveillance using consumer rotorcrafts and ground vehicles. Building such an autonomous system comes with two key challenges in persistent perception and obstacle avoidance. We begin with explaining two core algorithms to solve the challenges: an auto-landing algorithm that enables a quadrotor to land on a moving ground vehicle at a speed of 6.00 m/s, and an obstacle avoidance algorithm that ensures the safety of the quadrotor during searching process. On the basis of these algorithms, the architecture and infrastructure of Hawkeye framework are presented as well. Hawkeye is designed to be a generic platform with extensibility that allows integration of other domain applications. We demonstrate the potential of Hawkeye framework in a simulated agriculture monitoring mission and report its performance at the end of the paper. Delong Zhu 0001, Yegui Du, Chaoqun Wang 0009, Xun Xu 0001, Max Q.-H. Meng |
IROS | 1 |