Danping Zou

dblp:74/3124 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0003-3824-8741ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 2 first-author · 14 since 2021Systems, architecture and hardware · 17 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Head-Aware KV Cache Compression for Efficient Visual Autoregressive Modeling
abstract
Visual Autoregressive (VAR) models adopt a next-scale prediction paradigm, offering high-quality content generation with substantially fewer decoding steps. However, existing VAR models suffer from significant attention complexity and severe memory overhead due to the accumulation of key-value (KV) caches across scales. In this paper, we tackle this challenge by introducing KV cache compression into the next-scale generation paradigm. We begin with a crucial observation: attention heads in VAR models can be divided into two functionally distinct categories: Contextual Heads focus on maintaining semantic consistency, while Structural Heads are responsible for preserving spatial coherence. This structural divergence causes existing one-size-fits-all compression methods to perform poorly on VAR models. To address this, we propose HACK, a training-free Head-Aware KV cache Compression frameworK. HACK utilizes an offline classification scheme to separate head types, enabling it to apply pattern-specific compression strategies with asymmetric cache budgets for each category. By doing so, HACK effectively constrains the average KV cache length within a fixed budget B, reducing the theoretical attention complexity from O(n4) to O(Bn2). Extensive experiments on multiple VAR models across text-to-image and class-conditional tasks validate the effectiveness and generalizability of HACK. It achieves up to 70% KV cache compression without degrading output quality, resulting in memory savings and faster in- ference. For example, HACK provides a 1.75× memory reduction and a 1.57× speedup on Infinity-8B.
Ziran Qin, Youru Lv, Mingbao Lin, Hang Guo 0002, Zeren Zhang, Danping Zou, Weiyao Lin
AAAI6
2026 CoordAR: One-Reference 6D Pose Estimation of Novel Objects via Autoregressive Coordinate Map Generation
abstract
Object 6D pose estimation, a crucial task for robotics and augmented reality applications, becomes particularly challenging when dealing with novel objects whose 3D models are not readily available. To reduce dependency on 3D models, recent studies have explored one-reference-based pose estimation, which requires only a single reference view instead of a complete 3D model. However, existing methods that rely on real-valued coordinate regression suffer from limited global consistency due to the local nature of convolutional architectures and face challenges in symmetric or occluded scenarios owing to a lack of uncertainty modeling. We present CoordAR, a novel autoregressive framework for one-reference 6D pose estimation of unseen objects. CoordAR formulates 3D-3D correspondences between the reference and query views as a map of discrete tokens, which is obtained in an autoregressive and probabilistic manner. To enable accurate correspondence regression, CoordAR introduces 1) a novel coordinate map tokenization that enables probabilistic prediction over discretized 3D space; 2) a modality-decoupled encoding strategy that separately encodes RGB appearance and coordinate cues; and 3) an autoregressive transformer decoder conditioned on both position-aligned query features and the partially generated token sequence. With these novel mechanisms, CoordAR significantly outperforms existing methods on multiple benchmarks and demonstrates strong robustness to symmetry, occlusion, and other challenges in real-world tests.
Dexin Zuo, Wenxian Yu, Danping Zou
AAAI5
2025 ZQuery-LSM-Tree for ZNS SSD
Tao Cai 0003, Danping Zou, DeJiao Niu, Zihao Yinyi, Qiujing Huang, Yikang Deng
ICA3PP (5)2
2025 ThermoStereoRT: Thermal Stereo Matching in Real Time via Knowledge Distillation and Attention-Based Refinement
Anning Hu, Ang Li 0029, Xirui Jin, Danping Zou
ICRA4
2025 VisFly: An Efficient and Versatile Simulator for Training Vision-Based Flight
abstract
We present VisFly, a quadrotor simulator designed to efficiently train vision-based flight policies using reinforcement learning algorithms. VisFly offers a user-friendly framework and interfaces, leveraging Habitat-Sim's rendering engines to achieve frame rates exceeding 10,000 frames per second for rendering motion and sensor data. The simulator incorporates differentiable physics and is seamlessly wrapped with the Gym environment, facilitating the straightforward implementation of various learning algorithms. It supports the directly importing open-source scene datasets compatible with Habitat-Sim, enabling training on diverse real-world environments simultaneously. To validate our simulator, we also make three reinforcement learning examples for typical flight tasks relying on visual observations. The simulator is now available at [https://github.com/SJTU-ViSYS-team/VisFly].
Fanxing Li, Fangyu Sun, Tianbao Zhang, Danping Zou
ICRA4
2025 Mapless Collision-Free Flight via MPC using Dual KD-Trees in Cluttered Environments
abstract
Collision-free flight in cluttered environments is a critical capability for autonomous quadrotors. Traditional methods often rely on detailed 3D map construction, trajectory generation, and tracking. However, this cascade pipeline can introduce accumulated errors and computational delays, limiting flight agility and safety. In this paper, we propose a novel method for enabling collision-free flight in cluttered environments without explicitly constructing 3D maps or generating and tracking collision-free trajectories. Instead, we leverage Model Predictive Control (MPC) to directly produce safe actions from sparse waypoints and point clouds from a depth camera. These sparse waypoints are dynamically adjusted online based on nearby obstacles detected from point clouds. To achieve this, we introduce a dual KD-Tree mechanism: the Obstacle KD-Tree quickly identifies the nearest obstacle for avoidance, while the Edge KD-Tree provides a robust initial guess for the MPC solver, preventing it from getting stuck in local minima during obstacle avoidance. We validate our approach through extensive simulations and real-world experiments. The results show that our approach significantly outperforms the mapping-based methods and is also superior to imitation learning-based methods, demonstrating reliable obstacle avoidance at up to 12 m/s in simulations and 6 m/s in real-world tests. Our method provides a simple and robust alternative to existing methods. The code is publicly available at https://github.com/SJTU-ViSYS-team/avoid-mpc.
Linzuo Zhang, Yu Hu 0019, Feng Yu 0024, Danping Zou
IROS5
2025 PlanarGS: High-Fidelity Indoor 3D Gaussian Splatting Guided by Vision-Language Planar Priors
abstract
Three-dimensional Gaussian Splatting (3DGS) has recently emerged as an efficient representation for novel-view synthesis, achieving impressive visual quality. However, in scenes dominated by large and low-texture regions, common in indoor environments, the photometric loss used to optimize 3DGS yields ambiguous geometry and fails to recover high-fidelity 3D surfaces. To overcome this limitation, we introduce PlanarGS, a 3DGS-based framework tailored for indoor scene reconstruction. Specifically, we design a pipeline for Language-Prompted Planar Priors (LP3) that employs a pretrained vision-language segmentation model and refines its region proposals via cross-view fusion and inspection with geometric priors. 3D Gaussians in our framework are optimized with two additional terms: a planar prior supervision term that enforces planar consistency, and a geometric prior supervision term that steers the Gaussians toward the depth and normal cues. We have conducted extensive experiments on standard indoor benchmarks. The results show that PlanarGS reconstructs accurate and detailed 3D surfaces, consistently outperforming state-of-the-art methods by a large margin. Project page: https://planargs.github.io
Xirui Jin, Renbiao Jin, Boying Li, Danping Zou, Wenxian Yu
NeurIPS4
2025 Sparse-to-Dense Hint Guided Stereo-LiDAR Fusion
abstract
One challenge in stereo-LiDAR fusion arises from the sparsity and non-uniform distribution of LiDAR data. Existing methods expand sparse LiDAR data to produce semi-dense hints as guidance for fusion. However, the absence of depth cues beyond the expanded areas may still limit performance. To address this challenge, we propose a novel sparse-to-dense hint guided stereo-LiDAR fusion method. The key idea is to use a dense hint map generated by a lightweight network as guidance, with sparse LiDAR points and a monocular image as inputs. The dense hints are then employed to construct and explicitly regularize a multi-modal cost volume via integrating the geometric cues from the hints and the visual information from the images to produce better stereo prediction. The construction and aggregation of cost volume follow a well-designed coarse-to-fine strategy along with a pixel-wise search range adjustment module, facilitating fast computation while preserving fine details. Finally, a confidence-based fusion module is performed to adaptively produce the ultimate prediction based on the monocular and stereo estimations. The experimental results show that our method significantly outperforms existing methods with high inference efficiency across multiple benchmark datasets. To contribute to the community, we will release the code at: https://github.com/LiAngLA66/DG-Fusion.
Ang Li 0029, Dexin Zuo, Anning Hu, Wenxian Yu, Danping Zou
IEEE Trans. Circuits Syst. Video Technol.5
2024 Stereo-LiDAR Depth Estimation with Deformable Propagation and Learned Disparity-Depth Conversion
abstract
Accurate and dense depth estimation with stereo cameras and LiDAR is an important task for automatic driving and robotic perception. While sparse hints from LiDAR points have improved cost aggregation in stereo matching, their effectiveness is limited by the low density and non-uniform distribution. To address this issue, we propose a novel stereo-LiDAR depth estimation network with Semi-Dense hint Guidance, named SDG-Depth. Our network includes a deformable propagation module for generating a semi-dense hint map and a confidence map by propagating sparse hints using a learned deformable window. These maps then guide cost aggregation in stereo matching. To reduce the triangulation error in depth recovery from disparity, especially in distant regions, we introduce a disparity-depth conversion module. Our method is both accurate and efficient. The experimental results on benchmark tests show its superior performance. Our code is available at https://github.com/SJTU-ViSYS/SDG-Depth.
Anning Hu, Wenxian Yu, Danping Zou
ICRA5
2024 Ground-Fusion: A Low-cost Ground SLAM System Robust to Corner Cases
abstract
We introduce Ground-Fusion, a low-cost sensor fusion simultaneous localization and mapping (SLAM) system for ground vehicles. Our system features efficient initialization, effective sensor anomaly detection and handling, real-time dense color mapping, and robust localization in diverse environments. We tightly integrate RGB-D images, inertial measurements, wheel odometer and GNSS signals within a factor graph to achieve accurate and reliable localization both indoors and outdoors. To ensure successful initialization, we propose an efficient strategy that comprises three different methods: stationary, visual, and dynamic, tailored to handle diverse cases. Furthermore, we develop mechanisms to detect sensor anomalies and degradation, handling them adeptly to maintain system accuracy. Our experimental results on both public and self-collected datasets demonstrate that Ground-Fusion outperforms existing low-cost SLAM systems in corner cases. We release the code and datasets at https://github.com/SJTU-ViSYS/Ground-Fusion.
Wenxian Yu, Danping Zou
ICRA5
2024 DNZ-LSM-Tree for Hybrid Storage Systems
abstract
The Log-Structured Merge tree (LSM-tree) can transform random writes into sequential writes, adapting to the high sequential write performance of external storage devices. This has led to its widespread application in various Key-Value (KV) storage systems. However, the LSM-tree has issues such as write amplification and periodic sharp declines in performance. Non-Volatile Memory (NVM) and Zoned Namespace (ZNS) Solid State Drive (SSD) are emerging storage devices that differ significantly from traditional SSDs. Directly applying LSM-tree to NVM and ZNS SSD is challenging due to their unique advantages and characteristics. This paper proposes the DNZ-LSM-Tree, tailored for hybrid external storage constructed using NVM and ZNS SSD. Initially, we present the structure of the DNZ-LSM-Tree, which reconstructs the LSM-tree by utilizing the distinct features of memory, NVM, and ZNS SSD. Subsequently, to address the cascading compaction issue that significantly affects the performance of LSM-tree, we construct a Compaction Cache in NVM and design a layered distribution strategy for Sorted String Tables (SSTables) with cascading compactions. In addition, the zone allocation strategy based on key overlap ratio estimation and the wear leveling strategy for DNZ-LSM-Tree are designed to manage the zones of ZNS SSD. Finally, a prototype of KV storage system based on hybrid storage devices called HNZMS is implemented by employing DNZ-LSM-Tree and tested by YCSB. The results indicate that compared to the storage system called ListDB based on LSM-tree, HNZMS can increase the write throughput by 36.2% and reduce the write amplification by 51.9%.
Tao Cai 0003, DeJiao Niu, Qiangqiang Ni, Zihao Yinyi, Danping Zou
ISPA6
2024 Adaptive Hint Propagation for Iterative Stereo Matching
abstract
While sparse depth hints from LiDAR points have been utilized as guidance to enhance stereo matching, the improvement is hindered by the low density and uneven distribution of those points. To deal with these challenges, the sparse LiDAR hints are usually expanded for further processing. However, existing methods use only the local information of a fixed window surrounding the sparse hint, leading to inaccurate propagation that ultimately deteriorates stereo matching results. We introduce a new adaptive LiDAR propagation recurrent network by incorporating global context and local information, propagating the hints with an adaptive deformable window, and iteratively updating a disparity field through a recurrent unit. We have conducted comprehensive experiments on various public datasets. The results show that our method produces better matching quality than existing methods.
Anning Hu, Ang Li 0029, Danping Zou
VCIP3
2024 Transformer framework for depth-assisted UDA semantic segmentation
Yunna Song, Danping Zou, Caisheng Liu, Suqin Bai, Yunhan Sun
Eng. Appl. Artif. Intell.3
2024 TextSLAM: Visual SLAM With Semantic Planar Text Features
abstract
We propose a novel visual SLAM method that integrates text objects tightly by treating them as semantic features via fully exploring their geometric and semantic prior. The text object is modeled as a texture-rich planar patch whose semantic meaning is extracted and updated on the fly for better data association. With the full exploration of locally planar characteristics and semantic meaning of text objects, the SLAM system becomes more accurate and robust even under challenging conditions such as image blurring, large viewpoint changes, and significant illumination variations (day and night). We tested our method in various scenes with the ground truth data. The results show that integrating texture features leads to a more superior SLAM system that can match images across day and night. The reconstructed semantic 3D text map could be useful for navigation and scene understanding in robotic and mixed reality applications.
Boying Li, Danping Zou, Yuan Huang 0011, Xinghan Niu, Ling Pei, Wenxian Yu
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation
abstract
Estimating the 6-D poses of objects from RGB-D images holds great potential for several applications. However, given that the 6-D pose estimation accuracy is significantly affected by occlusion and noise between the objects in an image, this paper proposes a novel 6-D pose estimation method based on Efficient feature extraction and Point-pair feature matching. Specifically, we develop the Efficient channel attention Convolutional Neural Network (ECNN) and SO(3)-Encoder modules to extract 2-D features from the RGB image and SO(3)-equivariant features from the depth image, respectively. These features are fused in the DenseFusion module to obtain 3-D features in the camera space. Meanwhile, we exploit CAD model priors to obtain 3-D features in the model space through the model feature encoder, and then we globally regress the 3-D features in the camera and model space. According to these features, we generate oriented point clouds in each space, and then conduct point-pair feature matching to obtain pose information. Finally, we perform direct pose regression on the 3-D features in the camera and model space, and then resulting point-pair feature matching pose information is combined with the direct point-wise pose regression information to enhance pose prediction accuracy. Experimental results on three widely used benchmarking datasets demonstrate that our method achieves state-of-the-art performance, particularly for severe occluded scenes.
Danping Zou, Xin Shu 0001, Suqin Bai, Haowei Zhu, Yunhan Sun
IEEE Trans. Multim.3
2023 FeatureBooster: Boosting Feature Descriptors with a Lightweight Neural Network
abstract
We introduce a lightweight network to improve descriptors of keypoints within the same image. The network takes the original descriptors and the geometric properties of keypoints as the input, and uses an MLP-based self-boosting stage and a Transformer-based cross-boosting stage to enhance the descriptors. The boosted descriptors can be either real-valued or binary ones. We use the proposed network to boost both hand-crafted (ORB [34], SIFT [24]) and the state-of-the-art learning-based descriptors (SuperPoint [10], ALIKE [53]) and evaluate them on image matching, visual localization, and structure-from-motion tasks. The results show that our method significantly improves the performance of each task, particularly in challenging cases such as large illumination changes or repetitive patterns. Our method requires only 3.2ms on desktop GPU and 27ms on embedded GPU to process 2000 features, which is fast enough to be applied to a practical system. The code and trained weights are publicly available at github.com/SJTU-ViSYS/FeatureBooster.
Xinjiang Wang, Zeyu Liu 0001, Yu Hu 0019, Wenxian Yu, Danping Zou
CVPR6
2023 UAV Navigation With Monocular Visual Inertial Odometry Under GNSS-Denied Environment
abstract
In GNSS-denied environments, unmanned aerial vehicle (UAV) navigation based on visual inertial odometry has been widely studied. However, existing visual-inertial odometry methods still suffer from some practical problems such as image enhancement oversaturation and unreasonable weighting in backend optimization. Therefore, this paper presents monocular visual-inertial odometry with point-line fusion and backend adaptive optimization to improve the positioning accuracy and robustness of UAV navigation system. In the frontend, we proposed an adaptive gamma image correction algorithm for image preprocessing to avoid image oversaturation, which is more conducive to image extraction and matching. Instead of the traditional LSD line feature extraction algorithm, we employed an improved EDLines algorithm to enhance the efficiency of line feature extraction, better meeting the high dynamic real-time requirements of UAV. In the backend, we proposed a tightly coupled nonlinear adaptive optimization method based on a two-step approach to address the issue of unreasonable static weights. In the first step, we established factor graph model and performed the first nonlinear optimization based on a priori visual weights. In the second step, we calculated the reprojection error and established a functional model that examines the relationship between the reprojection error and the information matrix. We updated the information matrix using the reprojection error to adaptively adjust the weights of the point features and line features in real time. Finally, we performed a second nonlinear re-optimization. The proposed method was compared with the VINS-MONO [1] and PL-VINS [2] methods, the experimental results showing that the positioning accuracy of the proposed method on the public EuRoc dataset [3] improved by an average of 32.3% compared with the PL-VINS method, and by an average of 33.8% in three real-world scenarios under changing illumination, weak texture, and large-scale complex scenarios. The results demonstrated that the proposed method exhibited better robustness and higher positioning accuracy in various complex environments.
Haolong Luo, Guangyun Li, Danping Zou, Kailin Li 0003, Zidi Yang
IEEE Trans. Geosci. Remote. Sens.3
2022 Multi-instance Point Cloud Registration by Efficient Correspondence Clustering
abstract
We address the problem of estimating the poses of multiple instances of the source point cloud within a target point cloud. Existing solutions require sampling a lot of hypotheses to detect possible instances and reject the outliers, whose robustness and efficiency degrade notably when the number of instances and outliers increase. We propose to directly group the set of noisy correspondences into different clusters based on a distance invariance matrix. The instances and outliers are automatically identified through clustering. Our method is robust and fast. We evaluated our method on both synthetic and real-world datasets. The results show that our approach can correctly register up to 20 instances with an F1 score of 90.46% in the presence of 70% outliers, which performs significantly better and at least 10x faster than existing methods. (Source code: https://github.com/SITU-ViSYSlmulti-instant-reg).
Weixuan Tang 0001, Danping Zou
CVPR2
2021 StructDepth: Leveraging the structural regularities for self-supervised indoor depth estimation
abstract
Self-supervised monocular depth estimation has achieved impressive performance on outdoor datasets. Its performance however degrades notably in indoor environments because of the lack of textures. Without rich textures, the photometric consistency is too weak to train a good depth network. Inspired by the early works on indoor modeling, we leverage the structural regularities exhibited in indoor scenes, to train a better depth network. Specifically, we adopt two extra supervisory signals for self-supervised training: 1) the Manhattan normal constraint and 2) the co-planar constraint. The Manhattan normal constraint enforces the major surfaces (the floor, ceiling, and walls) to be aligned with dominant directions. The co-planar constraint states that the 3D points be well fitted by a plane if they are located within the same planar region. To generate the supervisory signals, we adopt two components to classify the major surface normal into dominant directions and detect the planar regions on the fly during training. As the predicted depth becomes more accurate after more training epochs, the supervisory signals also improve and in turn feedback to obtain a better depth model. Through extensive experiments on indoor benchmark datasets, the results show that our network outperforms the state-of-the-art methods. The source code is available at https://github.com/SJTU-ViSYS/StructDepth.
Boying Li, Yuan Huang 0011, Zeyu Liu 0001, Danping Zou, Wenxian Yu
ICCV4
2021 Robust Initialization of Multi-camera SLAM with Limited View Overlaps and Inaccurate Extrinsic Calibration
abstract
This paper proposes a robust initialization method for a multi-camera visual SLAM system where cameras have only a limited common field of views and inaccurate extrinsic calibration. The limited common field of views leads to only a few common features that can be matched between cameras. Inaccurate extrinsic poses, caused by vibrations or misplacement of cameras after offline calibration, make it even harder for triangulating the seed 3D points to initialize the SLAM system successfully. Instead of taking the extrinsic parameters as constants for feature matching and 3D point triangulation as most multi-camera systems did, we propose to take the inaccurate extrinsic poses as soft constraints to accommodate the calibration errors. Our initialization method consists of two stages by matching across different cameras and between two key frames. Both stages involve optimizing the cost functions that contain the extrinsic pose priors from inaccurate calibration parameters. By incorporating those soft pose constraints, we may avoid false feature matching and triangulation caused by inaccurate extrinsic parameters, while keeping solution space limited when only a few feature correspondences exist. The results in real-world tests show that such a simple solution can improve the success rate of SLAM initialization notably, even when the pose priors from the offline calibration differ significantly from the real ones.
Ang Li 0029, Danping Zou, Wenxian Yu
IROS2
2021 A novel framework for UAV returning based on FPGA
Qunfang He, Danping Zou, ZhiLei Chai
J. Supercomput.3
2020 TextSLAM: Visual SLAM with Planar Text Features
abstract
We propose to integrate text objects in man-made scenes tightly into the visual SLAM pipeline. The key idea of our novel text-based visual SLAM is to treat each detected text as a planar feature which is rich of textures and semantic meanings. The text feature is compactly represented by three parameters and integrated into visual SLAM by adopting the illumination-invariant photometric error. We also describe important details involved in implementing a full pipeline of text-based visual SLAM. To our best knowledge, this is the first visual SLAM method tightly coupled with the text features. We tested our method in both indoor and outdoor environments. The results show that with text features, the visual SLAM system becomes more robust and produces much more accurate 3D text maps that could be useful for navigation and scene understanding in robotic or augmented reality applications.
Boying Li, Danping Zou, Daniele Sartori, Ling Pei, Wenxian Yu
ICRA2
2019 A Novel Wireless Positioning Approach Based on Distributed Stochastic-Resonance-Enhanced Power Spectrum Fusion Technique
abstract
In the wireless positioning, the performances of most traditional methods are influenced by the low signal-to-noise (SNR) and none-line-of-sight (NLOS) problems seriously. So in this study, a kind of nonlinear stochastic-resonance (SR) signal enhancement technique combined with the distributed receiving array power spectrum fusion approach is proposed. By utilizing the signal power improvement property, the fitness of distributed SR system with the array structure, and the information fusion of the spectrum, the corresponding wireless positioning can be realized satisfactorily. Computer simulations also show the advantages over conventional methods in the direction of arrival (DoA) estimation and positioning estimation.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS3
2019 Positioning-Aided Scheme for Image Sensor Communication using Single-View Geometry
abstract
In this study, we investigate a positioning-aided scheme using computer vision techniques for image sensor communication (ISC), generally referred to as a visible light communication (VLC) system that utilizes an image sensor (camera). The image sensor in ISC is frequently used as a typical receiver because it has the ability to measure light intensity and distinguish multiple light sources. Additionally, thanks to its ability to detect the angle of arrival of light, the image sensor can estimate its own position using computer vision techniques. The positioning accuracy of ISC is reported to be in sub-meter levels that render visible light positioning (VLP). VLP is one of the most excellent positioning applications in indoor settings as global positioning system cannot provide satisfying indoor positioning services. To apply this significant positioning performance of VLP to our visible light system, we propose a positioning-aided scheme for ISC using computer vision techniques here. In addition, we evaluate the demodulation performance and positioning accuracy of our proposed scheme.
Zhengqiang Tang, Di He 0002, Shintaro Arai, Danping Zou
ISCAS4
2019 A Method of 2D Semantic Map Generation for Autonomous Flight of MAVs
abstract
In many applications, MAVs(Micro Aerial Vehicle) need to recognize the ground targets and locate them precisely, or in other words, to generate a semantic map of those targets, to guide the MAVs to complete mission automatically. In this paper, we introduce a method of fast 2D semantic map generation using only the images captured by the downward-looking camera in an unknown environment. The map is firstly stitched from each individual image by feature matching. Then the contour and area of each target are extracted through color detection in the global map. Finally, neural network is utilized for recognition of ground targets marked by different printed numbers. We have tested our method for automatic task performing as an MAV competition challenge in a virtual environment. By taking the generated 2D semantic map into the control loop, the MAV can localize itself, realize autonomous flight, detect and explore the environment.
Ruochen Yao, Danping Zou, Daniele Sartori, Ling Pei, Ling Gong
ISCAS2
2019 StructVIO: Visual-Inertial Odometry With Structural Regularity of Man-Made Environments
abstract
In this paper, we propose a novel visual-inertial odometry (VIO) approach that adopts structural regularity in man-made environments. Instead of using Manhattan world assumption, we use Atlanta world model to describe such regularity. An Atlanta world is a world that contains multiple local Manhattan worlds with different heading directions. Each local Manhattan world is detected on the fly, and their headings are gradually refined by the state estimator when new observations are received. With full exploration of structural lines that aligned with each local Manhattan worlds, our VIO method becomes more accurate and robust, as well as more flexible to different kinds of complex man-made environments. Through benchmark tests and real-world tests, the results show that the proposed approach outperforms existing visual-inertial systems in large-scale man-made environments.
Danping Zou, Yuanxin Wu, Ling Pei, Haibin Ling, Wenxian Yu
IEEE Trans. Robotics1
2019 Collaborative visual SLAM for multiple agents: A brief survey
abstract
This article presents a brief survey to visual simultaneous localization and mapping (SLAM) systems applied to multiple independently moving agents, such as a team of ground or aerial vehicles, a group of users holding augmented or virtual reality devices. Such visual SLAM system, name as collaborative visual SLAM, is different from a typical visual SLAM deployed on a single agent in that information is exchanged or shared among different agents to achieve better robustness, efficiency, and accuracy. We review the representative works on this topic proposed in the past ten years and describe the key components involved in designing such a system including collaborative pose estimation and mapping tasks, as well as the emerging topic of decentralized architecture. We believe this brief survey could be helpful to someone who are working on this topic or developing multi-agent applications, particularly micro-aerial vehicle swarm or collaborative augmented/virtual reality.
Danping Zou, Wenxian Yu
Virtual Real. Intell. Hardw.1
2018 Active Image-Based Modeling with a Toy Drone
abstract
Image-based modeling techniques [1]-[3] can now generate photo-realistic 3D models from images. But it is up to users to provide high quality images with good coverage and view overlap, which makes the data capturing process tedious and time consuming. We seek to automate data capturing for image-based modeling. The core of our system is an iterative linear method to solve the multi-view stereo (MVS) problem quickly and plan the Next-Best-View (NBV) effectively. Our fast MVS algorithm enables online model reconstruction and quality assessment to determine the NBVs on the fly. We test our system with a toy unmanned aerial vehicle (UAV) in simulated, indoor and outdoor experiments. Results show that our system improves the efficiency of data acquisition and ensures the completeness of the final model.
Rui Huang 0001, Danping Zou, Richard Vaughan 0001, Ping Tan 0002
ICRA2
2018 An Improved Kernel Clustering Algorithm Used in Computer Network Intrusion Detection
abstract
In the computer network intrusion detection system, data objects, mapped from original space, are analyzed based on kernel clustering algorithm. During the process of kernel clustering, some representative points are introduced to represent a cluster. In just one iteration, the distance between a data object and representative points of a cluster is computed to partition the data objects. The clustering results contain normal data and abnormal data, which achieve the goal of intrusion detection. At the same time, KDD CUP 1999 dataset is used to make simulations. The results show that the proposed algorithm has higher detecting probability under the condition of low constant false alarm rate compared with K-Means clustering algorithm and SVM.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS3
2018 Uplink Power Control Approach Based on Adaptive Chaotic Simulated Annealing
abstract
A kind of uplink power control approach of asynchronous code division multiple access (CDMA) mobile communication system based on the adaptive chaotic simulated annealing (CSA) is proposed. The particular influence of near-far effect can be solved by using the global optimal searching function of CSA method, and the multiple access interference (MAI) problem can be overcome by introducing adaptive MAI cancellation network in corresponding close loop control strategy. Computer simulations show the comparison results between proposed approach and conventional method with heterogeneous signal-to-interference ratio (SIR) threshold decision. It shows that this new approach can achieve higher channel capability and lower average outrage probability with some small increase of iteration steps, which can improve the performance of CDMA mobile communication system effectively.
Di He 0002, Xin Chen 0017, Danping Zou, Ling Pei, Ling-ge Jiang
ISCAS3
2017 An aerodynamic model-aided state estimator for multi-rotor UAVs
abstract
A robust state estimator is presented by fusing the aerodynamic model of multi-rotor UAVs with measurements from optical flow and other low-cost sensors such as IMU, magnetometer, and ultrasonic sensor. Due to the particular aerodynamics of multi-rotor UAVs, the body velocity in the rotor plane is able to be measured by the accelerometer. We therefore propose a novel state estimator by fully exploring the characteristic of aerodynamics of multi-rotor UAV. Our state estimator is fast and easy to be implemented. We have tested our estimator with different platforms in different scenes. Experimental results show that our estimator performs robustly in low light conditions where existing methods usually fail.
Rongzhi Wang, Danping Zou, Ling Pei, Wenxian Yu
IROS2
2015 Simultaneous video defogging and stereo reconstruction
abstract
We present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We first improve the photo-consistency term to explicitly model the appearance change due to the scattering effects. The prior matting Laplacian constraint on fog transmission imposes a detail-preserving smoothness constraint on the scene depth. We further enforce the ordering consistency between scene depth and fog transmission at neighboring points. These novel constraints are formulated together in an MRF framework, which is optimized iteratively by introducing auxiliary variables. The experiment results on real videos demonstrate the strength of our method.
Zhuwen Li, Robby T. Tan, Danping Zou, Steven Zhiying Zhou, Loong Fah Cheong
CVPR4
2013 CoSLAM: Collaborative Visual SLAM in Dynamic Environments
abstract
This paper studies the problem of vision-based simultaneous localization and mapping (SLAM) in dynamic environments with multiple cameras. These cameras move independently and can be mounted on different platforms. All cameras work together to build a global map, including 3D positions of static background points and trajectories of moving foreground points. We introduce intercamera pose estimation and intercamera mapping to deal with dynamic objects in the localization and mapping process. To further enhance the system robustness, we maintain the position uncertainty of each map point. To facilitate intercamera operations, we cluster cameras into groups according to their view overlap, and manage the split and merge of camera groups in real time. Experimental results demonstrate that our system can work robustly in highly dynamic environments and produce more accurate results in static environments.
Danping Zou
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Reconstructing 3D motion trajectories of particle swarms by global correspondence selection
abstract
This paper addresses the problem of reconstructing the 3D motion trajectories of particle swarms using two temporally synchronized and geometrically calibrated cameras. The 3D trajectory reconstruction problem involves two challenging tasks - stereo matching and temporal tracking. Existing methods separate the two and process them one at a time sequentially, and suffer from frequent irresolvable ambiguities in stereo matching and in tracking. We unify the two tasks, and propose a Global Correspondence Selection scheme to solve stereo matching and temporal tracking simultaneously. It treats 3D trajectory acquisition problem as selecting appropriate stereo correspondences among all possible ones for each target by minimizing a cost function. Experiment results show that the proposed method has significant performance advantage over existing approaches.
Danping Zou, Hai Shan Wu, Yan Qiu Chen
ICCV1
2007 Relative Epipolar Motion of Tracked Features for Correspondence in Binocular Stereo
abstract
Most 3D reconstruction solutions focus on surfaces, and there has not been much research attention paid to the problem of reconstructing 3D scenes made up of large numbers of particles, while the ability to reconstruct such dynamic scenes is potentially very useful in many areas such as colony behavior research and visual modeling. This paper proposes an approach - relative epipolar motion (REM) - towards solving the correspondence problem in stereopsis by utilizing the motion clue. It matches feature trajectories instead of the features themselves as used by existing methods. The proposed method has the following new capabilities: (1) it supports reconstructing dynamic 3D scenes of large number of undistinguishable drifting particles; (2) It is applicable to correspondence establishment for dynamic surfaces made up of repetitive textures; (3) It offers an alternative way to project structured light in active mode for deforming surface reconstruction. Experiment results on both simulated and real-world scenes demonstrate its effectiveness.
Hao Du 0004, Danping Zou, Yan Qiu Chen
ICCV2