EDBT 2026 Demo / reviewers in the wild / expert
Junqiao Zhao
dblp:27/7856
· DBLP profile ↗
29ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-7864-3255ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 13 since 2021Systems, architecture and hardware · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LPRFusion: Asymmetric cascade RoI refinement with LiDAR-pseudo point cloud fusion for 3D object detection
Yan Wu 0011, Yujian Mo, Junqiao Zhao, Jun Yan 0009 |
Neurocomputing | 4 |
| 2025 | Scrutinize What We Ignore: Reining In Task Representation Shift Of Context-Based Offline Meta Reinforcement LearningabstractOffline meta reinforcement learning (OMRL) has emerged as a promising approach for interaction avoidance and strong generalization performance by leveraging pre-collected data and meta-learning techniques.
Previous context-based approaches predominantly rely on the intuition that alternating optimization between the context encoder and the policy can lead to performance improvements, as long as the context encoder follows the principle of maximizing the mutual information between the task variable $M$ and its latent representation $Z$ ($I(Z;M)$) while the policy adopts the standard offline reinforcement learning (RL) algorithms conditioning on the learned task representation.
Despite promising results, the theoretical justification of performance improvements for such intuition remains underexplored.
Inspired by the return discrepancy scheme in the model-based RL field, we find that the previous optimization framework can be linked with the general RL objective of maximizing the expected return, thereby explaining performance improvements.
Furthermore, after scrutinizing this optimization framework, we observe that the condition for monotonic performance improvements does not consider the variation of the task representation. When these variations are considered, the previously established condition may no longer be sufficient to ensure monotonicity, thereby impairing the optimization process.
We name this issue \underline{task representation shift} and theoretically prove that the monotonic performance improvements can be guaranteed with appropriate context encoder updates.
We use different settings to rein in the task representation shift on three widely adopted training objectives concerning maximizing $I(Z;M)$ across different data qualities.
Empirical results show that reining in the task representation shift can indeed improve performance.
Our work opens up a new avenue for OMRL, leading to a better understanding between the task representation and performance improvements. Tianying Ji, Jinhang Liu, Anqi Guo, Junqiao Zhao, Lanqing Li |
ICLR | 6 |
| 2025 | CERTAIN: Context Uncertainty-aware One-Shot Adaptation for Context-based Offline Meta Reinforcement LearningabstractExisting context-based offline meta-reinforcement learning (COMRL) methods primarily focus on task representation learning and given-context adaptation performance. They often assume that the adaptation context is collected using task-specific behavior policies or through multiple rounds of collection. However, in real applications, the context should be collected by a policy in a one-shot manner to ensure efficiency and safety. We find that intrinsic context ambiguity across multiple tasks and out-of-distribution (OOD) issues due to distribution shift significantly affect the performance of one-shot adaptation, which has been largely overlooked in most COMRL research. To address this problem, we propose using heteroscedastic uncertainty in representation learning to identify ambiguous and OOD contexts, and train an uncertainty-aware context collecting policy for effective one-shot online adaptation. The proposed method can be integrated into various COMRL frameworks, including classifier-based, reconstrution-based and contrastive learning-based approaches. Empirical evaluations on benchmark tasks show that our method can improve one-shot adaptation performance by up to 36% and zero-shot adaptation performance by up to 34% compared to existing baseline COMRL methods. Hongtu Zhou, Ruiling Yang, Yakun Zhu, Haoqi Zhao, Junqiao Zhao, Chen Ye 0002 |
ICML | 7 |
| 2025 | Convex Hull-based Algebraic Constraint for Visual Quadric SLAMabstractUsing Quadrics as the object representation has the benefits of both generality and closed-form projection derivation between image and world spaces. Although numerous constraints have been proposed for dual quadric reconstruction, we found that many of them are imprecise and provide minimal improvements to localization. After scrutinizing the existing constraints, we introduce a concise yet more precise convex hull-based algebraic constraint for object landmarks, which is applied to object reconstruction, frontend pose estimation, and backend bundle adjustment. This constraint is designed to fully leverage precise semantic segmentation, effectively mitigating mismatches between complex-shaped object contours and dual quadrics. Experiments on public datasets demonstrate that our approach is applicable to both monocular and RGB-D SLAM and achieves improved object mapping and localization than existing quadric SLAM methods. The implementation of our method is available at https://github.com/tiev-tongji/convexhull-based-algebraic-constraint. Junqiao Zhao, Shuangfu Song, Zhongyang Zhu, Zihan Yuan, Chen Ye 0002, Tiantian Feng, Qiankun Yu |
IROS | 2 |
| 2025 | MutualVPR: A Mutual Learning Framework for Resolving Supervision Inconsistencies via Adaptive ClusteringabstractVisual Place Recognition (VPR) enables robust localization through image retrieval based on learned descriptors.
However, drastic appearance variations of images at the same place caused by viewpoint changes can lead to inconsistent supervision signals, thereby degrading descriptor learning.
Existing methods either rely on manually defined cropping rules or labeled data for view differentiation, but they suffer from two major limitations:
(1) reliance on labels or handcrafted rules restricts generalization capability;
(2) even within the same view direction, occlusions can introduce feature ambiguity.
To address these issues, we propose MutualVPR, a mutual learning framework that integrates unsupervised view self-classification and descriptor learning.
We first group images by geographic coordinates, then iteratively refine the clusters using K-means to dynamically assign place categories without manual labeling.
Specifically, we adopt a DINOv2-based encoder to initialize the clustering.
During training, the encoder and clustering co-evolve, progressively separating drastic appearance variations of the same place and enabling consistent supervision.
Furthermore, we find that capturing fine-grained image differences at a place enhances robustness.
Experiments demonstrate that MutualVPR achieves state-of-the-art (SOTA) performance across multiple datasets, validating the effectiveness of our framework in improving view direction generalization, occlusion robustness. Qiwen Gu, Xufei Wang, Junqiao Zhao, Siyue Tao, Tiantian Feng, Guang Chen 0001 |
NeurIPS | 3 |
| 2025 | CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningabstractConstrained hybrid-action reinforcement learning (RL) promises to learn a safe policy within a parameterized action space, which is particularly valuable for safety-critical applications involving discrete-continuous hybrid action spaces. However, existing hybrid-action RL algorithms primarily focus on reward maximization, which faces significant challenges for tasks involving both cost constraints and hybrid action spaces. In this work, we propose a novel Constrained Hybrid-action Policy Optimization algorithm (CHPO) to address the problems of constrained hybrid-action RL. Concretely, we rethink the limitations of hybrid-action RL in handling safe tasks with parameterized action spaces and reframe the objective of constrained hybrid-action RL by introducing the concept of Constrained Parameterized-action Markov Decision Process (CPMDP). Subsequently, we present a constrained hybrid-action policy optimization algorithm to confront the constrained hybrid-action problems and conduct theoretical analyses demonstrating that the CHPO converges to the optimal solution while satisfying safety constraints. Finally, extensive experiments demonstrate that the CHPO achieves competitive performance across multiple experimental tasks. Ao Zhou 0005, Jiayi Guan, Li Shen 0008, Fan Lu 0001, Sanqing Qu, Junqiao Zhao, Guang Chen 0001 |
NeurIPS | 6 |
| 2024 | Sparse Query Dense: Enhancing 3D Object Detection with Pseudo PointsabstractCurrent LiDAR-only 3D detection methods are limited by the sparsity of point clouds. The previous method used pseudo points generated by depth completion to supplement the LiDAR point cloud, but the pseudo points sampling process was complex, and the distribution of pseudo points was uneven. Meanwhile, due to the imprecision of depth completion, the pseudo points suffer from noise and local structural ambiguity, which limit the further improvement of detection accuracy. This paper presents SQDNet, a novel framework designed to address these challenges. SQDNet incorporates two key components: the SQD, which achieves sparse-to-dense matching via grid position indices, allowing for rapid sampling of large-scale pseudo points on the dense depth map directly, thus streamlining the data preprocessing pipeline. And use the density of LiDAR points within these grids to alleviate the uneven distribution and noise problems of pseudo points. Meanwhile, the sparse 3D Backbone is designed to capture long-distance dependencies, thereby improving voxel feature extraction and mitigating local structural blur in pseudo points. The experimental results validate the effectiveness of SQD and achieve considerable detection performance for difficult-to-detect instances on the KITTI test. Yujian Mo, Yan Wu 0011, Junqiao Zhao, Zhenjie Hou, Weiquan Huang, Jun Yan 0009 |
ACM Multimedia | 3 |
| 2024 | Towards an Information Theoretic Framework of Context-Based Offline Meta-Reinforcement LearningabstractAs a marriage between offline RL and meta-RL, the advent of offline meta-reinforcement learning (OMRL) has shown great promise in enabling RL agents to multi-task and quickly adapt while acquiring knowledge safely. Among which, context-based OMRL (COMRL) as a popular paradigm, aims to learn a universal policy conditioned on effective task representations. In this work, by examining several key milestones in the field of COMRL, we propose to integrate these seemingly independent methodologies into a unified framework. Most importantly, we show that the pre-existing COMRL algorithms are essentially optimizing the same mutual information objective between the task variable $M$ and its latent representation $Z$ by implementing various approximate bounds. Such theoretical insight offers ample design freedom for novel algorithms. As demonstrations, we propose a supervised and a self-supervised implementation of $I(Z; M)$, and empirically show that the corresponding optimization algorithms exhibit remarkable generalization across a broad spectrum of RL benchmarks, context shift scenarios, data qualities and deep learning architectures. This work lays the information theoretic foundation for COMRL methods, leading to a better understanding of task representation learning in the context of reinforcement learning. Given its
generality, we envision our framework as a promising offline pre-training paradigm of foundation models for decision making. Lanqing Li, Shatong Zhu, Junqiao Zhao, Pheng-Ann Heng |
NeurIPS | 6 |
| 2024 | Focus On What Matters: Separated Models For Visual-Based RL GeneralizationabstractA primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliary tasks to enhance generalization, few adopt image reconstruction due to concerns about exacerbating overfitting to task-irrelevant features during training. Perceiving the pre-eminence of image reconstruction in representation learning, we propose SMG (\blue{S}eparated \blue{M}odels for \blue{G}eneralization), a novel approach that exploits image reconstruction for generalization. SMG introduces two model branches to extract task-relevant and task-irrelevant representations separately from visual observations via cooperatively reconstruction. Built upon this architecture, we further emphasize the importance of task-relevant features for generalization. Specifically, SMG incorporates two additional consistency losses to guide the agent's focus toward task-relevant areas across different scenarios, thereby achieving free from overfitting. Extensive experiments in DMC demonstrate the SOTA performance of SMG in generalization, particularly excelling in video-background settings. Evaluations on robotic manipulation tasks further confirm the robustness of SMG in real-world applications. Source code is available at \url{https://anonymous.4open.science/r/SMG/}. Bowen Lv, Junqiao Zhao, Chang Huang, Hongtu Zhou, Chen Ye 0002 |
NeurIPS | 5 |
| 2023 | Robust Traffic Light Recognition Pipeline Based on YOLOv8 for Autonomous Driving SystemsabstractTraffic Light Recognition (TLR) aims at detecting Traffic Lights (TLs) and then classifying the status of light signals, being an essential constituent of autonomous driving perception systems. However, it’s challenging for existing TLR methods to accurately distinguish both color and shape status of TLs due to small object sizes, illumination variations, close resemblance with other objects and varying weather conditions. Existing public datasets for TLR have three main drawbacks:(i) poor diversity, (ii) sample imbalance, and (iii) insufficient category labels, greatly hindering the development of TLR. To overcome the aforementioned problems, we propose a Robust Traffic Light Recognition Pipeline based on YOLOv8 (RTLRP-YOLO) that can recognize TLs accurately with strong robustness based on adaptively generated high-quality images. Specifically, we develop a Self-Adaptive Preprocessing Module (SAPM) which is designed to adaptively generate high-quality images under hostile conditions, followed by a Two-stage Traffic Light Recognition Model based on YOLOv8 (TTRM) to obtain both the location and status information of TLs. Moreover, We also provide our self-made Tongji Small Traffic Light Dataset (TSTLD), covering a variety of weather conditions, regions, light intensities and shooting angles. To the best of our knowledge, our proposed method is the first one to be able of simultaneously identifying three colors (i.e., red, yellow and green) and four shapes (i.e., circle, left arrow, right arrow and up arrow) of TLs, achieving 95.43% accuracy on TSTLD with the inference time of 26 ms for per image. Yan Wu 0011, Junqiao Zhao |
ICPADS | 3 |
| 2023 | Multi-agent Decision-making at Unsignalized Intersections with Reinforcement Learning from DemonstrationsabstractIntersections are key nodes and also bottlenecks of urban road networks, so improving the traffic efficiency at intersections is beneficial to improving overall traffic throughput and mitigating traffic congestion. Previous methods such as rule-based, planning-based, and single-agent reinforcement learning usually oversimplify the policies of the surrounding vehicles and thus have difficulty modeling the complex interaction behaviors between vehicles, which limits the performance of these methods to some extent. Instead, we adopt a multi-agent reinforcement learning (MARL) approach to train and coordinate the policies of all vehicles to handle unsignalized intersection scenarios. Nevertheless, due to complex interactions between multiple agents, it is challenging to efficiently explore the environment and obtain high-reward samples. We therefore propose to pre-train the policy using demonstration data consisting of expert data and interaction data to improve the initial performance of agents and improve exploration, as well as to reduce the distributional shift between the demonstration data and the environmental interaction data. We experimentally prove that using interaction data generated by the algorithm in the demonstration data improves training stability. The proposed method enables effective exploration and greatly speeds up the training process. Chang Huang, Junqiao Zhao, Hongtu Zhou, Chen Ye 0002 |
IV | 2 |
| 2023 | How to Fine-tune the Model: Unified Model Shift and Model Bias Policy OptimizationabstractDesigning and deriving effective model-based reinforcement learning (MBRL) algorithms with a performance improvement guarantee is challenging, mainly attributed to the high coupling between model learning and policy optimization. Many prior methods that rely on return discrepancy to guide model learning ignore the impacts of model shift, which can lead to performance deterioration due to excessive model updates. Other methods use performance difference bound to explicitly consider model shift. However, these methods rely on a fixed threshold to constrain model shift, resulting in a heavy dependence on the threshold and a lack of adaptability during the training process. In this paper, we theoretically derive an optimization objective that can unify model shift and model bias and then formulate a fine-tuning process. This process adaptively adjusts the model updates to get a performance improvement guarantee while avoiding model overfitting. Based on these, we develop a straightforward algorithm USB-PO (Unified model Shift and model Bias Policy Optimization). Empirical results show that USB-PO achieves state-of-the-art performance on several challenging benchmark tasks. Junqiao Zhao, Hongtu Zhou, Chang Huang, Chen Ye 0002 |
NeurIPS | 3 |
| 2022 | Scale Estimation with Dual Quadrics for Monocular Object SLAMabstractThe scale ambiguity problem is inherently unsolvable to monocular SLAM without the metric baseline between moving cameras. In this paper, we present a novel scale estimation approach based on an object-level SLAM system. To obtain the absolute scale of the reconstructed map, we formulate an optimization problem to make the scaled dimensions of objects conform to the distribution of their sizes in the physical world, without relying on any prior information about gravity direction. The dual quadric is adopted to represent objects for its ability to describe objects compactly and accurately, thus providing reliable dimensions for scale estimation. In the proposed monocular object-level SLAM system, semantic objects are initialized first from fitted 3-D oriented bounding boxes and then further optimized under constraints of 2-D detections and 3-D map points. Experiments on indoor and outdoor public datasets show that our approach outperforms existing methods in terms of accuracy and robustness. Shuangfu Song, Junqiao Zhao, Tiantian Feng, Chen Ye 0002, Lu Xiong 0001 |
IROS | 2 |
| 2022 | Efficient Adaptive Upsampling Module for Real-Time Semantic SegmentationabstractUpsampling operation is necessary for semantic segmentation and other pixel-level prediction tasks. Among the commonly used upsampling operations, some are too simple to effectively recover the spatial details lost during downsampling process, and some are too complex and have high computation complexity. In real-world applications, it is critical to achieve high accuracy and maintain real-time inference speed. Therefore, an efficient upsampling operation is essential for these tasks. In this paper, we introduce efficient adaptive upsampling module (EAUM) for real-time semantic segmentation. Inspired by dynamic filter networks, EAUM adaptively predicts the kernel weight of each point in the upsampled feature map according to the corresponding points in the input feature map. To reduce computational cost, EAUM decomposes the spatial information and channel information required for upsampling. The proposed EAUM shows impressive performance on Cityscapes and CamVid benchmarks. Specifically, DenseENet with EAUM outperforms the baseline by 1.4% [Formula: see text] and 1.6% [Formula: see text] in accuracy with a slight drop in inference speed on Cityscapes test dataset. Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu, Yujun Liao, Yujian Mo |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2022 | G-VIDO: A Vehicle Dynamics and Intermittent GNSS-Aided Visual-Inertial State Estimator for Autonomous DrivingabstractThis paper proposes G-VIDO, a vehicle dynamics, and intermittent Global Navigation Satellite System (GNSS)-aided visual-inertial state estimator, to address the state estimation problem of autonomous vehicle localization (i.e., position and orientation estimation in the global coordinate system) under various GNSS states. A dynamics pre-integration theory is proposed on the basis of a two-degree-of-freedom (DOF) vehicle dynamics model, and dynamics constraints are built in the optimization back-end, considering the unobservable problem of the monocular visual-inertial system under degenerate motions. The proposed highly nonlinear system can be robustly initialized by loosely aligning the monocular structure from motion (SfM) results, pre-integrated IMU measurements, and vehicle motion information. GNSS is used for reference frame transformation and constraint construction in the sliding window. The cumulative error can be corrected with the aid of GNSS, and the vehicle’s position in the global coordinate system can be determined. A GNSS anomaly detection algorithm is proposed to improve the system robustness under intermittent GNSS. Experiments have shown that G-VIDO can provide real-time, robust, and seamless localization in multiple GNSS states, with an RMSE of less than 30 cm (with GNSS). Moreover, we proved that the initialization and local odometry modules in G-VIDO outperform several state-of-the-art VIO systems and our preliminary work VINS-Vehicle. Lu Xiong 0001, Rong Kang, Junqiao Zhao, Peizhi Zhang, Ran Ju, Chen Ye 0002, Tiantian Feng |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | GPU-Efficient Dense Convolutional Network for Real-time Semantic SegmentationabstractReal-time semantic segmentation is a challenging task as both accuracy and inference speed need to be considered simultaneously. In real-world applications, it is usually achieved by deploying a deep neural network in modern GPU device. However, most of the work focused on real-time semantic segmentation is designed by significantly reducing computation complexity and model size. There are other factors that have a significant impact on inference speed are overlooked, especially when the network is running in modern GPU device. In this paper, we focus on designing a GPU-efficient network as backbone for real-time semantic segmentation. Dense connectivity can preserve and accumulate feature maps of multiple receptive fields and is therefore ideal for semantic segmentation. Therefore, we design a GPU-efficient network (DenseENet) with dense connectivity. The proposed DenseENet shows an obvious advantage in balancing accuracy and inference speed in modern GPU device. Specifically, on Cityscapes test set, DenseENet with a simple FCN decoder achieves 75.2% mIoU with 83.6 FPS for an input of 1024 × 2048 resolution and 73.6% mIoU with 132 FPS for an input of 768 × 1536 resolution on a single GTX 1080Ti card. Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu |
ICRA | 3 |
| 2020 | Dense Dual-Path Network for Real-Time Semantic Segmentation
Xinneng Yang, Yan Wu 0011, Junqiao Zhao, Feilin Liu |
ACCV (1) | 3 |
| 2020 | Occlusion aware unsupervised learning of optical flow from videoabstractIn this paper, we proposed an unsupervised learning method for estimating the optical flow between video frames, especially to solve the occlusion problem. Occlusion is caused by the movement of an object or the movement of the camera, defined as when certain pixels are visible in one video frame but not in adjacent frames. Due to the lack of pixel correspondence between frames in the occluded area, incorrect photometric loss calculation can mislead the optical flow training process. In the video sequence, we found that the occlusion in the forward (t→t+1) and backward (t→t-1) frame pairs are usually complementary. That is, pixels that are occluded in subsequent frames are often not occluded in the previous frame and vice versa. Therefore, by using this complementarity, a new weighted loss is proposed to solve the occlusion problem. Our method achieves competitive optical flow accuracy compared to the baseline and some supervised methods on KITTI and Sintel benchmarks. Junqiao Zhao, Tiantian Feng |
ICMV | 2 |
| 2019 | DFNet: Semantic Segmentation on Panoramic Images with Dynamic Loss Weights and Residual Fusion BlockabstractFor the domain of self-driving and automatic parking, perception is a basic and critical technique, moreover, the detection of lane markings and parking slots is an important part of visual perception. Compared with front sight images, panoramic images(PI) can capture more comprehensive pavement information. However, the imbalance of different classes in PI is even more serious. Additionally, the judgment of boundary information between areas is a hard problem in deep models. Therefore, we propose a new model named DFNet to solve these problems. The proposed model has two main contributions, one is dynamic loss weights, and the other is residual fusion block(RFB). DFNet use dynamic loss weights to overcome the negative effect of imbalance dataset, which are calculated according to the pixel number of each class in a batch. RFB is composed of several convolutional layers, a pooling layer, and a fusion layer to combine the feature maps by pixel multiplication, which can reduce boundary information loss. We evaluate our method on PSV dataset, and the achieved advanced results demonstrate the effectiveness of the proposed model. Yan Wu 0011, Linting Guan, Junqiao Zhao |
ICRA | 4 |
| 2019 | Super Resolution Reconstruction Technique in Passive Microwave Images of Arctic Sea IceabstractPolar sea ice is one of the key parameters of cryosphere and polar environmental change, which plays an important role in the study of global climate change. High-resolution monitoring of polar sea ice relies mainly on optical satellite imagery and synthetic aperture radar (SAR) data, with limited spatial and temporal coverages for many applications. Passive microwave data is an important data source for continuous observations of polar sea ice, thanks to its working ability in all-sky conditions and its wide coverage. However, it is difficult to achieve high-resolution monitoring of polar sea ice using passive microwave data due to its coarse resolution. In order to solve this problem, super resolution (SR) reconstruction technique is adopted in this paper to improve the spatial resolution of passive microwave images. SR reconstruction technique based on both single-image and multi-image are attempted. AMSR2 level 3 (L3) Products of Brightness Temperatures (BTs) for Arctic sea ice are used as experimental data, and the reconstruction results obtained from different SR methods are compared and discussed. Tiantian Feng, Junqiao Zhao, Rongxing Li |
IGARSS | 3 |
| 2019 | DL-SLAM: Direct 2.5D LiDAR SLAM for Autonomous DrivingabstractPrecisely localizing a vehicle in the GNSS-denied urban area is crucial for autonomous driving. The occupancy grid-based 2D LiDAR SLAM methods scale poorly to outdoor road scenarios, while the 3D point cloud-based LiDAR SLAM methods suffer from huge computation and storage costs. Aiming at the precise real-time LiDAR SLAM for both indoor and outdoor, this paper proposed a direct 2.5D heightmap-based SLAM system. This system extended our previously proposed DLO (the direct 2.5D LiDAR odometry) method by introducing the 2.5D segment features for efficient loop closure detection. We experimented our SLAM method on the KITTI datasets and shown it superior performance compared with the existing LiDAR SLAM methods. Junqiao Zhao, Yuchen Kang, Chen Ye 0002, Lu Sun 0002 |
IV | 2 |
| 2019 | A Novel Robust Lane Change Trajectory Planning Method for Autonomous VehicleabstractA novel trajectory planning method is proposed in this paper for lane change of autonomous vehicle. Since it is difficult to accurately capture the trajectory of other vehicles, which means the trajectory for autonomous vehicle couldn't always easy to generate quickly. Moreover, the motion planning, as a kind of high-dimensional optimization problem with multiple nonlinear constraints, requires lots of resources to find a right solution. Therefore, we present a trajectory monitoring strategy to keep robust in lane change scenario, which generates the lane change and monitoring trajectory at the same time. If the former does not produce a safe trajectory or is time out, the monitoring trajectory will be taken as the result output. To meet the constraints of vehicle's motion and real-time requirements, B-spline-based method will be employed to plan a continuous curvature path. And RRT-based method works as a supplement for keeping algorithm completeness. Then the monitory trajectory mainly obeys collision-free requirements, which computes deceleration that keeps vehicle stability. The results illustrate that both B-spline-based method and RRT-based could generate curvature continuous and meet the limitation for motion, however, both have the possibility of timeout. Especially, there are challenge to the success rate as environment becomes more complex. Dequan Zeng, Zhuoping Yu, Lu Xiong 0001, Junqiao Zhao, Peizhi Zhang, Zhiqiang Fu |
IV | 4 |
| 2018 | Learn to Detect Objects IncrementallyabstractIntelligent vehicles need to detect new classes of traffic objects while keeping the performance of old ones. Deep convolution neural network (DCNN) based detector has shown superior performance, however, DCNN is ill-equipped for incremental learning, i.e., a DCNN based vehicle detector trained on traffic sign dataset will catastrophic forget how to detect vehicles. In this paper, we propose a novel method to alleviate this problem, our key insight is that the original class of objects also appears in new task data, by utilizing these objects, our method effectively keeps the detection accuracy of original models while incremental learning to detect new classes of objects. Detailed experiments on PASCAL VOC dataset and TSD-max database verified the effectiveness of our method. Linting Guan, Yan Wu 0011, Junqiao Zhao, Chen Ye 0002 |
Intelligent Vehicles Symposium | 3 |
| 2018 | Vision-based Semantic Mapping and Localization for Autonomous Indoor ParkingabstractIn this paper, we proposed a novel and practical solution for the real-time indoor localization of autonomous driving in parking lots. High-level landmarks, the parking slots, are extracted and enriched with labels to avoid the aliasing of low-level visual features. We then proposed a robust method for detecting incorrect data associations between parking slots and further extended the optimization framework by dynamically eliminating suboptimal data associations. Visual fiducial markers are introduced to improve the overall precision. As a result, a semantic map of the parking lot can be established fully automatically and robustly. We experimented the performance of real-time localization based on the map using our autonomous driving platform TiEV, and the average accuracy of 0.3m track tracing can be achieved at a speed of 10kph. Yewei Huang 0001, Junqiao Zhao, Shaoming Zhang, Tiantian Feng |
Intelligent Vehicles Symposium | 2 |
| 2018 | DLO: Direct LiDAR Odometry for 2.5D Outdoor EnvironmentabstractFor autonomous vehicles, high-precision real-time localization is the guarantee of stable driving. Compared with the visual odometry (VO), the LiDAR odometry (LO) has the advantages of higher accuracy and better stability. However, 2D LO is only suitable for the indoor environment, and 3D LO has less efficiency in general. Both are not suitable for the online localization of an autonomous vehicle in an outdoor driving environment. In this paper, a direct LO method based on the 2.5D grid map is proposed. The fast semi-dense direct method proposed for VO is employed to register two 2.5D maps. Experiments show that this method is superior to both the 3D-NDT and LOAM in the outdoor environment. Lu Sun 0002, Junqiao Zhao, Chen Ye 0002 |
Intelligent Vehicles Symposium | 2 |
| 2018 | VH-HFCN based Parking Slot and Lane Markings Segmentation on Panoramic Surround ViewabstractThe automatic parking is being massively developed by car manufacturers and providers. Until now, there are two problems with the automatic parking. First, there is no openly-available segmentation labels of parking slot on panoramic surround view (PSV) dataset. Second, how to detect parking slot and road structure robustly. Therefore, in this paper, we build up a public PSV dataset. At the same time, we proposed a highly fused convolutional network (HFCN) based segmentation method for parking slot and lane markings based on the PSV dataset. A surround-view image is made of four calibrated images captured from four fisheye cameras. We collect and label more than 4,200 surround view images for this task, which contain various illuminated scenes of different types of parking slots. A VH-HFCN network is proposed, which adopts an HFCN as the base, with an extra efficient VH-stage for better segmenting various markings. The VH-stage consists of two independent linear convolution paths with vertical and horizontal convolution kernels respectively. This modification enables the network to robustly and precisely extract linear features. We evaluated our model on the PSV dataset and the results showed outstanding performance in ground markings segmentation. Based on the segmented markings, parking slots and lanes are acquired by skeletonization, hough line transform and line arrangement. Yan Wu 0011, Tao Yang 0044, Junqiao Zhao, Linting Guan |
Intelligent Vehicles Symposium | 3 |
| 2017 | Fully Combined Convolutional Network with Soft Cost Function for Traffic Scene Parsing
Yan Wu 0011, Tao Yang 0044, Junqiao Zhao, Linting Guan, Jiqian Li |
ICIC (1) | 3 |
| 2017 | Pedestrian detection with dilated convolution, region proposal network and boosted decision treesabstractWith the rapid development of driverless cars, pedestrian detection has been a canonical instance of object detection. Although recent deep learning detectors such as RPN+BF and MS-CNN have shown excellent performance for pedestrian detection, they have limited success for detecting pedestrian, and the importance of final feature receptive field has been awared by previous leading deep learning pedestrian detectors. Applying the dilated convolution to the feature learning of pedestrian detection, we constructed a pedestrian detection framework along with the region proposal network and boosted decision trees. Pipeline of our proposed framework can be briefly generalized as follows: firstly, the fine-tuned RPN with specified aspect ratio is used to get boxes and scores. Secondly, the designed dilated convolution feature extraction model is used to get features. As different dilation factors provide different receptive field scales, we concat the features of different layers with the dilated convolutional features to get the final features. Finally, the candidate boxes are sent to the boosted decision trees to be classified using the scores and features. We evaluated our method on the Caltech Pedestrian Detection Benchmark. Comparing with other state-of-the-art detection methods, the proposed framework with dilated convolution has better performance. Jiqian Li, Yan Wu 0011, Junqiao Zhao, Linting Guan, Chen Ye 0002, Tao Yang 0044 |
IJCNN | 3 |
| 2010 | Quantitative analysis of discrete 3D geometrical detail levels based on perceptual metric
Qing Zhu 0012, Junqiao Zhao, Yeting Zhang |
Comput. Graph. | 2 |