VLDB 2026 Research / reviewers in the wild / expert
Long Chen 0005
dblp:64/5725-5
· DBLP profile ↗
98ranked-venue papers
19as first author
49since 2021 · last 2026
0000-0003-4925-0572ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 8 since 2021Systems, architecture and hardware · 13 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 6 since 2021Computer networks · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy RobustnessabstractThe safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events. Qimao Chen, Shaoqing Xu, Zhiyi Lai, Zixun Xie, Yuechen Luo, Shengyin Jiang, Hanbing Li, Long Chen 0005 |
AAAI | 9 |
| 2026 | Enhancing Diffusion Policies with Distribution-Matching Generator in Offline Reinforcement LearningabstractOffline reinforcement learning (RL) can learn policies from pre-collected offline datasets without interacting with the environment, but it suffers from the issue of out-of-distribution (OOD). Recent methods use the generative adversarial paradigm to learn policies, but easily fail to handle the conflict of fooling the discriminator and maximizing expected returns. In this paper, we propose a novel offline RL method named Distribution-Matching Generator-based Diffusion Policies (DMGDP). A distribution matching-based policy learning method is first developed, where the diffusion serves as the policy generator, to handle the conflict of fooling the discriminator and maximizing expected returns. Furthermore, a policy confidence mechanism based on discriminator regularization is designed to prevent the agent from taking OOD actions, with the aim of robust generative adversarial learning. We conducted extensive experiments on the D4RL benchmarks, and the results demonstrate that DMGDP outperforms state-of-the-art methods. Xuemin Hu, Yingfen Xu, Bo Tang 0011, Long Chen 0005 |
AAAI | 5 |
| 2026 | Think before Go: Hierarchical Reasoning for Image-goal NavigationabstractPengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Lin Zhao, Long Chen, Zhi-Xin Yang, Nanning Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengna Li, Kangyi Wu, Shaoqing Xu, Long Chen 0005, Nanning Zheng 0001 |
ACL (1) | 6 |
| 2026 | ConsistentID: Portrait Generation With Multimodal Fine-Grained Identity PreservingabstractDiffusion-based technologies have made significant strides, particularly in personalized and customized facial generation. However, existing methods struggle to achieve high-fidelity and detailed identity (ID) consistency. This is mainly due to two challenges: insufficient fine-grained control over specific facial areas and the absence of a comprehensive strategy for ID preservation that accounts for both intricate facial details and the overall facial structure. To address these limitations, we introduce ConsistentID, an innovative method crafted for diverse identity-preserving portrait generation under fine-grained multimodal facial prompts, utilizing only a single reference image. ConsistentID comprises two core components: a multimodal facial prompt generator and an ID-preservation network. The facial prompt generator combines localized facial features, facial feature descriptions, and overall facial descriptions to enhance the precision of facial detail reconstruction. The ID-preservation network, optimized with a facial attention localization strategy, ensures consistent identity preservation across facial regions. Together, these components leverage fine-grained multimodal identity information to improve identity preservation accuracy significantly. To drive ConsistentID's training, we propose a fine-grained portrait dataset, FGID, with over 500,000 facial images, offering greater diversity and comprehensiveness than existing public facial datasets. Experimental results substantiate that our ConsistentID achieves exceptional precision and diversity in personalized facial generation, surpassing existing methods in the MyStyle dataset. In addition, although ConsistentID introduces more multimodal ID information, it still maintains rapid inference speed during the generation process. Jiehui Huang, Wenhui Song, Zheng Chong, Zhenchao Tang, Yuhao Cheng, Long Chen 0005, Yiqiang Yan, Shengcai Liao, Xiaodan Liang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | A Visual Benchmark for Autonomous Driving in Open-Pit MinesabstractIn recent years, intelligent vehicles operating in urban environments have demonstrated the capability to autonomously execute various tasks, such as object detection, lane detection, segmentation, etc. This advancement is facilitated by the extensive datasets accumulated by researchers, alongside advancements in intelligent algorithms, as well as significant breakthroughs in software and hardware. However, within the autonomous driving community, there is a scarcity of data regarding scenarios encountered in mining environments. This scarcity presents challenges and bottlenecks for the advancement of comprehensive autonomous driving systems and autonomousoperations. Although we previously released our dataset, AutoMine, which includes over 18 hours of driving data in open-pit mines, its scope is limited to two specific tasks. This scope limitation impedes the training and validation of the majority of algorithms for different tasks in this particular scenario. To broaden the scope of autonomous driving visual tasks in mining environments, we have curated a diverse collection encompassing multiple tasks, including detection, segmentation, tracking, etc. Additionally, we have established benchmarks and set up baselines for the aforementioned multiple tasks. By comparing the performance differences of visual algorithms between mining areas and other scenarios, we demonstrate the distinctive characteristics of mining regions in an intuitive manner. We have developed a suite of tools for converting annotated data into the standardized format used in existing driving datasets. Our aspiration is to establish data and benchmark foundations, supporting research endeavors in intelligent transportation within mining environments and autonomous driving in comprehensive scenarios. Our project website can be seen in AutoMine, and the dataset can be downloaded via AutoMine-Benchmark. Yuchen Li 0004, Luxi Li, Zhenshan Bing, Libo Sun 0002, Alois C. Knoll, Fei-Yue Wang 0001, Long Chen 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | Panoramic Active Visual System for Distant Traffic Sign RecognitionabstractTraffic sign recognition (TSR) is one of the most important visual perception tasks for autonomous vehicles. Traffic signs situated far away occupy only a few pixels in the images captured by the cameras of the autonomous vehicle, which presents a challenge for TSR methods to give accurate and reliable results. In this paper, we propose a panoramic active visual system (PAVS) for distant traffic signs recognition. It combines the advantages of the active pan-tilt-zoom(PTZ) camera for long-distance viewing and the advantage of the panoramic camera to have a large field of view. The traffic signs will be preliminarily detected and tracked in the panoramic image and the objects with confidence lower than the threshold will be further detected by the active PTZ camera to get more reliable results. The proposed PAVS are tested with different state-of-the-art (SoTA) TSR networks in real traffic scenes and the experimental results show that, compared to the passive visual system, the performances of the proposed PAVS improve 8.5% in F1-score, 5.3% in mAP, and 60.2% in mean first detection distance (mFD) in traffic sign recognition tasks. Xuan Yuwen, Ziwang Lu, Jicheng Chen 0001, Long Chen 0005, Hui Zhang 0019 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | Long- and Short-Term Constraint-Driven Safe Reinforcement Learning for Autonomous DrivingabstractSafe reinforcement learning (RL) is developed to handle high-risk decision-making tasks, such as autonomous driving (AD), by constraining expected safety violation costs as a training objective. However, existing safe RL methods only consider the long-term objective but ignore the short-term state safety of exploration in the training process. In addition, it is difficult to achieve a balance between cost and return expectations, leading to deterioration of learning performance. Unlike these methods, we propose a novel algorithm named long-and short-term constraints (LSTCs) for safe RL. The short-term constraint is proposed to enhance the short-term state safety that the vehicle explores, while the long-term constraint enhances the overall safety of the vehicle throughout the decision-making process, both of which are jointly used to enhance vehicle safety in the training process. Furthermore, we develop a safe RL method with dual-constraint optimization based on the Lagrange multiplier to optimize the training process for end-to-end AD, balancing the cost and return expectations. Comprehensive experiments were conducted on the MetaDrive simulator. The experimental results demonstrate that the success rate increases by 13% and the episode cost decreases by 0.26 compared to the best results of the comparative methods, showing that the proposed method has better safety in continuous control tasks and exhibits a higher exploration performance in long-distance decision-making tasks compared to SOTA methods. Xuemin Hu, Yijun Wen, Bo Tang 0011, Long Chen 0005 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2025 | 3D Annotation-Free Learning by Distilling 2D Open-Vocabulary Segmentation Models for Autonomous DrivingabstractPoint cloud data labeling is considered a time-consuming and expensive task in autonomous driving, whereas annotation-free learning training can avoid it by learning point cloud representations from unannotated data. In this paper, we propose AFOV, a novel 3D Annotation-Free framework assisted by 2D Open-Vocabulary segmentation models. It consists of two stages: In the first stage, we innovatively integrate high-quality textual and image features of 2D open-vocabulary models and propose the Tri-Modal contrastive Pre-training (TMP). In the second stage, spatial mapping between point clouds and images is utilized to generate pseudo-labels, enabling cross-modal knowledge distillation. Besides, we introduce the Approximate Flat Interaction (AFI) to address the noise during alignment and label confusion. To validate the superiority of AFOV, extensive experiments are conducted on multiple related datasets. We achieved a record-breaking 47.73% mIoU on the annotation-free 3D segmentation task in nuScenes, surpassing the previous best model by 3.13% mIoU. Meanwhile, the performance of fine-tuning with 1% data on nuScenes and SemanticKITTI reached a remarkable 51.75% mIoU and 48.14% mIoU, outperforming all previous pre-trained models. Boyi Sun, Xingxia Wang, Bin Tian 0003, Long Chen 0005, Fei-Yue Wang 0001 |
AAAI | 5 |
| 2025 | Lightstereo: Channel Boost is All You Need for Efficient 2D Cost AggregationabstractWe present LightStereo, a cutting-edge stereomatching network crafted to accelerate the matching process. Departing from conventional methodologies that rely on aggregating computationally intensive 4D costs, LightStereo adopts the 3D cost volume as a lightweight alternative. While similar approaches have been explored previously, our breakthrough lies in enhancing performance through a dedicated focus on the channel dimension of the 3D cost volume, where the distribution of matching costs is encapsulated. Our exhaustive exploration has yielded plenty of strategies to amplify the capacity of the pivotal dimension, ensuring both precision and efficiency. We compare the proposed LightStereo with existing state-of-the-art methods across various benchmarks, which demonstrate its superior performance in speed, accuracy, and resource utilization. LightStereo achieves a competitive EPE metric in the SceneFlow datasets while demanding a minimum of only 22 GFLOPs and 17 ms of runtime, and ranks 1st on KITTI 2015 among real-time models. Our comprehensive analysis reveals the effect of 2 D cost aggregation for stereo matching, paving the way for realworld applications of efficient stereo systems. Code is available at https://github.com/XiandaGuo/OpenStereo. Xianda Guo, Chenming Zhang, Youmin Zhang 0008, Wenzhao Zheng, Dujun Nie, Matteo Poggi, Long Chen 0005 |
ICRA | 7 |
| 2025 | Adjacent-view Transformers for Supervised Surround-view Depth EstimationabstractDepth estimation has been widely studied and serves as the fundamental step of 3D perception for robotics and autonomous driving. Though significant progress has been made in monocular depth estimation in the past decades, these attempts are mainly conducted on the KITTI benchmark with only front-view cameras, which ignores the correlations across surround-view cameras. In this paper, we propose an Adjacent-View Transformer for Supervised Surround-view Depth estimation (AVT-SSDepth), to jointly predict the depth maps across multiple surrounding cameras. Specifically, we employ a global-to-local feature extraction module that combines CNN with transformer layers for enriched representations. Further, the adjacent-view attention mechanism is proposed to enable the intra-view and inter-view feature propagation. The former is achieved by the self-attention module within each view, while the latter is realized by the adjacent attention module, which computes the attention across multi-cameras to exchange the multi-scale representations across surround-view feature maps. In addition, AVT-SSDepth has strong cross-dataset generalization. Extensive experiments show that our method achieves superior performance over existing state-of-the-art methods on both DDAD and nuScenes datasets. Code is available at https://github.com/XiandaGuo/SSDepth. Xianda Guo, Wenjie Yuan 0004, Chenming Zhang, Qin Zou 0001, Long Chen 0005 |
IROS | 8 |
| 2025 | EANS: Reducing Energy Consumption for UAV with an Environmental Adaptive Navigation StrategyabstractUnmanned Aerial Vehicles (Uavs) are limited by the onboard energy. Refinement of the navigation strategy directly affects both the flight velocity and the trajectory based on the adjustment of key parameters in the Uavs pipeline, thus reducing energy consumption. However, existing techniques tend to adopt static and conservative strategies in dynamic scenarios, leading to inefficient energy reduction. Dynamically adjusting the navigation strategy requires overcoming the challenges including the task pipeline interdependencies, the environmental-strategy correlations, and the selecting parameters. To solve the aforementioned problems, this paper proposes a method to dynamically adjust the navigation strategy of the Uavs by analyzing its dynamic characteristics and the temporal characteristics of the autonomous navigation pipeline, thereby reducing Uavs energy consumption in response to environmental changes. We compare our method with the baseline through hardware-in-the-loop (HIL) simulation and real-world experiments, showing our method 3.2X and 2.6X improvements in mission time, 2.4X and 1.6X improvements in energy, respectively. Boyang Li 0009, Long Chen 0005, Kai Huang 0001 |
IROS | 4 |
| 2025 | WMNav: Integrating Vision-Language Models into World Models for Object Goal NavigationabstractObject Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although recent progress in Vision-Language Model (VLM)-based agents has demonstrated promising perception and decision-making abilities through prompting, none has yet established a fully modular world model design that reduces risky and costly interactions with the environment by predicting the future state of the world. We introduce WMNav, a novel World Model-based Navigation framework powered by Vision-Language Models (VLMs). It predicts possible outcomes of decisions and builds memories to provide feedback to the policy module. To retain the predicted state of the environment, WMNav proposes the online maintained Curiosity Value Map as part of the world model memory to provide dynamic configuration for navigation policy. By decomposing according to a human-like thinking process, WMNav effectively alleviates the impact of model hallucination by making decisions based on the feedback difference between the world model plan and observation. To further boost efficiency, we implement a two-stage action proposer strategy: broad exploration followed by precise localization. Extensive evaluation on HM3D and MP3D validates WMNav surpasses existing zero-shot benchmarks in both success rate and exploration efficiency (absolute improvement: +3.2% SR and +3.2% SPL on HM3D, +13.5% SR and +1.1% SPL on MP3D). Project page: https://b0b8k1ng.github.io/WMNav/. Dujun Nie, Xianda Guo, Yiqun Duan, Ruijun Zhang, Long Chen 0005 |
IROS | 5 |
| 2025 | SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language ModelsabstractAccurate spatial reasoning in outdoor environments—covering geometry, object pose, and inter-object relationships—is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematically evaluate the spatial reasoning capabilities of vision language models (VLMs). Built on the nuScenes dataset, SURDS comprises 41,080 vision–question–answer training instances and 9,250 evaluation samples, spanning six spatial categories: orientation, depth estimation, pixel-level localization, pairwise distance, lateral ordering, and front–behind relations. We benchmark leading general-purpose VLMs, including GPT, Gemini, and Qwen, revealing persistent limitations in fine-grained spatial understanding. To address these deficiencies, we go beyond static evaluation and explore whether alignment techniques can improve spatial reasoning performance. Specifically, we propose a reinforcement learning–based alignment scheme leveraging spatially grounded reward signals—capturing both perception-level accuracy (location) and reasoning consistency (logic). We further incorporate final-answer correctness and output-format rewards to guide fine-grained policy adaptation. Our GRPO-aligned variant achieves overall score of 40.80 in SURDS benchmark. Notably, it outperforms proprietary systems such as GPT-4o (13.30) and Gemini-2.0-flash (35.71). To our best knowledge, this is the first study to demonstrate that reinforcement learning–based alignment can significantly and consistently enhance the spatial reasoning capabilities of VLMs in real-world driving contexts. We release the SURDS benchmark, evaluation toolkit, and GRPO alignment code through: https://github.com/XiandaGuo/Drive-MLLM. Xianda Guo, Ruijun Zhang, Yiqun Duan, Dujun Nie, Wenke Huang 0003, Chenming Zhang, Shuai Liu 0009, Hao Zhao 0002, Long Chen 0005 |
NeurIPS | 10 |
| 2025 | ReSim: Reliable World Simulation for Autonomous DrivingabstractHow can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusively on real-world driving data composed mainly of safe expert trajectories, struggle to follow hazardous or non-expert behaviors, which are rare in such data. This limitation restricts their applicability to tasks such as policy evaluation. In this work, we address this challenge by enriching real-world human demonstrations with diverse non-expert data collected from a driving simulator (e.g., CARLA), and building a controllable world model trained on this heterogeneous corpus. Starting with a video generator featuring diffusion transformer architecture, we devise several strategies to effectively integrate conditioning signals and improve prediction controllability and fidelity. The resulting model, ReSim, enables Reliable Simulation of diverse open-world driving scenarios under various actions, including hazardous non-expert ones. To close the gap between high-fidelity simulation and applications that require reward signals to judge different actions, we introduce a Video2Reward module that estimates reward from ReSim’s simulated future. Our ReSim paradigm achieves up to 44% higher visual fidelity, improves controllability for both expert and non-expert actions by over 50%, and boosts planning and policy selection performance on NAVSIM by 2% and 25%, respectively. Jiazhi Yang, Kashyap Chitta, Shenyuan Gao, Long Chen 0005, Yuqian Shao, Xiaosong Jia, Hongyang Li 0001, Andreas Geiger 0001, Xiangyu Yue 0001, Li Chen 0008 |
NeurIPS | 4 |
| 2025 | Leveraging MAB for Efficient Neighbor Discovery with Directional Antennas in Wireless NetworksabstractDirectional antennas, with their ability to enhance signal strength and mitigate interference, are integral to modern wireless communication, especially in ground-air-space networks. However, the precise alignment requirement of directional antennas in communication pose challenges for efficient neighbor discovery (ND). Aiming at optimizing discovery time by dynamically adjusting sector selection probabilities, a novel Multi-Armed Bandit (MAB)-based algorithm is proposed in this paper. By analyzing the feedback that can accelerate the ND time and the theory of MAB, an important conclusion is drawn that MAB is only applicable to non-uniform distributions where there are significant differences in neighbor distributions. Based on analysis, an appropriate MAB-based ND algorithm is designed. Numerical simulations have demonstrated the correctness of the analysis and the performance advantages of the proposed algorithm both in uniform and non-uniform topologies. Huiyuan Xu, Yu Liu 0055, Long Chen 0005, Dalong Yang, You-Jiang Liu |
WCNC | 3 |
| 2025 | TemPrompt: Multi-task prompt learning for temporal relation extraction in RAG-based crowdsourcing systems
Jing Yang 0044, Linyao Yang, Xiao Wang 0002, Long Chen 0005, Fei-Yue Wang 0001 |
Neurocomputing | 5 |
| 2025 | The ParallelWorkforce: A Framework for Synergistic Collaboration in Digital, Robotic, and Biological Workers of Industry 5.0abstractAiming to boost production efficiency and reduce human workload, human-centricity has emerged as the core concept of Industry 5.0 (I5.0). However, current works have not established a unified automation and autonomous framework for human-centric smart manufacturing across various real world applications. Addressing this gap, this research introduces an innovative automated framework, ParallelWorkforce, which integrates blockchain intelligence and decentralized autonomous organizations and operations (DAOs) to drive the evolution from digital twins to parallel intelligence. First, this research conducts a comprehensive investigation into smart manufacturing in I5.0, summarizing the ongoing evolution. Next, a detailed exploration of ParallelWorkforce is provided to offer customized strategies for managing different levels of out-of-distribution events, significantly alleviating the workload on biological workers and maximizing the potential of both digital and robotic workers. Finally, the development of ParallelWorkforce across various key applications of smart manufacturing is demonstrated, including autonomous transportation, task assignment, and worker management. This research provides a viable solution for the further development of human-centered smart manufacturing and paves the way for the realization of “6S” goals in I5.0. Siyu Teng, Yutong Wang 0001, Xingxia Wang, Juanjuan Li, Yuchen Li 0004, Xiaotong Zhang 0007, Lingxi Li 0001, Long Chen 0005, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2024 | SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural NetworkabstractIn this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting of a dyadic decomposition front-end and backbone network), and quantifying the difficulty level of counting depending on sound polyphonicity. The dyadic decomposition front-end progressively decomposes the raw waveform dyadically along the frequency axis to obtain time-frequency representation in multi-stage, coarse-to-fine manner. Each intermediate waveform convolved by a parent filter is further processed by a pair of child filters that evenly split the parent filter's carried frequency response, with the higher-half child filter encoding the detail and lower-half child filter encoding the approximation. We further introduce an energy gain normalization to normalize sound loudness variance and spectrum overlap, and apply it to each intermediate parent waveform before feeding it to the two child filters. To better quantify sound counting difficulty level, we further design three polyphony-aware metrics: polyphony ratio, max polyphony and mean polyphony. We test DyDecNet on various datasets to show its superiority, and we further show dyadic decomposition network can be used as a general front-end to tackle other acoustic tasks. Zhuangzhuang Dai, Agathoniki Trigoni, Long Chen 0005, Andrew Markham |
AAAI | 4 |
| 2024 | GenAD: Generative End-to-End Autonomous Driving
Wenzhao Zheng, Xianda Guo, Chenming Zhang, Long Chen 0005 |
ECCV (65) | 5 |
| 2024 | AdvDiffuser: Generating Adversarial Safety-Critical Driving Scenarios via Guided DiffusionabstractSafety-critical scenarios are infrequent in natural driving environments but hold significant importance for the training and testing of autonomous driving systems. The prevailing approach involves generating safety-critical scenarios automatically in simulation by introducing adversarial adjustments to natural environments. These adjustments are often tailored to specific tested systems, thereby disregarding their transferability across different systems. In this paper, we propose AdvDiffuser, an adversarial framework for generating safety-critical driving scenarios through guided diffusion. By incorporating a diffusion model to capture plausible collective behaviors of background vehicles and a lightweight guide model to effectively handle adversarial scenarios, AdvDiffuser facilitates transferability. Experimental results on the nuScenes dataset demonstrate that AdvDiffuser, trained on offline driving logs, can be applied to various tested systems with minimal warm-up episode data and outperform other existing methods in terms of realism, diversity, and adversarial performance. Xianda Guo, Kunhua Liu, Long Chen 0005 |
IROS | 5 |
| 2024 | V2X-BGN: Camera-based V2X-Collaborative 3D Object Detection with BEV Global Non-Maximum SuppressionabstractIn recent years, research on Vehicle-to-Everything (V2X) cooperative perception algorithms mainly focuses on the fusion of intermediate features from LiDAR point clouds. Since the emergence of excellent single-vehicle visual perception models like BEVFormer, collaborative perception schemes based on camera and late-fusion have become feasible approaches. This paper proposes a V2X-collaborative 3D object detection structure in Bird's Eye View (BEV) space, based on global non-maximum suppression and late-fusion (V2X-BGN), and conducts experiments on the V2X-Set dataset. Focusing on complex road conditions with extreme occlusion, the paper compares the camera-based algorithm with the LiDAR-based algorithm, validating the effectiveness of pure visual solutions in the collaborative 3D object detection task. Additionally, this paper highlights the complementary potential of camera-based and LiDAR-based approaches and the importance of object-to-ego vehicle distance in the collaborative 3D object detection task. Caiji Zhang, Bin Tian 0003, Shi Meng, Shuangying Qi, Yunfeng Ai, Long Chen 0005 |
IV | 7 |
| 2024 | Conditional visibility aware view synthesis via parallel light fields
Yutong Wang 0001, Long Chen 0005, Fei-Yue Wang 0001 |
Neurocomputing | 5 |
| 2024 | Deep motion estimation through adversarial learning for gait recognition
Yuanhao Yue, Laixiang Shi, Long Chen 0005, Zhongyuan Wang 0001, Qin Zou 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | iMAPeM: A New Paradigm for Implementing Intelligent Mining With Humans in the LoopabstractSince the 1990s, the automated process within open-pit mines has been rapidly advanced. However, mineral transportation, as the most costly and dangerous production process, has not witnessed efficient achievements. Adverse weather conditions and complex work environments are two critical bottlenecks that impede the development and deployment of autonomous transportation of open-pit mines. To alleviate this issue, this research proposes a novel paradigm, named IMAPeM, designed to enable safe and efficient autonomous transportation with humans in the loop. IMAPeM includes three categories of miners: 1) biological miners; 2) digital miners; and 3) robotic miners, as well as three operational modes: 1) autonomous model (AM); 2) parallel model (PM); and 3) expert/emergency model (EM). IMAPeM employs these miners and modes depend on the complexity of the task, optimizing the utilization of digital and robotic miners while reducing the workload for biological miners. In addition, we developed the YUGONG system based on IMAPeM. The empirical implementation demonstrates the exceptional performance of the YUGONG system in autonomous transportation across diverse open-pit mines. This system contributes to the advancement of sustainable mining practices, which also carries profound significance for achieving long-term environmentally responsible mining operations. Yunfeng Ai, Siyu Teng, Yuchen Li 0004, Yu Gao 0011, Feng Meng, Shengli Yang, Bin Tian 0003, Long Chen 0005, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 10 |
| 2023 | Developing a multiview spatiotemporal model based on deep graph neural networks to predict the travel demand by busabstractThe accurate prediction of travel demand by bus is crucial for effective urban mobility demand management. However, most models of travel demand prediction by bus tend to focus on the bus’s spatiotemporal dependencies, while ignoring the interactions between buses and other transportation modes, such as metros and taxis. We propose a Multiview Spatiotemporal Graph Neural Network (MSTGNN) model to predict short-term travel demand by bus. It emphasizes the ability to capture the interaction dependencies among the travel demand of buses, metros, and taxis. Firstly, a multiview graph consisting of bus, metro, and taxi views is constructed, with each view containing both a local and global graph. Secondly, a multiview attention-based temporal graph convolution module is developed to capture spatiotemporal and cross-view interaction dependencies among different transport modes. Especially, to address the uneven spatial distributions of features in multiview learning, the cross-view spatial feature consistency loss is introduced as an auxiliary loss. Finally, we conduct intensive experiments using a real-world dataset from Shenzhen, China. The results demonstrate that our proposed MSTGNN model performs better than the existing models. Ablation experiments validate the contributions of various modes of transportation to the improvement of the model’s performance. Tianhong Zhao, Zhengdong Huang, Wei Tu 0001, Filip Biljecki, Long Chen 0005 |
Int. J. Geogr. Inf. Sci. | 5 |
| 2023 | Extreme Low-Resolution Action Recognition with Confident Spatial-Temporal Attention Transfer
Yucai Bai, Qin Zou 0001, Xieyuanli Chen, Lingxi Li 0001, Zhengming Ding, Long Chen 0005 |
Int. J. Comput. Vis. | 6 |
| 2023 | Improved interpolation with sub-pixel relocation method for strong barrel distortion
Xuan Yuwen, Silong Zhang, Long Chen 0005, Hui Zhang 0019 |
Signal Process. | 3 |
| 2023 | Learning Dynamic Graph for Overtaking Strategy in Autonomous DrivingabstractAutomatic overtaking is a challenging task for self-driving vehicles. Traditional rule-based methods for overtaking in autonomous driving heavily rely on many predefined rules and are difficult to apply in complex driving scenarios. Learning-based methods usually use convolutional networks, recurrent networks, and multilayer perceptrons, etc., to extract features from environments, but they fail to effectively represent geometric and interactive information among traffic participants. Classic graph convolutional networks (GCNs) have the ability of represent graph-structural information but are limited to stable relationship representation due to the fixed adjacency matrix when applied in autonomous driving. In this paper, we propose a novel dynamic graph learning method based on a graph convolutional network with a trainable adjacency matrix (TAM-GCN) to enable the learning of dynamic relationships among different nodes in an ever-changing driving scene. In addition, we develop a planning method for overtaking strategy in autonomous driving, where the proposed TAM-GCN is used to extract the spatial graph-structural features, select appropriate overtaking time, and generate efficient overtaking actions. The proposed model is trained using the imitation learning method. We conduct comprehensive experiments in both closed-loop and open-loop testing in the CARLA simulator and compare our method with state-of-the-art methods. Experimental results demonstrate the proposed method achieves better accuracy, safety and overtaking performance than existing methods. Xuemin Hu, Bo Tang 0011, Junchi Yan, Long Chen 0005 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Milestones in Autonomous Driving and Intelligent Vehicles - Part I: Control, Computing System Design, Communication, HD Map, Testing, and Human BehaviorsabstractInterest in autonomous driving (AD) and intelligent vehicles (IVs) is growing at a rapid pace due to the convenience, safety, and economic benefits. Although a number of surveys have reviewed research achievements in this field, they are still limited in specific tasks and lack systematic summaries and research directions in the future. Our work is divided into three independent articles and the first part is a survey of surveys (SoS) for total technologies of AD and IVs that involves the history, summarizes the milestones, and provides the perspectives, ethics, and future research directions. This is the second part (Part I for this technical survey) to review the development of control, computing system design, communication, high-definition map (HD map), testing, and human behaviors in IVs. In addition, the third part (Part II for this technical survey) is to review the perception and planning sections. The objective of this article is to involve all the sections of AD, summarize the latest technical milestones, and guide abecedarians to quickly understand the development of AD and IVs. Combining the SoS and Part II, we anticipate that this work will bring novel and diverse insights to researchers and abecedarians, and serve as a bridge between past and future. Long Chen 0005, Yuchen Li 0004, Chao Huang 0006, Yang Xing 0002, Daxin Tian, Li Li 0013, Zhongxu Hu, Siyu Teng, Chen Lv 0001, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | Milestones in Autonomous Driving and Intelligent Vehicles - Part II: Perception and PlanningabstractA growing interest in autonomous driving (AD) and intelligent vehicles (IVs) is fueled by their promise for enhanced safety, efficiency, and economic benefits. While previous surveys have captured progress in this field, a comprehensive and forward-looking summary is needed. Our work fills this gap through three distinct articles. The first part, a “survey of surveys” (SoS), outlines the history, surveys, ethics, and future directions of AD and IV technologies. The second part, “Milestones in AD and IVs Part I: Control, Computing System Design, Communication, high-definition map (HD map), Testing, and Human Behaviors” delves into the development of control, computing system, communication, HD map, testing, and human behaviors in IVs. This part, the third part, reviews perception and planning in the context of IVs. Aiming to provide a comprehensive overview of the latest advancements in AD and IVs, this work caters to both newcomers and seasoned researchers. By integrating the SoS and Part I, we offer unique insights and strive to serve as a bridge between past achievements and future possibilities in this dynamic field. Long Chen 0005, Siyu Teng, Bai Li 0002, Xiaoxiang Na, Yuchen Li 0004, Jinjun Wang, Dongpu Cao, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2023 | MetaMining: Mining in the MetaverseabstractMines are one of the important energy sources in the world. Due to mining areas are often affected by adverse weather and environmental conditions (sand, dust, extreme cold, heavy snow, etc.), the efficiency and security of mining are low. As a new, efficient, prospective mode, Metaverse has been studied and successfully applied in various industries. However, there is still no research about its usage in mines. In this article, we apply Metaverse to mining and propose MetaMining. We present its definition and analyze the functions it could perform. Furthermore, we propose its development phases and architecture. Specifically, MetaMining is a mode, which aims to achieve high-efficiency, high-security mining in the physical world through the interaction between the physical world and the virtual world. Its development phases involve three steps: 1) digital twins; 2) digital natives; and 3) surreality. Its architecture consists of three components: 1) human world; 2) virtual mining system; and 3) physical mining system. In addition, we analyze key technologies and possible challenges in the construction of MetaMining. Kunhua Liu, Long Chen 0005, Lingxi Li 0001, Huaiwei Ren, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | RadarVerses in Metaverses: A CPSI-Based Architecture for 6S Radar Systems in CPSSabstractMetaverses have caused significant changes in the industry and their academic foundation can be traced back to the term cyber–physical–social systems (CPSS), which was proposed in 2010. Radar is an important sensor in sensing systems that are widely applied in many fields, especially, in autonomous driving. To deal with the complex environment, smart radars with real-time information processing capabilities are required. Human factors play a critical role in the operation and management of radar systems, thus, digital twins’ radars in cyber–physical systems (CPS) are unable to achieve intelligence in CPSS due to an incomplete consideration of human involvement. For this consideration, we propose a novel framework of RadarVerses for smart radars in metaverses based on ACP-based parallel intelligence, which is also known as cyber–physical–social intelligence (CPSI). RadarVerses consist of five main parts which are physical radars, descriptive radars, predictive radars, prescriptive radars, and deep radars. To construct RadarVerses at the technical level, we introduce four main technical foundations: 1) communication technology; 2) scenarios engineering; 3) foundation models; and 4) digital workers. In addition, we also provide a case study about LiDARs’ predictive maintenance of accumulated snow in RadarVerses. Yonglin Tian, Yunfeng Ai, Bin Tian 0003, Er Wu, Long Chen 0005 |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2022 | AutoMine: An Unmanned Mine DatasetabstractAutonomous driving datasets have played an important role in validating the advancement of intelligent vehicle algorithms including localization, perception and prediction in academic areas. However, current existing datasets pay more attention to the structured urban road, which hampers the exploration on unstructured special scenarios. Moreover, the open-pit mine is one of the typical representatives for them. Therefore, we introduce the Autonomous driving dataset on the Mining scene (AutoMine) for positioning and perception tasks in this paper. The AutoMine is collected by multiple acquisition platforms including an SUV, a wide-body mining truck and an ordinary mining truck, depending on the actual mine operation scenarios. The dataset consists of 18+ driving hours, 18K annotated lidar and image frames for 3D perception with various mines, time-of-the-day and weather conditions. The main contributions of the AutoMine dataset are as follows: I.The first autonomous driving dataset for perception and localization in mine scenarios. 2.There are abundant dynamic obstacles of 9 degrees of freedom with large dimension difference (mining trucks and pedestrians) and extreme climatic conditions (the dust and snow) in the mining area. 3.Multi-platform acquisition strategies could capture mining data from multiple perspectives that fit the actual operation. More details can be found in our website(https://automine.cc). Yuchen Li 0004, Siyu Teng, Yu Zhang 0109, Yuchang Zhu, Dongpu Cao, Bin Tian 0003, Yunfeng Ai, Zhe Xuanyuan, Long Chen 0005 |
CVPR | 11 |
| 2022 | ULSM: Underground Localization and Semantic Mapping with Salient Region Loop Closure under Perceptually-Degraded EnvironmentabstractSimultaneous Localization and Mapping (SLAM) has greatly assisted in exploring perceptually-degraded underground environments, such as human-made tunnels, mine tunnels, and caves. However, the recurring sensor failures and spurious loop closures in these scenes bring significant challenges to applying SLAM. This paper proposes an architecture for underground localization and semantic mapping (ULSM) that promotes the robustness of odometry estimation and map-building. In this architecture, a two-stage robust motion compensation method is proposed to adapt to sensor-failure situations. The proposed salient region loop closure detection contributes to avoiding spurious loop closures. Meanwhile, the 2D pose as the initial value for point cloud registration is estimated without additional input. We also design a multi-robot cooperative mapping scheme based on descriptors of the salient region. Extensive experiments are conducted on datasets collected in the Tunnel Circuit of DARPA Subterranean Challenge. Bin Tian 0003, Long Chen 0005 |
IROS | 4 |
| 2022 | A Comparative Analysis of LiDAR SLAM-Based Indoor Navigation for Autonomous VehiclesabstractSimultaneous localization and mapping (SLAM) is a fundamental technique block in the indoor-navigation system for most autonomous vehicles and robots. SLAM aims at building a global consistent map of the environment while simultaneously determining the position and orientation of the robot in this map. Significant advances have been made in visual SLAM techniques in the past several years. However, due to the fragile performance in tracking feature points in environments that lack texture, e.g., a warehouse with blank white walls, visual SLAM can hardly provide a reliable localization. Compared with visual SLAM, LiDAR SLAM can often provide more robust localization in indoor environments by using 3D spatial information directly captured by LiDAR point clouds. Thus, LiDAR SLAM techniques are often employed in industrial applications such as automated guided vehicles (AGVs). In the past decades, a number of LiDAR SLAM methods have been proposed. However, the strength and weakness points of various LiDAR SLAMs are not clear, which may perplex the researchers and engineers. In this article, analysis and comparisons are made on different LiDAR SLAM-based indoor navigation methods, and extensive experiments are conducted to evaluate their performances in real environments. The comparative analysis and results can help researchers in academia and industry in constructing a suitable LiDAR SLAM system for indoor navigation for their own usage scenarios. Qin Zou 0001, Qin Sun, Long Chen 0005, Bu Nie, Qingquan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Three Principles to Determine the Right-of-Way for AVs: Safe Interaction With HumansabstractAutonomous vehicles (AVs) are widely believed to be good for improving transportation safety and efficiency. However, recent fatal accidents by some of their prototypes remind us that there are no operationalizable and quantitative safe driving strategies available for an AV in a wide range of situations to avoid collisions. In contrast with many recent studies that focused on ethical considerations when AVs are facing unavoidable harms, we study how to proactively prevent collisions by setting up a set of decision rules for AVs to determine the right-of-way efficiently. Notably, we summarize three essential principles for AVs designing to increase driving safety, and establish a rule-based nine-step communication-decision model to implement them. Our method is constructed by analyzing how human drivers solve potential conflicts. The decision rules are designed to be ambiguity-free and readily computable with the least communication so that human drivers and AVs could easily understand each other in terms of their behaviors and intentions of. We have demonstrated the effectiveness of our method by comparing it with some alternative approaches. Li Li 0013, Can Zhao 0004, Xiao Wang 0002, Zhiheng Li 0001, Long Chen 0005, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Conditional DQN-Based Motion Planning With Fuzzy Logic for Autonomous DrivingabstractMotion planning is one of the most significant part in autonomous driving. Learning-based motion planning methods attract many researchers’ attention due to the abilities of learning from the environment and directly making decisions from the perception. The deep Q-network, as a popular reinforcement learning method, has achieved great progress in autonomous driving, but these methods seldom use the global path information to handle the issue of directional planning such as making a turning at an intersection since the agent usually learns driving strategies only by the designed reward function, which is difficult to adapt to the driving scenarios of urban roads. Moreover, different motion commands such as the steering wheel and accelerator are associated with each other from classic Q-networks, which easily leads to an unstable prediction of the motion commands since they are independently controlled in a practical driving system. In this paper, a conditional deep Q-network for directional planning is proposed and applied in end-to-end autonomous driving, where the global path is used to guide the vehicle to drive from the origination to the destination. To handle the dependency of different motion commands in Q-networks, we take use of the idea of fuzzy control and develop a defuzzification method to improve the stability of predicting the values of different motion commands. We conduct comprehensive experiments in the CARLA simulator and compare our method with the state-of-the-art methods. Experimental results demonstrate the proposed method achieves better learning performance and driving stability performance than other methods. Long Chen 0005, Xuemin Hu, Bo Tang 0011 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Deep Learning for Image and Point Cloud Fusion in Autonomous Driving: A ReviewabstractAutonomous vehicles were experiencing rapid development in the past few years. However, achieving full autonomy is not a trivial task, due to the nature of the complex and dynamic driving environment. Therefore, autonomous vehicles are equipped with a suite of different sensors to ensure robust, accurate environmental perception. In particular, the camera-LiDAR fusion is becoming an emerging research theme. However, so far there has been no critical review that focuses on deep-learning-based camera-LiDAR fusion methods. To bridge this gap and motivate future research, this article devotes to review recent deep-learning-based data fusion approaches that leverage both image and point cloud. This review gives a brief overview of deep learning on image and point cloud data processing. Followed by in-depth reviews of camera-LiDAR fusion methods in depth completion, object detection, semantic segmentation, tracking and online cross-sensor calibration, which are organized based on their respective fusion levels. Furthermore, we compare these methods on publicly available datasets. Finally, we identified gaps and over-looked challenges between current academic researches and real-world applications. Based on these observations, we provide our insights and point out promising research directions. Yaodong Cui, Ren Chen, Wenbo Chu, Long Chen 0005, Daxin Tian, Ying Li 0036, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | CenterNet3D: An Anchor Free Object Detector for Point CloudabstractAccurate and fast 3D object detection from point clouds is a key task in autonomous driving. Existing one-stage 3D object detection methods can achieve real-time performance, however, they are dominated by anchor-based detectors which are inefficient and require additional post-processing. In this paper, we eliminate anchors and model an object as a single point—the center point of its bounding box. Based on the center point, we propose an anchor-free CenterNet3D network that performs 3D object detection without anchors. Our CenterNet3D uses keypoint estimation to find center points and directly regresses 3D bounding boxes. However, because inherent sparsity of point clouds, 3D object center points are likely to be in empty space which makes it difficult to estimate accurate boundaries. To solve this issue, we propose an extra corner attention module to enforce the CNN backbone to pay more attention to object boundaries. Besides, considering that one-stage detectors suffer from the discordance between the predicted bounding boxes and corresponding classification confidences, we develop an efficient keypoint-sensitive warping operation to align the confidences to the predicted bounding boxes. Our proposed CenterNet3D is non-maximum suppression free which makes it more efficient and simpler. We evaluate CenterNet3D on the widely used KITTI dataset and more challenging nuScenes dataset. Our method outperforms all state-of-the-art anchor-based one-stage methods and has comparable performance to two-stage methods as well. It has an inference speed of 20 FPS and achieves the best speed and accuracy trade-off. Our source code will be released athttps://github.com/wangguojun2018/CenterNet3d. Guojun Wang 0002, Jian Wu 0024, Bin Tian 0003, Siyu Teng, Long Chen 0005, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | RDC-SLAM: A Real-Time Distributed Cooperative SLAM System Based on 3D LiDARabstractTo improve the accuracy and efficiency of 3D LiDAR mapping, real-time cooperative SLAM has been considered to explore large and complex areas. To merge the individual maps from multiple robots, it is crucial to identify the common areas and obtain alternative matches between them. However, data transmission, especially in sparse networks with narrow bandwidth and limited range, is a challenging issue for the above problem. Since the distribution manner is suitable for limited communication, we proposed a common framework of 3D real-time distributed cooperative SLAM to fill the community gap. Assuming that each robot can communicate with others, the presented framework consists of four key modules: place recognition, relative pose estimation, distributed graph optimization, and communication. Meanwhile, we developed a complete real-time distributed cooperative SLAM system, called RDC-SLAM, by integrating state-of-the-art components into the framework. For computation and data transmission efficiency, descriptor-based registration is used instead of the conventional point cloud matching. An intensity-based descriptor is developed to perform the place recognition and obtain the alternative matches, while an eigenvalue-based segment descriptor is applied to further refine the relative pose estimations between these alternative matches. A distributed graph optimization method is utilized to obtain the maximum likelihood of multi-trajectory estimation. A communication protocol is also designed to associate data among robots that are easy to deploy and have low network requirements. The RDC-SLAM is validated by real-world experiments and exhibits superior performance concerning accuracy, computation efficiency, and data efficiency. Yachen Zhang, Long Chen 0005, Hui Cheng 0002, Wei Tu 0001, Dongpu Cao, Qingquan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Road-Model-Based Road Boundary Extraction for High Definition Map via LIDARabstractHigh definition (HD) map has become indispensable to autonomous driving systems, but its automated construction is challenging. Since road boundary lines are basic elements of the map, accurate and complete extraction of the lines is one of key to build HD map. We present a method to automatically extract the boundary lines based on the road model using multi-beam LIDAR. First, the road model based on B-spline surfaces is constructed to fit the ground, and the multiRANSAC is proposed to remove the surface points. Second, horizontal distance feature and strong straight constraint are used to extract the road boundary lines. Finally, the lines collected from multiple frames are fused to approximate the curves and complete the boundary information. In the matter, the proposed method extracts the lines on the curved and straight road accurately and completely. We verified the feasibility and robustness of the proposed method on KITTI dataset. The average precision, recall and F1 are 0.9599, 0.9066 and 0.9277 respectively, and the average runtime is 28$ms/frame$. Results show that the comprehensive performance of the proposed method outperforms the compared approaches. Huiyuan Xiong, Taohong Zhu, Yu Liu 0055, Yuelong Pan, Shaofang Wu, Long Chen 0005 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | Improved Vehicle LiDAR Calibration With Trajectory-Based Hand-Eye MethodabstractIn unmanned vehicles, LiDARs and GPS/INSs are the most popular sensors to achieve perception and positioning. Precise calibration of the extrinsic parameters between the LiDAR and the GPS/INS is necessary for a successful implementation of sensor fusion. The extrinsic transformation between the LiDAR and GPS/INS is 6D ($x$,$y$,$z$,$yaw$,$pitch$,$roll$), but the motion of a vehicle is mainly 3D ($x$,$y$,$yaw$). The problem is to calculate the 6D extrinsic parameters with the limitation of 3D motion (plane constraint). The solution to this problem has been breaking the plane constraint by designing specific vehicle motions. This paper proposes a new method, a trajectory-based hand-eye calibration method, which makes full use of the large range of unmanned vehicles. The trajectories with large and small ranges are used to solve the rotation and translation, respectively. It is proved that the extrinsic parameters can be solved when the trajectory range of the unmanned vehicle is sufficiently large. The method proposed is tested with simulation, custom and KITTI datasets, and compared with the state-of-the-art methods. The results demonstrate that the accuracy and efficiency of the method proposed are comparable to the state of the art methods. Xuan Yuwen, Long Chen 0005, Fengjun Yan, Hui Zhang 0019, Jianlin Tang, Bin Tian 0003, Yunfeng Ai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Transductive Zero-Shot Hashing for Multilabel Image RetrievalabstractHash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Given semantic annotations such as class labels and pairwise similarities of the training data, hashing methods can learn and generate effective and compact binary codes. While some newly introduced images may contain undefined semantic labels, which we call unseen images, zero-shot hashing (ZSH) techniques have been studied for retrieval. However, existing ZSH methods mainly focus on the retrieval of single-label images and cannot handle multilabel ones. In this article, for the first time, a novel transductive ZSH method is proposed for multilabel unseen image retrieval. In order to predict the labels of the unseen/target data, a visual-semantic bridge is built via instance-concept coherence ranking on the seen/source data. Then, pairwise similarity loss and focal quantization loss are constructed for training a hashing model using both the seen/source and unseen/target data. Extensive evaluations on three popular multilabel data sets demonstrate that the proposed hashing method achieves significantly better results than the comparison methods. Qin Zou 0001, Ling Cao, Zheng Zhang 0036, Long Chen 0005, Song Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Semasuperpixel: A Multi-Channel Probability-Driven Superpixel Segmentation MethodabstractSuperpixel, an efficient image segmentation approach, aggregates a group of similar pixels into the same cluster. Existing superpixel algorithms still mainly focus on the color information while ignoring the semantic distribution knowledge. In this paper, we propose a semantic information-driven method that adopts multi-channel semantic probabilities for superpixel segmentation. By conducting statistical analysis on the semantic output and then formulating the distance measure, the prior knowledge of the semantic with a dynamic confidence value could be utilized by our method during the global update effectively. Extensive experimental evaluations show that our method achieves a leading segmentation quality and convergence speed, compared to other five state-of-the-art algorithms, as measured by boundary recall, undersegmentation error, and explained variation. Qingyun Zhao, Lei Fan 0005, Yuzhi Zhao, Qiong Yan, Long Chen 0005 |
ICIP | 7 |
| 2021 | Toward Location-Enabled IoT (LE-IoT): IoT Positioning Techniques, Error Sources, and Error MitigationabstractLocalization techniques are becoming key to add location context to the Internet-of-Things (IoT) data without human perception and intervention. Meanwhile, the newly emerged low-power wide-area network (LPWAN) and 5G technologies have become strong candidates for mass-market localization applications. However, various error sources have limited localization performance by using such IoT signals. This article reviews the IoT localization system through the following sequence: IoT localization system review, localization data sources, localization algorithms, localization error sources and mitigation, and localization performance evaluation. Compared to the related surveys, this article has a more comprehensive and state-of-the-art review on IoT localization methods, an original review on IoT localization error sources and mitigation, an original review on IoT localization performance evaluation, and a more comprehensive review of IoT localization applications, opportunities, and challenges. Thus, this survey provides comprehensive guidance for peers who are interested in enabling localization ability in the existing IoT systems, using IoT systems for localization, or integrating IoT signals with the existing localization sensors. You Li 0001, Yuan Zhuang 0001, Xin Hu 0006, Zhouzheng Gao, Jia Hu 0001, Long Chen 0005, Zhe He 0002, Ling Pei, Kejie Chen, Maosong Wang, Xiaoji Niu, Ruizhi Chen, John S. Thompson, Fadhel M. Ghannouchi, Naser El-Sheimy |
IEEE Internet Things J. | 6 |
| 2021 | Real-Time Route Recommendations for E-Taxies Leveraging GPS TrajectoriesabstractElectric vehicles (EVs) currently face formidable challenges in promotion, i.e., short driving ranges, long charging times, and few charging stations, thereby limiting their acceptability to taxi drivers. Leveraging massive-scale taxi GPS trajectory data, we present a novel real-time route recommendation system for electric taxi (ET) drivers. Taxi travel knowledge, including the probability of picking up passengers and the distribution of destinations, is learned from the raw GPS trajectories. Considering the cascading effect of route decision making, consecutive ET actions are modeled with an action tree. The corresponding expected net revenue is estimated based on the learned knowledge. A prototype online system is developed for providing route recommendations, e.g., when to go to a charging station or cruise on certain roads. An experiment in Shenzhen demonstrates that the average daily net revenue of ET drivers is better than those of 76.2% of gasoline taxi drivers. The presented approach not only increases the revenue of ET drivers in the short term but also improves the viability of EVs in the long run. Wei Tu 0001, Ke Mai, Yatao Zhang, Yang Xu 0002, Jincai Huang 0002, Long Chen 0005, Qingquan Li 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2021 | Deep Neural Network Based Vehicle and Pedestrian Detection for Autonomous Driving: A SurveyabstractVehicle and pedestrian detection is one of the critical tasks in autonomous driving. Since heterogeneous techniques have been proposed, the selection of a detection system with an appropriate balance among detection accuracy, speed and memory consumption for a specific task has become very challenging. To deal with this issue and to provide guidance for model selection, this paper analyzes several mainstream object detection architectures, including Faster R-CNN, R-FCN, and SSD, along with several typical feature extractors, such as ResNet50, ResNet101, MobileNet_V1, MobileNet_V2, Inception_V2 and Inception_ResNet_V2. By conducting extensive experiments using the KITTI benchmark, which is a commonly used street dataset, we demonstrate that Faster R-CNN ResNet50 obtains the best average precision (AP) (58%) for vehicle and pedestrian detection, with a speed of 8.6 FPS. Faster R-CNN Inception_V2 performs best for detecting cars and detecting pedestrians respectively (74.5% and 47.3%). ResNet101 consumes the highest memory (9907 MB) and has the largest number of parameters (64.42 millions), and Inception_ResNet_V2 is the slowest model (3.05 FPS). SSD MobileNet_V2 is the fastest model (70 FPS), and SSD MobileNet_V1 is the lightest model in terms of memory usage (875 MB), both of which are suitable for applications on mobile and embedded devices. Long Chen 0005, Shaobo Lin, Xiankai Lu, Dongpu Cao, Hangbin Wu, Chi Guo, Chun Liu 0003, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Visualization Analysis of Intelligent Vehicles Research Field Based on Mapping Knowledge DomainabstractThis study combines applied mathematics, visual analysis technology, information science with an approach of Scientometrics to systematically analyze the development status, research distribution and future trend of intelligent vehicles research. A total number of 3933 published paper index by SCIE and SSCI from 2000 to 2019 are researched based on Mapping Knowledge Domain (MKD) and Scientometrics approaches. Firstly, this paper analyzes the literature content in the field of intelligent vehicles by including the literature number, literature productive countries, research organization, co-authorship of main research groups and the journals from which the articles are mainly sourced. Then, co-citation analysis is used to obtain five major research directions in the field of intelligent vehicles, which include “system framework”, “internet of vehicles”, “intersection control algorithms”, “influence on traffic flow”, and “policies and barriers”, respectively. The keyword co-occurrence analysis is applied to identify four dominant clusters: “planning and control system”, “autonomous vehicle questionnaire”, “sensor and vision”, and “connected vehicles”. Finally, we divide burst keywords into three phases according to the publication date to show more clearly the change of research focus and direction over time. Yi He 0013, Ching-Yao Chan, Long Chen 0005, Chaozhong Wu |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Learning a Deep Cascaded Neural Network for Multiple Motion Commands Prediction in Autonomous DrivingabstractIn autonomous driving, many learning-based methods for motion planing have been proposed in literature, which can predict motion commands directly from the sensory data of the environment, but these methods can neither predict multiple motion commands, such as steering angle, accelerator and brake, nor balance errors among different motion commands. In this paper, we propose a deep cascaded neural network for predicting multiple motion commands which can be trained in an end-to-end manner for autonomous driving. The proposed deep cascaded neural network consists of a convolutional neural network (CNN) and three long short-term memory (LSTM) units, fed with images from a front-facing camera installed at the vehicle. As the outputs, the proposed model can predict thee motion planning commands simultaneously including steering angle, acceleration, and brake to enable the autonomous driving. In order to balance errors among different motion commands and improve prediction accuracy, we propose a new network training algorithm, where three independent loss functions are designed to separately update the weights in the three LSTMs connected to three motion commands. We conduct comprehensive experiments using the data from a driving simulator and compare our method with the state-of-the-art methods. Simulation results demonstrate the proposed motion planning model achieves better accuracy performance than other models. Xuemin Hu, Bo Tang 0011, Long Chen 0005, Sheng Song, Xiuchi Tong |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | Lightweight Single-Image Super-Resolution Network with Attentive Auxiliary Feature Learning
Qing Wang 0025, Yuzhi Zhao, Junchi Yan, Lei Fan 0005, Long Chen 0005 |
ACCV (2) | 6 |
| 2020 | Imitative Reinforcement Learning Fusing Vision and Pure Pursuit for Self-drivingabstractAutonomous urban driving navigation is still an open problem and has ample room for improvement in unknown complex environments and terrible weather conditions. In this paper, we propose a two-stage framework, called IPP-RL, to handle these problems. IPP means an Imitation learning method fusing visual information with the additional steering angle calculated by Pure-Pursuit (PP) method, and RL means using Reinforcement Learning for further training. In our IPP model, the visual information captured by camera can be compensated by the calculated steering angle, thus it could perform well under bad weather conditions. However, imitation learning performance is limited by the driving data severely. Thus we use a reinforcement learning method-Deep Deterministic Policy Gradient (DDPG)-in the second stage training, which shares the learned weights from pretrained IPP model. In this way, our IPP-RL can lower the dependency of imitation learning on demonstration data and solve the problem of low exploration efficiency caused by randomly initialized weights in reinforcement learning. Moreover, we design a more reasonable reward function and use the n-step return to update the critic-network in DDPG. Our experiments on CARLA driving benchmark demonstrate that our IPP-RL is robust to lousy weather conditions and shows remarkable generalization capability in unknown environments on navigation task. Mingxing Peng, Zhihao Gong, Chen Sun 0008, Long Chen 0005, Dongpu Cao |
ICRA | 4 |
| 2020 | GOSMatch: Graph-of-Semantics Matching for Detecting Loop Closures in 3D LiDAR dataabstractDetecting loop closures in 3D Light Detection and Ranging (LiDAR) data is a challenging task since point-level methods always suffer from instability. This paper presents a semantic-level approach named GOSMatch to perform reliable place recognition. Our method leverages novel descriptors, which are generated from the spatial relationship between semantics, to perform frame description and data association. We also propose a coarse-to-fine strategy to efficiently search for loop closures. Besides, GOSMatch can give an accurate 6-DOF initial pose estimation once a loop closure is confirmed. Extensive experiments have been conducted on the KITTI odometry dataset and the results show that GOSMatch can achieve robust loop closure detection performance and outperform existing methods. Yachen Zhu, Yanyang Ma, Long Chen 0005, Maosheng Ye, Lingxi Li 0001 |
IROS | 3 |
| 2020 | Special Issue on Internet of Things for Connected Automated DrivingabstractInternet of Things (IoT) is becoming increasingly prevalent in transportation systems. The traffic system depends on safer, faster, and more intelligent vehicles. Vehicular networks [vehicle-to-vehicle (V2V) and vehicle-to- Infrastructure (V2I)] and automated driving technique are two of the cornerstone technologies enabling the construction of the future-generation highly functional and intelligent transportation system. The IoT-based transportation system can provide enormous connections of devices and sensors for the networked automated vehicles. The capacity of connected autonomous vehicles (CAVs) is expected to be dramatically enhanced by employing IoT techniques. Dongpu Cao, Li Li 0013, Clara Marina Martinez, Long Chen 0005, Yang Xing 0002, Weihua Zhuang |
IEEE Internet Things J. | 4 |
| 2020 | Parallel End-to-End Autonomous Mining: An IoT-Oriented ApproachabstractThis article proposes a new solution for end-to-end autonomous mining operations: Internet of Things (IoT)-based parallel mining, consisting of the concept definition, the solution given, and the concrete realization. The proposed parallel mining is inspired by the artificial societies (A) for modeling, computational experiments (C) for analysis, and parallel execution (P) for control (ACP) approach. The basic framework of parallel mining is given and its advantages are expounded. Then, the solution of parallel mining is proposed, which is mainly composed of four parts: 1) the management and control center for autonomous mining; 2) the autonomous transportation platform of truck; 3) the semiautonomous mining/shovel platform; and 4) the remote takeover platform. Key technologies of IoT-based parallel mining are discussed in detail, namely, network communication, virtual parallel mining construction, mining environment perception over-the-horizon for the moving area and obstacle detection, collaborative decision making, planning, and control for unmanned mining equipment, and parallel taking-over and remote control. Finally, the performance of IoT-based parallel mining, including fusion perception, collaborative decision making, planning, and control, is evaluated. The realization of parallel mining can fundamentally improve the safety of personnel and equipment, reduce the cost of mining operation, and increase the production rate. Yu Gao 0011, Yunfeng Ai, Bin Tian 0003, Long Chen 0005, Jian Wang 0003, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Learning Driving Models From Parallel End-to-End Driving Data SetabstractParallel end-to-end driving aims to improve the performance of end-to-end driving models using both simulated- and real-world data. However, how to efficiently utilize the data from both the simulated world and the real world remains a difficult issue, since these data are usually not well aligned. In this article, we build a data set called the parallel end-to-end driving data set (PED) for parallel end-to-end driving research. PED consists of 13 000 images from the simulated world and 13 000 images from the real world that are used to train the model, as well as 2700 images from the real world that are used to test the model. The simulated-world data in PED are constructed according to the real world, and each simulated-world image corresponds to a real-world image. PED also contains the vehicle measurement data (GPS, speed, steering angle, and heading direction of the vehicle) related to both the simulated- and real-world images, which are not available in some other data sets. We conduct two types of experiments to illustrate the effectiveness and the superiority of PED and explore some ways to mix the simulated-world data with the real-world data to improve the performance of end-to-end driving models. Long Chen 0005, Qing Wang 0025, Xiankai Lu, Dongpu Cao, Fei-Yue Wang 0001 |
Proc. IEEE | 1 |
| 2020 | A survey on deep learning methods for scene flow estimation
Ruihong Wu, Qingyun Zhao, Yiyou Guo, Long Chen 0005 |
Pattern Recognit. | 6 |
| 2020 | Low-Rank Tensor Regularized Fuzzy Clustering for Multiview DataabstractSince data are collected from a range of sources via different techniques, multiview clustering has become an emerging technique for unsupervised data classification. However, most existing soft multiview clustering methods only consider the pairwise correlations and ignore high-order correlations among multiple views. To integrate more comprehensive information from different views, this article innovates a fuzzy clustering model using the low-rank tensor to address the multiview data clustering problem. Our method first conducts a standard fuzzy clustering on different views of the data separately. Then, the obtained soft partition results are aggregated as the new data to be handled by a Kullback-Leibler (KL) divergence-based fuzzy model with low-rank tensor constraints. The KL divergence function, which replaces the traditional minimized Euclidean distance, can enhance the robustness of the model. More importantly, we formulate fuzzy partition matrices of different views as a third-order tensor. So, a low-rank tensor is introduced as a norm constraint in the KL divergence-based fuzzy clustering to obtain dexterously high-order correlations of different views. The minimization of the final model is convex and we present an efficient augmented Lagrangian alternating direction method to handle this problem. Specially, the global membership is derived by using tensor factorization. The efficiency and superiority of the proposed approach are demonstrated by the comparison with state-of-the-art multiview clustering algorithms on many multiple-view data sets. Huiqin Wei, Long Chen 0001, Keyu Ruan, Lingxi Li 0001, Long Chen 0005 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2020 | Surrounding Vehicle Detection Using an FPGA Panoramic Camera and Deep CNNsabstractSurrounding vehicle detection is one of the most important modules for a vision-based driver assistance system (VB-DAS) or an autonomous vehicle. In this paper, we put forward a wireless panoramic camera system for real-time and seamless imaging of the 360-degree driving scene. Using an embedded FPGA design, the proposed panoramic camera system can perform fast image stitching and produce panoramic videos in real-time, which greatly relives the computation and storage burden of a traditional multi-camera-based panoramic system. For surrounding vehicle detection, we present a novel deep convolutional neural network - EZ-Net, which perceives the potential vehicles by using 13 convolutional layers and locates the vehicles by a local non-maximum suppression process. Experimental results demonstrate that, the proposed EZ-Net performs vehicle detection on the panoramic video at a speed of 140 fps while holding a competing accuracy with the state-of-the-art detectors. Long Chen 0005, Qin Zou 0001, Ziyu Pan, Danyu Lai, Liwei Zhu, Zhoufan Hou, Jun Wang 0015, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Online Multi-Object Tracking Using Joint Domain Information in Traffic ScenariosabstractVisual tracking of multiple objects is an essential component for a perception system in autonomous driving vehicles. One of the favorable approaches is the tracking-by-detection paradigm, which links current detection hypotheses to previously estimated object trajectories (also known as tracks) by searching appearance or motion similarities between them. As this search operation is usually based on a very limited spatial or temporal locality, the association can fail in cases of motion noise or long-term occlusion. In this paper, we propose a novel tracking method that solves this problem by putting together information from both enlarged structural and temporal domain. For efficiency without loss of optimality, this approach is decomposed in to three stages, with each dealing with only one constrained association task, and thus, it follows the alternating optimization fashion. In our approach, detections are first assembled into small tracklets based on meta-measurements of object affinity. The association task for tracklets-to-tracks is solved by structural information based on a motion pattern between them. Here, we propose new rules to decouple the processing time from the tracklet length. Furthermore, constraints from temporal domain are introduced to recover objects, which are long-time disappearing due to failed detection or long-term occlusion. By putting together the heterogeneous domain information, our approach exhibits an improved state-of-the-art performance on standard benchmarks. With relatively little processing time, an online and real-time tracking is also permitted in our approach. Wei Tian 0001, Martin Lauer, Long Chen 0005 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | A Collaborative Visual Tracking Architecture for Correlation Filter and Convolutional Neural Network LearningabstractVisual object tracking has achieved remarkable progress in recent years and has been broadly applied in intelligent transportation systems such as autonomous vehicles and drones to monitor and analyze the behavior of specific targets. One typical tracking approach is the discriminative tracker, which branches into two main categories: the correlation filter (CF) and the convolutional neural network (CNN). However, most of the current researches consider both categories as two separate techniques and only rely on one of them. Thus, a dense cooperation between the CF and the CNN still remains less discovered and the question of how to effectively join both techniques to further boost the tracking performance is still open. To address this issue, in this paper, we propose a collaborative architecture which incorporates models constructed with both techniques and dynamically aggregates their response maps for target inference. By an alternating optimization, both models are learned on each other's errors to persistently improve the classification power of the whole tracker. For further efficiency, we present a faster solver for our utilized CF and an analytical solution for dynamic model weighting. Through experiments on standard benchmarks, we reveal the influence of key factors on the joint learning architecture and show that it outperforms the state-of-the-art approaches. Wei Tian 0001, Niels Ole Salscheider, Yunxiao Shan, Long Chen 0005, Martin Lauer |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Toward the Ghosting Phenomenon in a Stereo-Based Map With a Collaborative RGB-D RepairabstractAlthough 3-D reconstruction of dynamic road environment by moving cameras has been broadly applied in recognition and navigation systems, this task is still considered challenging, especially under circumstances with moving objects, where the reconstruction precision is strongly harassed by the ghosting problem. To address this issue, in this paper, we propose a novel approach for reconstructing 3-D maps of complete static scenes, based on a combination of an elaborately designed moving-object filtering mechanism and a map repairing and blank refilling procedure, where both plausible color and depth information from stereo image pairs are utilized. In this approach, first, we employ the planarity knowledge into the initial depth map based on the simple linear iterative cluster (SLIC) superpixel segmentation. The dynamic area in the image is determined under the supervision of odometry calculation. After wiping off moving objects, by collaboratively repairing color and depth information, the final 3-D map containing only static scene is obtained. The experimental results on extensive challenging real-world scenarios demonstrate the effectiveness and robustness of our approach. Jiasong Zhu, Lei Fan 0005, Wei Tian 0001, Long Chen 0005, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Improved Deep Hashing With Soft Pairwise Similarity for Multi-Label Image RetrievalabstractHash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Recently, many deep hashing methods have been proposed and shown largely improved performance over traditional feature-learning methods. Most of these methods examine the pairwise similarity on the semantic-level labels, where the pairwise similarity is generally defined in a hard-assignment way. That is, the pairwise similarity is “1” if they share no less than one class label and “0” if they do not share any. However, such similarity definition cannot reflect the similarity ranking for pairwise images that hold multiple labels. In this paper, an improved deep hashing method is proposed to enhance the ability of multi-label image retrieval. We introduce a pairwise quantified similarity calculated on the normalized semantic labels. Based on this, we divide the pairwise similarity into two situations-“hard similarity” and “soft similarity,” where cross-entropy loss and mean square error loss are adapted respectively for more robust feature learning and hash coding. Experiments on four popular datasets demonstrate that the proposed method outperforms the competing methods and achieves the state-of-the-art performance in multi-label image retrieval. Zheng Zhang 0036, Qin Zou 0001, Yuewei Lin, Long Chen 0005, Song Wang 0002 |
IEEE Trans. Multim. | 4 |
| 2020 | A Full Density Stereo Matching System Based on the Combination of CNNs and Slanted-PlanesabstractStereo matching methods consist of matching cost computation and several post processing steps. Deep learning methods have greatly raised the accuracy of matching cost and achieved the lowest error rate on several public datasets. However, their generality capabilities are not the best due to potential overfitting, which is the common problem of supervised learning approaches. This paper proposes a convolutional neural network (CNN) based cost estimation method for computing the similarity of image patches. In consideration of accuracy and generalization capability, small size convolution kernels are chosen in the convolution layer and dropout in the decision layer is used for preventing overfitting. After obtaining stereo matching cost from the output of the CNN, several post-processing operations are adopted for disparity optimization, which includes semi-global matching in 1-D from different directions, a left-right consistency check, and the slanted plane smoothing method. The method is evaluated on KITTI 2012, KITTI 2015, and Middlebury stereo datasets and the experimental results on the KITTI benchmark demonstrate the competitive accuracy performance of the approach. Additionally, to test the generalization of the method, a series of extended crossover experiments are conducted in which the training samples and testing samples come from different datasets. The results indicate superior generalization capability of our method than other supervised learning methods. Long Chen 0005, Lei Fan 0005, Jianda Chen, Dongpu Cao, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2019 | High-Resolution Driving Scene Synthesis Using Stacked Conditional Gans and Spectral NormalizationabstractLarge-scale dataset plays a key role in the driving scene understanding for deep learning based-autonomous driving tasks. Due to the fact that the annotation for a large number of images is extremely labor-intensive and time-consuming, many researchers turn to using image-synthesis techniques for automatic construction of training data. However, traditional methods often have difficulties in producing high-definition driving scene images. To tackle this problem, in this paper, we propose a novel deep model - hdCGAN - for high-definition image-to-image translation. The hdCGAN is built on a conditional GAN in combination with a spectral normalization. Moreover, we improve the hdCGAN by using a stacked network architecture and the enhanced model is called stack-hdCGAN. With the guidance of multi-scale discriminators and the constraint of spectral normalization in the training procedure, the learned models can generate high-resolution and high-quality driving scene images from corresponding semantic segmentation maps. Quantitative and qualitative evaluations on the Cityscapes dataset demonstrate the effectiveness of the proposed models. Shaobo Lin, Long Chen 0005, Qin Zou 0001, Wei Tian 0001 |
ICME | 2 |
| 2019 | Real-Time Vehicle Detection from Short-range Aerial Image with Compressed MobileNetabstractVehicle detection from short-range aerial image faces challenges including vehicle blocking, irrelevant object interference, motion blurring, color variation etc., leading to the difficulty to achieve high detection accuracy and real-time detection speed. In this paper, benefiting from the recent development in MobileNet family network engineering, we propose a compressed MobileNet which is not only internally resistant to the above listed challenges but also gains the best detection accuracy/speed tradeoff when comparing with the original MobileNet. In a nutshell, we reduce the bottleneck architecture number during the feature map downsampling stage but add more bottlenecks during the feature map plateau stage, neither extra FLOPs nor parameters are thus involved but reduced inference time and better accuracy are expected. We conduct experiment on our collected 5-k short-range aerial images, containing six vehicle categories: truck, car, bus, bicycle, motorcycle, crowded bicycles and crowded motorcycles. Our proposed compressed MobileNet achieves 110 FPS (GPU), 31 FPS (CPU) and 15 FPS (mobile phone), 1.2 times faster and 2% more accurate (mAP) than the original MobileNet. Ziyu Pan, Lingxi Li 0001, Yunxiao Shan, Dongpu Cao, Long Chen 0005 |
ICRA | 6 |
| 2019 | DSNet: Joint Learning for Scene Segmentation and Disparity EstimationabstractRecently, research works have attempted the joint prediction of scene semantics and optical flow estimation, which demonstrate the mutual improvement between both tasks. Besides, the depth information is also indispensable for the scene understanding, and disparity estimation is necessary for outputting dense depth maps. Such task shares a great similarity with the optical flow estimation since they can all be cast into a problem of capturing the difference at a location of two image frames. However, as far as we know, currently there are few networks for the joint learning of semantic and disparity. Moreover, since deep semantic information and disparity feature maps can learn from each other, we find it unnecessary with two independent encoding modules to separately extract semantic and disparity features. Therefore, we propose a unified multi-tasking architecture DSNet, for the simultaneous estimation of semantic and disparity information. In our model, semantic features, extracted by the encoding module ResNet from the left and right images, are used to obtain the deep disparity features via a novel matching module which performs pixel-to-pixel matching. In addition, we also use the disparity map to perform warp operation on deep features of the right image to deal with the problem of lacking of semantic labels. The effectiveness of our method is demonstrated by extensive experiments. Wujing Zhan, Xinqi Ou, Yunyi Yang, Long Chen 0005 |
ICRA | 4 |
| 2019 | DeepSqueezeNet-CRF: A Lightweight Deep Model for Semantic Image SegmentationabstractDeep convolutional neural networks (DCNNs) has shown the powerful capability in image semantic segmentation, along with the huge model sizes and massive parameters. Our key insight is to build a new neural network for image semantic segmentation with fewer parameters and smaller model size under the premise of maintaining state-of-the-art performance. In this paper, we propose a lightweight deep model named DeepSqueezeNet, which introduces the SqueezeNet to the fullyconvolutional networks, replacing parts of convolutional layers by Fire Module layers. Our DeepSqueezeNet contains fewer parameters, smaller model size and higher accuracy than the original fully-convolutional networks. Further more, we propose DeepSqueezeNet-CRF by combining the responses at the final DCNN layer with a fully-connected Conditional Random Field (CRF) to improve localization accuracy. Both quantitative and qualitative results on PASAL VOC 2012 demonstrate the priority and feasibility of our methods, comparing with several other state-of-the-art methods. Especially, our methods can reduce the number of network parameters, model size and training time of the original fully-connected network, with even better performance. Danyu Lai, Yique Deng, Long Chen 0005 |
IJCNN | 3 |
| 2019 | Monocular Outdoor Semantic Mapping with a Multi-task NetworkabstractIn many robotic applications, especially for the autonomous driving, understanding the semantic information and the geometric structure of surroundings are both essential. Semantic 3D maps, as a carrier of the environmental knowledge, are then intensively studied for their abilities and applications. However, it is still challenging to produce a dense outdoor semantic map from a monocular image stream. Motivated by this target, in this paper, we propose a method for large-scale 3D reconstruction from consecutive monocular images. First, with the correlation of underlying information between depth and semantic prediction, a novel multi-task Convolutional Neural Network (CNN) is designed for joint prediction. Given a single image, the network learns low-level information with a shared encoder and separately predicts with decoders containing additional Atrous Spatial Pyramid Pooling (ASPP) layers and the residual connection which merits disparities and semantic mutually. To overcome the inconsistency of monocular depth prediction for reconstruction, post-processing steps with the superpixelization and the effective 3D representation approach are obtained to give the final semantic map. Experiments are compared with other methods on both semantic labeling and depth prediction. We also qualitatively demonstrate the map reconstructed from large-scale, difficult monocular image sequences to prove the effectiveness and superiority. Yucai Bai, Lei Fan 0005, Ziyu Pan, Long Chen 0005 |
IROS | 4 |
| 2019 | An Adaptive Path Tracking Controller Based on Reinforcement Learning with Urban Driving ApplicationabstractUrban driving requires the autonomous vehicles to drive with smooth control and track the planned path accurately. However, most of the existing path-tracking controllers pay more attention to the tracking errors than the smoothness because of the difficulties to balance them. This paper proposes a learning-based method to achieve the trade-off between the smooth control and the tracking-error control. An Reinforcement Learning algorithm, which is called Proximal Policy Optimization, is used to train a neural model to tune the weights of a designed controller PP_PID (Pure-Pursuit_Proportional Integral Derivative). The successfully trained model will adaptively select the optimal weights for the Pure-Pursuit and PID to guarantee the control smoothness and accuracy. Finally, the proposed controller will be tested in two path tracking scenarios. The results show that the proposed controller can change the weights adaptively to maintain a balance in the tracking error and lateral acceleration under the 35km/h. Longsheng Chen, Yuanpeng Chen, Xiangtong Yao, Yunxiao Shan, Long Chen 0005 |
IV | 5 |
| 2019 | Improving classification with semi-supervised and fine-grained learning
Danyu Lai, Wei Tian 0001, Long Chen 0005 |
Pattern Recognit. | 3 |
| 2019 | High-Speed Scene Flow on Embedded Commercial Off-the-Shelf SystemsabstractScene flow is an essential part of a stereo-based perception system for autonomous driving and mobile robotics. As in most of these platforms, the computing resource is limited but the computing requirement is high, embedded and parallelized algorithms are of vital importance for real-time tasks. This paper develops a cross-platform embedded scene flow algorithm by using an OpenCL (Open Computing Language) programming. Meanwhile, we propose a method to achieve a good performance by using a novel coarse-grained software pipeline for the embedded stream application. Experimental results show that the proposed algorithm can boost the average processing speed to 50 fps for different commercial off-the-shelf (COTS) hardware, including desktop graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and mobile phone platforms. For certain GPUs, the peak frame rates can also reach 1000 fps. By comparing the efficiency among the serial platform, we illustrate that with the help of OpenCL programming, COTS platforms can provide enough computing resources for the stereo-based perception algorithm. Long Chen 0005, Mingyue Cui, Feihu Zhang, Biao Hu 0001, Kai Huang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | Deep Integration: A Multi-Label Architecture for Road Scene RecognitionabstractDeep convolutional neural networks have been applied by automobile industries, Internet giants, and academic institutes to boost autonomous driving technologies; while progress has been witnessed in environmental perception tasks, such as object detection and driver state recognition, the scene-centric understanding and identification still remain a virgin land. This mainly encompasses two key issues: 1) the lack of shared large datasets with comprehensively annotated road scene information and 2) the difficulty to find effective ways to train networks concerning the bias of category samples, image resolutions, scene dynamics, and capturing conditions. In this paper, we make two contributions: 1) we introduce a large-scale dataset with over 110 k images, dubbed DrivingScene, covering traffic scenarios under different weather conditions, road structures, and environmental instances and driving places, which is the first large-scale dataset for multi-class traffic scenes classification and 2) we propose a multi-label neural network for road scene recognition, which incorporates both single- and multi-class classification modes into a multi-level cost function for training with imbalanced categories and utilizes a deep data integration strategy to improve the classification ability on hard samples. The experimental results on DrivingScene and PASCAL VOC demonstrate the effectiveness of the proposed approach in handling the challenge of data imbalance. Long Chen 0005, Wujing Zhan, Wei Tian 0001, Qin Zou 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Feature Selection Based on Tensor Decomposition and Object Proposal for Night-Time Multiclass Vehicle DetectionabstractNight-time vehicle detection is essential in building intelligent transportation systems (ITS) for road safety. Most of current night-time vehicle detection approaches focus on one or two classes of vehicles. In this paper, we present a novel multiclass vehicle detection system based on tensor decomposition and object proposal. Commonly used features such as histogram of oriented gradients and local binary pattern often produce useless image blocks (regions), which can result in unsatisfactory detection performance. Thus, we select blocks via feature ranking after tensor decomposition and only extract features from these selected blocks. To generate windows that contain all vehicles, we propose a novel object-proposal approach based on a state-of-the-art object-proposal method, local features, and image region similarity. The three terms are summed with learned weights to compute the reliability score of each proposal. A bio-inspired image enhancement method is used to enhance the brightness and contrast of input images. We have built a Hong Kong night-time multiclass vehicle dataset for evaluation. Our proposed vehicle detection approach can successfully detect four types of vehicles: 1) car; 2) taxi; 3) bus; and 4) minibus. Occluded vehicles and vehicles in the rain can also be detected. Our proposed method obtains 95.82% detection rate at 0.05 false positives per image, and it outperforms several state-of-the-art night-time vehicle detection approaches. Hulin Kuang, Long Chen 0005, Leanne Lai Chan, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | Dress Fashionably: Learn Fashion Collocation With Deep Mixed-Category Metric LearningabstractIn this paper, we seek to enable machine to answer questions like, given a clutch bag, what kind of skirt, heel and even accessory best fashionably collocate with it ? This problem, dubbed fashion collocation, has almost been neglected by researchers due to the large uncertainty lies in fashion collocation and professional expertise required to address it. In this paper, we narrow down the well-collocated samples to be fashion images shared on fashion websites, with which we propose an end-to-end trainable deep mixed-category metric learning method to project well-collocated clothing items to lie close but items violating well-collocation far apart in the deep embedding space. Specifically, we simultaneously model the intra-category exclusiveness and cross-category inclusiveness of fashion collocation by feeding a set of well-collocated clothing items and corresponding bad-collocated clothing items to the deep neural network, further a hard-aware online exemplar mining strategy is designed to force the whole neural network to be trainable and learn discriminative features at the early and later training stages respectively. To motivate more research in fashion collocation, we collect a dataset of 0.2 million fashionably well-collocated images consisting of either on-body or off-body clothing items or accessories. Extensive experimental results show the feasibility and superiority of our method. Long Chen 0005 |
AAAI | 1 |
| 2018 | A Robust Look-ahead Distance Tuning Strategy for the Geometric Path Tracking ControllersabstractGeometric path tracking controllers are very popular to be implemented in the autonomous vehicles because of its simple implementation and validation in the practical applications. However, it also suffered from the sensitivity to the selections of the look-ahead distance. Therefore, reasonable tuning strategy of the look-ahead distance is very important to the application of the geometric controllers. This paper proposed a more robust tuning strategy for all the geometric controllers based on the careful analysis of the previous methods. Instead of utilizing as much as possible elements, we just want to make use of only two variables, the distance from the current location to the reference path and the changing rate of it. Fuzzy logic was utilized to combine these two variables to tune the look-ahead distance. We will show that we just use two variables, but we actually consider another two variables' affections indirectly. Lastly, we will add our strategy to two different classical geometric controllers and compare the relative benefits of previous tuning strategies. Longsheng Chen, Yunxiao Shan, Long Chen 0005 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Planecell: Representing Structural Space with Plane ElementsabstractReconstruction based on the stereo camera has received considerable attention recently, but two particular challenges still remain. The first concerns the need to present and compress data in an effective way, and the second is to maintain as much of the available information as possible while ensuring sufficient accuracy. To overcome these issues, we propose a new 3D representation method, namely, planecell, that extracts planarity from the depth-assisted image segmentation and then directly projects these depth planes into the 3D world. The proposed method demonstrates its advancement especially dealing with large-scale structural environment, such as autonomous driving scene. The reconstruction result of our method achieves equal accuracy compared to dense point clouds and compresses the output file 200 times. To further obtain global surfaces, an energy function formulated from Conditional Random Field that generalizes the planar relationships is maximized. We evaluate our method with reconstruction baselines on the KITTI outdoor scene dataset, and the results indicate the superiorities compared to other 3D space representation methods in accuracy, memory requirements and the scope of applications. Lei Fan 0005, Long Chen 0005, Kai Huang 0001, Dongpu Cao |
Intelligent Vehicles Symposium | 2 |
| 2018 | A Reliable Road Segmentation and Edge Extraction for Sparse 3D Lidar DataabstractPrecise segmentation of road areas using cheap Lidar is a tough and critical task due to data sparsity problem. With sparse point clouds, reliable perception of environment is difficult due to the lack of available information and loss of object features. This paper presents a new approach to use sparse 3D Lidar data for road segmentation by fusing multiple frames of point cloud. With registration of multiple frames into a same coordinate system, reliable data can be provided for later ground segmentation and edge extraction. The accuracies of extensive experiments on three kinds of roads demonstrate that the proposed approach obtains high precision and reliability. Jianfeng Gu 0001, Yuehui Wang, Long Chen 0005, Zhe Xuanyuan, Kai Huang 0001 |
Intelligent Vehicles Symposium | 3 |
| 2018 | Learning a Deep Motion Planning Model for Autonomous DrivingabstractTo deal with the issue of computational complexity and robustness of traditional motion planning methods for autonomous driving, an end-to-end motion planning model based on a deep cascaded neural network is proposed in this paper. The model can directly predict the driving parameters from the input sequence images. We combine two classical deep learning models including the convolution neural network (CNN) and the long short-term memory (LSTM) which are used to extract spatial and temporary features of the input images, respectively. The proposed model can fit the nonlinear relationship between the input sequence images and the output motion parameters for making the end-to-end planning. The experiments are conducted using the data collected from a driving simulator. Experimental results show that the proposed method can efficiently learn humans' driving behaviors, adapt to different roads, and has a better robustness performance than some existing methods. Sheng Song, Xuemin Hu, Liyun Bai, Long Chen 0005 |
Intelligent Vehicles Symposium | 5 |
| 2018 | Online Cooperative 3D Mapping for Autonomous DrivingabstractAutonomous driving requires 3D representations of the environments as high definition maps. In many cases, it is not efficient for a single vehicle to map the entire large environment. Therefore, a group of vehicles could cooperate to build maps. In this paper, we propose an approach for cooperative 3D mapping by multiple vehicles working simultaneously as a team. Each vehicle uses 3D LIDAR sensor and local mapping algorithms to build local map and the global map can be obtained by merging all the local maps in an consistent manner. The challenges in cooperative mapping lie in both accuracy and efficiency. We show that our cooperative mapping approach can save mapping time as well as reduce the accumulated error often suffered by single vehicle mapping algorithms. Meanwhile, real world experiments results indicate that our mapping algorithm can be implemented online with minimum burden imposed on communication channel and computation resources on each vehicle. Zhe Xuanyuan, Boyang Li 0009, Long Chen 0005, Kai Huang 0001 |
Intelligent Vehicles Symposium | 4 |
| 2018 | Adas on Cots with OpenCL: A Case Study with Lane DetectionabstractThe concept of autonomous cars is driving a boost for car electronics and the size of automotive electronics market is foreseen to double by 2025. How to benefit from this boost is an interesting question. This article presents a case study to test the feasibility of using OpenCL as the programming language and Cots components as the underlying computing platforms for Adas development. For representative Adas applications, a scalable lane detection is developed that can tune the trade-off between detection accuracy and speed. Our OpenCL implementation is tested on 14 video streams from different data-sets with different road scenarios on 5 Cots platforms. We demonstrate that the Cots platforms can provide more than sufficient computing power for the lane detection in the meanwhile our OpenCL implementation can exploit the massive parallelism provided by the Cots platforms. Kai Huang 0001, Biao Hu 0001, Long Chen 0005, Alois C. Knoll, Zhihua Wang 0001 |
IEEE Trans. Computers | 3 |
| 2018 | Bayes Saliency-Based Object Proposal Generator for Nighttime Traffic ImagesabstractObject proposal is one of the most key pre-processing steps for nighttime vehicle detection systems in intelligent transportation systems. However, most current object proposal methods are developed on daytime data sets, and these methods demonstrate unsatisfactory results when they are used on nighttime images. Therefore, this paper presents a novel Bayes saliency-based object proposal generator for nighttime RGB traffic images to generate a modest and accurate set of proposals, which are more likely to be vehicles for preceding vehicle detection. First, we propose a new Bayes saliency detection approach in which prior estimation, feature extraction, weight estimation, and Bayes rule are used to compute saliency maps. Then, we propose a simple but effective object proposal generator based on the Bayes saliency map. Multi-scale sliding window, proposal rejecting, scoring, and non-maximum suppression are combined to generate a modest and effective set of proposals. Experimental results demonstrate that our proposed approach generates a modest set of proposals and outperforms some state-of-the-art methods on nighttime images in terms of various evaluation metrics. Furthermore, our proposed object proposal approach can improve the detection performance and the speed of several state-of-the-art vehicle detection approaches. Hulin Kuang, Kaifu Yang, Long Chen 0005, Yongjie Li 0001, Leanne Lai Chan, Hong Yan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Vehicle Tracking at Nighttime by Kernelized Experts With Channel-Wise and Temporal Reliability EstimationabstractDespite the fact that in recent years, vision-based tracking approaches have made significant progress, the task of tracking vehicles at night still remains challenging. Visual information is strongly deteriorated or at least degraded due to poor illumination conditions. This reduces the perceptive ability of vision systems significantly and can even lead to target loss, resulting in false estimation and/or false prediction of object behavior. In this paper, we propose a novel online-learning method to track vehicles at night. Our method is based on the kernelized correlation filter and assembles different feature channels to kernelized experts. By estimating their reliabilities, we force the appearance model to focus on the most discriminative visual features to accomplish the classification. In addition, a temporal optimization step in conjunction with a memory model is used to remove outliers and keep the most reliable samples to train the tracker models. Experiments over various daytime and weather conditions show that our approach outperforms existing trackers at night and in case of bad weather while offering state-of-the-art performance in more favorable situations. As our tracker has only little computational cost, it is appropriate for use cases with real-time requirements like in automotive or industrial applications. Wei Tian 0001, Long Chen 0005, Ke Zou, Martin Lauer |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | RGB-T SLAM: A flexible SLAM framework by combining appearance and thermal informationabstractVisual SLAM in low illumination scenes remains a considerably challenging task since the available amount of appearance information frequently stays insufficient. To tackle with this problem, we propose a novel SLAM framework by using both appearance information and thermal information, which possesses illumination-free recognizable contents, in a flexible manner. The key idea is to continuously update a RGB-T map, which contains both RGB and thermal map points to implement location and mapping. More specifically, in our SLAM system, we detect features in both RGB and thermal images and combine them together to update the RGB-T map and implement simultaneous location and mapping. Both quantitative and qualitative results demonstrate the effectiveness of our framework, especially under low illumination environments. Long Chen 0005, Libo Sun 0002, Lei Fan 0005, Kai Huang 0001, Zhe Xuanyuan |
ICRA | 1 |
| 2017 | Real-time scene flow on COTS embedded systems by coarse-grained software pipelineabstractScene flow is a key function of stereo-based environment perception system for mobile robotics and autonomous vehicle. Due to the heavy computing requirement and the limited computing resource, parallelized and embedded algorithms become quite important for the application of the mobile robotics. This paper develops a cross-platform embedded scene flow algorithm by using a coarse-grained software pipeline and OpenCL programming language. Our OpenCL algorithm is tested on 10 video streams from different datasets with different scenarios on different commercial-off-the-shelf (COTS) hardware. The average frame rates for the 10 videos can reach about 50 fps on both GPU and mobile device. The peak frame rates for certain videos on GPU can reach almost 450 fps. We also demonstrate that the COTS platform can provide sufficient computing power for stereo-based perception algorithm potentially by using OpenCL programming. Long Chen 0005, Mingyue Cui, Kai Huang 0001, Zhe Xuanyuan |
Intelligent Vehicles Symposium | 1 |
| 2017 | Compressive perceptual hashing tracking
Long Chen 0005, Jian-Fei Yang |
Neurocomputing | 1 |
| 2017 | Let the robot tell: Describe car image with natural language via LSTM
Long Chen 0005, Lei Fan 0005 |
Pattern Recognit. Lett. | 1 |
| 2017 | Moving-Object Detection From Consecutive Stereo Pairs Using Slanted Plane SmoothingabstractDetecting moving objects is of great importance for autonomous unmanned vehicle systems, and a challenging task especially in complex dynamic environments. This paper proposes a novel approach for the detection of moving objects and the estimation of their motion states using consecutive stereo image pairs on mobile platforms. First, we use a variant of the semi-global matching algorithm to compute initial disparity maps. Second, assisted by the initial disparities, boundaries in the image segmentation produced by simple linear iterative clustering are classified into coplanar, hinge, and occlusion. Moving points are obtained during ego-motion estimation by a modified random sample consensus) algorithm without resorting to time-consuming dense optical flow. Finally, the moving objects are extracted by merging superpixels according to the boundary types and their movements. The proposed method is accelerated on the GPU at 20 frames per second. The data which we use for testing and benchmarking is released, thus completing similar data sets. It includes 812 image pairs and 924 moving objects with ground truth for better algorithms evaluation. Experimental results demonstrate that the proposed method achieves competitive results in terms of moving-object detection and their motion state estimation in challenging urban scenarios. Long Chen 0005, Lei Fan 0005, Guodong Xie, Kai Huang 0001, Andreas Nüchter |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Transforming a 3-D LiDAR Point Cloud Into a 2-D Dense Depth Map Through a Parameter Self-Adaptive FrameworkabstractThe 3-D LiDAR scanner and the 2-D charge-coupled device (CCD) camera are two typical types of sensors for surrounding-environment perceiving in robotics or autonomous driving. Commonly, they are jointly used to improve perception accuracy by simultaneously recording the distances of surrounding objects, as well as the color and shape information. In this paper, we use the correspondence between a 3-D LiDAR scanner and a CCD camera to rearrange the captured LiDAR point cloud into a dense depth map, in which each 3-D point corresponds to a pixel at the same location in the RGB image. In this paper, we assume that the LiDAR scanner and the CCD camera are accurately calibrated and synchronized beforehand so that each 3-D LiDAR point cloud is aligned with its corresponding RGB image. Each frame of the LiDAR point cloud is then projected onto the RGB image plane to form a sparse depth map. Then, a self-adaptive method is proposed to upsample the sparse depth map into a dense depth map, in which the RGB image and the anisotropic diffusion tensor are exploited to guide upsampling by reinforcing the RGB-depth compactness. Finally, convex optimization is applied on the dense depth map for global enhancement. Experiments on the KITTI and Middlebury data sets demonstrate that the proposed method outperforms several other relevant state-of-the-art methods in terms of visual comparison and root-mean-square error measurement. Long Chen 0005, Jianda Chen, Qingquan Li 0001, Qin Zou 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Turn Signal Detection During Nighttime by CNN Detector and Perceptual Hashing TrackingabstractDetecting vehicle turn signals at night is critical for both assistant driving systems and autonomous driving systems. In this paper, we propose a novel method that consists of detection and tracking modules to achieve a high level of robustness. For nighttime vehicle detection, a Nakagami-image-based method is used to locate the regions containing vehicle lights. At the same time, a set of vehicle object proposals is generated using a region proposal network based on convolutional neural network (CNN) feature maps. Then, the light regions and proposals are combined to generate the regions of interest (ROIs) for the further detection. Vehicle candidates are extracted from the ROIs using a softmax classifier with CNN-based features. For the tracking module, we propose a perceptional hashing algorithm to track these vehicle candidates. During the tracking, turn signals are detected by analyzing the continuous intensity variation of the vehicle box sequences. Experimental results for typical sequences show that the proposed method can robustly detect and track a vehicle in front with over 95% accuracy and recognize the turning signals in night scenes with a detection rate of over 90%. The vehicle detection method improves the miss rate of state-of-the-art systems by more than 20%. In addition, the proposed vehicle tracking method outperforms other state-of-the-art systems. Long Chen 0005, Xuemin Hu, Tong Xu 0014, Hulin Kuang, Qingquan Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2016 | Fast Fashion Guided Clothing Image Retrieval: Delving Deeper into What Feature Makes Fashion
Long Chen 0005 |
ACCV (5) | 2 |
| 2016 | Multi-task Relative Attribute Prediction by Incorporating Local Context and Global Style Information
Long Chen 0005, Jianda Chen |
BMVC | 2 |
| 2016 | Geodesic-based pavement shadow removal revisitedabstractShadows often incur uneven illumination to pavement images, which brings great challenges to image-based pavement crack detection. Thus, it is desired to remove pavement shadows before detecting pavement cracks. However, due to the large penumbras cast by trees, light poles, etc., it is difficult to locate shadows in a pavement image. In this paper, an automatic pavement shadow removal method is proposed based on geodesic analysis. First, a geodesic shadow model is used to partition a pavement shadow into a number of geodesic regions. Then, an optimal background region is selected for reference by statistic analysis. Finally, a texture-balanced illuminance compensation is applied on all geodesic regions over the image. Experiments demonstrate the effectiveness of the proposed method. Qin Zou 0001, Zhongwen Hu, Long Chen 0005, Qian Wang 0002, Qingquan Li 0001 |
ICASSP | 3 |
| 2015 | Sparse depth map upsampling with RGB image and anisotropic diffusion tensorabstractThis paper proposes a novel algorithm for up-sampling 2D sparse depth map projected by laser scanner (i.e. Velodyne HDL-64E) using its synchronised RGB image and Anisotropic Diffusion Tensor. We assume each depth-unknown pixel's depth value derives from all depth-known pixels and their affinity can be measured in a geodesic manner. Specifically, for each depth-unknown point, we compute its geodesic distance to all depth-known points, the cost of each geodesic path is calculated by spatial distance, color aberrance and tensor discrepancy, which compacts with the assumption that color homogenous region corresponds to close or smooth depth value, while depth discontinuity often happens in image edges. To mitigate computation complexity, we further introduced a flexible approximation algorithm in which the complexity is linear to image's size and can be further reduced w.r.t. different accuracy requirement. Finally, we evaluate our algorithm on both KITTI visual benchmark suite and Middlebury dataset. Experiment shows that, while generating smooth and dense upsampling result, our algorithm retains sharp depth discontinuity even in edges of objects lies very far and few laser scanner points cover it. Besides, our algorithm is parameter-nonsensitive, which frees us from laboriously finding appropriate parameters to get fine result. We hope our work would motivate more research on laser point upsampling and their combination with RGB image, especially in outdoor scenes for autonomous driving. Long Chen 0005 |
Intelligent Vehicles Symposium | 2 |
| 2015 | Reinforcement learning control for coordinated manipulation of multi-robots
Yanan Li 0001, Long Chen 0005, Keng Peng Tee, Qingquan Li 0001 |
Neurocomputing | 2 |
| 2015 | ALIMC: Activity Landmark-Based Indoor Mapping via CrowdsourcingabstractIndoor maps are integral to pedestrian navigation systems, an essential element of intelligent transportation systems (ITS). In this paper, we propose ALIMC, i.e., Activity Landmark-based Indoor Mapping system via Crowdsourcing. ALIMC can automatically construct indoor maps for anonymous buildings without any prior knowledge using crowdsourcing data collected by smartphones. ALIMC abstracts the indoor map using a link-node model in which the pathways are the links and the intersections of the pathways are the nodes, such as corners, elevators, and stairs. When passing through the nodes, pedestrians do the corresponding activities, which are detected by smartphones. After activity detection, ALIMC extracts the activity landmarks from the crowdsourcing data and clusters the activity landmarks into different clusters, each of which is treated as a node of the indoor map. ALIMC then estimates the relative distances between all the nodes and obtains a distance matrix. Based on the distance matrix, ALIMC generates a relative indoor map using the multidimensional scaling technique. Finally, ALIMC converts the relative indoor map into an absolute one based on several reference points. To evaluate ALIMC, we implement ALIMC in an office building. Experiment results show that the 80th percentile error of the mapping accuracy is about 0.8-1.5 m. Baoding Zhou, Qingquan Li 0001, Qingzhou Mao, Wei Tu 0001, Xing Zhang 0003, Long Chen 0005 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2012 | 3D LIDAR point cloud based intersection recognition for autonomous drivingabstractFinding road intersections in advance is crucial for navigation and path planning of moving autonomous vehicles, especially when there is no position or geographic auxiliary information available. In this paper, we investigate the use of a 3D point cloud based solution for intersection and road segment classification in front of an autonomous vehicle. It is based on the analysis of the features from the designed beam model. First, we build a grid map of the point cloud and clear the cells which belong to other vehicles. Then, the proposed beam model is applied with a specified distance in front of autonomous vehicle. A feature set based on the length distribution of the beam is extracted from the current frame and combined with a trained classifier to solve the road-type classification problem, i.e., segment and intersection. In addition, we also make the distinction between +-shaped and T-shaped intersections. The results are reported over a series of real-world data. A performance of above 80% correct classification is reported at a real-time classification rate of 5 Hz. Quanwen Zhu, Long Chen 0005, Qingquan Li 0001, Andreas Nüchter, Jian Wang 0086 |
Intelligent Vehicles Symposium | 2 |
| 2011 | Traffic sign detection and recognition for intelligent vehicleabstractIn this paper, we propose a computer vision based system for real-time robust traffic sign detection and recognition, especially developed for intelligent vehicle. In detection phase, a color-based segmentation method is used to scan the scene in order to quickly establish regions of interest (ROI). Sign candidates within ROIs are detected by a set of Haar wavelet features obtained from AdaBoost training. Then, the Speeded Up Robust Features (SURF) is applied for the sign recognition. SURF finds local invariant features in a candidate sign and matches these features to the features of template images that exist in data set. The recognition is performed by finding out the template image that gives the maximum number of matches. We have evaluated the proposed system on our intelligent vehicle SmartVII. A recognition accuracy of over 90% in real-time has been achieved. Long Chen 0005, Qingquan Li 0001, Qingzhou Mao |
Intelligent Vehicles Symposium | 1 |
| 2010 | Block-constraint line scanning method for lane detectionabstractConsidering the plentiful road markings in China, we present a Block-Constraint Line scanning (BCLS) method for lane detection in this paper. In this method, images are firstly pre-processed by a morphological top-hat transform, and then an imaging model is created for building relationship between lane parameters of the image coordinate and the WGS coordinate, from which target points on lane lines could be retained by a block-constraint line scanning algorithm. Finally, lanes could be extracted by a Progressive Probabilistic Hough Transform (PPHT) and the number of lanes is figured out through clustering. Our method is fast enough to meet real-time requirement. Experiments were carried out on the intelligent vehicle SmartV (Fig.1) on the Wuhan urban roads in China and the results show that this method can efficiently and accurately extract lanes in complex environments, even with the presence of non-lane road markings. Long Chen 0005, Qingquan Li 0001, Qingzhou Mao, Qin Zou 0001 |
Intelligent Vehicles Symposium | 1 |