VLDB 2026 Research / reviewers in the wild / expert
Bin Dai 0001
dblp:79/292-1
· DBLP profile ↗
35ranked-venue papers
1as first author
13since 2021 · last 2025
0000-0001-9405-2626ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TPKD: Teacher-Pruned Knowledge Distillation for Point Cloud-Based 3D Object Detection
Liang Xiao 0007, Dawei Zhao 0003, Qi Zhu 0004, Yiming Nie, Bin Dai 0001 |
ICIC (22) | 6 |
| 2025 | Spatiotemporal Context Adapting Framework for Visual Object TrackingabstractABSTRACT Visual object tracking is widely applied in intelligent transportation systems and visual surveillance systems that serve smart cities, as well as in autonomous vehicles. Existing methods usually utilise a relation‐modelling framework to model the visual object tracking problem, with auxiliary spatial context and temporal information. The spatial context is often extracted by enlarging the target template, which can introduce more background and positional information. The temporal correlation is obtained by associating the search image with previous images. However, due to noise interference, existing methods often partially exploit auxiliary data, leading to underutilisation of spatiotemporal information. To address these issues, we propose a novel and concise tracking framework, uniformly encoding all auxiliary data, including the enlarged target template, previous images, and corresponding target bounding boxes. Specifically, to mitigate the unstable factors introduced by these raw inputs, we propose a spatiotemporal context adaptive encoder, which can adaptively select appropriate information in noisy data. Extensive experiments show that the proposed method achieves state‐of‐the‐art performance on various benchmarks, demonstrating its superiority. Kunlong Zhao, Dawei Zhao 0003, Xu Wang 0043, Liang Xiao 0007, Yulong Huang 0003, Yiming Nie, Yonggang Zhang 0001, Bin Dai 0001 |
IET Image Process. | 8 |
| 2025 | Efficient Distillation Using Channel Pruning for Point Cloud-Based 3D Object DetectionabstractAlthough point cloud-based 3D object detectors have advanced significantly in recent years, they are frequently hindered by substantial computational overheads. Lightweight model techniques, such as knowledge distillation, have recently been proven effective for 3D object detector compression. However, neural network pruning’s complementary role in knowledge distillation is often overlooked. In this paper, we propose an efficient distillation using channel pruning for point cloud-based 3D object detection. Firstly, given the complete teacher model, we introduce random and magnitude channel pruning methods to generate several compact student models and investigate the effects of different combinations on 3D and 2D layers. Secondly, we introduce model compression scores to explore the impact of channel compression ratios and input resolutions, enabling us to select suitable pruned models for distillation from the given set. Furthermore, we employ multi-source knowledge distillation to facilitate more effective spatial and semantic knowledge transfer. To highlight the features of the foreground regions during distillation, we then propose a soft pivotal position selection mask. Extensive evaluations on various datasets using both pillar-and voxel-based 3D detectors validate the efficiency of our method in compressing point cloud-based 3D detectors. Codes are publicly available at https://github.com/lifuyang-1919/Efficient-Distillation.git Juan Wang 0033, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | Contrastive Label Disambiguation for Self-Supervised Terrain Traversability Learning in Off-Road EnvironmentsabstractDiscriminating terrain traversability stands as a pivotal challenge for autonomous driving in off-road environments. The complexity arises from the diverse and ambiguous nature of off-road conditions, coupled with the specific characteristics of the driving platform. To address this challenge, we introduce a novel self-supervised learning framework for terrain traversability analysis, incorporating a contrastive label disambiguation mechanism. The proposed framework integrates traversability learning with real-time scene reconstruction. By projecting actual driving experience onto the terrain models, weakly labeled training samples with pseudo-labels can be automatically generated. Furthermore, a prototype-based contrastive representation learning method with the aid of a local window-based transformer encoder is designed to learn distinguishable embeddings, facilitating the self-supervised updating of those pseudo labels. Through the iterative interaction between representation learning and pseudo label updating, the inherent ambiguities associated with those pseudo labels are gradually eliminated. This enables the acquisition of fine-grained and platform-specific terrain traversability insights, eliminating the need for any human-provided annotations. Experimental results on the publicly available RELLIS-3D dataset and two self-collected datasets demonstrate the effectiveness of the proposed method. Hanzhang Xue, Liang Xiao 0007, Xiaochang Hu, Hao Fu 0001, Yiming Nie, Bin Dai 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingabstractVision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning. Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001 |
CVPR | 13 |
| 2024 | Pre-pruned Distillation for Point Cloud-based 3D Object DetectionabstractKnowledge distillation has recently been proven to be effective for model compression and acceleration of point cloud-based 3D object detection. However, the complementary network pruning is often overlooked during knowledge distillation. In this paper, we propose a pre-pruned distillation framework that combines network pruning and knowledge distillation to better transfer knowledge from the teacher to the student. To maintain the feature consistency between the student and the teacher, we train a teacher model and then generate a compact student model by structural channel pruning. Then, we employ multi-source knowledge distillation to transfer both mid-level and high-level information to the student model. Additionally, to improve the object detection performance of the student model, we propose a soft pivotal position selection mask to emphasize the features of the foreground regions during distillation. We conduct experiments on both pillarand voxel-based 3D object detectors on the Waymo datasets, demonstrating the effectiveness of our approach in compressing point cloud-based 3D detectors. Liang Xiao 0007, Dawei Zhao 0003, Shubin Si, Hanzhang Xue, Yiming Nie, Bin Dai 0001 |
IV | 8 |
| 2024 | A Two-Stage Active Domain Adaptation Framework for Vehicle Re-Identification
Linzhi Shang, Dawei Zhao 0003, Yiming Nie, Kunlong Zhao, Liang Xiao 0007, Bin Dai 0001 |
PRCV (1) | 6 |
| 2024 | Deep Reinforcement Learning: A SurveyabstractDeep reinforcement learning (DRL) integrates the feature representation ability of deep learning with the decision-making ability of reinforcement learning so that it can achieve powerful end-to-end learning control capabilities. In the past decade, DRL has made substantial advances in many tasks that require perceiving high-dimensional input and making optimal or near-optimal decisions. However, there are still many challenging problems in the theory and applications of DRL, especially in learning control tasks with limited samples, sparse rewards, and multiple agents. Researchers have proposed various solutions and new theories to solve these problems and promote the development of DRL. In addition, deep learning has stimulated the further development of many subfields of reinforcement learning, such as hierarchical reinforcement learning (HRL), multiagent reinforcement learning, and imitation learning. This article gives a comprehensive overview of the fundamental theories, key algorithms, and primary research domains of DRL. In addition to value-based and policy-based DRL algorithms, the advances in maximum entropy-based DRL are summarized. The future research topics of DRL are also analyzed and discussed. Xu Wang 0043, Xingxing Liang, Dawei Zhao 0003, Jincai Huang 0001, Xin Xu 0001, Bin Dai 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Fast and Accurate Deep Loop Closing and Relocalization for Reliable LiDAR SLAMabstractLoop closing and relocalization are crucial techniques to establish reliable and robust long-term SLAM by addressing pose estimation drift and degeneration. This article begins by formulating loop closing and relocalization within a unified framework. Then, we propose a novel multi-head network LCR-Net to tackle both tasks effectively. It exploits novel feature extraction and a pose-aware attention mechanism to precisely estimate similarities and 6-DoF poses between pairs of LiDAR scans. In the end, we integrate our LCR-Net into a SLAM system and achieve robust and accurate online LiDAR SLAM in outdoor driving environments. We thoroughly evaluate our LCR-Net through three setups derived from loop closing and relocalization, including candidate retrieval, closed-loop point cloud registration, and continuous relocalization using multiple datasets. The results demonstrate that LCR-Net excels in all three tasks, surpassing the state-of-the-art methods and exhibiting a remarkable generalization ability. Notably, our LCR-Net outperforms baseline methods without using a time-consuming robust pose estimator, rendering it suitable for online SLAM applications. To our best knowledge, the integration of LCR-Net yields the first LiDAR SLAM with the capability of deep loop closing and relocalization. The implementation of our methods is open-sourced athttps://github.com/nubot-nudt/LCR-Net. Chenghao Shi, Xieyuanli Chen, Junhao Xiao 0001, Bin Dai 0001, Huimin Lu 0002 |
IEEE Trans. Robotics | 4 |
| 2023 | RDMNet: Reliable Dense Matching Based Point Cloud Registration for Autonomous DrivingabstractPoint cloud registration is an important task in robotics and autonomous driving to estimate the ego-motion of the vehicle. Recent advances following the coarse-to-fine manner show promising potential in point cloud registration. However, existing methods rely on good superpoint correspondences, which are hard to be obtained reliably and efficiently, thus resulting in less robust and accurate point cloud registration. In this paper, we propose a novel network, named RDMNet, to find dense point correspondences coarse-to-fine and improve final pose estimation based on such reliable correspondences. Our RDMNet uses a devised 3D-RoFormer mechanism to first extract distinctive superpoints and generates reliable superpoints matches between two point clouds. The proposed 3D-RoFormer fuses 3D position information into the transformer network, efficiently exploiting point clouds’ contextual and geometric information to generate robust superpoint correspondences. RDMNet then propagates the sparse superpoints matches to dense point matches using the neighborhood information for accurate point cloud registration. We extensively evaluate our method on multiple datasets from different environments. The experimental results demonstrate that our method outperforms existing state-of-the-art approaches in all tested datasets with a strong generalization ability. Chenghao Shi, Xieyuanli Chen, Huimin Lu 0002, Wenbang Deng, Junhao Xiao 0001, Bin Dai 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2022 | ORFD: A Dataset and Benchmark for Off-Road Freespace DetectionabstractFreespace detection is an essential component of autonomous driving technology and plays an important role in trajectory planning. In the last decade, deep learning based freespace detection methods have been proved feasible. However, these efforts were focused on urban road environments and few deep learning based methods were specifically designed for off-road freespace detection due to the lack of off-road dataset and benchmark. In this paper, we present the ORFD dataset, which, to our knowledge, is the first off-road freespace detection dataset. The dataset was collected in different scenes (woodland, farmland, grassland and countryside), different weather conditions (sunny, rainy, foggy and snowy) and different light conditions (bright light, daylight, twilight, darkness), which totally contains 12,198 LiDAR point cloud and RGB image pairs with the traversable area, non-traversable area and unreachable area annotated in detail. We propose a novel network named OFF-Net, which unifies Transformer architecture to aggregate local and global information, to meet the requirement of large receptive fields for freespace detection task. We also propose the cross-attention to dynamically fuse LiDAR and RGB image information for accurate off-road freespace detection. Dataset and code are publicly available at https://github.com/chaytonmin/OFF-Net. Weizhong Jiang, Dawei Zhao 0003, Jiaolong Xu, Liang Xiao 0007, Yiming Nie, Bin Dai 0001 |
ICRA | 7 |
| 2022 | Trajectory Prediction for Autonomous Driving with Topometric MapabstractState-of-the-art autonomous driving systems rely on high definition (HD) maps for localization and navigation. However, building and maintaining HD maps is time-consuming and expensive. Furthermore, the HD maps assume structured environment such as the existence of major road and lanes, which are not present in rural areas. In this work, we propose an end-to-end transformer networks based approach for map-less autonomous driving. The proposed model takes raw LiDAR data and noisy topometric map as input and produces precise local trajectory for navigation. We demonstrate the effectiveness of our method in real-world driving data, including both urban and rural areas. The experimental results show that the proposed method outperforms state-of-the-art multimodal methods and is robust to the perturbations of the topometric map. The code of the proposed method is publicly available at https://github.com/Jiaolong/trajectory-prediction. Jiaolong Xu, Liang Xiao 0007, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001 |
ICRA | 5 |
| 2021 | LiDAR-based Drivable Region Detection for Autonomous DrivingabstractFor autonomous driving, drivable region detection is one of the most basic and essential tasks. In this paper, a novel LiDAR-based drivable region detection algorithm which could output a complete, accurate and stable result is proposed. To promote the completeness of the detection result, the Bayesian generalized kernel inference and bilateral filtering are utilized to estimate the attribute of those unobserved cells. To ensure the traversability, a region growing operator is performed on the normal vector map which reflects the slope of the terrain, thus closely related to the traversability of the vehicle. To improve the result’s stability, information from multiple frames are fused together in the Kalman Filter framework. Experiments are performed both on public dataset and our own dataset. Experimental results show that the proposed algorithm could run in real-time and outperforms state-of-the-art approaches. Hanzhang Xue, Hao Fu 0001, Ruike Ren, Bokai Liu, Bin Dai 0001 |
IROS | 7 |
| 2020 | Self-Supervised Domain Adaptation with Consistency TrainingabstractWe consider the problem of unsupervised domain adaptation for image classification. To learn target-domain-aware features from the unlabeled data, we create a self-supervised pretext task by augmenting the unlabeled data with a certain type of transformation (specifically, image rotation) and ask the learner to predict the properties of the transformation. However, the obtained feature representation may contain a large amount of irrelevant information with respect to the main task. To provide further guidance, we force the feature representation of the augmented data to be consistent with that of the original data. Intuitively, the consistency introduces additional constraints to representation learning, therefore, the learned representation is more likely to focus on the right information about the main task. Our experimental results validate the proposed method and demonstrate state-of-the-art performance on classical domain adaptation benchmarks. Code is available at https://github.com/Jiaolong/ss-da-consistency. Liang Xiao 0007, Jiaolong Xu, Dawei Zhao 0003, Yiming Nie, Bin Dai 0001 |
ICPR | 7 |
| 2020 | Drosophila-inspired 3D moving object detection based on point clouds
Dawei Zhao 0003, Tao Wu 0001, Hao Fu 0001, Liang Xiao 0007, Xin Xu 0001, Bin Dai 0001 |
Inf. Sci. | 8 |
| 2019 | Augmenting cascaded correlation filters with spatial-temporal saliency for visual tracking
Dawei Zhao 0003, Liang Xiao 0007, Hao Fu 0001, Tao Wu 0001, Xin Xu 0001, Bin Dai 0001 |
Inf. Sci. | 6 |
| 2018 | Toward Autonomous Driving in Highway and Urban Environment: HQ3 and IVFC 2017abstractThe 2017 Intelligent Vehicle Future Challenge of China (IVFC) was held in Changshu between 24th November and 26th November, 2017. As the ninth series of this event, last year's competition has introduced many new features and has attracted 21 teams to join this competition. The HQ3 autonomous vehicle, jointly developed by National University of Defense Technology, Jilin University and Central South University, took part in this competition. This paper mainly describes the key modules of HQ3, including GPS-free localization, environment perception and behavior planning. All of these modules together enable HQ3 to perform well during the competition. Lilin Qian, Hao Fu 0001, Xiaohui Li 0007, Bang Cheng, Tingbo Hu, Zengping Sun, Tao Wu 0001, Bin Dai 0001, Xin Xu 0001 |
Intelligent Vehicles Symposium | 8 |
| 2018 | Hybrid conditional random field based camera-LIDAR fusion for road detection
Liang Xiao 0007, Ruili Wang 0001, Bin Dai 0001, Yuqiang Fang, Daxue Liu, Tao Wu 0001 |
Inf. Sci. | 3 |
| 2017 | Accurate extrinsic calibration between monocular camera and sparse 3D Lidar points without markersabstractIt is of practical interest to automatically calibrate the multiple sensors in autonomous vehicles. In this paper, we deal with an interesting case when used low-resolution Lidar and present a practical approach to extrinsic calibration between monocular camera and Lidar with sparse 3D measurements. We formulate the problem as directly minimizing the feature error evaluated between frames following the way of image warping. To overcome the difficulties in the optimization problem, we propose to use the distance transform and further projection error model to obtain the key approximated edge points that are sensitive to the loss function. Finally, the loss minimization is solved by an efficient random selection algorithm. Experimental results on KITTI dataset show that our proposed method can achieve competitive results and an improvement in translation estimation particularly. Zhipeng Xiao, Hongdong Li, Dingfu Zhou, Yuchao Dai, Bin Dai 0001 |
Intelligent Vehicles Symposium | 5 |
| 2016 | Learning deep compact channel features for object detection in traffic scenesabstractIn this work, we present a new multiple channel feature called Deep Compact Channel Feature (DCCF), which generates a compact, discriminative feature representation by a pre-trained deep encoder-decoder. With the combination of DCCF and boosted decision trees, a new object detector is proposed which achieved outstanding performance on standard pedestrian dataset INRIA and Caltech. Furthermore, a large scale and challenging Chinese Traffic Sign Detection benchmark is constructed. DCCF and other related methods are evaluated on this dataset. The dataset and baselines are available online. Yuqiang Fang, Lin Sun 0004, Hao Fu 0001, Tao Wu 0001, Ruili Wang 0001, Bin Dai 0001 |
ICIP | 6 |
| 2016 | Likelihood-Field-Model-Based Dynamic Vehicle Detection and Tracking for Self-DrivingabstractDynamic vehicle detection and tracking is crucial for self-driving in urban environments. The main problem of the previous beam-model-based algorithms is that they cannot detect and track dynamic vehicles that are occluded by other objects. In this paper, we develop a novel dynamic vehicle detection and tracking algorithm to solve this problem for our autonomous land vehicle (ALV), which is equipped with a Velodyne LIDAR and a GPS-aid inertial navigation system. For detection, our improved two-dimensional virtual scan is presented to detect the potential dynamic vehicles with a scan differencing operation. Then, for each potential dynamic vehicle, a novel likelihood-field-based vehicle measurement model is proposed to weight its possible poses. Finally, our newly modified scaling series algorithm and the importance sampling technique are adopted to estimate the initial pose and the corresponding velocity for each vehicle, respectively. The scaling series algorithm coupled with a Bayesian filter (SSBF) was previously used to handle the tactile localization problem in static background scenes. For tracking dynamic vehicles, we improve the SSBF by adding the ego-motion compensation so that the improved algorithm is able to update the pose and velocity for each vehicle in dynamic background scenes. Both the quantitative and qualitative experimental results validate the performance of our dynamic vehicle detection and tracking algorithm on the KITTI datasets and the Velodyne data collected by our ALV in dynamic urban environments. Tongtong Chen, Ruili Wang 0001, Bin Dai 0001, Daxue Liu, Jinze Song |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Velodyne-based curb detection up to 50 meters awayabstractLong range curb detection is crucial for an Autonomous Land Vehicle (ALV) navigation in urban environments. This paper presents a novel curb detection algorithm which can detect the curbs up to 50 meters away with Velodyne LIDAR. Instead of building a Digital Elevation Map (DEM) and utilizing geometric features (like normal direction) to extract candidate curb points, we take each scan line of Velodyne LIDAR as a processing unite directly. Some feature points, which are extracted from individual scan lines, are selected as the initial curb points by the distance criterion and Hough Transform (HT). Eventually, iterative Gaussian Process Regression (GPR), which utilizes the above initial curb points as the initial seeds, is exploited to represent both the curved and straight-line curb model. In order to verify the effectiveness of our algorithm quantitatively, 2934 Velodyne scans are collected in various urban scenes with our ALV, and 566 of them are labelled manually1. Our algorithm is also compared with two other curb detection techniques. The experimental results on the dataset show promising performance. Tongtong Chen, Bin Dai 0001, Daxue Liu, Jinze Song |
Intelligent Vehicles Symposium | 2 |
| 2015 | CRF based road detection with multi-sensor fusionabstractIn this paper, we propose to fuse the LIDAR and monocular image in the framework of conditional random field to detect the road robustly in challenging scenarios. LIDAR points are aligned with pixels in image by cross calibration. Then boosted decision tree based classifiers are trained for image and point cloud respectively. The scores of the two kinds of classifiers are treated as the unary potentials of the corresponding pixel nodes of the random field. The fused conditional random field can be solved efficiently with graph cut. Extensive experiments tested on KITTI-Road benchmark show that our method reaches the state-of-the-art. Liang Xiao 0007, Bin Dai 0001, Daxue Liu, Tingbo Hu, Tao Wu 0001 |
Intelligent Vehicles Symposium | 2 |
| 2015 | Graph-Based Learning via Auto-Grouped Sparse Regularization and Kernelized ExtensionabstractThe key task in developing graph-based learning algorithms is constructing an informative graph to express the contextual information of a data manifold. Since traditional graph construction methods are sensitive to noise and less datum-adaptive to changes in density, a new method called$\ell^1$-graph was proposed recently. A graph construction needs to have two important properties: sparsity and locality. The$\ell^1$-graph has a strong sparsity property, but a weak locality property. Thus, we propose a new method of constructing an informative graph using auto-grouped sparse regularization based on the$\ell^1$-graph, which is called as Group Sparse graph (GS-graph). We also show how to efficiently construct a GS-graph in reproducing kernel Hilbert space with the kernel trick. The new methods, the GS-graph and its kernelized version (KGS-graph), have the same noise-insensitive property as that of$\ell^1$-graph and also can successively preserve the properties of sparsity and locality simultaneously. Furthermore, we integrate the proposed graph with several graph-based learning algorithms to demonstrate the effectiveness of our method. The empirical studies on benchmarks show that the proposed methods outperform the$\ell^1$-graph and other traditional graph construction methods in various learning tasks. Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2014 | Performance of global descriptors for velodyne-based urban object recognitionabstractObject Recognition is an essential component for Autonomous Land Vehicle (ALV) navigation in urban environments. This paper presents a thorough evaluation of the performance of some state of the art global descriptors on the public Sydney Urban Objects Dataset1, which was collected in the Central Business District of Sydney. These descriptors are Bounding Box descriptor, Histogram of Local Point Level descriptor, Hierarchy descriptor, and Spin Image (SI). We also propose a novel Global Fourier Histogram (GFH) descriptor. Experimental results on the public data set show that GFH descriptor turns out to be one of the best global descriptors for the object recognition in urban environments, and the results on the data collected by our own ALV in urban environments also demonstrate its usefulness. Tongtong Chen, Bin Dai 0001, Daxue Liu, Jinze Song |
Intelligent Vehicles Symposium | 2 |
| 2014 | Efficient Vehicle Localization Based on Road-Boundary Maps
Dawei Zhao 0003, Tao Wu 0001, Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001 |
PRICAI | 6 |
| 2014 | Decomposition and Extraction: A New Framework for Visual ClassificationabstractIn this paper, we present a novel framework for visual classification based on hierarchical image decomposition and hybrid midlevel feature extraction. Unlike most midlevel feature learning methods, which focus on the process of coding or pooling, we emphasize that the mechanism of image composition also strongly influences the feature extraction. To effectively explore the image content for the feature extraction, we model a multiplicity feature representation mechanism through meaningful hierarchical image decomposition followed by a fusion step. In particularly, we first propose a new hierarchical image decomposition approach in which each image is decomposed into a series of hierarchical semantical components, i.e, the structure and texture images. Then, different feature extraction schemes can be adopted to match the decomposed structure and texture processes in a dissociative manner. Here, two schemes are explored to produce property related feature representations. One is based on a single-stage network over hand-crafted features and the other is based on a multistage network, which can learn features from raw pixels automatically. Finally, those multiple midlevel features are incorporated by solving a multiple kernel learning task. Extensive experiments are conducted on several challenging data sets for visual classification, and experimental results demonstrate the effectiveness of the proposed method. Yuqiang Fang, Qiang Chen 0007, Lin Sun 0004, Bin Dai 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2012 | Graph-Oriented Learning via Automatic Group Sparsity for Data AnalysisabstractThe key task in graph-oriented learning is constructing an informative graph to model the geometrical and discriminant structure of a data manifold. Since traditional graph construction methods are sensitive to noise and less datum-adaptive to changes in density, a new graph construction method so-called ℓ1-Graph has been proposed [1] recently. A graph construction method needs to have two important properties: sparsity and locality. However, the ℓ1-Graph is strong in sparsity property, but weak in locality. In order to overcome such limitation, we propose a new method of constructing an informative graph using automatic group sparse regularization based on the work of ℓ1-Graph, which is called as group sparse graph (GroupSp-Graph). The newly developed GroupSp-Graph has the same noise-insensitive property as ℓ1-Graph, and also can successively preserve the group and local information in the graph. In other words, the proposed group sparse graph has both properties of sparsity and locality simultaneously. Furthermore, we integrate the proposed graph with several graph-oriented learning algorithms: spectral embedding, spectral clustering, subspace learning and manifold regularized non-negative matrix factorization. The empirical studies on benchmark data sets show that the proposed algorithms achieve considerable improvement over classic graph constructing methods and the ℓ1-Graph method in various learning task. Yuqiang Fang, Ruili Wang 0001, Bin Dai 0001 |
ICDM | 3 |
| 2011 | Adaptive sample collection using active learning for kernel-based approximate policy iterationabstractApproximate policy iteration (API) has been shown to be a class of reinforcement learning methods with stability and sample efficiency. However, sample collection is still an open problem which is critical to the performance of API methods. In this paper, a novel adaptive sample collection strategy using active learning-based exploration is proposed to enhance the performance of kernel-based API. In this strategy, an online kernel-based least squares policy iteration (KLSPI) method is adopted to construct nonlinear features and approximate the Q-function simultaneously. Therefore, more representative samples can be obtained for value function approximation. Simulation results on typical learning control problems illustrate that by using the proposed strategy, the performance of KLSPI can be improved remarkably. Chunming Liu, Xin Xu 0001, Haiyun Hu, Bin Dai 0001 |
ADPRL | 4 |
| 2011 | How Can Multipath Dissemination Help to Detect Prefix Hijacking?abstractMultiple path dissemination is recognized as an important feature to improve the reliability and efficiency of networks. However, most multipath routing proposals focus only on disseminating additional routes to increase the reliability of the Internet, and do not concern the security issues that the additional routes bring to the table. In this paper, we attempt to understand the feasibility of utilizing multipath dissemination to detect prefix hijacking attacks. We investigate the minimum level of security enhancement that would be achieved by multipath dissemination. We systematically analyze the effectiveness of two types of multipath advertisements, advertising the most disjoint path advertisement or the second best path advertisement with the best path, on detecting prefix hijacking. Our analysis and measurement results show that advertising the second best route with the best route is an efficient and effective way to disseminate multipath information with respect to inter-domain routing security. Feng Wang 0017, Bin Dai 0001, Jinshu Su |
ICCCN | 2 |
| 2011 | LIDAR-based Long Range Road Intersection DetectionabstractLong range road intersection detection is crucial for localization and local path planning of autonomous vehicle in urban environments. In this paper, a new long-range road intersection detection approach for autonomous vehicle equipped with 3D LIDAR is presented. The approach first analyzes the admissible space in front of the autonomous vehicle, and then a virtual 3D LIDAR is placed in the admissible space 20 meters away from the vehicle. Finally the beam model of range finders and an improved toe-finding algorithm for virtual 3D LIDAR is used to find the road intersection. Experiments are carried out at the autonomous vehicle in campus, and results show the promising performance of the presented method. Tongtong Chen, Bin Dai 0001, Daxue Liu |
ICIG | 2 |
| 2011 | Real-Time Long-Range Lane Detection and Tracking for Intelligent VehicleabstractThis paper presents a real-time long-range lane detection and tracking approach to meet the requirements of the high-speed intelligent vehicles running on highway roads. Based on a linear-parabolic two-lane highway road model and a novel strong lane marking feature named Lane Marking Segmentation, the maximal lane detection distance of this approach is up to 120 meters. Then the lane lines are selected and tracked by estimating the ego vehicle lateral offset with a Kalman filter. Experiment results with test dataset extracted from real traffic scenes on highway roads show that the approaches proposed in this paper can achieve a high detection rate with a low time cost. Bin Dai 0001, Jinze Song, Hangen He |
ICIG | 2 |
| 2011 | Fast detection of small infrared objects in maritime scenes using local minimum patternsabstractThis paper describes a novel approach for fast detecting small maritime objects in infrared (IR) images. It is based on the local minimum patterns (LMP), which are theoretically the approximations of some stationary wavelet transforms (SWT). Using LMP to estimate the background with a single image, we obtain an object-aware saliency map by background subtraction. Regions of potential objects are then segmented by an adaptive threshold based on the histogram of the saliency map. We finally propose a fast clustering algorithm for localizing objects from segmented regions. Extensive experiments on challenging data sets show a competitive performance. Baojun Qi, Tao Wu 0001, Bin Dai 0001, Hangen He |
ICIP | 3 |
| 2009 | MORT: A Technique to Improve Routing Efficiency in Fault-Tolerant Multipath RoutingabstractMultipath routing is thought of as a promising direction of the current routing system as it can improve the network performance in terms of reliability and throughput. However, there are some challenging problems to solve towards Internet-wide multipath routing. One of them is the dramatically increasing control message overhead caused by network dynamics. More message overhead will consume more computing resources and more storage. Meanwhile, more message overhead will lead to slower convergence process for routing protocols due to longer processing time. In this paper, we present MORT to solve the above problem. MORT is based on a technique called ¿information hiding¿. The ¿information hiding¿ technique allows routers in network to hide some routing information such as link failures and link cost changes to other routers without introducing any serious bad effect to the routing protocols. Multipath routing protocols embedded with MORT will have fewer routing message overhead and shorter routing convergence time when facing network events such as link failures and link recoveries. In the simulations, we apply MORT to a newly presented multipath protocol to show that MORT can reduce message overhead significantly as well as shortening the routing convergence time. Bin Dai 0001, Huabiao Lu, Zhigang Sun 0002, Ziming Song, Yanpeng Ma, Jinshu Su |
MSN | 1 |
| 2008 | Self-learning path-tracking control of autonomous vehicles using kernel-based approximate dynamic programmingabstractWith the fast development of robotics and intelligent vehicles, there has been much research work on modeling and motion control of autonomous vehicles. However, due to model complexity, and unknown disturbances from dynamic environment, the motion control of autonomous vehicles is still a difficult problem. In this paper, a novel self-learning path-tracking control method is proposed for a car-like robotic vehicle, where kernel-based approximate dynamic programming (ADP) is used to optimize the controller performance with little prior knowledge on vehicle dynamics. The kernel-based ADP method is a recently developed reinforcement learning algorithm called kernel least-squares policy iteration (KLSPI), which uses kernel methods with automatic feature selection in policy evaluation to get better generalization performance and learning efficiency. By using KLSPI, the lateral control performance of the robotic vehicle can be optimized in a self-learning and data-driven style. Compared with previous learning control methods, the proposed method has advantages in learning efficiency and automatic feature selection. Simulation results show that the proposed method can obtain an optimized path-tracking control policy only in a few iterations, which will be very practical for real applications. Xin Xu 0001, Bin Dai 0001, Hangen He |
IJCNN | 3 |