EDBT 2026 Demo / reviewers in the wild / expert
Yeqiang Qian
dblp:203/9959
· DBLP profile ↗
27ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0003-0831-8702ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DrivingEditor: 4D Composite Gaussian Splatting for Reconstruction and Edition of Dynamic Autonomous Driving ScenesabstractIn recent years, with the development of autonomous driving, 3D reconstruction for unbounded large-scale scenes has attracted researchers' attention. Existing methods have achieved outstanding reconstruction accuracy in autonomous driving scenes, but most of them lack the ability to edit scenes. Although some methods have the capability to edit scenarios, they are highly dependent on manually annotated 3D bounding boxes, leading to their poor scalability. To address the issues, we introduce a new Gaussian representation, called DrivingEditor, which decouples the scene into two parts and handles them by separate branches to individually model the dynamic foreground objects and the static background during the training process. By proposing a framework for decoupled modeling of scenarios, we can achieve accurate editing of any dynamic target, such as dynamic objects removal, adding and etc, meanwhile improving the reconstruction quality of autonomous driving scenes especially the dynamic foreground objects, without resorting to 3D bounding boxes. Extensive experiments on Waymo Open Dataset and KITTI benchmarks demonstrate the performance in 3D reconstruction for both dynamic and static scenes. Besides, we conduct extra experiments on unstructured large-scale scenarios, which can more convincingly demonstrate the performance and robustness of our proposed model when rendering the unstructured scenes. Our code is available at https://github.com/WangXu-xxx/DrivingEditor. Yeqiang Qian, Yun-Fu Liu, Lei Tuo, Huiyong Chen, Ming Yang 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | ParkOcc: A Novel Dataset and Benchmark for Surround-View Fisheye 3-D Semantic Occupancy Prediction in Automated Parking Scenarios
Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2026 | MPSS: A Model Pruning Method for Semantic Image Segmentation NetworksabstractThis paper proposes a model pruning named MPSS for semantic image segmentation networks, so that semantic image segmentation models can be deployed into embedded devices. Most existing model pruning methods aim at image classification models. Since semantic segmentation is a fine grained task, the directly use of traditional model pruning methods greatly reduces the model accuracy. The core problems of model pruning in the semantic segmentation task are to determine the appropriate pruning kernels and the pruning structure. We propose a new composite index that defines the similarity between convolution kernels to determine the pruning kernels. Furthermore, we propose a structure mending method based on the neural architecture search to determine the pruning structure. Compared with the method of manually defining the pruning rate, the proposed structure mending method obtains a better pruning structure. We conduct experiments based on two semantic segmentation networks, the FCN and the FASSD-Net. The experimental results show that the proposed model pruning method enables the pruned network to obtain higher accuracy under the same compression rate. In addition, we deploy the compressed models on an embedded platform, and the FASSD Net inference speed is twice as fast as the unpruned model on NVIDIA Xavier NX. Yeqiang Qian, Qihang Su, Ming Yang 0002 |
IEEE Trans. Multim. | 1 |
| 2026 | Structure-Guided Memory-Efficient 3D Gaussians for Large-Scale Reconstructionabstract3D reconstruction is a critical technology with significant implications for applications such as urban planning, autonomous driving, and virtual reality. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated impressive results in small-scale scenes, achieving high-quality reconstructions with real-time rendering capabilities. However, when applied to large-scale scenes, existing 3DGS methods face significant challenges due to the exponential growth of model size, often exceeding the memory capacity of consumer-grade GPUs and making training and rendering infeasible. In this paper, we propose a structure-guided memory-efficient 3DGS framework that uses only half the memory of current large-scale 3DGS methods while maintaining state-of-the-art reconstruction accuracy. Specifically, we introduce a structure-guided density control mechanism that uses a heuristic approach to split Gaussian ellipsoids in challenging regions and optimizes their attributes during densification, significantly reducing memory storage requirements while preserving structural details with fewer ellipsoids. Moreover, we propose a novel structure loss to supervise the learning of scene structural information, enabling the model to better capture and preserve geometric details such as straight lines and edges, further enhancing reconstruction accuracy. We also propose the largest known drone dataset for 3D reconstruction, comprising over 10,000 high-resolution images covering more than 2.5 million square meters. Extensive experiments on multiple benchmark datasets and our proposed dataset demonstrate that our new method is highly memory-efficient with high accuracy. We strongly recommend you to watch our demo at https://lvzinan.github.io/STGS.github.io/. Zinan Lv, Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | FP-TTC: Fast Prediction of Time-to-Collision Using Monocular ImagesabstractTime-to-Collision (TTC) is a measure of the time until an object collides with the observation plane which is a critical input indicator for obstacle avoidance and other downstream modules. Previous works have utilized deep neural networks to estimate TTC with monocular cameras in an end-to-end manner, which obtain the state-of-the-art (SOTA) accuracy performance. However, these models usually have deep layers and numerous parameters, resulting in long inference time and high computational overhead. Moreover, existing methods use two frames which are the current and future moments as input to calculate the TTC resulting in a delay during the calculation process. To solve these issues, we propose a novel fast TTC prediction model: FP-TTC. We first use an attention-based scale encoder to model the scale-matching process between images, which significantly reduces the computational overhead as well as improves the model’s accuracy. Meanwhile, a simple but powerful trick is introduced to the model, where we built a time-series decoder and predict the current TTC from RGB images in the past, avoiding the computational delay caused by the system time step interval, and further improved the TTC prediction speed. Our model achieves a parameter reduction of 89.1%, a 5.5-fold increase in inference speed, a 19.3% improvement in accuracy. We also provided a lightweight version of FP-TTC, which further optimized the inference speed and parameter count by 15%. Our code is available athttps://github.com/LChanglin/FP-TTC. Yeqiang Qian, Songan Zhang, Ming Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | AdaptiveOcc: Adaptive Octree-Based Network for Multi-Camera 3D Semantic Occupancy Prediction in Autonomous DrivingabstractMulti-camera 3D semantic occupancy prediction is a critical task for autonomous driving, playing a vital role in understanding the environment. Current methods mainly rely on uniform voxel representation to encode space, which greatly limits their resolution scalability. It causes most existing methods to struggle with scaling to finer granularities, as the cubic growth nature of uniform voxel leads to a significant increase in the demand for computational and storage resources when scaling. To address this, we propose a multi-level hierarchical model AdaptiveOcc. Using the octree structure, our model can adaptively represent different parts of space with varying voxel granularity. It can selectively extend resolution only for a small subset of voxels, thus mitigating the substantial computational and storage burden brought by scaling. To endow our model with adaptability, we propose a distance-adaptive octree construction rule for generating supervised labels. Considering that the voxel granularity requirements vary for different distance ranges in environmental perception, such a construction rule results in a higher likelihood of coarser granularity for distant regions and finer granularity for nearby regions. This ensures a more efficient and rational allocation of computational resources, further reducing the inference latency. Extensive experiments on nuScenes, SemanticKITTI and Waymo dataset validate that our method can scale to finer granularities with faster speed, and less training memory compared with other state-of-the-art methods. Our code is available athttps://github.com/yty-sky/AdaptiveOcc. Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | TA-TOS: Terrain-Aware Tiny Obstacle Segmentation Based on MRF Road Modeling Using 3-D LiDAR ScansabstractRobust obstacle segmentation remains critical for the safety of intelligent transportation systems (ITS), where LiDAR-based perception systems form the cornerstone of vehicle-environment interaction. Although state-of-the-art (SOTA) LiDAR-based approaches have demonstrated high performance in segmenting common obstacles, the results for tiny obstacle segmentation are still unsatisfactory. However, such tiny obstacles, e.g., curbs, gravel, and potholes, pose significant threats to ground vehicles, undermining ITS operational safety and surface transportation traffic efficiency. It is challenging for SOTA methods to distinguish tiny obstacles due to their inability to precisely model road surfaces, particularly bumpy road surfaces. To address this problem, this paper proposes a road modeling method based on the Markov random field (MRF), possessing stronger road surface modeling capability. A novel negative exponential energy function is introduced to simultaneously ensure the smoothness of the road model and the consistency with the road undulation. After the energy minimization of the MRF, the segmentation of obstacles (including both positive and negative obstacles) is achieved by computing the signed distance to the refined road model. Our proposed terrain-aware tiny obstacle segmentation (TA-TOS) method is compatible with different terrains and different LiDARs, without any prior data or pre-training. We evaluate the performance of TA-TOS on the SemanticKITTI dataset, and two self-built datasets containing tiny obstacles from actual urban mobility systems (road scenarios) and mining haulage systems (off-road scenarios), respectively. Our proposed TA-TOS method achieves much better performance than the SOTA LiDAR-based segmentation approaches, particularly on roads with pronounced undulation. The results show that the improvement is more significant for the segmentation of smaller obstacles. Our source code is publicly available at github.com/ryming2001/TA-TOS. Nan Ming, Yeqiang Qian, Chunyu Feng, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | Local Vectorized High Definition Map Construction for Autonomous Driving: A Comprehensive ReviewabstractWith the advancement of autonomous driving technology, high-definition (HD) maps are crucial for accurate vehicle positioning and safe navigation. Traditional HD map construction relies on offline processing and manual annotation, which are costly and inflexible for dynamic road environments. Local vectorized HD map construction (LV-HDMC) has emerged as a key technology to address these limitations. LV-HDMC uses advanced computer vision techniques to generate map elements from vehicle-mounted sensors in real time, meeting the demands of autonomous driving. This review provides a comprehensive analysis of LV-HDMC research, tracing the evolution of HD map generation and offering an overview of the LV-HDMC task. It explores methods for creating ground truth, including local region acquisition and map element representation, and classifies network structures related to LV-HDMC, emphasizing the role of computer vision in feature extraction and decoding. Evaluation metrics and benchmarks are introduced, comparing the performance of existing methods. The review also discusses future research directions, highlighting how advancements in computer vision could enhance the accuracy and efficiency of LV-HDMC. This review aims to offer valuable insights into LV-HDMC tasks and contribute to the advancement of autonomous driving technology. Yangrong Zhang, Yeqiang Qian, Hongjun Yi, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | AMFD: Distillation via Adaptive Multimodal Fusion for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection has been shown to be effective in improving performance in complex illumination scenarios. However, prevalent double-stream networks in multispectral detection employ two separate feature extraction branches for multi-modal data, leading to nearly double the inference time compared to single-stream networks utilizing only one feature extraction branch. This increased inference time has hindered the widespread employment of multispectral pedestrian detection in embedded devices for autonomous systems. To efficiently compress multispectral object detection networks, we propose a novel distillation method, the Adaptive Modal Fusion Distillation (AMFD) framework. Unlike traditional distillation methods, the AMFD framework fully leverages the original modal features from the teacher network, thereby significantly enhancing the performance of the student network. Specifically, a Modal Extraction Alignment (MEA) module is utilized to derive learning weights for student networks, integrating focal and global attention mechanisms. This methodology enables the student network to acquire optimal fusion strategies independent from that of teacher network without necessitating an additional feature fusion module. Furthermore, we present the SMOD dataset, a well-aligned challenging multispectral dataset for detection. Extensive experiments on the challenging KAIST, LLVIP, SUNRGB-D and SMOD datasets are conducted to validate the effectiveness of AMFD. The results demonstrate that our method outperforms existing state-of-the-art methods in both reducing log-average Miss Rate and improving mean Average Precision. The code is available athttps://github.com/bigD233/AMFD.git. Zizhao Chen, Yeqiang Qian, Xiaoxiao Yang, Ming Yang 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | LESS-Map: Lightweight and Evolving Semantic Map in Parking Lots for Long-term Self-LocalizationabstractPrecise and long-term stable localization is essential in parking lots for tasks like autonomous driving or autonomous valet parking, etc. Existing methods rely on a fixed and memory-inefficient map, which lacks robust data association approaches. And it is not suitable for precise localization or long-term map maintenance. In this paper, we propose a novel mapping, localization, and map update system based on ground semantic features, utilizing low-cost cameras. We present a precise and lightweight parameterization method to establish improved data association and achieve accurate localization at centimeter-level. Furthermore, we propose a novel map update approach by implementing high-quality data association for parameterized semantic features, allowing continuous map update and refinement during re-localization, while maintaining centimeter-level accuracy. We validate the performance of the proposed method in real-world experiments and compare it against state-of-the-art algorithms. The proposed method achieves an average accuracy improvement of 5cm during the registration process. The generated maps consume only a compact size of 450 KB/km and remain adaptable to evolving environments through continuous update. Xinyang Tang, Yeqiang Qian, Jiming Chen 0001, Liang Li 0010 |
ICRA | 3 |
| 2024 | Non-Repetitive: A Promising LiDAR Scanning PatternabstractLiDAR is an essential sensor for intelligent vehicles. Recently, LiDARs used in vehicles produced by different companies have significant differences in their scanning patterns. Some vehicles use mechanical and solid-state (repetitive) LiDARs, while others use prism-based (non-repetitive) LiDARs. The scanning pattern of a LiDAR has a profound impact on its scanning performance. To investigate the influence of LiDAR scanning patterns, we created the "Repetitive-or-not" dataset, which is collected simultaneously by LiDARs with both repetitive and non-repetitive scanning patterns in the CARLA simulation environment. Using this dataset, we conducted a comprehensive statistical analysis of the scanning ability of repetitive and non-repetitive LiDARs. Furthermore, we looked into the effects of these two LiDAR scanning patterns on the performance of various 3D object detection algorithms. Finally, we explored the domain gap in the point cloud data produced by repetitive and non-repetitive LiDARs. Through an in-depth investigation of the "Repetitive-or-not" dataset, we have discovered that non-repetitive LiDAR shows great promise. This conclusion is primarily supported by its superior object scanning capabilities. Angchen Xie, Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IROS | 2 |
| 2023 | TTC4MCP: Monocular Collision Prediction Based on Self-Supervised TTC EstimationabstractVision-based collision prediction for autonomous driving is a challenging task due to the dynamic movement of vehicles and diverse types of obstacles. Most existing methods rely on object detection algorithms, which only predict predefined collision targets, such as vehicles and pedestrians, and cannot anticipate emergencies caused by unknown obstacles. To address this limitation, we propose a novel approach using pixel-wise time-to-collision (TTC) estimation for monocular collision prediction (TTC4MCP). Our approach predicts TTC and optical flow from monocular images and identifies potential collision areas using feature clustering and motion analysis. To overcome the challenge of training TTC estimation models without ground truth data in new scenes, we propose a self-supervised TTC training method, enabling collision prediction in a wider range of scenarios. TTC4MCP is evaluated on multiple road conditions and demonstrates promising results in terms of accuracy and robustness. Yeqiang Qian, Weihao Yan 0001, Ming Yang 0002 |
IROS | 2 |
| 2023 | Threshold-Adaptive Unsupervised Focal Loss for Domain Adaptation of Semantic SegmentationabstractSemantic segmentation is an important task for intelligent vehicles to understand the environment. Current deep learning based methods require large amounts of labeled data for training. Manual annotation is expensive, while simulators can provide accurate annotations. However, the performance of the semantic segmentation model trained with synthetic datasets will significantly degenerate in the actual scenes. Unsupervised domain adaptation (UDA) for semantic segmentation is used to reduce the domain gap and improve the performance on the target domain. Existing adversarial-based and self-training methods usually involve complex training procedures, while entropy-based methods have recently received attention for their simplicity and effectiveness. However, entropy-based UDA methods have problems that they barely optimize hard samples and lack an explicit semantic connection between the source and target domains. In this paper, we propose a novel two-stage entropy-based UDA method for semantic segmentation. In stage one, we design a threshold-adaptative unsupervised focal loss to regularize the prediction in the target domain. It first introduces unsupervised focal loss into UDA for semantic segmentation, helping to optimize hard samples and avoiding generating unreliable pseudo-labels in the target domain. In stage two, we employ cross-domain image mixing (CIM) to bridge the semantic knowledge between two domains and incorporate long-tail class pasting to alleviate the class imbalance problem. Extensive experiments on synthetic-to-real and cross-city benchmarks demonstrate the effectiveness of our method. It achieves state-of-the-art performance using DeepLabV2, as well as competitive performance using the lightweight BiSeNet with great advantages in training and inference time. Weihao Yan 0001, Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | BAANet: Learning Bi-directional Adaptive Attention Gates for Multispectral Pedestrian DetectionabstractThermal infrared (TIR) image has proven effectiveness in providing temperature cues to the RGB features for multispectral pedestrian detection. Most existing methods directly inject the TIR modality into the RGB-based framework or simply ensemble the results of two modalities. This, however, could lead to inferior detection performance, as the RGB and TIR features generally have modality-specific noise, which might worsen the features along with the propagation of the network. Therefore, this work proposes an effective and efficient cross-modality fusion module called Bi-directional Adaptive Attention Gate (BAA-Gate). Based on the attention mechanism, the BAA-Gate is devised to distill the informative features and recalibrate the representations asymptotically. Concretely, a bi-direction multi-stage fusion strategy is adopted to progressively optimize features of two modalities and retain their specificity during the propagation. Moreover, an adaptive interaction of BAA-Gate is introduced by the illumination-based weighting strategy to adaptively adjust the recalibrating and aggregating strength in the BAA-Gate and enhance the robustness towards illumination changes. Considerable experiments on the challenging KAIST dataset demonstrate the superior performance of our method with satisfactory speed. Xiaoxiao Yang, Yeqiang Qian, Hui-Jie Zhu, Ming Yang 0002 |
ICRA | 2 |
| 2022 | Pedestrian Graph +: A Fast Pedestrian Crossing Prediction Model Based on Graph Convolutional NetworksabstractEstimating when pedestrians cross the street is essential for intelligent transportation systems. Accurate, real-time prediction is critical to ensure the safety of the most vulnerable road users while improving passenger comfort. In the present work, we developed a model called Pedestrian Graph +, an improvement of our previous work, Pedestrian Graph, which predicts pedestrian crossing action in urban areas based on a Graph Convolution Network. We integrated two convolutional modules in the new model that provide additional context information (cropped images, cropped segmentation maps, ego-vehicle velocity data) to the main Graph Convolutional module, thus increasing accuracy. Our model is faster and smaller than other state-of-the-art models, achieving equivalent accuracy. Our model is faster than state-of-the-art models, with an inference time of 6 ms (on a GTX 1080) and low memory consumption (0.3 MB). We tested our model on two datasets, Joint Attention in Autonomous Driving (JAAD) and Pedestrian Intention Estimation (PIE), achieving 86% and 89% accuracy, respectively. Another contribution of our work is the ability to dynamically process almost any input size in the time domain without significant loss of accuracy. It is possible due to the fully convolutional property of ConvNets. Our models and results are available athttps://github.com/RodrigoGantier/Pedestrian_graph_plus. Pablo Rodrigo Gantier Cadena, Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Gated-Residual Block for Semantic Segmentation Using RGB-D DataabstractSemantic segmentation is an important technique for scene understanding in the intelligent transportation system. RGB-D data shows great advantages over the unimodal data in this area, and it can be easily obtained from consumer sensors nowadays. How to design effective fusion structures to fuse RGB and depth signals in RGB-D data is a challenging problem. This paper proposes a novel gated-residual block to address this problem. The structure consists of two residual units and one gated fusion unit, the residual unit progressively aggregates modality-specific features from the modality-specific signals and the gate mechanism computes complementary features for them. Based on the gated-residual block, the paper presents the deep multimodal networks, named GRBNet, for RGB-D semantic segmentation. Experiments on ScanNet, Cityscapes and SUN RGB-D datasets verify the effectiveness of the proposed approach and demonstrate that the GRBNet achieved competitive performance. Yeqiang Qian, Liuyuan Deng, Tianyi Li 0003, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Survey on Fish-Eye Cameras and Their Applications in Intelligent VehiclesabstractFish-eye cameras have become essential sensors in intelligent vehicles. Due to its unique projection principle, a fish-eye camera can provide a large field of view. Benefiting from this special feature, fish-eye cameras have rich applications in intelligent vehicles. However, dataset and distortion problems are still challenges when applying fish-eye cameras in reality. This work introduces the projection principle of fish-eye cameras, and four classic fish-eye image representation models are presented. Then, the typical fish-eye datasets are presented, including real collected data and virtually generated data. Through the organization and summarization of the relevant studies, we demonstrate various applications of fish-eye cameras in intelligent vehicles, e.g., object detection and tracking, image segmentation, mapping and localization, and around-view monitoring. These works design various strategies to exploit the advantages of fish-eye cameras and prevent image distortion problems, showing the broad application prospects of such cameras. Finally, we discuss the development tendencies of intelligent vehicle applications involving fish-eye cameras. Yeqiang Qian, Ming Yang 0002, John M. Dolan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | SPADE-E2VID: Spatially-Adaptive Denormalization for Event-Based Video ReconstructionabstractEvent-based cameras have several advantages over traditional cameras that shoot videos in frames. Event cameras have a high temporal resolution, high dynamic range, and almost non-existence of blurriness. The data that is produced by event sensors forms a chain of events when a change in brightness is reported in each pixel. This feature makes it difficult to directly apply existing algorithms and take advantage of the event camera data. Due to the developments in neural networks, important advances were made in event-based image reconstruction. Even though these neural networks achieve precise reconstructions while preserving most of the properties of the event cameras, there is still an initialization time that needs to have the highest possible quality in the reconstructed frames. In this work, we present the SPADE-E2VID neural network model that improves the quality of early frames in an event-based reconstructed video, as well as the overall contrast. The SPADE-E2VID model improves the quality of the first reconstructed frames by 15.87% for MSE error, 4.15% for SSIM, and 2.5% in LPIPS. In addition, the SPADE layer in our model allows training our model to reconstruct videos without a temporal loss function. Another advantage of our model is that it has a faster training time. In a many-to-one training style, we avoid running the loss function at each step, executing the loss function at the end of each loop only once. In the present work, we also carried out experiments with event cameras that do not have polarity data. Our model produces quality video reconstructions with non-polarity events in HD resolution (1200 × 800). The Video, the code, and the datasets will be available at: https://github.com/RodrigoGantier/SPADE_E2VID. Pablo Rodrigo Gantier Cadena, Yeqiang Qian, Ming Yang 0002 |
IEEE Trans. Image Process. | 2 |
| 2021 | Neutral Cross-Entropy Loss Based Unsupervised Domain Adaptation for Semantic SegmentationabstractThe generalization performance for semantic segmentation remains a major challenge when the data distributions between the source and target domain mismatch. Unsupervised domain adaptation (UDA) approaches are proposed to mitigate the problem above, among which entropy-minimization-based methods have gained more and more attention. However, the methods merely follow the cluster assumption sharpening the prediction distribution, thus have limited performance improvement. Without additional priors, the entropy loss can easily over-sharpen the prediction distribution, which brings noisy information into the learning process. On the other hand, the gradient of the entropy loss is strongly biased toward easy samples, also leading to limited generalization advances. In this paper, we firstly propose a pixel-level consistency regularization method, which introduces the smoothness prior to the UDA problem. Furthermore, we propose the neutral cross-entropy loss based on the consistency regularization, and reveal that its internal neutralization mechanism mitigates the over-sharpness of entropy minimization via the flatness effect of consistency regularization. We also demonstrate that the gradient bias toward easy samples is inherently tackled via the neutral cross-entropy loss. The experiments show that the proposed method has outperformed state-of-the-art methods in two synthetic-to-real experiments, only using the lightweight network. Ming Yang 0002, Liuyuan Deng, Yeqiang Qian |
IEEE Trans. Image Process. | 4 |
| 2021 | Adversarial Training-Based Hard Example Mining for Pedestrian Detection in Fish-Eye ImagesabstractSince fish-eye cameras are popular in the intelligent transportation systems, accurate pedestrian detection in fish-eye images becomes more and more critical in low-speed scenarios. Usually, big data based training is the key for detectors to handle the distortion problem in fish-eye images. Especially, hard examples are more important for detectors. They have more complex features and are hard to recognize. In conventional methods, fish-eye images are collected and labeled manually. These methods are expensive and labor-intense. More importantly, these methods are still hard to collect abundant hard examples since they are rare in reality. This work proposes the Distortion Generation Network, which generates generous fish-eye images automatically using only small samples. Moreover, the Adversarial Distortion Generation Network is proposed to mine hard examples via adversarial training. These hard examples benefit detectors to be more robust to seriously distorted objects in fish-eye images. Experiments with the ETH, the KITTI, the GM-ATCI and real fish-eye datasets demonstrate that the proposed methods achieve higher accuracy than conventional methods in pedestrian detection in fish-eye images. Yeqiang Qian, Ming Yang 0002, Hao Li 0024, Bing Wang 0006 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | G2P: a new descriptor for pedestrian detection
Ming Yang 0002, Yeqiang Qian, Linji Xue, Hao Li 0024, Liuyuan Deng |
Neural Comput. Appl. | 2 |
| 2020 | Monocular pedestrian orientation estimation based on deep 2D-3D feedforward
Chenchen Zhao 0001, Yeqiang Qian, Ming Yang 0002 |
Pattern Recognit. | 2 |
| 2020 | DLT-Net: Joint Detection of Drivable Areas, Lane Lines, and Traffic ObjectsabstractPerception is an essential task for self-driving cars, but most perception tasks are usually handled independently. We propose a unified neural network named DLT-Net to detect drivable areas, lane lines, and traffic objects simultaneously. These three tasks are most important for autonomous driving, especially when a high-definition map and accurate localization are unavailable. Instead of separating tasks in the decoder, we construct context tensors between sub-task decoders to share designate influence among tasks. Therefore, each task can benefit from others during multi-task learning. Experiments show that our model outperforms the conventional multi-task network in terms of the task-wise accuracy and the overall computational efficiency, in the challenging BDD dataset. Yeqiang Qian, John M. Dolan, Ming Yang 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Oriented Spatial Transformer Network for Pedestrian Detection Using Fish-Eye CameraabstractPedestrian detection using fish-eye cameras is a principal research focus in computer vision. Lack of pedestrian datasets of fish-eye images and pedestrian distortion in fish-eye images are two primary challenges. In this paper, two approaches are proposed to deal with these two challenges, respectively. On the one hand, the projective model transformation (PMT) algorithm is proposed, which can transform normal images into fish-eye images. The PMT can be applied to most of the pedestrian datasets and generates corresponding fish-eye image datasets. In this way, enough training data can be provided through the PMT. On the other hand, the oriented spatial transformer network (OSTN) is designed to rectify warped pedestrian features using CNNs, so that pedestrians in fish-eye images are easier for detectors to recognize. The OSTN can be embedded into universal deep learning based detectors easily. Moreover, the new pedestrian detector, where the OSTN is embedded, can be trained end to end. Finally, the OSTN based fish-eye pedestrian detectors can be trained using fish-eye images, which are generated using the PMT. Experiments on ETH, KITTI, Citypersons, and real pedestrian datasets show the effectiveness of the PMT and accuracy improvement of pedestrian detection in fish-eye images using the OSTN. Yeqiang Qian, Ming Yang 0002, Xu Zhao 0001, Bing Wang 0006 |
IEEE Trans. Multim. | 1 |
| 2018 | Pedestrian Feature Generation in Fish-Eye Images via AdversaryabstractPedestrian detection in fish-eye images is always an important problem in advanced driver assistance systems (ADAS). In conventional methods, pedestrian detectors will be trained using fish-eye images. But it is hard to collect and label enough fish-eye images manually. Therefore, a new strategy for training fish-eye pedestrian detectors using images from normal pedestrian datasets is proposed in this work. Concretely, Fish-eye Spatial Transformer Network (FSTN) is designed to generate pedestrian features in fish-eye images. FSTN aims to simulate distorted pedestrian features on the feature maps. Then the entire network is trained via adversary. FSTN is trained to generate examples which are difficult for pedestrian detectors to classify. So that the detectors are more robust to the deformation. FSTN can be embedded into state-of-the-art detectors easily. And the entire pedestrian detector, where the FSTN embedded, can be trained end to end via adversary. Moreover, experiments on ETH and KITTI pedestrian datasets show the slight accuracy improvement of pedestrian detection in fish-eye images using adversarial network compared with conventional methods. Yeqiang Qian, Ming Yang 0002, Bing Wang 0006 |
ICRA | 1 |
| 2017 | CNN based semantic segmentation for urban traffic scenes using fisheye cameraabstractSemantic segmentation is an important step of visual scene understanding for autonomous driving. Recently, Convolutional Neural Network (CNN) based methods have successfully applied in semantic segmentation using narrow-angle or even wide-angle pinhole camera. However, in urban traffic environments, autonomous vehicles need wider field of view to perceive surrounding things and stuff, especially at intersections. This paper describes a CNN-based semantic segmentation solution using fisheye camera which covers a large field of view. To handle the complex scene in the fisheye image, Overlapping Pyramid Pooling (OPP) module is proposed to explore local, global and pyramid local region context information. Based on the OPP module, a network structure called OPP-net is proposed for semantic segmentation. The net is trained and evaluated on a fisheye image dataset for semantic segmentation which is generated from an existing dataset of urban traffic scenes. In addition, zoom augmentation, a novel data augmentation policy specially designed for fisheye image, is proposed to improve the net's generalization performance. Experiments demonstrate the outstanding performance of the OPP-net for urban traffic scenes and the effectiveness of the zoom augmentation. Liuyuan Deng, Ming Yang 0002, Yeqiang Qian, Bing Wang 0006 |
Intelligent Vehicles Symposium | 3 |
| 2017 | Self-adapting part-based pedestrian detection using a fish-eye cameraabstractNowadays, fish-eye cameras play an increasingly important role in intelligent vehicles because of its wide field of view. Using fish-eye camera, pedestrians around the vehicles could be monitored expediently, but the problem of pedestrian distortion has always existed. This paper creates a new warping pedestrian benchmark using imaging principle of the fish-eye camera based on ETH pedestrian benchmark. With this practical benchmark, warping pedestrians are trained differently according to the position in fish-eye images. A self-adapting part-based algorithm is proposed to detect pedestrian with different degrees of deformation. Moreover, GPU is used to accelerate the whole algorithm to guarantee the real-time performance. Experiments show that the algorithm has competitive accuracy. Yeqiang Qian, Ming Yang 0002, Bing Wang 0006 |
Intelligent Vehicles Symposium | 1 |