VLDB 2026 Research / reviewers in the wild / expert
Chi Zhang 0020
dblp:91/195-20
· DBLP profile ↗
32ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-9604-2800ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Perception-Failure-Induced Test Scenario Searching via Online Causal Reinforcement LearningabstractDespite compliance with safety standards such as ISO 26262, the perception systems of autonomous vehicles still face numerous challenges during real-world operation. In this work, we propose a framework to identify safety-critical scenarios specifically induced by perception failures, aiming to support more targeted and effective scenario-based testing. Importantly, the key challenge lies in disentangling whether perception failures directly lead to critical outcomes. To address this, we propose an Online Causal Reinforcement Searching (OCRS) framework that simultaneously performs scenario search and causal reasoning. Focusing on visual unawareness and perception degradation, OCRS employs an LSTM-RNN controller to identify accident-prone scenarios linked to perception failures, which are then verified in a counterfactual world to determine causal responsibility. To improve efficiency, an online scenario classifier with passive-aggressive updates is introduced to dynamically filter out non-critical cases. The experimental results demonstrate both the effectiveness and efficiency of the proposed approach, achieving a 33% reduction in execution time for searchingperception-failure-induced corner cases. Furthermore, we evaluate the method under adverse weather conditions such as snow and fog, confirming that OCRS remains effective in identifying various failure modes. Chi Zhang 0020, Tingting Long, Linhai Xu, Mingwen Bi, Xingyu Chen 0001, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Style Nursing with Spatial and Semantic Guidance for Zero-Shot Traffic Scene Style TransferabstractRecent advances in text-to-image diffusion models have shown an outstanding ability in zero-shot style transfer. However, existing methods often struggle to balance preserving the semantic content of the input image and faithfully transferring the target style in line with the edit prompt. Especially when applied to complex traffic scenes with diverse objects, layouts, and stylistic variations, current diffusion models tend to exhibit Style Neglection, i.e., failing to generate the required style in the prompt. To address this issue, we propose Style Nursing, which directs the model to focus on style subject tokens in the text prompt and excites their corresponding visual activations. Moreover, we introduce Spatial and Semantic Guidance to guide the preservation of content after editing, which utilizes spatial features from the DDIM sampling process together with attention maps from the semantic reconstruction. To evaluate the performance of zero-shot style transfer methods in traffic scenes, we present STREET-6K, a new benchmark dataset comprising 6000 images showcasing diverse traffic scenes and style transfer variations, accompanied by comprehensive annotations and evaluation metrics. Our approach beats state-of-the-art image translation methods in comprehensive quantitative metrics and human evaluations on traffic scene image synthesis while seamlessly generalizing to various other types of images without training or fine-tuning. Further experiments on detection and segmentation tasks show that fine-tuning perception models on our synthesized images improves Recall and mean Intersection over Union (mIoU) by over 10% and 3% respectively in rarely-seen traffic scenes. Zihang Lin, Yuehu Liu, Chi Zhang 0020 |
AAAI | 5 |
| 2025 | Mitigating Shortcut Learning in Online Action Detection and Anticipation via Cross-Modal Semantic AlignmentabstractOnline Action Detection (OAD) and Online Action Anticipation (OAA) are conventionally framed as multi-class classification tasks reliant on action representation learning. However, existing methods easily overfit to appearance features corre-lated with specific categories, neglecting semantically meaningful action features, i.e., shortcut learning. This results in a misinter-pretation of visually similar actions with distinct semantics. We argue that shortcut learning stems from the one-hot label supervision, which simplifies task objectives from semantic recognition to category differentiation. Inspired by advances in text-supervised visual representation learning (e.g., CLIP), we propose a CLIP for Online Action (CLIP40A), a unified model that formulates OAD and OAA as Video-Text Retrieval tasks. This approach mitigates shortcut learning by supervising the alignment between actions and their corresponding labels. Specifically, CLIP40A extracts action contexts from distant videos across multiple time scales, and then learns current and future action representations based on these contexts and recent videos. For labels, CLIP40A introduces a learnable prompt mechanism to compensate for the lack of label contexts. Subsequently, CLIP40A leverages the pretrained CLIP text encoder to convert labels into representations, preserving label semantics and semantic associations with visual data. By similarity comparisons, the labels most similar to current and future actions are the results of OAD and OAA. CLIP40A achieves superior performance on THUMOS'14 and TVSeries. Yuehu Liu, Chi Zhang 0020 |
CBMI | 3 |
| 2025 | Semantic Graph Embedded Energy Minimization Learning for Scene Graph GenerationabstractThe performance of current scene graph generation models is affected by training with cross-entropy loss, exacerbating the problem of prediction bias stemming from biased training data. Energy-based model adopts a learning method for joint image and scene graph to alleviate this challenge. However, this method only focuses on the visual features of images, neglecting the rich relation information contained in the semantic space. To address this issue, we innovatively employ powerful pre-trained large models to realize a simple yet effective semantic graph embedded energy minimization framework for the SGG task. Specifically, we use large models to generate image descriptions and extract relation triplets, which are then transformed into semantic graphs with entities as nodes and relations as edges. Moreover, by mapping these graphs into the same space using GNN to learn the minimal energy value, our approach enables SGG model to learn structural information in both semantic and visual spaces. We validate the effectiveness and efficiency of our method on the SGG benchmark Visual Genome dataset. Compared with prevailing models and EBM, we achieve a significant performance improvement of up to 2.99% and 2.24%, respectively. Jinghang Chen, Chi Zhang 0020, Yuehu Liu, Le Wang 0003 |
ICASSP | 2 |
| 2025 | Worst Perception Scenario Prediction for Testing Autonomous Driving PerceptionabstractRecent studies have suggested that potential short-comings of certain perception modules can be discovered by analyzing the performance of worst scenarios. However, finding the worst perception scenario (WPS) requires datasets with rich semantic annotations of the scenes in visual perception tasks and it is time-consuming to label all the scenario data. To address this, we proposed a method of prediction for WPS, which utilized prior information to predict the model performance under the absence of annotations. Specifically, this paper introduced a scenario matcher based on hybrid re-ranking, which combined the labeled and unlabeled data to generate the pseudo-sample set. In addition, we designed a sample reorganization module to update this sample set through the nearest neighbor retrieval. We also discussed the distribution relationship between labeled and unlabeled data, categorizing it into three cases, and validated the effectiveness of the proposed method on the KITTI and ApolloScape datasets. Liheng Xu, Chi Zhang 0020, Yuehu Liu, Li Li 0013 |
IV | 3 |
| 2025 | 3D Shape Adaptation Across Datasets for Weakly Supervised Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is a key yet challenging task that usually involves extensive and expensive manual annotation of 3D boxes. To eliminate the dependence on 3D box labels, weakly supervised M3D (WM3D) has been explored using only 2D annotations, which necessitate the use of extra resources, like LiDAR data, stereo images, and video sequences. However, the strict correspondence and complex calibration between the target image and additional resources limit their applicability. In this work, we propose a simple yet effective framework, 3D Shape Adaptation across datasets for Weakly supervised Monocular 3D Detection (SAWM3D). We observed that directly applying a source-dataset detector to the target dataset results in a significant domain gap, with the primary contribution coming from the 3D location, while orientation and dimensions have a smaller impact. This enables us to view WM3D as 3D shape adaptation optimization on the target dataset. Directly scaling the predicted shape results in a significant reduction of the adaptation gap; fine-tuning on the target dataset using only 2D supervision also yields impressive results. Experiments on the KITTI benchmark demonstrate the effectiveness of our strategies. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 4 |
| 2025 | Rethinking SSIM-Based Optimization in Neural Field TrainingabstractThe Structural Similarity (SSIM) index is a widely used metric for evaluating image quality, with broad applications in areas such as image restoration, 3D reconstruction, and novel view synthesis. A number of previous works have introduced SSIM-based optimization into neural field training to enhance the model's performance. Despite its widespread use, there has been limited research on how to effectively incorporate SSIM loss into the training process. In this work, we explore this gap and provide insights into the role of SSIM loss in neural field training. Our key finding is that SSIM loss is particularly beneficial during the early phase of training, before the model fully learns the luminance information. We show that SSIM loss acts as an effective “guidance” mechanism in the initial training phase, and removing it after the model has learned the luminance does not harm the final performance-in fact, it may improve it. Our experiments demonstrate the effectiveness of our strategy, offering new insights into how SSIM loss can be more efficiently used in neural field training. We believe these findings will not only enhance SSIM's application in neural field training but also inspire further research into more adaptive loss functions for deep learning models. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 4 |
| 2025 | 3D Shape Transfer Learning for Enhanced Monocular 3D Object DetectionabstractMonocular 3D object detection (M3D) is challenging due to the lack of depth information in the RGB image. Existing works resort to various additional resources to enhance detection performance, including LiDAR data, depth information, CAD models, stereo images, and others, where strict correspondence or synchronization may limit their applicability and scalability. In this work, we propose a simple yet effective framework, 3D Shape Transfer Learning for Enhanced Monocular 3D Object Detection (STLM3D). It views M3D as 3D shape reconstruction and leverages transfer learning across datasets to enhance shape reconstruction capability, thereby enhancing M3D performance. In addition, we design a plug-and-play 3D detection branch for 3D attributes prediction. Experimental results on the KITTI benchmark demonstrate that the proposed method achieves state-of-the-art performance compared to existing approaches. Yuanqi Su, Haoang Lu, Chi Zhang 0020, Yuehu Liu |
IV | 4 |
| 2025 | BiOMamba: Mamba-based Forward-Then-Backward Temporal Modeling for Online Action Detection and AnticipationabstractGiven that action evolution follows temporal progression, recent studies for Online Action Detection (OAD) and Online Action Anticipation (OAA) generally adopt forward temporal modeling to capture dependencies in observable video sequences. However, the strictly sequential nature of forward temporal modeling prevents subsequent frames from being used to enhance the earlier modeling process. In particular, the current frame, the last observable frame in the online video stream, serves as the direct visual cue for ongoing action recognition and the informative context for future action anticipation. As modeling errors accumulate over time, the resulting representations may progressively deviate from the actual semantics. Findings in cognitive neuroscience show that the hippocampus performs backward replay after observation to reinforce and correct the interpretation of previous observations. Inspired by this, we propose to incorporate backward temporal modeling following forward temporal modeling, enabling the model to leverage backward temporal modeling to enhance forward temporal modeling. Based on this idea, we propose a unified model for OAD and OAA, named Bidirectional Online Mamba (BiOMamba). Specifically, to address the excessive length and relevance imbalance in observable sequences, BiOMamba compresses distant long-term memory and preserves recent short-term memory. Then, BiOMamba sequentially model both forward and backward temporal dependencies in the whole memory. Finally, according to the temporal modeling result, BiOMamba generates representations for current and future actions. BiOMamba achieves state-of-the-art performance on THUMOS'14 (OAD: 73.3% mAP, OAA: 59.7% mAP) and TVSeries (OAD: 89.9% mcAP, OAA: 83.7% mcAP). Yuehu Liu, Chi Zhang 0020 |
ACM Multimedia | 3 |
| 2025 | Long and Short-Term Collaborative Decision-Making Transformer for Online Action Detection and Anticipation
Chi Zhang 0020, Le Wang 0003, Yuehu Liu |
Pattern Recognit. | 2 |
| 2025 | Flight Mastery in Turbulent Skies: Shared Control and Curriculum Reinforcement Learning for Crosswind LandingabstractLanding in crosswind conditions poses significant challenges for aircraft, as traditional control methods often fail to ensure stability in rapidly changing wind environments. While reinforcement learning offers a promising alternative, it typically suffers from low sample efficiency and limited generalization under stochastic wind fields. To address these challenges, we propose a Shared Control and Curriculum Reinforcement Learning framework. We model the crosswind landing task as a Markov Decision Process (MDP), explicitly defining the state space, action space, and wind field representation. To initialize learning, we decompose the multi-objective landing task into four sub-tasks—altitude, attitude, heading, and speed control—and train expert policies for each. These are then distilled into a shared control model via behavior cloning, providing a pre-trained policy with basic flight control capabilities. We further fine-tune this model using curriculum reinforcement learning, progressively increasing the complexity of wind conditions to enhance robustness and generalization. Experimental results across multiple aircraft and wind scenarios show that our method improves landing success rates and trajectory smoothness, while generalizing more effectively to unseen wind conditions, outperforming PID controllers, imitation learning, and mainstream RL baselines. Zechen Shi, Xingyu Chen 0001, Zeyang Liu 0001, Chi Zhang 0020, Yimeng Yu, Junbin You, Xuguang Lan |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | DDGPnP: Differential degree graph based PnP solution to handle outliers
Zhichao Cui, Zeqi Chen, Chi Zhang 0020, Gaofeng Meng, Yuehu Liu, Xiangmo Zhao |
Comput. Vis. Image Underst. | 3 |
| 2024 | Asynchronous Threshold ECDSA With Batch ProcessingabstractThreshold Elliptic Curve Digital Signature Algorithm (ECDSA) has attracted a lot of attention due to the wide applications of ECDSA in crypto asset. Although several variants of threshold signature protocols can provide functions, such as key generation and signing, they suffer from two shortfalls. First, these schemes only discuss a single signature computation task in a synchronous algorithm context, which is difficult to adapt to real crypto-asset applications, such as custody. Second, these schemes are computing intensive and not scalable, hence can hardly support large-scale processing operations in real life even after traditional optimization, such as multithreading, is applied. In this article, we propose an innovative computation method called asynchronous threshold ECDSA with batch processing, based on the interactive threshold signature protocols. The method provides a reliable solution for critical operational scenarios, such as threshold signing and distributed key generation (DKG) in crypto-asset custody, and can be a future reference in secure data distribution mechanisms. The performance and scalability of our methods are validated through a benchmark testing. Hongxin Zhang 0001, Guanghuan Xie, Chi Zhang 0020, Zhuo Li 0014, Rui Qin 0002, Gang Xiong 0001, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Worst Perception Scenario Search via Recurrent Neural Controller and K-Reciprocal Re-RankingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. A recent theoretical study shows that the generalization on the worst-group of test samples is far more difficult than others. Therefore, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address this, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as the discrete search on the Visual Operation Design Domain (ODD), namely scenario representation, by optimizing LSTM-RNN controller with the worst-performance reward. Moreover, a time-efficient K-reciprocal re-ranking technique is utilized to match the predicted scenario parameters with existing test data. The proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Furthermore, searching performances w.r.t different Visual ODDs are investigated and it is found that visual representations through generative adversarial network contribute to a better performance. Chi Zhang 0020, Xiaoning Ma, Liheng Xu, Haoang Lu, Le Wang 0003, Yuanqi Su, Yuehu Liu, Li Li 0013 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Reliable Boundary Samples-Based Proxy Pairs for Unsupervised Person Re-identification
Chang Zou, Zeqi Chen, Yuehu Liu, Chi Zhang 0020 |
PRCV (12) | 4 |
| 2023 | Dual Clustering Co-Teaching With Consistent Sample Mining for Unsupervised Person Re-IdentificationabstractIn unsupervised person Re-ID, peer-teaching strategy leveraging two networks to facilitate training has been proven to be an effective method to deal with the pseudo label noise. However, training two networks with a set of noisy pseudo labels reduces the complementarity of the two networks and results in label noise accumulation. To handle this issue, this paper proposes a novel Dual Clustering Co-teaching (DCCT) approach. DCCT mainly exploits the features extracted by two networks to generate two sets of pseudo labels separately by clustering with different parameters. Each network is trained with the pseudo labels generated by its peer network, which can increase the complementarity of the two networks to reduce the impact of noises. Furthermore, we propose dual clustering with dynamic parameters (DCDP) to make the network adaptive and robust to dynamically changing clustering parameters. Moreover, Consistent Sample Mining (CSM) is proposed to find the samples with unchanged pseudo labels during training for potential noisy sample removal. Extensive experiments demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art unsupervised person Re-ID methods by a considerable margin and surpasses most methods utilizing camera information. Zeqi Chen, Zhichao Cui, Chi Zhang 0020, Jiahuan Zhou, Yuehu Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Orthogonal multi-view tensor-based learning for clustering
Shuangxun Ma, Yuehu Liu, Guangcan Liu, Qinghai Zheng, Chi Zhang 0020 |
Neurocomputing | 5 |
| 2022 | An Improved Azimuth Signal Reconstruction Algorithm for Wide-Beam Distributed SARabstractDistributed multichannel synthetic aperture radar (MC-SAR) is a system in which transmitting or receiving arrays are distributed on multiple platforms or at different locations on one platform. The along-track component of the baseline makes distributed SAR promising in high-resolution wide-swath (HRWS) imaging such as azimuth MC-SAR. However, the additional channel mismatch introduced by the cross-track baseline (CTB) is considered for the distributed SAR. When the azimuth beam is wide, the azimuth-variant channel mismatch caused by the CTB must be compensated before SAR imaging. First, an improved azimuth signal reconstruction algorithm for distributed wide-beam SAR is proposed in this paper. The azimuth variance of the channel mismatch is considered in a reconstruction filter to further suppress the ambiguity, and the computational consumption is decreased by approximately decomposing the mismatch matrix. Second, the ambiguity suppression performance of the proposed method is analyzed quantitatively. Finally, a simulation and real data processing are provided to demonstrate the effectiveness of the proposed method. Chi Zhang 0020, Zegang Ding, Han Li 0006, Tianyi Zhang 0006 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Switching: understanding the class-reversed sampling in tail sample memorization
Chi Zhang 0020, Benyi Hu, Yuhang Liuzhang, Le Wang 0003, Yuehu Liu |
Mach. Learn. | 1 |
| 2022 | Density-Aware Haze Image Synthesis by Self-Supervised Content-Style DisentanglementabstractThe key procedure of haze image synthesis with adversarial training lies in the disentanglement of the feature involved only in haze synthesis, i.e.,the style feature, from the feature representing the invariant semantic content, i.e.,the content feature. Previous methods introduced a binary classifier to constrain the domain membership from being distinguished through the learned content feature during the training stage, thereby the style information is separated from the content feature. However, we find that these methods cannot achieve complete content-style disentanglement. The entanglement of the flawed style feature with content information inevitably leads to the inferior rendering of haze images. To address this issue, we propose a self-supervised style regression model with stochastic linear interpolation that can suppress the content information in the style feature. Ablative experiments demonstrate the disentangling completeness and its superiority in density-aware haze image synthesis. Moreover, the synthesized haze data are applied to test the generalization ability of vehicle detectors. Further study on the relation between haze density and detection performance shows that haze has an obvious impact on the generalization ability of vehicle detectors and that the degree of performance degradation is linearly correlated to the haze density, which in turn validates the effectiveness of the proposed method. Chi Zhang 0020, Zihang Lin, Liheng Xu, Zongliang Li, Wei Tang 0016, Yuehu Liu, Gaofeng Meng, Le Wang 0003, Li Li 0013 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Point Cloud Segmentation via Edge-fused Local Graph LearningabstractTraditional convolution for capturing local structures and relationships remains a key technical limit in 3D semantic segmentation, which neglects the certain influence of the adjacent points on the central point in the disordered local point clouds. In this paper, we propose a novel joint-edge graph convolution neural network (JEGCN), which can extract the dynamic features of each local area and transfer the edge information between the vertex pairs to the adjacent vertices. In the proposed graph convolution module, the adjacent vertices are selected with high classification confidence which can guide the central vertex, and then reweight these vertices. Considering the lack of texture features in 3D point clouds, we incorporate 2D image features to adjacent feature propagation to effectively extract the local and global features of point clouds. The experimental results based on ScanNet and S3DIS datasets demonstrate the effectiveness of the proposed method. Mengtao Han, Yaochen Li, Liangyu Zuo, Chi Zhang 0020, Yuanqi Su |
ICRA | 5 |
| 2020 | Multi-label X-Ray Imagery Classification via Bottom-Up Attention and Meta Fusion
Benyi Hu, Chi Zhang 0020, Le Wang 0003, Qilin Zhang 0004, Yuehu Liu |
ACCV (6) | 2 |
| 2020 | Calibrank: Effective Lidar-Camera Extrinsic Calibration By Multi-Modal Learning To RankabstractPrecise and online LiDAR-camera extrinsic calibration is one of the prerequisites of multi-modal data fusion for autonomous perception. The existing 6-DoF pose regression networks take majority effort on coarse-to-fine training strategy to gradually approach the global minimum. However, with limited computing resources, the optimal pose parameters seem unreachable. Moreover, recent research on neural network interpretability reveals that learning-based pose regression is nothing but the interpolation with most relevant samples. Motivated by this notion, we propose to solve the calibration problem in a retrieval way. Concretely, the learning-to-rank pipeline is introduced for ranking the top n relevant poses in the gallery set, which is then fused in to the final prediction. To better explore the pose relevance between ground truth samples, we further propose an exponential mapping from parametric space to the relevance space. The superiority of the proposed method is validated and demonstrated in the comparative and ablative experimental analysis. Xiannong Wu, Chi Zhang 0020, Yuehu Liu |
ICIP | 2 |
| 2020 | Worst Perception Scenario Search for Autonomous DrivingabstractAchieving excellent generalization on perceiving real traffic scenarios with diversity is the long-term goal for building robust autonomous driving systems. In this paper, we propose to discover potential shortness of certain perception module by analyzing its worst-scenario performance. However, with the benchmark datasets growing huge and tremendous, exhaustive searching for the worst perception scenario (WPS) seems to be time consuming and unnecessary. To address, we present an automatic searching scheme empowered by reinforcement learning. In this case, worst scenario mining is formulated as a discrete search problem. A single layer recurrent neural network with LSTM neurons is employed to predict WPS according to the searching reward, which is optimized by a vanilla policy gradient method. Moreover, to deal with the imbalanced distribution of real traffic scenarios, a KNN-like retrieval is utilized for searching the closest scenario samples. Effective yet efficient, the proposed method has been validated by finding the most challenging scenarios for various vehicle detectors on KITTI, BDD100k and our own benchmark set EVB. Further experiments reveal that detection networks with structural similarity share the similar WPS. Liheng Xu, Chi Zhang 0020, Yuehu Liu, Le Wang 0003, Li Li 0013 |
IV | 2 |
| 2020 | Coarse-to-fine 3D road model registration for traffic video augmentationabstractThis study addresses the problem of non‐perspective pose estimation from line correspondences in the traffic scenarios. A coarse‐to‐fine 3D road registration method is proposed for this problem in two stages. Firstly, the iterative closest point algorithm is exploited to estimate the pose coarsely. An objective function is then established to incorporate the feature correspondences for refining the coarse pose. Besides, the framework including road registration is employed for traffic video augmentation. The framework begins with the inputs of traffic videos, road information from Geographic Information Systems and 3D models of traffic elements (e.g. vehicles, pedestrians). Subsequently, 3D road model generation and point‐to‐line correspondence establishment are achieved in the preprossessing stage. After road and viewpoint registration, the 3D graphic engine is employed to simulate the traffic scene with the road, viewpoints and traffic elements. The augmented videos are generated by fusing the original frames and newly projected traffic elements. The authors demonstrate the superiority of the proposed registration method by the comparison to state‐of‐the‐arts in both quantitative and qualitative experiments. In addition, the frames of the augmented videos validate the proposed method in the application. Zhichao Cui, Yaochen Li, Chi Zhang 0020, Yuehu Liu, Fuji Ren |
IET Image Process. | 3 |
| 2019 | VIASEG: Visual Information Assisted Lightweight Point Cloud SegmentationabstractRapid and precise point cloud segmentation is one of the prerequisites for real-time and robust autonomous perception and environmental understanding, which requires a balance between speed and accuracy in architecture design. However, recent lightweight architectures, though fast enough, rely on domain adaptation from time-consuming-constructed synthetic dataset and sophisticated post-processing procedure to improve their performance, neglecting the rich visual information acquired by cameras aside from LiDAR sensors. In this paper, such color information is embedded at data-level to boost the performance of real-time point cloud segmentation. Furthermore, a multiscale lightweight fully convolutional network, VIASeg, is proposed based on the newly designed Super Squeeze Residual module and Semantic Connection from higher convolutional layers to lower layers, which improves the performance by feature denoising with high level semantic information. The superiority of the proposed method is validated and demonstrated in the comparative and ablative experimental analysis, while maintaining the real-time characteristic. Zhibin Zhong, Chi Zhang 0020, Yuehu Liu, Ying Wu 0001 |
ICIP | 2 |
| 2019 | SAR Ship Detection Based on Resnet and Transfer LearningabstractSynthetic Aperture Radar (SAR) ship detection has been a research hotspot and is significant for marine surveillance. Traditional constant false alarm rate (CFAR) detector has the disadvantages of high false alarm and poor adaptability. Deep learning provides a unique solution for SAR ship detection. However, the traditional deep network cannot reach very deep thus the accuracy is limited, and the training speed is slow. In this paper, a very deep network ResNet with higher accuracy and faster training speed is applied to train the SAR ship detection model. Moreover, transfer learning is applied to combat the small dataset. The proposed method is tested on a general SAR ship dataset and achieves 94.7% average precision. Comparative experiments show that our method has the best performance and which verifies the effectiveness of our method. Zegang Ding, Chi Zhang 0020, Yan Wang 0011, Jing Chen 0023 |
IGARSS | 3 |
| 2019 | Road Scene Layout Reconstruction based on CNN and its Application in Traffic SimulationabstractIn this paper, we propose a road scene prediction framework based on the control points of road boundaries using CNN. Firstly, the image features are extracted and the heatmaps are generated by CNN to locate the control points of road boundaries. The input images are then segmented to specify the scene layout based on the control points. Furthermore, the 3D traffic scene models are constructed. The applications for traffic simulation are then developed. The evaluations and comparisons based on TSD-max dataset prove the effectiveness of the proposed method. Yaochen Li, Yuehu Liu, Zhichao Cui, Chi Zhang 0020 |
IV | 6 |
| 2019 | Joint Task Difficulties Estimation and Testees Ranking for Intelligence EvaluationabstractIn this paper, we study the testing tasks evaluation and testees ranking problem, in which tasks have different difficulty levels, and testees have different capabilities.We assume that a testee may have a probability to pass a certain task so as to allow certain uncertainty. The goal of this problem is to simultaneously determine the relative difficulty level of each testing task and the relative capability of every testee, purely based on the test outcome. We design two models to solve this problem. The first one assumes that the test outcome follows a certain Bernoulli distribution; while the second one assumes that the test outcome follows a certain Bernoulli distribution with the beta distribution-type a priori knowledge. Then, we form the original problem into likelihood estimation problems and solve them by using coordinate descent algorithms. We show that the beta distribution-type a priori knowledge is needed, when we only carry out a limited number of tests due to time and financial budgets. All these findings are useful to intelligence tests. Finally, we discuss how to extend this statistical learning model for more general cases as well as in a specific case in the field of Computational Social Systems like artificial social cognition evaluation. Chi Zhang 0020, Yuehu Liu, Li Li 0013, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2018 | Traffic Sensory Data Classification by Quantifying Scenario ComplexityabstractFor unmanned ground vehicle (UGV) off-line testing and performance evaluation, massive amount of traffic scenario data is often required. The annotations in current off-line traffic sensory dataset typically include I) types of roadways II) scene types III) specific characteristics that are generally considered challenging for cognitive algorithms. While such annotations are helpful in manual selection of data, they are insufficient for comprehensive and quantitate measurement of per-roadway-segment scenario complexity. To resolve such limitations, we propose a traffic sensory data classification paradigm based on quantifying the scenario complexity for each roadway segment, where such quantification is jointly based on road semantic complexity and traffic element complexity. The road semantic complexity is a proposed measurement of the complexity incurred by the static elements such as curvy roads, intersections, merges and splits, which is predicted with a Support Vector Regression (SVR). The traffic element complexity is a measurement of complexity due to dynamic traffic elements, such as nearby vehicles and pedestrians. Experimental results and a case study verify the efficacy of the proposed method. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004 |
Intelligent Vehicles Symposium | 2 |
| 2018 | A Graded Offline Evaluation Framework for Intelligent Vehicle's Cognitive AbilityabstractCognitive ability evaluation in intelligent vehicles is conventionally evaluated by classical autonomous driving dataset, which lacks comprehensive annotations of driving difficulty. Realistically, different driving conditions require vast different level of cognitive ability, e.g., driving in highly congested traffic is much more challenging than driving on limited access highway; driving in a blizzard/hurricane requires much more robust environmental cognition abilities than driving under ordinary conditions. Different datasets contain different proportions of various driving conditions, rendering intelligent vehicle evaluation susceptible to dataset variations. To overcome such limitations, we propose to first benchmark the driving difficulty with the proposed “Cascaded Tanks Model” and obtain a fine-grained per-segment difficulty rating based on our proposed Semantic Descriptor. With the proposed Graded Offline Evaluation (GOE) framework, it is demonstrated that offline validation of the cognitive abilities in Intelligent Vehicles (IV) is more consistent regardless of dataset choice. Chi Zhang 0020, Yuehu Liu, Qilin Zhang 0004, Le Wang 0003 |
Intelligent Vehicles Symposium | 1 |
| 2015 | Autonomous Driving Simulation for Unmanned VehiclesabstractHuman can judge driver's driving ability by observing the vehicle motion in different traffic scenes. Identically, driving behavior can be the main basis for evaluating the performance of an unmanned vehicle in both field test and simulation test. Although simulation test avoids disadvantages of field test, existing simulation systems lack traffic scene data with perception granularity of visual sensors. In order for realizing vehicle-in-loop simulation, simulation technique of driving behaviors must be able to exhibit actual motion of unmanned vehicles. In this paper, we propose an automatic approach of simulating autonomous driving behaviors of vehicles in traffic scene represented by image sequences. Different from general simulation systems, we use actual traffic environment data to build the traffic scene and simulate the driving behaviors. After the proposed method was embedded in scene browser, a typical traffic scene including the intersections was chosen for virtual vehicle to execute the driving tasks of lane change, overtaking, slowing down and stop, right turn and U-Turn. The experimental results show that different driving behaviors of vehicles in typical traffic scene can be exhibited smoothly and realistically. Our method can also be used for generating simulation data of traffic scenes that are difficult to collect. Danchen Zhao, Yuehu Liu, Chi Zhang 0020, Yaochen Li |
WACV | 3 |