EDBT 2026 Demo / reviewers in the wild / expert
Jiawei Xu 0004
dblp:79/8798-4 · also Gary J. W. Xu
· DBLP profile ↗
30ranked-venue papers
12as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From honeybee to helicopter: A low-cost dual bionic model for flight control under hazardous situations
Shiwen Pan, Fengshuo Yan, Hong Cheng 0002, Kun Guo 0004, Jiawei Xu 0004, Zhao-Hui Sun, Xiaoru Wanyan, Edmond Q. Wu |
Neurocomputing | 7 |
| 2025 | FreqGAN: Infrared and Visible Image Fusion via Unified Frequency Adversarial LearningabstractTraditional fusion methods based on deep learning mainly employ convolutional or self-attention operations to model local or global dependencies, which often lead to the oversight of frequency-domain information. To address this deficiency, we introduce a unified frequency adversarial learning network, termed FreqGAN. Our method involves a frequency-compensated generator that employs discrete wavelet transformation to decompose encoded spatial features into multiple frequency bands. Leveraging skip connections, low and high-frequency components are respectively directed into the encoder and decoder, compensating for additional outline and detail. Moreover, we construct a hybrid frequency aggregation module, which enables a progressive optimization of activity levels across multiple scales and makes the various frequency bands correlated. Complementing our generative model, we devise dual frequency-constrained discriminators. These discriminators are tasked with dynamically adjusting weights for each input frequency band, thereby obligating the generator to accurately reconstruct salient frequency information from different modality images. Additionally, a frequency-supervised function is formulated to further safeguard against the loss of frequency information. Our comprehensive experimental evaluations, encompassing a wide range of fusion tasks and subsequent applications, distinctly highlight FreqGAN’s superior performance, establishing it as a frontrunner in comparison to existing state-of-the-art alternatives. The source codes are forthcoming at:https://github.com/Zhishe-Wang/FreqGAN. Zhishe Wang, Zhuoqun Zhang, Wuqiang Qi, Fengbao Yang, Jiawei Xu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MLP-Net: Multilayer Perceptron Fusion Network for Infrared Small Target DetectionabstractInfrared small target detection (IRSTD) faces various challenges such as long distances, weak features, and small scales. While methodologies based on convolutional neural networks (CNNs) have made strides, they are inherently hampered by a bias toward local reduction, limiting their global interpretive power. Conversely, Transformer-based approaches, though capable of capturing long-range dependencies, struggle with computational inefficiencies due to their quadratic complexity. To surmount these challenges, this article presents MLP-Net, a novel multilayer perceptron (MLP) fusion network for IRSTD. The architecture combines the advantages of CNNs and MLPs to capture global semantic information from local features and significantly enhance feature representation. Additionally, we develop a parallel token interaction mixer (PTIM) that processes the token representations with direction-specific interactive information across the height, width, and channel dimensions on MLPs, dynamically reinforcing the ability of long-range dependency modeling. Complementing this, we devise a contextual selection fusion module (CSFM) to gradually aggregate high-level semantics and low-level details from coarse to fine. This module integrates the complementary characteristics of different layers to promote detection accuracy. Finally, comprehensive experiments on the NUAA-SIRST, NUDT-SIRST, and IRSTD-1K benchmarks demonstrate that the proposed MLP-Net delivers promising detection performance, transcending other state-of-the-art alternatives. The relevant codes will be available athttps://github.com/Zhishe-Wang/MLP-Net. Zhishe Wang, Chunfa Wang, Chaoqun Xia, Jiawei Xu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | PKNet: Infrared Small Target Detection via Parallel Interactive Kolmogorov-Arnold NetworkabstractInfrared small target detection (IRSTD) remains challenging due to low signal-to-clutter ratios, unpredictable environmental interference, and weak spatial features. Convolutional neural network (CNN)-based methods excel at extracting local details but often fail to capture long-range dependencies, while Transformer-based approaches model global interactions through self-attention mechanisms but suffer from high computational costs. To overcome these challenges, we introduce a novel parallel interactive kolmogorov–arnold network, termed PKNet. In this architecture, the CNN branch is designed to extract fine-grained local features, while the KAN branch leverages learnable univariate functions to model contextual dependencies. To further strengthen global feature representation, we design a multi-grained KAN (MG-KAN) Block, which enhances context modeling by promoting nonlinear interactions across multiple token dimensions, enabling efficient feature extraction while preserving long-range dependencies. Moreover, we develop a cyclic interactive fusion module that facilitates bidirectional information refinement between the CNN and KAN branches. This module dynamically aligns and integrates multi-scale local and global features, significantly improving the network’s ability to distinguish small targets from complex backgrounds. Extensive experiments on three public benchmarks demonstrate that PKNet achieves superior performance in terms of both detection accuracy and efficiency, substantially outperforming existing state-of-the-art methods. The code will be available at:https://github.com/Zhishe-Wang/PKNet. Xiaomei Yan, Wang Ye, Chunfa Wang, Chaoqun Xia, Jiawei Xu 0004, Zhishe Wang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Human-Factors-in-Aviation-Loop: Multimodal Deep Learning for Pilot Situation Awareness Analysis Using Gaze Position and Flight Control DataabstractSituation awareness (SA) is a crucial factor affecting flight safety for pilots, yet few studies have focused specifically on modeling SA for pilots, resulting in limited success. In this paper, we propose a novel multimodal deep learning approach to monitor pilots’ SA. The approach combines handcrafted and deep features obtained from eye movement and flight control data collected from 27 novice pilots across different training phases using a flight simulator. Ground truth SA measurements were obtained using the Situation Awareness Global Assessment Technique (SAGAT). The handcrafted features included 13 eye movements and 22 flight control features, while deep features were extracted from time-series of gaze positions using a deep extractor based on Transformer. By fusing the handcrafted features of eye movement and flight control, along with one deep feature of eye movement, we predicted the final SA level. Through leave-one-flight-out cross-validation, our model achieved a higher accuracy of 92.04%. The results indicate that the multimodal model outperforms the unimodal models, with the eye movement modality demonstrating superiority over the flight control modality in predicting SA. This suggests our method provides an objective means of predicting pilot’s SA and offers new insights for SA assessment in aviation and other fields. Overall, our multimodal deep learning approach holds promise for enhancing pilot training and flight safety by facilitating a more comprehensive understanding of pilots’ SA during critical flight scenarios. Jiawei Xu 0004, Sicheng Pan, Zhao-Hui Sun, Kun Guo 0004, Seop Hyeong Park, Fengshuo Yan, Xiaoru Wanyan, Hong Cheng 0002, Qi Wu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | AITFuse: Infrared and visible image fusion via adaptive interactive transformer learning
Zhishe Wang, Jiawei Xu 0004, Fengbao Yang, Xiaomei Yan |
Knowl. Based Syst. | 4 |
| 2024 | Video-Based Engagement Estimation of Game Streamers: An Interpretable Multimodal Neural Network ApproachabstractIn this paper, we propose a non-intrusive and nonrestrictive multimodal deep learning model for estimating the engagement levels of game streamers. We incorporate three modalities from the streamers' videos (facial, pixel, and audio information) to train the multimodal neural network. Additionally, we introduce a novel interpretation technique that directly calculates the contribution of each modality to the model's classification performance without the need to retrain single modality models. Experimental results demonstrate that our model achieves an accuracy of 77.2% on the test set, with the sound modality identified as a key modality for engagement estimation. By utilizing the proposed interpretation technique, we further analyze the modality contributions of the model in handling different categories and samples from various players. This enhances the model's interpretability and reveals its limitations, as well as future directions for improvement. The proposed approach and findings have potential applications in the fields of game streaming and audience analysis, as well as in domains related to multimodal learning and affective computing. Sicheng Pan, Jiawei Xu 0004, Kun Guo 0004, Seop Hyeong Park, Hongliang Ding |
IEEE Trans. Games | 2 |
| 2024 | Cultural Insights in Souls-Like Games: Analyzing Player Behaviors, Perspectives, and Emotions Across a Multicultural ContextabstractSouls-like games are one of the most popular and emerging genres in the contemporary gaming world. This study compared the behavioral characteristics, perspectives, and emotional expressions of players in Souls-like games from different cultural backgrounds, specifically examining the distinctions and commonalities among them. Natural language processing techniques were employed to analyze English, Chinese, and Russian reviews of 17 Souls-like games to investigate players' gaming experiences, including gameplay behaviors, game evaluations, and emotional experiences. The findings revealed significant disparities among players from different cultures in all three aspects of their engagement with Souls-like games. Specifically, these players exhibited significant culture-related variations in their behavioral characteristics towards Souls-like games. In terms of perspectives, English-speaking players tended to focus more on game optimization, whereas Chinese and Russian players paid greater attention to game combat design. Regarding emotional expressions, Chinese players were more prone to exhibit emotions of anger and disgust, while English and Russian players displayed a more neutral emotional stance. These cultural insights provide valuable information for game developers to better meet the needs and expectations of players from different cultural backgrounds. This study not only broadens our understanding of player behaviors and cultural influences but also lends robust support to cross-cultural gaming research. Sicheng Pan, Jiawei Xu 0004, Kun Guo 0004, Seop Hyeong Park, Hongliang Ding |
IEEE Trans. Games | 2 |
| 2023 | A Cross-Scale Iterative Attentional Adversarial Fusion Network for Infrared and Visible ImagesabstractRecent existing methods generally adopt a simple concatenation or addition strategy to integrate features at the fusion layer, failing to adequately consider the intrinsic characteristics of different modal images and feature interaction of different scales, which may produce a limited fusion performance. Toward this end, we introduce a cross-scale iterative attentional adversarial fusion network, namely CrossFuse. More specifically, in the generator, we design a cross-modal attention integrated module to merge the intrinsic content of different modal images. The parallel spatial-independent and channel-independent pathways are proposed to calculate the attentional weights, which are assigned to measure the activity levels of source images at the same scale. Moreover, we construct a cross-scale iterative decoder framework to interact with different modality features at different scales, which can constantly optimize their activity levels. By this means, the generator learns to integrate their modality characteristics via attentional weights in an iterative manner, and the generated result characterizes competitive infrared radiant intensity and distinct visible detail description. Extensive experiments on three different benchmarks demonstrate that our CrossFuse outperforms other nine state-of-the-art methods in terms of fusion performance, generalization ability and computational efficiency. Our codes will be released athttps://github.com/Zhishe-Wang/CrossFuse. Zhishe Wang, Wenyu Shao, Jiawei Xu 0004, Lei Zhang 0168 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Human-Factors-in-Driving-Loop: Driver Identification and Verification via a Deep Learning Approach using Psychological Behavioral DataabstractDriver identification has been popular in the field of driving behavior analysis, which has a broad range of applications in anti-thief, driving style recognition, insurance strategy, and fleet management. However, most studies to date have only researched driver identification without a robust verification stage. This paper addresses driver identification and verification through a deep learning (DL) approach using psychological behavioral data, i.e., vehicle control operation data and eye movement data collected from a driving simulator and an eye tracker, respectively. We design an architecture that analyzes the segmentation windows of three-second data to capture unique driving characteristics and then differentiate drivers on that basis. The proposed model includes a fully convolutional network (FCN) and a squeeze-and-excitation (SE) block. Experimental results were obtained from 24 human participants driving in 12 different scenarios. The proposed driver identification system achieves an accuracy of 99.60% out of 15 drivers. To tackle driver verification, we combine the proposed architecture and a Siamese neural network, and then map all behavioral data into two embedding layers for similarity computation. The identification system achieves significant performance with average precision of 96.91%, recall of 95.80%, F1 score of 96.29%, and accuracy of 96.39%, respectively. Importantly, we scale out the verification system to imposter detection and achieve an average verification accuracy of 90.91%. These results imply the invariable characteristics from human factors rather than other traditional resources, which provides a superior solution for driving behavior authentication systems. Jiawei Xu 0004, Sicheng Pan, Zhao-Hui Sun, Seop Hyeong Park, Kun Guo 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | Infrared and Visible Image Fusion via Interactive Compensatory Attention Adversarial LearningabstractThe existing generative adversarial fusion methods generally concatenate source images or deep features, and extract local features through convolutional operations without considering their global characteristics, which tends to produce a limited fusion performance. Toward this end, we propose a novel interactive compensatory attention fusion network, termed ICAFusion. In particular, in the generator, we construct a multi-level encoder-decoder network with a triple path, and design infrared and visible paths to provide additional intensity and gradient information for the concatenating path. Moreover, we develop the interactive and compensatory attention modules to communicate their pathwise information, and model their long-range dependencies through a cascading channel-spatial model. The generated attention maps can more focus on infrared target perception and visible detail characterization, and are used to reconstruct the fusion image. Therefore, the generator takes full advantage of local and global features to further increase the representation ability of feature extraction and feature reconstruction. Extensive experiments illustrate that our ICAFusion obtains superior fusion performance and better generalization ability, which precedes other advanced methods in the subjective visual description and objective metric evaluation. Our codes will be public athttps://github.com/Zhishe-Wang/ICAFusion. Zhishe Wang, Wenyu Shao, Jiawei Xu 0004, Xiaoqin Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | UNFusion: A Unified Multi-Scale Densely Connected Network for Infrared and Visible Image FusionabstractInfrared image retains typical thermal targets while visible image preserves rich texture details, image fusion aims to reconstruct a synthesized image containing prominent targets and abundant texture details. Most of deep learning-based methods mainly focus on convolution operation to extract the local features, but do not fully consider their multi-scale characteristics and global dependencies, which may cause loss of target regions and texture details in the fused image. Towards this goal, we present a unified multi-scale densely connected fusion network in this paper, named as UNFusion. We carefully design a multi-scale encoder-decoder architecture that can efficiently extract and reconstruct multi-scale deep features. Dense skip connections are employed in both encoder and decoder sub-networks to reuse all the intermediate features of different layers and scales for fusion tasks. In the fusion layer,$L_{p} $normalized attention models, which include three kinds of different norms, are proposed to highlight and combine these deep features from spatial and channel dimensions, and the combined spatial and channel attention maps are used to reconstruct a final fused image. We conduct extensive experiments on the public TNO and Roadscene datasets, and the results demonstrate that our UNFusion can simultaneously preserve high brightness of typical thermal targets and abundant texture details to obtain superior scene representation and better visual perception. Besides, our UNFusion achieves better fusion performance and transcends other state-of-the-art methods in terms of qualitative and quantitative comparisons. Our code is available athttps://github.com/Zhishe-Wang/UNFusion. Zhishe Wang, Junyao Wang 0002, Jiawei Xu 0004, Xiaoqin Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | The Improvement of Road Driving Safety Guided by Visual Inattentional BlindnessabstractThe computational modeling of human visual attention has received much attention in recent decades. In advanced industrial applications, it has been demonstrated that computational visual attention models (CVAMs) can predict visual attention very similarly to human visual attention. However, it is controversial whether the driver’s eye fixation location (EFL) or the predicted eye fixation location of computational visual attention models is more reliable and helpful for actual driving. To address this issue, an open database of videos taken under the most common 18 driving conditions in everyday driving has been established. In experiments using this database, expert drivers found that it was not sufficient for drivers to rely on only one of the two EFLs. Based on this finding, a hybrid EFL recommendation strategy is proposed for improving driving safety. By extracting visual characteristics from human dynamic vision, the performance of the proposed recommendation method demonstrates its potential value in these collected driving tasks. In addition, the visual comfort of driving is further addressed to enhance the safety of driving. From the results of experiments on 108 driving video clips taken of the most common 18 real driving conditions, it is confirmed that the proposed EFL recommendation achieves an experience rating of driving comfort between 88.1 and 92.7 out of 100. Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002, Jie Hu 0041 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | The Alleviation of Perceptual Blindness During Driving in Urban Areas Guided by Saccades RecommendationabstractIn advanced industrial applications, computational visual attention models (CVAMs) could predict visual attention very similarly to actual human attention allocation. This has been used as a very important component of technology in advanced driver assistance systems (ADAS). Given that the biological inspiration of the driving-related CVAMs could be obtained from skilled drivers in complex driving conditions, in which the driver’s attention is constantly directed at various salient and informative visual stimuli by alternating the eye fixations via saccades to drive safely, this paper proposes a saccade recommendation strategy to enhance the driving safety under urban road environment, particularly when the driver’s vision is often impaired by the visual crowding. The altered and directed saccades are collected and optimized by extracting four innate features from human dynamic vision. A neural network isdesigned to classify preferable saccades to reduce perceptual blindness due to visual crowding under urban scenes. A state-of-the-art CVAM is firstly adopted to localize the predicted eye fixation locations (EFLs) in driving video clips. Besides, human subjects’ gaze at the recommended EFLs is measured via an eye-tracker. The time delays between the predicted EFLs and drivers’ EFLs are analyzed under different driving conditions, followed by the time delays between the predicted EFLs and the driver’s hand control. The visually safe margin is then measured by mediating the driving speed and the total delay. Experimental results demonstrate that the recommended saccades can effectively reduce the amount of perceptual blindness, which is known to be of help to further improve road driving safety. Jiawei Xu 0004, Xiaoqin Zhang 0002, Seop Hyeong Park, Kun Guo 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Haze concentration adaptive network for image dehazing
Tao Wang 0052, Li Zhao 0005, Pengcheng Huang 0002, Xiaoqin Zhang 0002, Jiawei Xu 0004 |
Neurocomputing | 5 |
| 2021 | Robust feature learning for adversarial defense via hierarchical feature alignment
Xiaoqin Zhang 0002, Tao Wang 0052, Runhua Jiang, Jiawei Xu 0004, Li Zhao 0005 |
Inf. Sci. | 5 |
| 2020 | Self-calibrated Attention Residual Network for Image Super-ResolutionabstractDeep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in single image super-resolution (SISR). However, most SR methods restore high resolution (HR) images from single-scale region in the low resolution (LR) input, which limits the ability of method to infer multi-scales of details for high resolution (HR) output. In this paper a novel basic building block called self-Calibrated residual block (SARB) is proposed to solve this problem. SARB consists of carefully designed multi-scale paths, which can capture rich structure information from different scale. In addition, self-Calibrated residual block is introduced to adaptively learn informatively context to make network generate more discriminative representations. These blocks are composed of self-calibrated attention residual network (SARN) for image super-resolution. Experiments results on five benchmark datasets demonstrate that the proposed SARN achieves comparable results compared with the previous most of the state-of-the-art methods. Anqi Rong, Li Zhao 0005, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 4 |
| 2020 | A Nonlocal Denoising Framework Based on Tensor Robust Principal Component Analysis with ℓp normabstractThis paper have given a nonlocal denoising framework based on tensor robust principal component analysis with ℓpnorm for color image and video (NDFCIV), which have following three features: (1) it is capable of processing zero-mean Gaussian noise, impulse noise and any other noise that is created by mixing the two for color image and video at same time. (2) Meanwhile, nonlocal denoising strategy is adopted to promote the effectiveness of the denoising framework. (3) Moreover, we present a non-convex constraint method which can get more exact low rank tensor recovery result and enhance the denoising effect of the framework further. The experimental results demonstrate the effectiveness of the proposed denoising framework. Mengqing Sun, Li Zhao 0005, Jiawei Xu 0004 |
IEEE BigData | 4 |
| 2020 | Feature Fusion Based on Sparse Block for Image Super-resolutionabstractRecently, deep neural networks have been widely used in the task of single image super-resolution. However, existing deep neural networks always take huge parameters to map low-resolution images to high-resolution ones. In addition, most of them only consider high-level features to reconstruct high-resolution images. These two methodologies not only cause the difficulty of practical applications but also the inefficiency of restoring image details. Therefore, in this work, the authors propose a novel sparse block to learn high-level features. Based on this block, a fusion method is proposed to fuse features from multiple levels. By incorporating these two approaches, a lightweight neural network, i.e. Sparse Block Fusion Network (SBFN), is proposed for end-to-end training. Through extensive experiments, it is demonstrated the proposed methods can achieve comparable performance with few parameters. By making comprehensive comparisons, effectiveness of SBFN is also verified in multiple benchmark datasets. Shengping Wang, Li Zhao 0005, Runhua Jiang, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 5 |
| 2020 | Multi-level Feature Fusion Network for Single Image Super-ResolutionabstractRecently, deep convolution neural networks have achieved remarkable performance in the task of single image super-resolution (SISR). However, effectiveness of existing networks highly relies on their receptive field, which always increases with the depth of the network. In this work, we propose a novel module, named as residual group, to effectively learn feature maps by using dynamic receptive field. This residual group firstly uses a selective kernel convolution layer to dynamically learn multi-scale information from its input features. Then, several residual blocks are employed to further refine the learned feature. In addition, we also propose a selective feature fusion module to fuse appearance information in multi-level features. Within this module, the low-level features and high-level features are selectively fused to complement the high-level ones. Finally, by combining these two methods, we introduce a multi-level feature fusion network (MLFFN) for single image super-resolution (SISR). Through comprehensive experiments, we demonstrate that the proposed MLFFN achieves state-of-the-art performance both quantitatively and qualitatively. Xinxia Zhang, Xiaoqin Zhang 0002, Li Zhao 0005, Runhua Jiang, Pengcheng Huang 0002, Jiawei Xu 0004 |
IEEE BigData | 6 |
| 2020 | Improvement of viewing experience on stereoscopic image guided by human stereo vision
Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
Multim. Tools Appl. | 1 |
| 2020 | A Temporally Irreversible Visual Attention Model Inspired by Motion Sensitive NeuronsabstractWhen a human perceives videos composed of the same images in various orders, such as normal order, reverse order, and random order, the human visual attention system perceives them as different visual inputs. This means that the temporal change in the image sequence exerts a considerable influence on the human visual system. However, most state-of-the-art computational visual attention models have not considered the temporal cues adequately. Motivated by this deficiency, we propose a novel temporally irreversible visual attention model considering the following three aspects. First, the central bias of human dynamic vision is incorporated into the model to manifest this tendency. Second, the depth and directional motion-sensitive neurons are fused to discern different motion patterns. Third, the rarity factor is integrated into the model to mimic the attention shift when human observer perceives new emerging motion cues. The proposed model demonstrates its competitiveness to select attentive events in our experiments, in both laboratory setting and real driving video clips. When compared with recent visual attention models, the proposed model achieves the highest score in similarity with human dynamic vision. The proposed model could be one of the fundamental building blocks for any visual attention systems coping with dynamic scenes. Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | A bio-inspired motion sensitive model and its application to estimating human gaze positions under classified driving conditions
Jiawei Xu 0004, Seop Hyeong Park, Xiaoqin Zhang 0002 |
Neurocomputing | 1 |
| 2017 | What has been missed for real life driving? an inspirational thinking from human innate biasesabstractNature is gorgeous for her imbalance. The innate bias from non-human to human results in a wonderful yet mysterious biological foundation towards the inspirational thinking for real application. The vision researchers evidenced that the left gaze bias in humans and non-humans. Nevertheless, the acousticians observed the right ear advantages in both non-humans and humans. Unlike the vision and acoustician researchers investigating the underlying mechanisms of human innate bias, we are more interested in mimicking these characteristics. In this paper, we propose two simple yet effective methods to generate the left eye gaze bias and the right ear advantage. We further discuss the potential applications, e.g., real life driving, from these inherent phenomena. We believe that this paper could bring an inspirational impact for future cognitive transportation, by implementing these human innate biases properly. Jiawei Xu 0004, Yu-An Chen, Kun Guo 0004, Jiheng Wang, Federica Menchinelli, Ling Shao 0001 |
AVSS | 1 |
| 2017 | Traffic Sign Detection Using a Cascade Method With Fast Feature Extraction and Saliency TestabstractAutomatic traffic sign detection is challenging due to the complexity of scene images, and fast detection is required in real applications such as driver assistance systems. In this paper, we propose a fast traffic sign detection method based on a cascade method with saliency test and neighboring scale awareness. In the cascade method, feature maps of several channels are extracted efficiently using approximation techniques. Sliding windows are pruned hierarchically using coarse-to-fine classifiers and the correlation between neighboring scales. The cascade system has only one free parameter, while the multiple thresholds are selected by a data-driven approach. To further increase speed, we also use a novel saliency test based on mid-level features to pre-prune background windows. Experiments on two public traffic sign data sets show that the proposed method achieves competing performance and runs 2~7 times as fast as most of the state-of-the-art methods. Xinwen Hou, Jiawei Xu 0004, Shigang Yue, Cheng-Lin Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | A saliency-based cascade method for fast traffic sign detectionabstractWe propose a cascade method for fast and accurate traffic sign detection. The main feature of the method is that mid-level saliency test is used to efficiently and reliably eliminate background windows. Fast feature extraction is adopted in the subsequent stages for rejecting more negatives. Combining with neighbor scales awareness in window search, the proposed method runs at 3~5 fps for high resolution (1360×800) images, 2~7 times as fast as most state-of-the-art methods. Compared with them, the proposed method yields competitive performance on prohibitory signs while sacrifices performance moderately on danger and mandatory signs. Shigang Yue, Jiawei Xu 0004, Xinwen Hou, Cheng-Lin Liu 0001 |
Intelligent Vehicles Symposium | 3 |
| 2015 | Building up a Bio-Inspired Visual Attention Model by Integrating Top-Down Shape Bias and Improved Mean Shift Adaptive SegmentationabstractThe driver-assistance system (DAS) becomes quite necessary in-vehicle equipment nowadays due to the large number of road traffic accidents worldwide. An efficient DAS detecting hazardous situations robustly is key to reduce road accidents. The core of a DAS is to identify salient regions or regions of interest relevant to visual attended objects in real visual scenes for further process. In order to achieve this goal, we present a method to locate regions of interest automatically based on a novel adaptive mean shift segmentation algorithm to obtain saliency objects. In the proposed mean shift algorithm, we use adaptive Bayesian bandwidth to find the convergence of all data points by iterations and the k-nearest neighborhood queries. Experiments showed that the proposed algorithm is efficient, and yields better visual salient regions comparing with ground-truth benchmark. The proposed algorithm continuously outperformed other known visual saliency methods, generated higher precision and better recall rates, when challenged with natural scenes collected locally and one of the largest publicly available data sets. The proposed algorithm can also be extended naturally to detect moving vehicles in dynamic scenes once integrated with top-down shape biased cues, as demonstrated in our experiments. Jiawei Xu 0004, Shigang Yue |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2014 | Mimicking visual searching with integrated top down cues and low-level features
Jiawei Xu 0004, Shigang Yue |
Neurocomputing | 1 |
| 2013 | A Motion Attention Model Based on Rarity Weighting and Motion cues in Dynamic ScenesabstractNowadays, motion attention model is a controversial topic in the biological computer vision area. The computational attention model can be decomposed into a set of features via predefined channels. Here we designed a bio-inspired vision attention model, and added the rarity measurement onto it. The priority of rarity is emphasized under the assumption of weighting effect upon the features logic fusion. At this stage, a final saliency map at each frame is adjusted by the spatiotemporal and rarity values. By doing this, the process of mimicking human vision attention becomes more realistic and logical to the real circumstance. The experiments are conducted on the benchmark dataset of static images and video sequences. We simulated the attention shift based on several dataset. Most importantly, our dynamic scenes are mostly selected from the objects moving on the highway and dynamic scenes. The former one can be developed on the detection of car collision and will be a useful tool for further application in robotics. We also conduct experiment on the other video clips to prove the rationality of rarity factor and feature cues fusion methods. Finally, the evaluation results indicate our visual attention model outperforms several state-of-the-art motion attention models. Jiawei Xu 0004, Shigang Yue, Yuchao Tang |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | Visual Based Contour Detection by Using the Improved Short Path Finding
Jiawei Xu 0004, Shigang Yue |
EANN | 1 |