Jianwei Zhang 0013

dblp:144/1628-13 · also Jian-wei Zhang 0013 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
31since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 An improved differential evolutionary algorithm integrating neighborhood search variation mechanism and multi-strategies for airport gate allocation
Lirong Zhang, Edmond Q. Wu, Wu Deng 0001, Jianwei Zhang 0013, Yi Lin 0006
Expert Syst. Appl.5
2026 Enhancing unsupervised unified anomaly detection via semi-supervised learning with multi-source uncertainty mining
Borui Kang, Yuzhong Zhong, Lin Deng 0003, Maoning Wang, Jianwei Zhang 0013
Neurocomputing5
2026 Joint scheduling of runway and taxiway considering uncertain taxiing time based on an improved ACO algorithm
Lirong Zhang, Yi Lin 0006, Suwan Yin, Wu Deng 0001, Hongyu Yang 0002, Jianwei Zhang 0013
Inf. Sci.6
2026 C3aptioner: Improving change captioning by leveraging momentum cross-view and cross-modality contrastive learning
abstract
The primary goal of change captioning is to identify subtle visual differences between two similar images and express them in natural language. Existing research has been significantly influenced by the task of vision change detection and has mainly concentrated on the identification and description of visual changes. However, we contend that an effective change captioner should go beyond mere detection and description of what has changed. Two additional aspects are crucial: 1) retaining significant and unique semantic elements that persist across both images, and 2) forging a robust link between visual cues and their concomitant descriptive linguistic elements. This paper addresses these challenges by presenting the C 3 aptioner, which seamlessly incorporates dual momentum contrastive learning objectives into change captioning. Our model architecture consists of intra-image and inter-image Transformer encoders for visual feature extraction, complemented by unimodal language and multimodal decoders. Specifically, we introduce a cross-view contrastive learning objective to capture essential invariant features by aligning cross-view representations with a momentum-updated queue of negative samples, addressing the challenge of viewpoint variations. Additionally, our cross-modality contrastive learning objective aligns and interacts visual and textual modalities using a separate momentum-maintained queue, resolving the modality gap that hampers existing methods. This dual contrastive approach enables C 3 aptioner to model both changed and unchanged elements while establishing strong vision-language correspondence, resulting in more contextually rich and human-like descriptions. Extensive experiments across five distinct datasets confirm that our approach achieves state-of-the-art performance, with particularly significant improvements in challenging scenarios involving extreme viewpoint changes. Source code is available at https://github.com/DenglinGo/C-3aptioner .
Lin Deng 0003, Borui Kang, Yuzhong Zhong, Maoning Wang, Jianwei Zhang 0013
Neural Networks5
2026 DWSF-Net: A Dynamic Wavelet-Based Spatial-Frequency Fusion Network for Multispectral Object Detection
abstract
Multispectral object detection aims to identify tar gets under diverse illumination conditions by leveraging complementary information from multiple spectral modalities. A major challenge in this field lies in effectively fusing multispectral features while accounting for both spatial and frequency domain characteristics. Existing methods primarily focus on spatial fusion, often neglecting critical frequency domain cues and treating all spectral channels equally, despite their distinct properties. In particular, RGB images capture high-frequency texture and color, whereas infrared (IR) images focus on low frequency thermal signatures—rendering conventional spatial only fusion suboptimal. To address these challenges, we propose a novel Dynamic Wavelet-based Spatial-Frequency Fusion Net work (DWSF-Net) that integrates both spatial and frequency information for enhanced multispectral representation. DWSF Net introduces a learnable wavelet encoder to adaptively extract frequency-aware features, a wavelet modulation fusion module to selectively combine informative sub-bands across spectra, and a frequency-domain sub-band fusion scheme with adaptive weight learning to refine cross-spectral integration. Finally, modulated spatial features and adaptively fused frequency components are aggregated to form the final representation. Extensive experiments conducted on three public datasets demonstrate that the proposed DWSF-Net achieves state-of-the-art performance, highlighting its effectiveness and potential for improving the accuracy of multispectral object detection.
Fan Yang 0104, Wei Li 0075, Lei Li 0020, Jianwei Zhang 0013
IEEE Trans. Multim.5
2025 Improved YOLOv7 for small object detection in airports: Task-oriented feature learning with Gaussian Wasserstein loss and attention mechanisms
Ruijie Peng, Chuanlin Liao, Weijun Pan, Xiaolin Gou, Jianwei Zhang 0013, Yi Lin 0006
Neurocomputing5
2025 Multidimensional Fusion Network for Multispectral Object Detection
abstract
Multispectral object detection has attracted increasing attention recently due to its superior detection capacity under various illumination conditions. The key challenge lies in the effective aggregation of multi-spectral features to derive highly discriminative representations. To address this challenge, we propose a novel Multidimensional Fusion Network (MMFN) to explore multi-modal information from local, global, and channel perspectives. Specifically, at the local level, local features of different modalities and their inter-relationships are captured by a window-shifted fusion. As a complement to the local information, we designed a global interaction module that facilitates the fusion of holistic, high-level semantic information spanning the entire image. We distillate the channel dependencies and complementarities between different modalities through cross-channel learning and generate the final fused representation. Comprehensive experiments conducted on three publicly available datasets provide compelling evidence validating the superiority of the proposed methodology. The results exhibit notable performance gains over state-of-the-art multispectral object detectors. Our code will be released.
Fan Yang 0104, Binbin Liang, Wei Li 0075, Jianwei Zhang 0013
IEEE Trans. Circuits Syst. Video Technol.4
2025 A Non-Autoregressive Multi-Horizon Flight Trajectory Prediction Framework With Gray Code Representation
abstract
Flight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing methods generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error accumulation and low-efficiency problems. In this paper, a novel framework, called FlightBERT++, is proposed to i) forecast multi-horizon flight trajectories directly in a non-autoregressive way, and ii) improve the limitation of the binary encoding (BE) representation in the FlightBERT framework. Specifically, the proposed framework is implemented by a generalized encoder-decoder architecture, in which the encoder learns the temporal-spatial patterns from historical observations and the decoder predicts the flight status for the future horizons. Compared to conventional architecture, an innovative horizon-aware context generator is dedicatedly designed to consider the prior horizon information, which further enables non-autoregressive multi-horizon prediction. Additionally, the Gray code representation and the differential prediction paradigm are designed to cope with the high-bit misclassifications of the BE representation, which significantly reduces the outliers in the predictions. Moreover, a differential prompted decoder is proposed to enhance the capability of the differential predictions by leveraging the stationarity of the differential sequence. Extensive experiments are conducted to validate the proposed framework on a real-world flight trajectory dataset. The experimental results demonstrated that the proposed framework outperformed the competitive baselines in both FTP performance and computational efficiency. The code is publicly available at: https://github.com/gdy-scu/FlightBERT_PP_V2
Dongyue Guo, Fengshuo Ye, Jianwei Zhang 0013, Hongyu Yang 0002, Yi Lin 0006
IEEE Trans. Intell. Transp. Syst.5
2025 Exploring Contextual Knowledge-Enhanced Speech Recognition in Air Traffic Control Communication: A Comparative Study
abstract
Accurate recognition of named entities from spoken instructions remains a significant challenge for automatic speech recognition (ASR) techniques in air traffic control (ATC), which limits the reliability of ASR-based applications. A promising solution to overcome this challenge is to integrate prior contextual knowledge into ASR since it contains rich named entities used in ATC communications. Although existing studies have investigated ATC-related contextual ASR techniques, there is a lack of benchmarks to evaluate the advantages of different approaches. In this article, a comprehensive comparative study is presented to explore effective contextual ASR approaches for the ATC domain. Specifically, several typical contextual ASR approaches are introduced in ATC to conduct a comprehensive comparison. Moreover, a novel contextual ASR model, denoted CATCNet, is presented to dedicatedly address the domain-specific problems in ATC, such as limited resources, fast speech, and volatile noise. Several evaluation metrics are proposed to validate the performance of comparison approaches based on the practical requirements of ATC efforts. Extensive experiments are conducted across two real-world ATC speech corpora to build the benchmark. The experimental results demonstrated that integrating context knowledge is effective in improving the recognition performance of named entities. Crucially, the proposed CATCNet outperforms other baseline models by confirming all technical improvements, achieving 80.0% and 86.54% instruction recognition accuracy (IRA) on the ATCSpeech and C-ATCSpeech corpora, respectively. It is believed that this work not only overcomes the bottleneck of ASR performance in the ATC domain, but also provides an applicable solution for ATC-related ASR applications.
Dongyue Guo, Jianwei Zhang 0013, Bo Yang 0063, Yi Lin 0006
IEEE Trans. Neural Networks Learn. Syst.3
2025 Inserting Objects into Any Background Images via Implicit Parametric Representation
abstract
Inserting an object into a background scene has wide applications in image editing and mixed reality. However, existing methods still struggle to seamlessly adapt the object to the background while maintaining its individual characteristics. In this article, we propose to fine-tune a pre-trained diffusion-based insertion model such that it learns to establish a unique correspondence between a few weights and the target object, given as input few-shot images of an object. A novel individualized feature extraction (IFE) module is designed to extract the individual detail features from few-shot object images. Then, the individual features of the target object, together with the semantic features of the target object and the background context features extracted by the pre-trained image encoders are injected into the cross-attention modules of the latent diffusion model, enabling it to learn the correlation information of the target object and the background scene through the attention mechanism. The weights obtained by fine-tuning implicitly serve as an alternative representation of the target object, with which the object can be easily inserted into any background images. Extensive comparative experiments validate the superiority of the proposed method to the state-of-the-art insertion methods in maintaining the individual details of the inserted object and adapting it to background scenes, including allowing the interaction between the inserted object and the background scene, correctly handling their occlusion relationship, maintaining the consistency of their viewpoints and poses.
Qi Zhang 0139, Guanyu Xing, Mengting Luo, Jianwei Zhang 0013, Yanli Liu 0002
IEEE Trans. Vis. Comput. Graph.4
2024 FlightBERT++: A Non-autoregressive Multi-Horizon Flight Trajectory Prediction Framework
abstract
Flight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing approaches generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error accumulation and low-efficiency problems. In this paper, a novel framework, called FlightBERT++, is proposed to i) forecast multi-horizon flight trajectories directly in a non-autoregressive way, and ii) improve the limitation of the binary encoding (BE) representation in the FlightBERT. Specifically, the FlightBERT++ is implemented by a generalized encoder-decoder architecture, in which the encoder learns the temporal-spatial patterns from historical observations and the decoder predicts the flight status for the future horizons. Compared with conventional architecture, an innovative horizon-aware contexts generator is dedicatedly designed to consider the prior horizon information, which further enables non-autoregressive multi-horizon prediction. Moreover, a differential prompted decoder is proposed to enhance the capability of the differential predictions by leveraging the stationarity of the differential sequence. The experimental results on a real-world dataset demonstrated that the FlightBERT++ outperformed the competitive baselines in both FTP performance and computational efficiency.
Dongyue Guo, Jianwei Zhang 0013, Yi Lin 0006
AAAI4
2024 GenUDC: High Quality 3D Mesh Generation With Unsigned Dual Contouring Representation
Ruowei Wang, Dan Zeng 0002, Xueqi Ma, Zixiang Xu, Jianwei Zhang 0013, Qijun Zhao
ACM Multimedia6
2024 2D human skeleton action recognition with spatial constraints
abstract
Abstract Human actions are predominantly presented in 2D format in video surveillance scenarios, which hinders the accurate determination of action details not apparent in 2D data. Depth estimation can aid human action recognition tasks, enhancing accuracy with neural networks. However, reliance on images for depth estimation requires extensive computational resources and cannot utilise the connectivity between human body structures. Besides, the depth information may not accurately reflect actual depth ranges, necessitating improved reliability. Therefore, a 2D human skeleton action recognition method with spatial constraints (2D‐SCHAR) is introduced. 2D‐SCHAR employs graph convolution networks to process graph‐structured human action skeleton data comprising three parts: depth estimation, spatial transformation, and action recognition. The initial two components, which infer 3D information from 2D human skeleton actions and generate spatial transformation parameters to correct abnormal deviations in action data, support the latter in the model to enhance the accuracy of action recognition. The model is designed in an end‐to‐end, multitasking manner, allowing parameter sharing among these three components to boost performance. The experimental results validate the model's effectiveness and superiority in human skeleton action recognition.
Lei Wang 0225, Jianwei Zhang 0013, Wenbing Yang, Song Gu, Shanmin Yang
IET Comput. Vis.2
2024 MSTAD: A masked subspace-like transformer for multi-class anomaly detection
Borui Kang, Yuzhong Zhong, Zhimin Sun, Lin Deng 0003, Maoning Wang, Jianwei Zhang 0013
Knowl. Based Syst.6
2024 ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
abstract
Radio speech echo is a specific phenomenon in the air traffic control (ATC) domain, which degrades speech quality and further impacts automatic speech recognition (ASR) accuracy. In this work, a time-domain recognition-oriented speech enhancement (ROSE) framework is proposed to improve speech intelligibility and also advance ASR accuracy based on convolutional encoder-decoder-based U-Net framework, which serves as a plug-and-play tool in ATC scenarios and does not require additional retraining of the ASR model. Specifically, 1) In the U-Net architecture, an attention-based skip-fusion (ABSF) module is applied to mine shared features from encoders using an attention mask, which enables the model to effectively fuse the hierarchical features. 2) A channel and sequence attention (CSAtt) module is innovatively designed to guide the model to focus on informative features in dual parallel attention paths, aiming to enhance the effective representations and suppress the interference noises. 3) Based on the handcrafted features, ASR-oriented optimization targets are designed to improve recognition performance in the ATC environment by learning robust feature representations. By incorporating both the SE-oriented and ASR-oriented losses, ROSE is implemented in a multi-objective learning manner by optimizing shared representations across the two task objectives. The experimental results show that the ROSE significantly outperforms other state-of-the-art methods for both the SE and ASR tasks, in which all the proposed improvements are confirmed by designed experiments. In addition, the proposed approach can contribute to the desired performance improvements on public datasets.
Xincheng Yu, Dongyue Guo, Jianwei Zhang 0013, Yi Lin 0006
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Research on Clustering Detection Method for Security Attack Behaviors Based on Air Traffic Control Network
abstract
The problem of high similarity in attack data leading to unsatisfactory detection results of air traffic control network security attack behavior is addressed. This article designs a new clustering detection method for air traffic control network security attack behavior. Set the characteristic state of air traffic control network security attack behavior, obtain the set of air traffic control network security attack behavior characteristics through recursive feature addition method, and extract the characteristics of air traffic control network security attack behavior by determining the degree of feature criticality. Calculate the expected information gain and entropy value of feature data, determine the information gain of feature data, and reduce the interference of similar feature data. Introduce an automatic encoder in artificial intelligence algorithms to encode and decode the characteristics of air traffic control network security attack behavior, and achieve dimensionality reduction processing of air traffic control network security attack behavior data. Based on the above processing, a Unsupervised learning algorithm for clustering detection of air traffic control network security attacks is designed. Firstly, determine the distance between clustering clusters of air traffic control network security attack behavior characteristics, calculate the clustering threshold, and construct the initial clustering center. Then, recalculate the new mean of all feature objects in each cluster as the new cluster center point. Secondly, traverse all objects in the clustering cluster of air traffic control network security attack behavior feature data. Finally, clustering detection of air traffic control network security attack behavior is completed through the calculation of the objective function. The experiment takes three sets of experimental attack behavior datasets as the test subjects, with detection rate, false detection rate, and recall rate as the test indicators, and selects three similar methods for comparative testing. The experimental results show that the detection rate of the proposed method remains around 98%, the false detection rate remains below 1%, and the recall rate is above 97%. It has been proven that the proposed method can improve the detection performance of air traffic control network security attack behavior.
Ruchun Jia, Jianwei Zhang 0013, Zekun Jiang
HealthCom2
2023 3D Semantic Subspace Traverser: Empowering 3D Generative Model with Shape Editing Capability
abstract
Shape generation is the practice of producing 3D shapes as various representations for 3D content creation. Previous studies on 3D shape generation have focused on shape quality and structure, without or less considering the importance of semantic information. Consequently, such generative models often fail to preserve the semantic consistency of shape structure or enable manipulation of the semantic attributes of shapes during generation. In this paper, we proposed a novel semantic generative model named 3D Semantic Subspace Traverser that utilizes semantic attributes for category-specific 3D shape generation and editing. Our method utilizes implicit functions as the 3D shape representation and combines a novel latent-space GAN with a linear subspace model to discover semantic dimensions in the local latent space of 3D shapes. Each dimension of the subspace corresponds to a particular semantic attribute, and we can edit the attributes of generated shapes by traversing the coefficients of those dimensions. Experimental results demonstrate that our method can produce plausible shapes with complex structures and enable the editing of semantic attributes. The code and trained models are available at https://github.com/TrepangCat/3D Semantic Subspace Tra verser
Ruowei Wang, Pei Su, Jianwei Zhang 0013, Qijun Zhao
ICCV4
2023 CONICA: A Contrastive Image Captioning Framework with Robust Similarity Learning
abstract
Contrastive Language Image Pre-training (CLIP) has recently made significant advancements in image captioning by providing effective multi-modal representation learning capabilities. However, previous studies primarily rely on the language-aligned visual semantics as input for the captioning model, leaving the learned robust vision-language relevance under-exploited. In this paper, we propose CONICA, a unified CONtrastive Image CAptioning framework that investigates how contrastive learning can further enhance image captioning from three aspects. Firstly, we introduce contrastive learning objectives into the typical image captioning training pipeline with minimal overhead. Secondly, we construct fine-grained contrastive samples to obtain image-text similarities that correlate with the evaluation metric of image captioning. Finally, we incorporate the learned contrastive knowledge into the captioning decoding strategy to search for better captions. Experimental results demonstrate that CONICA significantly improves performance over standard captioning baselines and achieves new state-of-the-art results on the MSCOCO and Flikr30K. Source code is available at https://github.com/DenglinGo/CONICA.
Lin Deng 0003, Yuzhong Zhong, Maoning Wang, Jianwei Zhang 0013
ACM Multimedia4
2023 M2ATS: A Real-world Multimodal Air Traffic Situation Benchmark Dataset and Beyond
abstract
Air Traffic Control (ATC) is a complicated, time-evolving, and real-time procedure to direct flight operations in a safer and ordered manner. Although enormous data storages are available during air traffic operations for over 40 years, data-driven intelligent application in aviation is still an emerging task due to the safety-critical issue. With the prevalence of the Next Generation ATC system, artificial intelligence (AI) -empowered research topics are attracting increasing attention from both industrial and academic domains and a high-quality dataset naturally becomes the prerequisite for such practices. However, almost all ATC-related datasets are only unimodal for certain tasks, which fails to comprehensively illustrate the traffic situation to further support real-world studies. To address this gap, a multimodal air traffic situation (M2ATS) dataset is constructed to advance AI-related research in the ATC domain, including airspace information, flight plan, trajectory, and speech. M2ATS covers 10362 flights ATC situation data, involving 110000+ utterances (104 hours) with diversity golden text annotations, 16 intents, and 51 slots. Considering the real-world ATC requirements, a total of 10 multimedia-related tasks (24 baselines) are designed to validate the proposed dataset, covering automatic speech recognition, natural language processing, and spatial-temporal data processing. New ATC-related metrics corresponding to ATC applications are proposed in addition to the common metrics to evaluate task performance. Extensive experiment results demonstrate that the selective baselines can achieve designed tasks on this new dataset, and further investigations are also required to address task and data specificities. It is believed that the proposed new dataset is a new practice to advance AI applications to an industrial scene, which not only promotes ATC-related applications but also provides diverse research topics in the common multimedia community.
Dongyue Guo, Yi Lin 0006, Xuehang You, Zhongping Yang, Jizhe Zhou 0001, Bo Yang 0063, Jianwei Zhang 0013, Shasha Hu
ACM Multimedia7
2023 Spatiotemporal Interaction Transformer Network for Video-Based Person Reidentification in Internet of Things
abstract
Video-based person reidentification, which is a significant application in the Internet of Things, aims to identify the same person in different video sequences across nonoverlapping cameras. Existing methods usually utilize temporal cues to enhance spatial features. However, these methods learn the temporal and spatial information separately, which breaks the relationship between them and ignores the positive role of temporal information for learning frame-level spatial representation in the process of spatial representation learning. In this article, we propose a novel spatiotemporal interaction transformer network (SITN) to solve this problem. To model the temporal information and the relationship between frames, we introduce a temporal interaction module (TIM) to interact between frame information. Meanwhile, we combine TIM with spatial transformer encoder to explore the positive role of temporal information in the learning procedure of the frame-level spatial feature. Moreover, we propose a transformer local learning scheme by reconstructing the 2-D spatial information of the frame patch sequences and extracting local features in a striped manner to strengthen the discriminative capability of our model. Extensive experiments are conducted on four public benchmarks. The results show that our model is superior compared with state-of-the-art methods.
Fan Yang 0104, Wei Li 0075, Binbin Liang, Jianwei Zhang 0013
IEEE Internet Things J.4
2023 Key Frame Mechanism for Efficient Conformer Based End-to-End Speech Recognition
abstract
Recently, Conformer as a backbone network for end-to-end automatic speech recognition achieved state-of-the-art performance. The Conformer block leverages a self-attention mechanism to capture global information, along with a convolutional neural network to capture local information, resulting in improved performance. However, the Conformer-based model encounters an issue with the self-attention mechanism, as computational complexity grows quadratically with the length of the input sequence. Inspired by previous Connectionist Temporal Classification (CTC) guided blank skipping during decoding, we introduce intermediate CTC outputs as guidance into the downsampling procedure of the Conformer encoder. We define the frame with non-blank output as key frame. Specifically, we introduce the key frame-based self-attention (KFSA) mechanism, a novel method to reduce the computation of the self-attention mechanism using key frames. The structure of our proposed approach comprises two encoders. Following the initial encoder, we introduce an intermediate CTC loss function to compute the label frame, enabling us to extract the key frames and blank frames for KFSA. Furthermore, we introduce the key frame-based downsampling (KFDS) mechanism to operate on high-dimensional acoustic features directly and drop the frames corresponding to blank labels, which results in new acoustic feature sequences as input to the second encoder. By using the proposed method, which achieves comparable or higher performance than vanilla Conformer and other similar work such as Efficient Conformer. Meantime, our proposed method can discard more than 60% useless frames during model training and inference, which will accelerate the inference speed significantly.
Changhao Shan, Sining Sun, Qing Yang 0033, Jianwei Zhang 0013
IEEE Signal Process. Lett.5
2023 A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches
abstract
Automatic spoken instruction understanding (SIU) of the controller-pilot conversations in the air traffic control (ATC) requires not only recognizing the words and semantics of the speech but also determining the role of the speaker. However, few of the published works on the automatic understanding systems in air traffic communication focus on speaker role identification (SRI). In this article, we formulate the SRI task of controller-pilot communication as a binary classification problem. Furthermore, the text-based, speech-based, and speech-and-text-based multi-modal methods are proposed to achieve a comprehensive comparison of the SRI task. To ablate the impacts of the comparative approaches, various advanced neural network architectures are applied to optimize the implementation of text-based and speech-based methods. Most importantly, a multi-modal speaker role identification network (MMSRINet) is designed to achieve the SRI task by considering both the speech and textual modality features. To aggregate modality features, the modal fusion module is proposed to fuse and squeeze acoustic and textual representations by modal attention mechanism and self-attention pooling layer, respectively. Finally, the comparative approaches are validated on the ATCSpeech corpus collected from a real-world ATC environment. The experimental results demonstrate that all the comparative approaches worked for the SRI task, and the proposed MMSRINet shows competitive performance and robustness compared with the other methods on both seen and unseen data, achieving 98.56% and 98.08% accuracy, respectively.
Dongyue Guo, Jianwei Zhang 0013, Bo Yang 0063, Yi Lin 0006
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 FlightBERT: Binary Encoding Representation for Flight Trajectory Prediction
abstract
Flight Trajectory Prediction (TP) is an essential task in Air Traffic Control (ATC). Currently, the TP task is usually achieved by regression approaches, which concatenates several scalar attributes of the observation into a low-dimensional vector as the inputs. However, it is difficult to accurately model aircraft motion patterns using low-dimensional features in complex and time-varying ATC environments. To improve the performance of the TP task, in this paper, a novel framework, called FlightBERT, is proposed based on Binary Encoding (BE) representation, which enables us to tackle the TP task as a multi binary classification problem. Specifically, the scalar attributes of the flight trajectory are encoded into binary codes and transformed into a high-dimensional representation by the attribute embedding module. Considering the prior knowledge among flight attributes, an Attribute Correlation Attention (ACoAtt) block is designed to explicitly capture the correlations among the specific attributes. A stacked Transformer block is applied to serve as the backbone network, which is followed by the predictor to generate the outputs. Considering the nature of flight trajectory, a hybrid constrained loss, i.e., combining the mean square error loss with the binary cross-entropy loss, is innovatively designed to optimize the proposed framework. The proposed method is validated on a large-scale dataset, which is collected from the real-world ATC environment. The experimental results demonstrate that the proposed method outperforms other baselines by quantitative and qualitative evaluations.
Dongyue Guo, Qi Wu 0003, Jianwei Zhang 0013, Rob Law 0001, Yi Lin 0006
IEEE Trans. Intell. Transp. Syst.4
2023 DHI-GAN: Improving Dental-Based Human Identification Using Generative Adversarial Networks
abstract
In this work, a novel semisupervised framework is proposed to tackle the small-sample problem of dental-based human identification (DHI), achieving enhanced performance via a "classifying while generating" paradigm. A generative adversarial network (GAN), called the DHI-GAN, is presented to implement this idea, in which an extra classifier is also dedicatedly proposed to achieve an efficient training procedure. Considering the complex specificities of this problem, except for the noise input of the generator, an identity embedding-guided architecture is proposed to retain informative features for each individual. A parallel spatial and channel fusion attention block is innovatively designed to encourage the model to learn discriminative and informative features by focusing on different regional details and abstract concepts. The attention block is also widely applied to the overall classifier to learn identity-dependent information. A loss combination of the ArcFace and focal loss is utilized to address the small-sample problem. Two parameters are proposed to control the generated samples that are fed into the classifier during the optimization procedure. The proposed DHI-GAN framework is finally validated on a real-world dataset, and the experimental results demonstrate that it outperforms other baselines, achieving a 92.5% top-one accuracy rate. Most importantly, the proposed GAN-based semisupervised training strategy is able to reduce the required number of training samples (individuals) and can also be incorporated into other classification models. Our code will be available at https://github.com/sculyi/MedicalImages/.
Yi Lin 0006, Jianwei Zhang 0013, Jizhe Zhou 0001, Peixi Liao, Hu Chen 0002, Zhenhua Deng, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.3
2022 Sub-AVG: Overestimation reduction for cooperative multi-agent reinforcement learning
Jianwei Zhang 0013, Yi Lin 0006
Neurocomputing2
2022 Automatic repetition instruction generation for air traffic control training using multi-task learning with an improved copy network
Jianwei Zhang 0013, Dongyue Guo, Yang Zhou 0019, Bo Yang 0063, Yi Lin 0006
Knowl. Based Syst.1
2022 Scale-Adaptive Deep Model for Bacterial Raman Spectra Identification
abstract
The combination of Raman spectroscopy and deep learning technology provides an automatic, rapid, and accurate scheme for the clinical diagnosis of pathogenic bacteria. However, the accuracy of existing deep learning methods is still limited because of the single and fixed scales of deep neural networks. We propose a deep neural network that can learn multi-scale features of Raman spectra by using the automatic combination of multi-receptive fields of convolutional layers. This model is based on the expert knowledge that the discrimination information of Raman spectra is composed of multi-scale spectral peaks. We enhance the interpretability of the model by visualizing the activated wavenumbers of the bacterial spectrum that can be used for reference in related work. Compared with existing state-of-the-art methods, the proposed method achieves higher accuracy and efficiency for bacterial identification on isolate-level, empiric-treatment-level, and antibiotic-resistance-level tasks. The clinical bacterial identification task requires significantly fewer patient samples to achieve similar accuracy. Therefore, this method has tremendous potential for the identification of clinical pathogenic bacteria, antibiotic susceptibility testing, and prescription guidance.
Lin Deng 0003, Yuzhong Zhong, Maoning Wang, Xiujuan Zheng, Jianwei Zhang 0013
IEEE J. Biomed. Health Informatics5
2021 Heterogeneous Face Recognition with Attention-guided Feature Disentangling
abstract
This paper proposes an attention-guided feature disentangling framework (AgFD) to eliminate the large cross-modality discrepancy for Heterogeneous Face Recognition (HFR). Existing HFR methods either focus only on extracting identity features or impose linear/no independence constraints on the decomposed components. Instead, our AgFD disentangles the facial representation and forces intrinsic independence between identity features and identity-irrelevant variations. To this end, an Attention-based Residual Decomposition Module (AbRDM) and an Adversarial Decorrelation Module (ADM) are presented. AbRDM provides hierarchical complementary feature disentanglement, while ADM is introduced for decorrelation learning. Extensive experiments on the challenging CASIA NIR-VIS 2.0 Database, Oulu-CASIA NIR&VIS Database, BUAA-VisNir Database, and IIIT-D Viewed Sketch Database demonstrate the generalization ability and competitive performance of the proposed method.
Shanmin Yang, Xiao Yang 0029, Yi Lin 0006, Peng Cheng 0006, Yi Zhang 0018, Jianwei Zhang 0013
ACM Multimedia6
2021 GPU-based multi-slice per pass algorithm in interactive volume illumination rendering
abstract
Volume rendering plays a significant role in medical imaging and engineering applications. To obtain an improved three-dimensional shape perception of volumetric datasets, realistic volume illumination has been considerably studied in recent years. However, the calculation overhead associated with interactive volume rendering is unusually high, and the solvability of the problem is adversely affected when the data size and algorithm complexity are increased. In this study, a scalable and GPU-based multi-slice per pass (MSPP) volume rendering algorithm is proposed which can quickly generate global volume shadow and achieve a translucent effect based on the transfer function, so as to improve perception of the shape and depth of volumetric datasets. In our real-world data tests, MSPP significantly outperforms some complex volume shadow algorithms without losing the illumination effects, for example, half-angle slicing. Furthermore, the MSPP can be easily integrated into the parallel rendering frameworks based on sort-first or sort-last algorithms to accelerate volume rendering. In addition, its scalable slice-based volume rendering framework can be combined with several traditional volume rendering frameworks.
Dening Luo, Yi Lin 0006, Jianwei Zhang 0013
Frontiers Inf. Technol. Electron. Eng.3
2021 A Deep Learning Framework of Autonomous Pilot Agent for Air Traffic Controller Training
abstract
In this work, a deep learning-based framework is proposed to implement an autonomous pilot agent (APA), which serves as a human pseudo-pilot to assist air traffic controller (ATCO) training. A novel paradigm, including speech recognition, language understanding, pilot repetition generation (PRG), and text-to-speech (TTS), is designed to formulate the framework pipeline, which also incorporates a simulation system interface. We mainly focus on the PRG and TTS models to address the ATC specificities in this work. The neural architecture is proposed to generate the text repetition instruction by using a sequence-to-sequence text mapping. The Transformer block is improved to implement a high-efficient TTS model, in which the nonautoregressive mechanism is applied to achieve the parallel synthesis. A dedicated phoneme vocabulary is designed to cope with the multilingual issue in the ATC domain and address the out-of-vocabulary problem. With the APA framework, a virtual training mode is proposed to complete the training task without the limitation of time and location. Experimental results on a real-world dataset show that the proposed APA framework replaces the human pilot with considerable high confidence in a real-time manner during the simulation training. Most importantly, the APA framework and the virtual training system are able to cope with the dilemma of physical attendance (like COVID-19) and improve the equipment utilization capacity for the ATCO training.
Yi Lin 0006, Dongyue Guo, Changyu Yin, Bo Yang 0063, Jianwei Zhang 0013
IEEE Trans. Hum. Mach. Syst.7
2021 A Unified Framework for Multilingual Speech Recognition in Air Traffic Control Systems
abstract
This work focuses on robust speech recognition in air traffic control (ATC) by designing a novel processing paradigm to integrate multilingual speech recognition into a single framework using three cascaded modules: an acoustic model (AM), a pronunciation model (PM), and a language model (LM). The AM converts ATC speech into phoneme-based text sequences that the PM then translates into a word-based sequence, which is the ultimate goal of this research. The LM corrects both phoneme- and word-based errors in the decoding results. The AM, including the convolutional neural network (CNN) and recurrent neural network (RNN), considers the spatial and temporal dependences of the speech features and is trained by the connectionist temporal classification loss. To cope with radio transmission noise and diversity among speakers, a multiscale CNN architecture is proposed to fit the diverse data distributions and improve the performance. Phoneme-to-word translation is addressed via a proposed machine translation PM with an encoder-decoder architecture. RNN-based LMs are trained to consider the code-switching specificity of the ATC speech by building dependences with common words. We validate the proposed approach using large amounts of real Chinese and English ATC recordings and achieve a 3.95% label error rate on Chinese characters and English words, outperforming other popular approaches. The decoding efficiency is also comparable to that of the end-to-end model, and its generalizability is validated on several open corpora, making it suitable for real-time approaches to further support ATC applications, such as ATC prediction and safety checking.
Yi Lin 0006, Dongyue Guo, Jianwei Zhang 0013, Zhengmao Chen, Bo Yang 0063
IEEE Trans. Neural Networks Learn. Syst.3
2020 A Real-Time ATC Safety Monitoring Framework Using a Deep Learning Approach
abstract
A deep learning-based safety monitoring framework for air traffic control (ATC) systems is proposed in this paper to reduce human errors and relieve the controllers' workload by regulating the controlling procedure, eliminating communication misunderstanding, monitoring flight conformance, and detecting potential conflicts. The framework comprises automatic speech recognition (ASR), controlling intent inference (CII), and control safety monitoring (CSM) subsystems. The pipeline of the proposed framework can be described as follows: the ASR subsystem translates the pilot-controller voice communications (PCVCs) into texts, which are then converted to the predefined data structure by the CII subsystem. Three types of air traffic safety measures, including repetition check, flight conformance verification, and potential conflict detection, are finally validated by the CSM subsystem. An improved end-to-end ASR model with convolutional, bidirectional long short-term memory (BLSTM) and fully connected (FC) layers is trained using the connectionist temporal classification loss function. The BLSTM and FC combined CII model is designed to infer the controlling intent and slot filling. A language model is also trained in this subsystem to improve the overall performance of the framework. After converting the PCVCs to ATC data, the CSM subsystem checks the given safety monitoring tasks and sends warnings to the current system. The experimental results show that the proposed ASR model obtains a better performance than that of other approaches, and the tasks in the CII subsystem are fulfilled with a high classification precision. The CSM subsystem is also tested to confirm its safety monitoring function by playing back the data and several simulated instructions. To the best of our knowledge, this is pioneering work in the safety monitoring of flight control by recognizing the PCVCs with deep learning-based methods.
Yi Lin 0006, Linjie Deng, Zhengmao Chen, Xiping Wu, Jianwei Zhang 0013, Bo Yang 0063
IEEE Trans. Intell. Transp. Syst.5
2019 An efficient stereo matching based on fragment matching
Yingjiang Li, Jianwei Zhang 0013, Yuzhong Zhong, Maoning Wang
Vis. Comput.2
2018 An algorithm for trajectory prediction of flight plan based on relative motion between positions
abstract
Traditional methods for plan path prediction have low accuracy and stability. In this paper, we propose a novel approach for plan path prediction based on relative motion between positions (RMBP) by mining historical flight trajectories. A probability statistical model is introduced to model the stochastic factors during the whole flight process. The model object is the sequence of velocity vectors in the three-dimensional Earth space. First, we model the moving trend of aircraft including the speed (constant, acceleration, or deceleration), yaw (left, right, or straight), and pitch (climb, descent, or cruise) using a hidden Markov model (HMM) under the restrictions of aircraft performance parameters. Then, several Gaussian mixture models (GMMs) are used to describe the conditional distribution of each moving trend. Once the models are built, machine learning algorithms are applied to obtain the optimal parameters of the model from the historical training data. After completing the learning process, the velocity vector sequence of the flight is predicted by the proposed model under the Bayesian framework, so that we can use kinematic equations, depending on the moving patterns, to calculate the flight position at every radar acquisition cycle. To obtain higher prediction accuracy, a uniform interpolation method is used to correct the predicted position each second. Finally, a plan trajectory is concatenated by the predicted discrete points. Results of simulations with collected data demonstrate that this approach not only fulfils the goals of traditional methods, such as the prediction of fly-over time and altitude of waypoints along the planned route, but also can be used to plan a complete path for an aircraft with high accuracy. Experiments are conducted to demonstrate the superiority of this approach to some existing methods.
Yi Lin 0006, Jianwei Zhang 0013
Frontiers Inf. Technol. Electron. Eng.2