Yi Lin 0006

dblp:42/5120-6 · DBLP profile ↗
← Back
39ranked-venue papers
7as first author
35since 2021 · last 2027
0000-0002-7194-5023ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 An improved differential evolutionary algorithm integrating neighborhood search variation mechanism and multi-strategies for airport gate allocation
Lirong Zhang, Edmond Q. Wu, Wu Deng 0001, Jianwei Zhang 0013, Yi Lin 0006
Expert Syst. Appl.6
2026 An effective UNet using feature interaction and fusion for organ segmentation in medical image
Xiaolin Gou, Chuanlin Liao, Jizhe Zhou 0001, Fengshuo Ye, Yi Lin 0006
Eng. Appl. Artif. Intell.5
2026 Joint scheduling of runway and taxiway considering uncertain taxiing time based on an improved ACO algorithm
Lirong Zhang, Yi Lin 0006, Suwan Yin, Wu Deng 0001, Hongyu Yang 0002, Jianwei Zhang 0013
Inf. Sci.2
2026 PaTH: Patch-wise temporal hierarchical modeling for high-resolution flight trajectory prediction
Guoxin Huang, Fengshuo Ye, Yi Lin 0006, Hongyu Yang 0002, Yunxiang Han, Dongyue Guo
Pattern Recognit.3
2026 DSPNet: Dual-Path Saliency Prediction With Reconstructed Feature Injection and Discrepancy-Aware KL Loss
abstract
Visual saliency prediction (VSP) is a fundamental task in computer vision, which enables efficient sensing and computation by prioritizing informative regions, such as object detection, scene understanding, and video compression. However, the VSP task still faces three challenges: (i) fine-grained texture is irreversibly lost during repeated down-sampling in single-stream convolutional networks; (ii) traditional attention modules apply fixed, simplistic weights that adapt poorly to complex scenes; and (iii) standard Kullback-Leibler (KL) divergence fails to sufficiently amplify gradients in weakly activated regions, thereby hindering the learning of hard samples. A dual-path saliency prediction network (DSPNet) is proposed to address the aforementioned challenges. A vector-quantized variational autoencoder (VQ-VAE) branch discretizes local textures and reinjects them into a U-Net backbone to recover lost details caused by down-sampling. The restored texture is integrated with high-level semantics via a bidirectional mutual transpose attention (BMTA) decoder, enabling nonlinear cross-layer interactions and enhancing spatial discrimination in complex scenes. A discrepancy-aware focal KL (F-KL) loss is designed to optimize the model training, which adaptively emphasizes discrepant regions to improve attention to areas with prediction gaps. Evaluations on four public benchmarks (MIT300, SALICON, PASCAL-S, and TORONTO) show state-of-the-art results across metrics such as Similarity (SIM), KL divergence, and Normalized Scanpath Saliency (NSS), confirming that DSPNet effectively balances local texture fidelity with global semantic understanding.
Chuanlin Liao, Xiaolin Gou, Junyu Fan, Yi Lin 0006
IEEE Trans. Multim.4
2026 MSG-Net: Structure-Guided Enhancement for Underwater Images Based on Multi-View Feature Interaction
abstract
Underwater images often suffer from blurring, low contrast, and color distortion. Although current multi-view methods show better potential than single-input methods in mitigating these degradations, they still struggle to fuse and extract features across views. Moreover, most existing methods struggle to capture fine texture details and preserve key structural information, resulting in images with insufficient detail and unclear structures. To address these issues, we propose a Multi-View Input and Structure-Guided Network, MSG-Net, to achieve underwater image enhancement. Specifically, White Balance (WB) and Contrast Limited Adaptive Histogram Equalization (CLAHE) are employed to construct multi-view input, enabling more effective handling of diverse underwater degradation factors. A Cross-Channel Attention Fusion (CCAF) module is designed to effectively fuse these multi-view features via a channel-wise attention mechanism. Additionally, a Channel-Group Axial Transformer (CGAT) is introduced to extract axial features within channel groups, enhancing cross-channel interactions while alleviating the patch boundary discontinuities in Transformer-based methods. Furthermore, a Structure-Guided Transformer (SGT) is proposed to incorporate structural information into the enhancement process, enriching texture details and highlighting key structures. Extensive experiments demonstrate that MSG-Net outperforms or matches state-of-the-art underwater image enhancement methods across various benchmarks. Moreover, MSG-Net exhibits outstanding performance on other enhancement tasks.
Junyu Fan, Chuanlin Liao, Yi Lin 0006
IEEE Trans. Multim.5
2025 Enhancing Multivariate Time Series Forecasting with Multi-scale Moving Transformation
abstract
Recently the performance of long-term time series forecasting has been greatly improved by deep models. In this paper, we propose a general framework called Multi-scale Moving Transformation (MMT) that can be applied to the state-of-the-art models of time series forecasting tasks. Via comparative experiments, we demonstrate the effectiveness of the MMT framework against the models with the moving average. By incorporating the MMT framework with state-of-the-art backbone models based on different methods, our experimental results on various public datasets demonstrate that the improvements outperform their corresponding baseline counterparts.
Wenjie Ou, Hongmin Du, Dongyue Guo, Yi Lin 0006
ICASSP5
2025 Adversarial Feature Disentanglement Framework for Voice Pathology Detection
abstract
Voice pathology detection plays an important role in diagnosis and medical intervention. Existing methods suffer from inferior performance with limited and imbalanced training samples since the pathological information is coupled with other linguistic and paralinguistic attributes. In this work, a novel adversarial feature disentanglement framework is proposed to achieve the voice pathology detection task by extracting task-oriented features and optimizing feature space. Specifically, to suppress noise features in the coupled representations, an adversarial feature disentanglement mechanism is proposed to decouple pathological and non-pathological information, in which a mutual information discriminator is introduced to prevent information leakage. A classification and contrastive learning (CCL) module is designed to cluster intra-class embeddings in high-dimensional space. Experiments on the open-source SVD and FEMH datasets demonstrate that the proposed model outperforms other competitive baselines, achieving 87.94% and 92.06% accuracy, respectively. The visualization also validates the effectiveness in distinguishing pathological and non-pathological features.
Dongyue Guo, Lipeng Shen, Wei Mo, Yi Lin 0006
ICASSP6
2025 MFA-Net: A Multi-Stage Network for Facial Acupoint Localization with Global-Local Feature Fusion and Acupoint Encoding
abstract
Automatic acupoint localization (AAL) combines Traditional Chinese Medicine (TCM) with modern scientific technology, enhancing the popularization of TCM and promoting its modernization. However, current AAL models have low accuracy, which limits their widespread application. To address this issue, an image-based global-to-local multi-stage facial acupoint localization network (MFA-Net) is proposed, which contains a global localization stage (GLS) and a local multi-stage localization stage (LMLS). GLS determines the approximate position of an acupoint on an image, while LMLS adjusts the offset of the acupoint within the local region. Moreover, to validate the effectiveness of MFA-Net and resolve the lack of baselines in selected facial acupoint dataset, several recent state-of-the-art (SOTA) models from the field of facial landmark detection are applied as baselines for facial acupoint localization. Experiments show that MFA-Net performs best compared to all baselines.
Chuanlin Liao, Yi Lin 0006
ICME4
2025 AV-RISE: Hierarchical Cross-Modal Denoising for Learning Robust Audio-Visual Speech Representation
abstract
Audio-visual speech recognition (AVSR) leverages complementary visual cues to improve speech recognition. However, in real-world scenarios, both modalities may suffer from noise or occlusion. In such scenarios, most existing fusion strategies overlook the variation in modality-specific quality under different degradation conditions. This limitation may lead to dominance of corrupted modality in the fusion process, resulting in worse AVSR performance than unimodal systems, termed as Corrupted Modality Bias (CMB) in this work. To address this, a self-supervised speech representation learning framework, called AV-RISE, is proposed to employ teacher-student self-distillation to robustly reconstruct clean speech representations from corrupted audio-visual inputs. A hierarchical fusion mechanism is designed to progressively refine audio and visual representations by integrating the Suppression and Enhancement Interaction (SEI) module into each layer of the pre-trained encoder. In the SEI module, cross-modal suppression and modality-oriented enhancement are performed to mitigate noise-induced feature inconsistencies, which strengthens the modeling of complementary semantic representations. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that AV-RISE outperforms SOTA AVSR models, especially under extreme degradation. Most importantly, the hierarchical SEI-based fusion effectively enhances reliable semantic representations to mitigate CMB, by evaluating feature similarities between clean and noise samples.
Zhishuo Zhao, Yi Lin 0006, Dongyue Guo, Junyu Fan
ACM Multimedia2
2025 Enhancing air traffic control: A transparent deep reinforcement learning framework for autonomous conflict resolution
Hongyu Yang 0002, Yi Lin 0006, Suwan Yin
Expert Syst. Appl.3
2025 FAcupoint: The first dense facial acupoint localization dataset and baselines
Jizhe Zhou 0001, Hongyu Yang 0002, Yi Lin 0006
Expert Syst. Appl.5
2025 MLFINet : A multi-level feature interaction 3D medical image segmentation network
Chuanlin Liao, Xiaolin Gou, Kemal Polat, Jingchun Zhou, Yi Lin 0006
Neurocomputing5
2025 Improved YOLOv7 for small object detection in airports: Task-oriented feature learning with Gaussian Wasserstein loss and attention mechanisms
Ruijie Peng, Chuanlin Liao, Weijun Pan, Xiaolin Gou, Jianwei Zhang 0013, Yi Lin 0006
Neurocomputing6
2025 See Through Water: Heuristic Modeling Toward Color Correction for Underwater Image Enhancement
abstract
Color cast is one of the main degradations in underwater images. Existing data-driven methods, while capable of learning color correction rules from large datasets, often overlook the imaging characteristics and light behavior in underwater environments, making them unable to accurately restore colors in complex water bodies. To address this, we use color constancy and an underwater imaging model to heuristically model the underwater environment for accurate color restoration. On one hand, we propose a multi-scale joint prior network architecture to fully explore the rich feature-level information at different scales in underwater images. This is used to fit the complex parameters of the underwater imaging model, deriving high-quality potential undegraded images. On the other hand, to tackle the challenges of color distortion caused by complex imaging factors in different water environments, we estimate the background light of the water body through the color constancy of underwater objects and dynamically incorporate it into the underwater imaging model as a prior. This not only guides the learning process more effectively but also allows the model to consider key aspects of underwater optical propagation, making it adaptable to different water environments and improving the color accuracy of the enhanced images. We have also conducted extensive experiments to demonstrate the effectiveness of the proposed method, which not only achieves the best overall performance in qualitative analysis and quantitative comparison but also boasts the best color accuracy and the fastest inference speed. The code is available athttps://github.com/JunyuFan/MJPNet.
Junyu Fan, Jingchun Zhou, Danling Meng, Yi Lin 0006
IEEE Trans. Circuits Syst. Video Technol.5
2025 Multi-Strategy Quantum Differential Evolution Algorithm With Cooperative Co-Evolution and Hybrid Search for Capacitated Vehicle Routing
abstract
Capacitated Vehicle Routing Problem (CVRP) is a critical challenge in logistics optimization, which directly impact operational costs and service efficiency. While quantum differential evolution (QDE) algorithm offers potential advantages in solving combinatorial optimization problems, its application in CVRP is still limited due to the premature convergence, poor search capability and stagnation. To address these limitations, a novel multi-strategy QDE algorithm with cooperative co-evolution (CC) framework and hybrid local search strategy, namely MSCFLQDE is proposed to effectively solve the CVRP. Firstly, a new multi-population strategy with CC framework is designed to solve each sub-CVRP for enabling parallel optimization and preserving global constraints. Then an adaptive differential mutation mechanism is developed to balance the exploration and exploitation and accelerate the convergence. Thirdly, a new quantum rotation mode with the sorting coding rule is designed to adjust the search direction and reduce stagnation. In the later stage, a hybrid local search strategy is proposed to dynamically eliminate the redundant nodes and intersections. Finally, the experiment results on the five CVRPs (set A, set B, set P, set E, and set G) demonstrate that the MSCFLQDE has better search ability, higher convergence and stronger stability by comparing with the state-of-the-art algorithms(such as CCDE, CCDE-D, CCDE-R, CCDE-S, HGS, BILA, AGA-ES and TAMLS and so on), which achieves 5.53% shorter distances for P51_K10 by comparing with AGA-ES.
Wu Deng 0001, Shifan Shang, Lirong Zhang, Yi Lin 0006, Huimin Zhao 0002, Xiaojuan Ran, Xiangbing Zhou, Huiling Chen 0001
IEEE Trans. Intell. Transp. Syst.4
2025 A Non-Autoregressive Multi-Horizon Flight Trajectory Prediction Framework With Gray Code Representation
abstract
Flight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing methods generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error accumulation and low-efficiency problems. In this paper, a novel framework, called FlightBERT++, is proposed to i) forecast multi-horizon flight trajectories directly in a non-autoregressive way, and ii) improve the limitation of the binary encoding (BE) representation in the FlightBERT framework. Specifically, the proposed framework is implemented by a generalized encoder-decoder architecture, in which the encoder learns the temporal-spatial patterns from historical observations and the decoder predicts the flight status for the future horizons. Compared to conventional architecture, an innovative horizon-aware context generator is dedicatedly designed to consider the prior horizon information, which further enables non-autoregressive multi-horizon prediction. Additionally, the Gray code representation and the differential prediction paradigm are designed to cope with the high-bit misclassifications of the BE representation, which significantly reduces the outliers in the predictions. Moreover, a differential prompted decoder is proposed to enhance the capability of the differential predictions by leveraging the stationarity of the differential sequence. Extensive experiments are conducted to validate the proposed framework on a real-world flight trajectory dataset. The experimental results demonstrated that the proposed framework outperformed the competitive baselines in both FTP performance and computational efficiency. The code is publicly available at: https://github.com/gdy-scu/FlightBERT_PP_V2
Dongyue Guo, Fengshuo Ye, Jianwei Zhang 0013, Hongyu Yang 0002, Yi Lin 0006
IEEE Trans. Intell. Transp. Syst.7
2025 Exploring Contextual Knowledge-Enhanced Speech Recognition in Air Traffic Control Communication: A Comparative Study
abstract
Accurate recognition of named entities from spoken instructions remains a significant challenge for automatic speech recognition (ASR) techniques in air traffic control (ATC), which limits the reliability of ASR-based applications. A promising solution to overcome this challenge is to integrate prior contextual knowledge into ASR since it contains rich named entities used in ATC communications. Although existing studies have investigated ATC-related contextual ASR techniques, there is a lack of benchmarks to evaluate the advantages of different approaches. In this article, a comprehensive comparative study is presented to explore effective contextual ASR approaches for the ATC domain. Specifically, several typical contextual ASR approaches are introduced in ATC to conduct a comprehensive comparison. Moreover, a novel contextual ASR model, denoted CATCNet, is presented to dedicatedly address the domain-specific problems in ATC, such as limited resources, fast speech, and volatile noise. Several evaluation metrics are proposed to validate the performance of comparison approaches based on the practical requirements of ATC efforts. Extensive experiments are conducted across two real-world ATC speech corpora to build the benchmark. The experimental results demonstrated that integrating context knowledge is effective in improving the recognition performance of named entities. Crucially, the proposed CATCNet outperforms other baseline models by confirming all technical improvements, achieving 80.0% and 86.54% instruction recognition accuracy (IRA) on the ATCSpeech and C-ATCSpeech corpora, respectively. It is believed that this work not only overcomes the bottleneck of ASR performance in the ATC domain, but also provides an applicable solution for ATC-related ASR applications.
Dongyue Guo, Jianwei Zhang 0013, Bo Yang 0063, Yi Lin 0006
IEEE Trans. Neural Networks Learn. Syst.5
2024 FlightBERT++: A Non-autoregressive Multi-Horizon Flight Trajectory Prediction Framework
abstract
Flight Trajectory Prediction (FTP) is an essential task in Air Traffic Control (ATC), which can assist air traffic controllers in managing airspace more safely and efficiently. Existing approaches generally perform multi-horizon FTP tasks in an autoregressive manner, thereby suffering from error accumulation and low-efficiency problems. In this paper, a novel framework, called FlightBERT++, is proposed to i) forecast multi-horizon flight trajectories directly in a non-autoregressive way, and ii) improve the limitation of the binary encoding (BE) representation in the FlightBERT. Specifically, the FlightBERT++ is implemented by a generalized encoder-decoder architecture, in which the encoder learns the temporal-spatial patterns from historical observations and the decoder predicts the flight status for the future horizons. Compared with conventional architecture, an innovative horizon-aware contexts generator is dedicatedly designed to consider the prior horizon information, which further enables non-autoregressive multi-horizon prediction. Moreover, a differential prompted decoder is proposed to enhance the capability of the differential predictions by leveraging the stationarity of the differential sequence. The experimental results on a real-world dataset demonstrated that the FlightBERT++ outperformed the competitive baselines in both FTP performance and computational efficiency.
Dongyue Guo, Jianwei Zhang 0013, Yi Lin 0006
AAAI5
2024 AMG-AVSR: Adaptive Modality Guidance for Audio-Visual Speech Recognition via Progressive Feature Enhancement
Zhishuo Zhao, Dongyue Guo, Wenjie Ou, Yi Lin 0006
ACML5
2024 ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
abstract
Radio speech echo is a specific phenomenon in the air traffic control (ATC) domain, which degrades speech quality and further impacts automatic speech recognition (ASR) accuracy. In this work, a time-domain recognition-oriented speech enhancement (ROSE) framework is proposed to improve speech intelligibility and also advance ASR accuracy based on convolutional encoder-decoder-based U-Net framework, which serves as a plug-and-play tool in ATC scenarios and does not require additional retraining of the ASR model. Specifically, 1) In the U-Net architecture, an attention-based skip-fusion (ABSF) module is applied to mine shared features from encoders using an attention mask, which enables the model to effectively fuse the hierarchical features. 2) A channel and sequence attention (CSAtt) module is innovatively designed to guide the model to focus on informative features in dual parallel attention paths, aiming to enhance the effective representations and suppress the interference noises. 3) Based on the handcrafted features, ASR-oriented optimization targets are designed to improve recognition performance in the ATC environment by learning robust feature representations. By incorporating both the SE-oriented and ASR-oriented losses, ROSE is implemented in a multi-objective learning manner by optimizing shared representations across the two task objectives. The experimental results show that the ROSE significantly outperforms other state-of-the-art methods for both the SE and ASR tasks, in which all the proposed improvements are confirmed by designed experiments. In addition, the proposed approach can contribute to the desired performance improvements on public datasets.
Xincheng Yu, Dongyue Guo, Jianwei Zhang 0013, Yi Lin 0006
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Long-Term Airport Network Performance Forecasting With Linear Diffusion Graph Networks
abstract
Precise forecasting of airport performances, such as landing rates and delays, is essential for the smooth operation of air traffic management systems and for improving the passenger experience. While current efforts predominantly address short-term predictions, the imperative for long-term forecasting is undeniable, particularly for strategic operational planning and resource management. Equally important is the explainability of these forecasts, which is critical for effective decision-making. To meet these needs, our study introduces an innovative approach to airport performance forecasting with the Linear-Diffusion Graph Network (LDGN), an explainable and probabilistic model. The LDGN is intricately structured, comprising stacked temporal linear layers and graph diffusion layers that harness the clarity of linear time series models. This configuration adeptly captures the nuanced interactions between graph-based diffusion processes and the dynamic spread of conditions across airport performances. Departing from conventional point forecasts, the LDGN produces a probabilistic output, prioritizing predictability and a strong capacity for generalization. The model’s pre-training is enhanced with stochastic mask reconstruction, a technique that significantly improves its ability to generalize. Through rigorous testing on real-world datasets, we have validated the LDGN’s superior performance in both long-term and very long-term forecasting. Our results demonstrate not only high accuracy and explainability but also a robust capacity for uncertainty quantification.
Jing Yang 0017, Yi Lin 0006, Hongyu Yang 0002
IEEE Trans. Intell. Transp. Syst.4
2024 Spatiotemporal Propagation Learning for Network-Wide Flight Delay Prediction
abstract
Accurate and interpretable delay predictions are vital for decision-making in the aviation industry. However, effectively incorporating spatiotemporal dependencies and external factors related to delay propagation remains a challenge. To address this challenge, we propose the SpatioTemporal Propagation Network (STPN), a novel space-time separable graph convolutional network that models delay propagation by considering both spatial and temporal factors. STPN uses a multi-graph convolution model that considers both geographic proximity and airline schedules from a spatial perspective, while employing a multi-head self-attention mechanism that can be learned end-to-end and explicitly accounts for various types of temporal dependencies in delay time series from a temporal perspective. Experiments on two real-world delay datasets show that STPN outperforms state-of-the-art methods for multi-step ahead arrival and departure delay prediction in large-scale airport networks. Additionally, the counterfactuals generated by STPN provide evidence of its ability to learn explainable delay propagation patterns. Comprehensive experiments also demonstrate that STPN sets a robust benchmark for general spatiotemporal forecasting. The code for STPN is available athttps://github.com/Kaimaoge/STPN.
Hongyu Yang 0002, Yi Lin 0006
IEEE Trans. Knowl. Data Eng.3
2023 M2ATS: A Real-world Multimodal Air Traffic Situation Benchmark Dataset and Beyond
abstract
Air Traffic Control (ATC) is a complicated, time-evolving, and real-time procedure to direct flight operations in a safer and ordered manner. Although enormous data storages are available during air traffic operations for over 40 years, data-driven intelligent application in aviation is still an emerging task due to the safety-critical issue. With the prevalence of the Next Generation ATC system, artificial intelligence (AI) -empowered research topics are attracting increasing attention from both industrial and academic domains and a high-quality dataset naturally becomes the prerequisite for such practices. However, almost all ATC-related datasets are only unimodal for certain tasks, which fails to comprehensively illustrate the traffic situation to further support real-world studies. To address this gap, a multimodal air traffic situation (M2ATS) dataset is constructed to advance AI-related research in the ATC domain, including airspace information, flight plan, trajectory, and speech. M2ATS covers 10362 flights ATC situation data, involving 110000+ utterances (104 hours) with diversity golden text annotations, 16 intents, and 51 slots. Considering the real-world ATC requirements, a total of 10 multimedia-related tasks (24 baselines) are designed to validate the proposed dataset, covering automatic speech recognition, natural language processing, and spatial-temporal data processing. New ATC-related metrics corresponding to ATC applications are proposed in addition to the common metrics to evaluate task performance. Extensive experiment results demonstrate that the selective baselines can achieve designed tasks on this new dataset, and further investigations are also required to address task and data specificities. It is believed that the proposed new dataset is a new practice to advance AI applications to an industrial scene, which not only promotes ATC-related applications but also provides diverse research topics in the common multimedia community.
Dongyue Guo, Yi Lin 0006, Xuehang You, Zhongping Yang, Jizhe Zhou 0001, Bo Yang 0063, Jianwei Zhang 0013, Shasha Hu
ACM Multimedia2
2023 A Comparative Study of Speaker Role Identification in Air Traffic Communication Using Deep Learning Approaches
abstract
Automatic spoken instruction understanding (SIU) of the controller-pilot conversations in the air traffic control (ATC) requires not only recognizing the words and semantics of the speech but also determining the role of the speaker. However, few of the published works on the automatic understanding systems in air traffic communication focus on speaker role identification (SRI). In this article, we formulate the SRI task of controller-pilot communication as a binary classification problem. Furthermore, the text-based, speech-based, and speech-and-text-based multi-modal methods are proposed to achieve a comprehensive comparison of the SRI task. To ablate the impacts of the comparative approaches, various advanced neural network architectures are applied to optimize the implementation of text-based and speech-based methods. Most importantly, a multi-modal speaker role identification network (MMSRINet) is designed to achieve the SRI task by considering both the speech and textual modality features. To aggregate modality features, the modal fusion module is proposed to fuse and squeeze acoustic and textual representations by modal attention mechanism and self-attention pooling layer, respectively. Finally, the comparative approaches are validated on the ATCSpeech corpus collected from a real-world ATC environment. The experimental results demonstrate that all the comparative approaches worked for the SRI task, and the proposed MMSRINet shows competitive performance and robustness compared with the other methods on both seen and unseen data, achieving 98.56% and 98.08% accuracy, respectively.
Dongyue Guo, Jianwei Zhang 0013, Bo Yang 0063, Yi Lin 0006
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2023 Towards Recognition for Radio-Echo Speech in Air Traffic Control: Dataset and a Contrastive Learning Approach
abstract
In the air traffic control (ATC) domain, automatic speech recognition (ASR) suffers from radio speech echo, which cannot be addressed by existing echo cancellation due to auditory-oriented optimization and poor generalization ability caused by volatile radio transmission. In this work, a contrastive learning-based framework is proposed to tackle the radio-echo speech for the ASR task based on convolution networks with multiple paths and recurrent neural networks. 1) By analyzing the communication mechanism of the ATC speech, a novel transmission method is designed to collect clean and noisy speech samples (with the same texts) via a bypass device in a real-world ATC environment. 2) To enhance the model capacity, a temporal and frequency attention block is innovatively designed to guide the model to focus on informative frames and frequencies, aiming at learning shared representations between the clean and noisy speech signals with the same texts. 3) By incorporating contrastive loss, the proposed approach is implemented by a multi-objective optimization, in which the loss weights are dynamically determined to enhance the ASR performance in a learnable manner. With the proposed transmission method, a real-world dataset is collected and annotated to validate the proposed approach. Experimental results demonstrate that the proposed approach outperforms other comparative baselines with different technical frameworks, achieving a 6.76% character error rate on the test dataset. Most importantly, all the proposed improvements are confirmed by designed experiments, in which contrastive learning with learnable multi-objective loss weights contributes to the primary performance improvement.
Yi Lin 0006, Xincheng Yu, Zichen Zhang 0020, Dongyue Guo, Jizhe Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 FlightBERT: Binary Encoding Representation for Flight Trajectory Prediction
abstract
Flight Trajectory Prediction (TP) is an essential task in Air Traffic Control (ATC). Currently, the TP task is usually achieved by regression approaches, which concatenates several scalar attributes of the observation into a low-dimensional vector as the inputs. However, it is difficult to accurately model aircraft motion patterns using low-dimensional features in complex and time-varying ATC environments. To improve the performance of the TP task, in this paper, a novel framework, called FlightBERT, is proposed based on Binary Encoding (BE) representation, which enables us to tackle the TP task as a multi binary classification problem. Specifically, the scalar attributes of the flight trajectory are encoded into binary codes and transformed into a high-dimensional representation by the attribute embedding module. Considering the prior knowledge among flight attributes, an Attribute Correlation Attention (ACoAtt) block is designed to explicitly capture the correlations among the specific attributes. A stacked Transformer block is applied to serve as the backbone network, which is followed by the predictor to generate the outputs. Considering the nature of flight trajectory, a hybrid constrained loss, i.e., combining the mean square error loss with the binary cross-entropy loss, is innovatively designed to optimize the proposed framework. The proposed method is validated on a large-scale dataset, which is collected from the real-world ATC environment. The experimental results demonstrate that the proposed method outperforms other baselines by quantitative and qualitative evaluations.
Dongyue Guo, Qi Wu 0003, Jianwei Zhang 0013, Rob Law 0001, Yi Lin 0006
IEEE Trans. Intell. Transp. Syst.6
2023 DHI-GAN: Improving Dental-Based Human Identification Using Generative Adversarial Networks
abstract
In this work, a novel semisupervised framework is proposed to tackle the small-sample problem of dental-based human identification (DHI), achieving enhanced performance via a "classifying while generating" paradigm. A generative adversarial network (GAN), called the DHI-GAN, is presented to implement this idea, in which an extra classifier is also dedicatedly proposed to achieve an efficient training procedure. Considering the complex specificities of this problem, except for the noise input of the generator, an identity embedding-guided architecture is proposed to retain informative features for each individual. A parallel spatial and channel fusion attention block is innovatively designed to encourage the model to learn discriminative and informative features by focusing on different regional details and abstract concepts. The attention block is also widely applied to the overall classifier to learn identity-dependent information. A loss combination of the ArcFace and focal loss is utilized to address the small-sample problem. Two parameters are proposed to control the generated samples that are fed into the classifier during the optimization procedure. The proposed DHI-GAN framework is finally validated on a real-world dataset, and the experimental results demonstrate that it outperforms other baselines, achieving a 92.5% top-one accuracy rate. Most importantly, the proposed GAN-based semisupervised training strategy is able to reduce the required number of training samples (individuals) and can also be incorporated into other classification models. Our code will be available at https://github.com/sculyi/MedicalImages/.
Yi Lin 0006, Jianwei Zhang 0013, Jizhe Zhou 0001, Peixi Liao, Hu Chen 0002, Zhenhua Deng, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.1
2022 Sub-AVG: Overestimation reduction for cooperative multi-agent reinforcement learning
Jianwei Zhang 0013, Yi Lin 0006
Neurocomputing4
2022 Automatic repetition instruction generation for air traffic control training using multi-task learning with an improved copy network
Jianwei Zhang 0013, Dongyue Guo, Yang Zhou 0019, Bo Yang 0063, Yi Lin 0006
Knowl. Based Syst.7
2021 Heterogeneous Face Recognition with Attention-guided Feature Disentangling
abstract
This paper proposes an attention-guided feature disentangling framework (AgFD) to eliminate the large cross-modality discrepancy for Heterogeneous Face Recognition (HFR). Existing HFR methods either focus only on extracting identity features or impose linear/no independence constraints on the decomposed components. Instead, our AgFD disentangles the facial representation and forces intrinsic independence between identity features and identity-irrelevant variations. To this end, an Attention-based Residual Decomposition Module (AbRDM) and an Adversarial Decorrelation Module (ADM) are presented. AbRDM provides hierarchical complementary feature disentanglement, while ADM is introduced for decorrelation learning. Extensive experiments on the challenging CASIA NIR-VIS 2.0 Database, Oulu-CASIA NIR&VIS Database, BUAA-VisNir Database, and IIIT-D Viewed Sketch Database demonstrate the generalization ability and competitive performance of the proposed method.
Shanmin Yang, Xiao Yang 0029, Yi Lin 0006, Peng Cheng 0006, Yi Zhang 0018, Jianwei Zhang 0013
ACM Multimedia3
2021 Improving speech recognition models with small samples for air traffic control systems
Yi Lin 0006, Bo Yang 0063, Huachun Tan, Zhengmao Chen
Neurocomputing1
2021 GPU-based multi-slice per pass algorithm in interactive volume illumination rendering
abstract
Volume rendering plays a significant role in medical imaging and engineering applications. To obtain an improved three-dimensional shape perception of volumetric datasets, realistic volume illumination has been considerably studied in recent years. However, the calculation overhead associated with interactive volume rendering is unusually high, and the solvability of the problem is adversely affected when the data size and algorithm complexity are increased. In this study, a scalable and GPU-based multi-slice per pass (MSPP) volume rendering algorithm is proposed which can quickly generate global volume shadow and achieve a translucent effect based on the transfer function, so as to improve perception of the shape and depth of volumetric datasets. In our real-world data tests, MSPP significantly outperforms some complex volume shadow algorithms without losing the illumination effects, for example, half-angle slicing. Furthermore, the MSPP can be easily integrated into the parallel rendering frameworks based on sort-first or sort-last algorithms to accelerate volume rendering. In addition, its scalable slice-based volume rendering framework can be combined with several traditional volume rendering frameworks.
Dening Luo, Yi Lin 0006, Jianwei Zhang 0013
Frontiers Inf. Technol. Electron. Eng.2
2021 A Deep Learning Framework of Autonomous Pilot Agent for Air Traffic Controller Training
abstract
In this work, a deep learning-based framework is proposed to implement an autonomous pilot agent (APA), which serves as a human pseudo-pilot to assist air traffic controller (ATCO) training. A novel paradigm, including speech recognition, language understanding, pilot repetition generation (PRG), and text-to-speech (TTS), is designed to formulate the framework pipeline, which also incorporates a simulation system interface. We mainly focus on the PRG and TTS models to address the ATC specificities in this work. The neural architecture is proposed to generate the text repetition instruction by using a sequence-to-sequence text mapping. The Transformer block is improved to implement a high-efficient TTS model, in which the nonautoregressive mechanism is applied to achieve the parallel synthesis. A dedicated phoneme vocabulary is designed to cope with the multilingual issue in the ATC domain and address the out-of-vocabulary problem. With the APA framework, a virtual training mode is proposed to complete the training task without the limitation of time and location. Experimental results on a real-world dataset show that the proposed APA framework replaces the human pilot with considerable high confidence in a real-time manner during the simulation training. Most importantly, the APA framework and the virtual training system are able to cope with the dilemma of physical attendance (like COVID-19) and improve the equipment utilization capacity for the ATCO training.
Yi Lin 0006, Dongyue Guo, Changyu Yin, Bo Yang 0063, Jianwei Zhang 0013
IEEE Trans. Hum. Mach. Syst.1
2021 A Unified Framework for Multilingual Speech Recognition in Air Traffic Control Systems
abstract
This work focuses on robust speech recognition in air traffic control (ATC) by designing a novel processing paradigm to integrate multilingual speech recognition into a single framework using three cascaded modules: an acoustic model (AM), a pronunciation model (PM), and a language model (LM). The AM converts ATC speech into phoneme-based text sequences that the PM then translates into a word-based sequence, which is the ultimate goal of this research. The LM corrects both phoneme- and word-based errors in the decoding results. The AM, including the convolutional neural network (CNN) and recurrent neural network (RNN), considers the spatial and temporal dependences of the speech features and is trained by the connectionist temporal classification loss. To cope with radio transmission noise and diversity among speakers, a multiscale CNN architecture is proposed to fit the diverse data distributions and improve the performance. Phoneme-to-word translation is addressed via a proposed machine translation PM with an encoder-decoder architecture. RNN-based LMs are trained to consider the code-switching specificity of the ATC speech by building dependences with common words. We validate the proposed approach using large amounts of real Chinese and English ATC recordings and achieve a 3.95% label error rate on Chinese characters and English words, outperforming other popular approaches. The decoding efficiency is also comparable to that of the end-to-end model, and its generalizability is validated on several open corpora, making it suitable for real-time approaches to further support ATC applications, such as ATC prediction and safety checking.
Yi Lin 0006, Dongyue Guo, Jianwei Zhang 0013, Zhengmao Chen, Bo Yang 0063
IEEE Trans. Neural Networks Learn. Syst.1
2020 ATCSpeech: A Multilingual Pilot-Controller Speech Corpus from Real Air Traffic Control Environment
abstract
Automatic Speech Recognition (ASR) is greatly developed in recent years, which expedites many applications on other fields. For the ASR research, speech corpus is always an essential foundation, especially for the vertical industry, such as Air Traffic Control (ATC). There are some speech corpora for common applications, public or paid. However, for the ATC, it is difficult to collect raw speeches from real systems due to safety issues. More importantly, for a supervised learning task like ASR, annotating the transcription is a more laborious work, which hugely restricts the prospect of ASR application. In this paper, a multilingual speech corpus (ATCSpeech) from real ATC systems, including accented Mandarin Chinese and English, is built and released to encourage the non-commercial ASR research in ATC domain. The corpus is detailly introduced from the perspective of data amount, speaker gender and role, speech quality and other attributions. In addition, the performance of our baseline ASR models is also reported. A community edition for our speech database can be applied and used under a special contrast. To our best knowledge, this is the first work that aims at building a real and multilingual ASR corpus for the air traffic related research.
Bo Yang 0063, Xianlong Tan, Zhengmao Chen, Min Ruan, Zhongping Yang, Xiping Wu, Yi Lin 0006
INTERSPEECH9
2020 A Real-Time ATC Safety Monitoring Framework Using a Deep Learning Approach
abstract
A deep learning-based safety monitoring framework for air traffic control (ATC) systems is proposed in this paper to reduce human errors and relieve the controllers' workload by regulating the controlling procedure, eliminating communication misunderstanding, monitoring flight conformance, and detecting potential conflicts. The framework comprises automatic speech recognition (ASR), controlling intent inference (CII), and control safety monitoring (CSM) subsystems. The pipeline of the proposed framework can be described as follows: the ASR subsystem translates the pilot-controller voice communications (PCVCs) into texts, which are then converted to the predefined data structure by the CII subsystem. Three types of air traffic safety measures, including repetition check, flight conformance verification, and potential conflict detection, are finally validated by the CSM subsystem. An improved end-to-end ASR model with convolutional, bidirectional long short-term memory (BLSTM) and fully connected (FC) layers is trained using the connectionist temporal classification loss function. The BLSTM and FC combined CII model is designed to infer the controlling intent and slot filling. A language model is also trained in this subsystem to improve the overall performance of the framework. After converting the PCVCs to ATC data, the CSM subsystem checks the given safety monitoring tasks and sends warnings to the current system. The experimental results show that the proposed ASR model obtains a better performance than that of other approaches, and the tasks in the CII subsystem are fulfilled with a high classification precision. The CSM subsystem is also tested to confirm its safety monitoring function by playing back the data and several simulated instructions. To the best of our knowledge, this is pioneering work in the safety monitoring of flight control by recognizing the PCVCs with deep learning-based methods.
Yi Lin 0006, Linjie Deng, Zhengmao Chen, Xiping Wu, Jianwei Zhang 0013, Bo Yang 0063
IEEE Trans. Intell. Transp. Syst.1
2019 Detecting multi-oriented text with corner-based region proposals
Linjie Deng, Yanxiang Gong, Yi Lin 0006, Jingwen Shuai, Xiaoguang Tu, Yuefei Zhang, Zheng Ma 0005, Mei Xie
Neurocomputing3
2018 An algorithm for trajectory prediction of flight plan based on relative motion between positions
abstract
Traditional methods for plan path prediction have low accuracy and stability. In this paper, we propose a novel approach for plan path prediction based on relative motion between positions (RMBP) by mining historical flight trajectories. A probability statistical model is introduced to model the stochastic factors during the whole flight process. The model object is the sequence of velocity vectors in the three-dimensional Earth space. First, we model the moving trend of aircraft including the speed (constant, acceleration, or deceleration), yaw (left, right, or straight), and pitch (climb, descent, or cruise) using a hidden Markov model (HMM) under the restrictions of aircraft performance parameters. Then, several Gaussian mixture models (GMMs) are used to describe the conditional distribution of each moving trend. Once the models are built, machine learning algorithms are applied to obtain the optimal parameters of the model from the historical training data. After completing the learning process, the velocity vector sequence of the flight is predicted by the proposed model under the Bayesian framework, so that we can use kinematic equations, depending on the moving patterns, to calculate the flight position at every radar acquisition cycle. To obtain higher prediction accuracy, a uniform interpolation method is used to correct the predicted position each second. Finally, a plan trajectory is concatenated by the predicted discrete points. Results of simulations with collected data demonstrate that this approach not only fulfils the goals of traditional methods, such as the prediction of fly-over time and altitude of waypoints along the planned route, but also can be used to plan a complete path for an aircraft with high accuracy. Experiments are conducted to demonstrate the superiority of this approach to some existing methods.
Yi Lin 0006, Jianwei Zhang 0013
Frontiers Inf. Technol. Electron. Eng.1