Tengfei Yu

dblp:308/3627 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 A new mineral quantification method via experiment-enhanced transfer learning of linear mixed mid-infrared spectra data
Tengfei Yu, Zhenhao Xu
Eng. Appl. Artif. Intell.2
2026 Toward Large-Scale and Robust Indoor Positioning: Deep Learning-Augmented VLP/INS Fusion With Efficient Anchor Calibration
abstract
Visible Light Positioning (VLP) has emerged as a promising indoor localization technology owing to its high accuracy, low power consumption, and lighting compatibility. The Received Signal Strength (RSS)-based multi-anchor VLP pre-serves these advantages while having drawn considerable research attention due to its simple implementation and high reliability. However, its large-scale deployment encounters challenges at every stage: inefficient anchor calibration, limited model-based ranging performance, and robustness reduction from undetected gross errors. To address these issues, we propose a VLP and inertial navigation system fusion framework comprising an anchor position estimation module, a distance estimation module, and a fusion positioning module. For anchor calibration, a LiDAR-inertial odometry-based calibration scheme enhanced by a twolayer optimization strategy is introduced, which provides prior knowledge of anchor positions for the whole system. To improve the performance of the RSS-based ranging method, a hybrid Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN)-Bidirectional Long Short-Term Memory (Bi-LSTM) network with an embedded distance quality assessor is developed, achieving over 50% higher accuracy than model-based baselines. It delivers precise distance estimates and their validity labels for measurement updates in subsequent fusion positioning. Additionally, a two-stage error detection mechanism filters low quality observations by combining network-generated usability labels with prior-state estimates. The system consistently attains decimeter-level positioning across various trajectories, meeting the needs of diverse Internet of Things applications.
Xiaoxiang Cao, Xuan Wang 0015, Tengfei Yu, Zhenghua Zhang, Jingxue Bi, Yue Yu 0003, Yulin Hu
IEEE Trans. Mob. Comput.3
2025 A Fast and Robust Calibration Method for Lambert Model in VLP System Without Any Geometric Measurement
abstract
Visible light positioning (VLP) is one of the most promising technologies for providing high-precision, low-cost indoor positioning and navigation services. However, the traditional calibration methods of VLP are complex and unfriendly to users, which hinders the large-scale commercial deployment of VLP systems. In this paper, a fast and robust calibration method for received signal strength (RSS)-based VLP system is proposed, which dispenses with geometric measurement and greatly simplifies the calibration procedures. The proposed method calibrates the Lambert model through two steps, firstly using a ratio method to calibrate the Lambert order, and then estimating the constant term. The actual processes only require moving a robot equipped with a photo-detector (PD) along a rectangular trajectory once, and then the program will automatically estimate the required parameters by analyzing the RSS during this period. Experimental results show a good stability of the calibrated parameters, as well as a excellent distance measuring accuracy within 12 cm. The proposed method is 3.7 times more efficient than traditional methods while the positioning accuracy is close, which will greatly reduce the time and labor costs of large-scale deployment.
Tianming Huang, Yuan Zhuang 0001, Xiansheng Yang, Xiao Sun 0009, Tengfei Yu, Xiaoxiang Cao
IEEE Internet Things J.5
2025 VLP-BERT: BERT-Enhanced IMU and Visible Light Tightly Coupled Integration Positioning System
abstract
Visible Light Positioning (VLP) has emerged as a promising indoor localization technology due to its high accuracy, low cost. However, it still faces challenges such as environmental interference, signal noise, and occlusion. To address the above issues, a Bidirectional Encoder Representation from Transformer (BERT)-enhanced VLP and inertial navigation fusion positioning system is developed. Firstly, to tackle the problem of inaccurate ranging caused by signal noise, we propose a Transformer-based network, VLP-BERT, which leverages long-sequence masking to enhance the network’s feature extraction capabilities from visible light signals. Moreover, the VLP-BERT is integrated into an autoencoder-decoder architecture for signal denoising. Secondly, to overcome the limitations of traditional ranging models in complex environments, a deep learning-based centralized VLP ranging model is proposed. Finally, to enhance the system’s reliability under varying conditions, a tightly coupled fusion method integrating VLP with Pedestrian Dead Reckoning (PDR) is proposed, incorporating error detection and state-constrained strategies. Extensive experimental evaluations demonstrate the effectiveness of VLP-BERT in both denoising and accurate ranging. The system was compared with nine different methods, the results show that the proposed tightly coupled approach not only achieves sub-meter-level accuracy but also significantly enhances the system’s robustness, even in challenging scenarios such as signal blockage and poor signal quality.
Xuan Wang 0015, Xiaoxiang Cao, Tengfei Yu, Zhenqi Zheng, Zhenghua Zhang, Yue Yu 0003
IEEE Internet Things J.3
2025 DIO-VL: Deep Learning-Based Inertial Odometry and Visible Light Fusion Localization
abstract
Visible Light Positioning (VLP) has attracted significant attention due to its low cost, low power consumption and high accuracy. However, challenges such as noise interference and limited coverage still affect the system’s performance. On the other hand, Inertial Measurement Units (IMUs) can provide position measurements that are unaffected by environmental factors, but traditional methods tend to diverge over time. To solve these issues, a method for robust fusion positioning of a deep learning-enhanced pedestrian trolley odometer and VLP has been proposed. Firstly, a deep learning-based inertial odometry (DIO) is introduced to capture gyroscope noise and hidden motion features, which can achieve precise displacement estimation within a specified window size. Secondly, a tightly coupled fusion filter that integrates IMU-based odometry with received signal strength (RSS)-based VLP is designed, it significantly enhances the system’s robustness under conditions of signal sparsity and serious signal noise interference. Finally, a method for calculating the observation error covariance based on RSS that considers measurement uncertainty has been proposed, it ensures that the system remains robust even in the presence of weak signal strength or environmental noise. Experimental evaluations demonstrate that the proposed DIO improves accuracy by 33.4% compared to existing methods, with only an 0.76% increase in average computation time. Furthermore, compared to the DIO, VLP, and particle filter fusion systems, the proposed fusion-based localization system achieves accuracy improvements of 78.1%, 23.4%, and 9.6%, respectively, while maintaining significantly lower average computational time than the particle filter system. These results indicate that the proposed fusion method significantly outperforms existing techniques and effectively addresses the limitations of individual positioning methods.
Tengfei Yu, Yuan Zhuang 0001, Xuan Wang 0015, Xiaoxiang Cao
IEEE Internet Things J.1
2025 Four Typical Variation Patterns of Mid-Infrared Spectra of the Felsic Mineral Anomalies for Fault Zone Identification
abstract
Mid-infrared spectroscopy has the advantage of being fast and non-destructive for mineral identification and geological analysis. Mineral anomalies in fault zones are geological markers for fault identification. However, fault rock spectra acquired by remote sensing techniques are often a mixture of multiple endmembers. The mid-infrared spectral variability of mineral anomaly patterns in fault zones is still unclear, causing difficulties in identifying faults by analyzing mineral anomaly characteristics through mid-infrared spectroscopy. From the perspective of remote detection of faults, we constructed 14 groups of binary, ternary, and quaternary mineral assemblages of the felsic anomaly pattern in fault zones, tested the mid-infrared spectra, analyzed the influences of mineral component and content changes on spectral characteristics, and summarized four categories of the variation patterns of the mid-infrared spectra of mixed minerals: selfstabilization (SS), superimposed interference (SI), camouflage enhancement (CE), and annihilation (AH). Further, we proposed the spectral identification criteria for anomalous mineral assemblages and clarified the identification characteristics and detection limits of endmember minerals in mixed minerals. Besides, we discussed the nonlinear mixing effect and quantitative potential of the mixed spectra. The results show that the variation effect weakens the linear relationship between spectral peak intensity and mineral abundance in the mixed spectrum, making it difficult to calibrate the mineral content by the reflection peak intensity. Our new finding of the spectral reflection peak shift has a strong linear correlation with changes in mineral content, which provides new ideas for the quantitative interpretation of mid-infrared spectra of mixed minerals in remote sensing applications.
Tengfei Yu, Zhenhao Xu, Ruiqi Shao
IEEE Trans. Geosci. Remote. Sens.1
2024 Detecting and Classifying Invasive Breast Carcinoma by the Faster R-CNN with Resnet 50
abstract
In recent statistical data, compared with other types of cancer, breast cancer has been the most commonly diagnosed cancer in the world for females, and invasive breast carcinoma occupies a significant ratio in breast cancer. Therefore, in order to short the detection time consumption, we need a computer-aided, deep-learning-based invasive breast carcinoma detection and classification solution. The dataset includes around two thousand images with two unbalanced classes: benign and malignant. The accuracy of the Faster-RCNN with Resnet-50 in this dataset is 85.5% for malignant and 84.5% for benign. As for the patients, the mean AP is 99.9% that 99.1% in malignant patients and 99.0% in benign patients.
Tengfei Yu, Zhongbao Chen
ICIS2
2024 Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation
abstract
End-to-end speech translation (ST) presents notable disambiguation challenges as it necessitates simultaneous cross-modal and crosslingual transformations.While word sense disambiguation is an extensively investigated topic in textual machine translation, the exploration of disambiguation strategies for ST models remains limited.Addressing this gap, this paper introduces the concept of speech sense disambiguation (SSD), specifically emphasizing homophones -words pronounced identically but with different meanings.To facilitate this, we first create a comprehensive homophone dictionary and an annotated dataset rich with homophone information established based on speech-text alignment.Building on this unique dictionary, we introduce AmbigST, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy.Our experiments on different MuST-C and CoVoST ST benchmarks demonstrate that AmbigST sets new performance standards.Specifically, it achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity.
Tengfei Yu, Xuebo Liu 0002, Liang Ding 0006, Kehai Chen, Dacheng Tao, Min Zhang 0005
ACL (1)1
2024 Curriculum Consistency Learning for Conditional Sentence Generation
abstract
Consistency learning (CL) has proven to be a valuable technique for improving the robustness of models in conditional sentence generation (CSG) tasks by ensuring stable predictions across various input data forms.However, models augmented with CL often face challenges in optimizing consistency features, which can detract from their efficiency and effectiveness.To address these challenges, we introduce Curriculum Consistency Learning (CCL), a novel strategy that guides models to learn consistency in alignment with their current capacity to differentiate between features.CCL is designed around the inherent aspects of CL-related losses, promoting task independence and simplifying implementation.Implemented across four representative CSG tasks, including instruction tuning (IT) for large language models and machine translation (MT) in three modalities (text, speech, and vision), CCL demonstrates marked improvements.Specifically, it delivers +2.0 average accuracy point improvement compared with vanilla IT and an average increase of +0.7 in COMET scores over traditional CL methods in MT tasks.Our comprehensive analysis further indicates that models utilizing CCL are particularly adept at managing complex instances, showcasing the effectiveness and efficiency of CCL in improving CSG models.Code and scripts are available at https://github.com/xinxinxing/ Curriculum-Consistency-Learning.
Liangxin Liu, Xuebo Liu 0002, Lian Lian, Shengjun Cheng, Jun Rao, Tengfei Yu, Hexuan Deng, Min Zhang 0005
EMNLP6
2024 Self-Powered LLM Modality Expansion for Large Speech-Text Models
abstract
Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities.Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable benefits, these methods generally entail significant resource demands and tend to overfit specific tasks.This study aims to refine the use of speech datasets for LSM training by addressing the limitations of vanilla instruction tuning.We explore the instruction-following dynamics within LSMs, identifying a critical issue termed speech anchor bias-a tendency for LSMs to over-rely on speech inputs, mistakenly interpreting the entire speech modality as directives, thereby neglecting textual instructions.To counteract this bias, we introduce a self-powered LSM that leverages augmented automatic speech recognition data generated by the model itself for more effective instruction tuning.Our experiments across a range of speech-based tasks demonstrate that selfpowered LSM mitigates speech anchor bias and improves the fusion of speech and text modalities in LSMs.
Tengfei Yu, Xuebo Liu 0002, Zhiyi Hou, Liang Ding 0006, Dacheng Tao, Min Zhang 0005
EMNLP1
2024 Deep-Learning-Enhanced Visible Light Positioning System Based on the LED Array
abstract
The escalating demand for indoor location-based services has propelled advancements in indoor positioning technologies in which visible light positioning (VLP) standing out for its accuracy and eco-friendliness. Current VLP systems are typically categorized into multi-anchor-based and single-anchor-based methods. The former encounters challenges like high wiring costs, limited adaptability in narrow spaces, and the segregation of lighting and positioning. Current single-anchor methods using cameras or photodiode (PD) arrays face complex hardware design costs and application limitations. To address these issues, a deep learning-enhanced VLP method with a light-emitting diode (LED) array is proposed. Firstly, a cost-effective single-anchor VLP system is designed, utilizing a LED array and a single PD, which dramatically reduces the costs. Secondly, to mitigate the impact of various noises on the received signal, a transformer-based signal-denoising network is designed, resulting in a significant reduction in the large ranging errors. Furthermore, a distance error estimation network based on graph neural network (GNN) is introduced to address the issue of poor anchor configuration caused by LED arrays. The GNN aggregates inherent correlations among LEDs, significantly reducing ranging errors. Finally, a weighted particle swarm optimization (PSO)-based positioning algorithm is introduced to further optimize results, and GNN’s outputs are converted into weights for PSO using a self-designed weight function, enhancing accuracy and robustness. Experimental evaluations of the hardware and proposed methods demonstrate sub-meter-level positioning accuracy. Moreover, the effective coverage range is comparable to that of a regular LED fixture of the same size, suggesting its broad applications prospects in the Internet of Things (IoT) after it integrates lighting and positioning capabilities.
Xiaoxiang Cao, Yuan Zhuang 0001, Xuan Wang 0015, Tengfei Yu, Jiale Jiang
IEEE Internet Things J.4
2023 PromptST: Abstract Prompt Learning for End-to-End Speech Translation
abstract
An end-to-end speech-to-text (S2T) translation model is usually initialized from a pretrained speech recognition encoder and a pretrained text-to-text (T2T) translation decoder.Although this straightforward setting has been shown empirically successful, there do not exist clear answers to the research questions: 1) how are speech and text modalities fused in S2T model and 2) how to better fuse the two modalities?In this paper, we take the first step toward understanding the fusion of speech and text features in S2T model.We first design and release a 10GB linguistic probing benchmark, namely Speech-Senteval, to investigate the acoustic and linguistic behaviors of S2T models.Preliminary analysis reveals that the uppermost encoder layers of the S2T model can not learn linguistic knowledge efficiently, which is crucial for accurate translation.Based on the finding, we further propose a simple plug-in prompt-learning strategy on the uppermost encoder layers to broaden the abstract representation power of the encoder of S2T models.We call such a promptenhanced S2T model PromptST.Experimental results on four widely-used S2T datasets show that PromptST can deliver significant improvements over a strong baseline by capturing richer linguistic knowledge.Benchmarks,
Tengfei Yu, Liang Ding 0006, Xuebo Liu 0002, Kehai Chen, Meishan Zhang, Dacheng Tao, Min Zhang 0005
EMNLP1