Yifei Ding

dblp:194/4297 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment
abstract
Yixuan Wang, Yue Huang, Hong Qian, Yunzhao Wei, Yifei Ding, Wenkai Wang, Zhi Liu, Zhongjing Huang, Aimin Zhou, Jiajun Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hong Qian, Yunzhao Wei, Yifei Ding, Zhongjing Huang, Aimin Zhou, Jiajun Guo
ACL (1)5
2026 Deep learning-driven digital twin system for pedestrian tracking and evacuation load assessment in public spaces
abstract
Real-time pedestrian localization is essential for effective emergency evacuation in large indoor public spaces. This study presents an intelligent digital twin system for evacuation monitoring, integrating deep learning and computer vision. The system includes four components: (1) Internet of Things sensor network, (2) cloud computing server, (3) Artificial Intelligence processing engine, and (4) interactive user interface. The Artificial Intelligence engine introduces three innovations: automated detection and tracking of pedestrian coordinates using You Only Look Once-Pose (YOLO-Pose) and Deep Simple Online and Realtime Tracking (DeepSORT); transformation of multi-camera data into a unified world coordinate system; and the Multi-Object Matching Operation (MOMO) algorithm for identity association. These enable accurate detection, localization, and counting while minimizing identifiability. The system was validated in controlled experiments and a high-speed rail station waiting hall with dense, dynamic pedestrian flow. It achieves high localization precision, with a root mean square error of 5.3 cm, a mean absolute error of 4.8 cm, and a people counting accuracy of 92.34% while processing 30 frames per second video at 27.8 ms per frame. These results demonstrate the potential of the digital twin framework in intelligent evacuation management. The main contribution in Artificial Intelligence is the Multi-Object Matching Operation algorithm, and the engineering contribution is the realization of a real-time digital twin system in a large public facility. • Proposes a digital twin system integrating YOLO-Pose and DeepSORT for real-time pedestrian tracking. • Introduces a novel multi-camera calibration method for global coordinate unification. • Achieves 92.34% people counting accuracy in complex public infrastructure environments.
Huakai Sun, Yifei Ding, Ruiwen Fan, Tianhang Zhang
Eng. Appl. Artif. Intell.2
2026 Autonomous navigation of non-perception firefighting robot through CCTV-informed vision sharing
Saizhe Ding, Yifei Ding, Asif Sohail Usmani
Expert Syst. Appl.3
2025 Paper-Level Computerized Adaptive Testing for High-Stakes Examination via Multi-Objective Optimization
abstract
Computerized Adaptive Testing (CAT) is a testing technique that accurately infers students' proficiency levels using a relatively small number of questions.Most existing CAT systems operate on a question-level adaptive paradigm, which is suitable for practice scenarios.However, in computerized standardized high-stakes examinations such as the GRE and GMAT, this paradigm faces several challenges: (1) the lack of comparability in exam results, (2) high implementation costs due to the reliance on real-time interactions and the financial burden of maintaining CAT testing system, and (3) the difficulty in balancing multiple factors of diagnosis quality, attribute coverage, and question exposure.To address these challenges, we propose a Paper-level Computerized Adaptive Testing (PCAT) and its corresponding evaluation method.PCAT divides an exam into multiple testing stages, where examinees adaptively receive test papers of varying difficulty based on their performance in previous stages.The paper assembly problem in PCAT is solved using a population-based multi-objective optimization (MOO) approach.PCAT offers several advantages: First, the paper-level adaptive mechanism ensures that the questions faced by examinees depend solely on their performance in the earlier stages, maintaining adaptability while enhancing the comparability of results across different examinees.Second, PCAT replaces the selection strategy module in traditional CAT with an assembly module, allowing computationally intensive tasks such as cognitive diagnosis and paper assembly to be completed offline before the exam, eliminating the need for real-time interactions.Additionally, the population-based MOO approach generates a set of high-quality solutions in one run, meeting the demands of frequent administration of standardized high-stakes exams like the GRE and reducing the financial burden of maintaining a large-scale CAT system.Finally, MOO naturally
Mingjia Li 0002, Junkai Tong 0002, Yifei Ding, Hong Qian, Aimin Zhou
KDD (2)4
2025 Smart building evacuation by tracking multi-camera network and explainable Re-identification model
abstract
Real-time crowd data from surveillance devices is essential for the emergency decision-making and management inside complex buildings. Traditional evacuation monitoring with single-camera tracking often leads to erratic information, so multi-camera tracking for building occupants is critical to enhance evacuation safety and emergency response. This research proposes a novel real-time multi-camera tracking framework for the detection, tracking and re-identification (Re-ID) of evacuees across multi-camera. The framework consists of (1) a multi-camera network, (2) human detection model, (3) tracking model, (4) an explainable attention-aided Re-ID (AAR) model, and (5) a module of feature matching and re-distribution algorithm. The attention-aided Re-ID model presents outstanding performance on both the standard big benchmarks and our custom dataset. Moreover, a simple evacuation drill is conducted to demonstrate real-time multi-camera tracking, showing good accuracy in Re-ID and personnel counting, where the overall Re-ID tracking accuracy exceeds 75% and the personnel counting accuracy is approaching 100%. Lastly, the class activation map (CAM) illustrates the model explainability and limitations. The proposed multi-camera tracking framework helps develop a more automated monitoring system and an intelligent digital twin for building emergency safety management. • Establish a multi-camera re-identification (Re-ID) framework for monitoring building evacuation. • Propose an explainable attention-aided Re-ID network to enhance human feature extraction ability. • Develop a novel Re-ID dataset annotation tool for building emergency management and digital twin. • Achieve multi-camera monitoring walking evacuees with constant identities and personnel counting.
Yifei Ding, Xinghao Chen 0009
Eng. Appl. Artif. Intell.1
2025 Log-Cumulative feature alignment for enhanced Prognosis of Aero-Engine remaining Useful life
Xingxing Jiang, Benlian Xu, Yifei Ding
Expert Syst. Appl.4
2024 An autoregressive model-based degradation trend prognosis considering health indicators with multiscale attention information
Jichao Zhuang, Yifei Ding, Minping Jia, Ke Feng 0004
Eng. Appl. Artif. Intell.3
2024 Deep temporal-spectral domain adaptation for bearing fault diagnosis
Yifei Ding, Minping Jia, Peng Ding 0002, Xiaoli Zhao 0002, Chi-Guhn Lee
Knowl. Based Syst.1
2024 Unsupervised Fault Detection With Deep One-Class Classification and Manifold Distribution Alignment
abstract
Fault detection or anomaly detection relies heavily on learning from datasets where only normal samples are available, resulting in the emergence of numerous one-class classification (OCC) methods. However, learning discriminative deep representatives with good generalization from cross-domain positive samples remains challenging. Therefore, this work proposes an end-to-end framework, deep transfer one-class classification (DTOCC) for unsupervised fault detection, which combines adversarial generative OCC and distribution alignment from the perspective of manifold learning. Specifically, pseudo-negative samples are generated outside the positive manifold, facilitating the model to learn discrimination with respect to normal and anomaly. Further, cross-domain positive samples are aligned in log-Euclidean manifold space to enhance representation learning. Then, we provide the specific implementations for fault detection and validate its superiority through case studies on multiclass and run-to-failure datasets, simulating both offline and online scenarios.
Yifei Ding, Minping Jia, Xiaoan Yan, Xiaoli Zhao 0002, Chi-Guhn Lee
IEEE Trans. Ind. Informatics1
2023 Domain generalization via adversarial out-domain augmentation for remaining useful life prediction of bearings under unseen conditions
Yifei Ding, Minping Jia, Peng Ding 0002, Xiaoli Zhao 0002, Chi-Guhn Lee
Knowl. Based Syst.1
2023 Incremental Learning for Remaining Useful Life Prediction via Temporal Cascade Broad Learning System With Newly Acquired Data
abstract
Deep neural networks have promoted the technology development of fault classification and remaining useful life (RUL) prediction for mechanical equipment due to their powerful nonlinear feature extraction capability. However, the performance of traditional deep learning models is limited by the depth of networks, which is directly related to the training consumption. In addition, the parameters of networks can only be updated by retraining when faced with newly acquired data. To address the above problems, an incremental learning method based on a temporal cascade broad learning system (TCBLS) is proposed for the RUL prediction of machinery with newly acquired data. Specifically, linear and nonlinear feature information is first learned by the TCBLS. The ridge regression method is developed to calculate the weights of the network and establish an end-to-end mapping between the feature information layer and the prediction layer. Finally, the incremental learning of new data and the incremental learning of nodes are proposed for adaptively updating the weights of the network in the face of newly acquired data and insufficient prediction accuracy. The effectiveness of the proposed method is verified by four run-to-failure datasets. The comparison results with classical deep learning models show that the proposed method is promising for RUL prediction as it achieves high prediction accuracy while saving training time consumption across orders of magnitude and effectively handling newly acquired data without retraining.
Minping Jia, Peng Ding 0002, Xiaoli Zhao 0002, Yifei Ding
IEEE Trans. Ind. Informatics5
2023 Intelligent Fault Diagnosis of Gearbox Under Variable Working Conditions With Adaptive Intraclass and Interclass Convolutional Neural Network
abstract
The industrial gearboxes usually work in harsh and variable conditions, which results in partial failure of gears or bearings. Accordingly, the continuous irregular fluctuations of gearbox under variable conditions maybe increase the intraclass difference and reduce the interclass difference for the monitored samples. To this end, a new intelligent fault diagnosis method of gearbox based on adaptive intraclass and interclass convolutional neural network (AIICNN) under variable working conditions is proposed. The core of the proposed algorithm is to apply the designed intraclass and interclass constraints to improve the distribution differences of samples. Meanwhile, the adaptive activation function is added into the 1-D convolutional neural network (1dCNN) to enlarge the heterogeneous distance and narrow the homogeneous distance of samples. Specifically, the training sample subset with intraclass and interclass spacing fluctuations under variable conditions is first converted into frequency domain through the fast Fourier transform (FFT), and the designed AIICNN algorithm is employed for model training. Afterward, the testing subset is provided to the trained AIICNN algorithm for fault diagnosis. The experimental data of the planetary gearbox test rig verify the feasibility of the proposed diagnosis method and algorithm. Compared with other methods, this method can eliminate the difference of sample distribution under variable conditions and improve its diagnostic generalization.
Xiaoli Zhao 0002, Jianyong Yao, Wenxiang Deng, Peng Ding 0002, Yifei Ding, Minping Jia, Zheng Liu 0002
IEEE Trans. Neural Networks Learn. Syst.5
2022 Intelligent machinery health prognostics under variable operation conditions with limited and variable-length data
Peng Ding 0002, Minping Jia, Yifei Ding, Xiaoli Zhao 0002
Adv. Eng. Informatics3
2021 A High-Availability K-modes Clustering Method Based on Differential Privacy
Shaobo Zhang 0001, Liujie Yuan, Yifei Ding
ICA3PP (2)5
2019 Domain agnostic online semantic segmentation for multi-dimensional time series
abstract
Unsupervised semantic segmentation in the time series domain is a much studied problem due to its potential to detect unexpected regularities and regimes in poorly understood data. However, the current techniques have several shortcomings, which have limited the adoption of time series semantic segmentation beyond academic settings for four primary reasons. First, most methods require setting/learning many parameters and thus may have problems generalizing to novel situations. Second, most methods implicitly assume that all the data is segmentable and have difficulty when that assumption is unwarranted. Thirdly, many algorithms are only defined for the single dimensional case, despite the ubiquity of multi-dimensional data. Finally, most research efforts have been confined to the batch case, but online segmentation is clearly more useful and actionable. To address these issues, we present a multi-dimensional algorithm, which is domain agnostic, has only one, easily-determined parameter, and can handle data streaming at a high rate. In this context, we test the algorithm on the largest and most diverse collection of time series datasets ever considered for this task and demonstrate the algorithm's superiority over current solutions.
Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding 0003, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn J. Keogh
Data Min. Knowl. Discov.3
2019 Correction to: Domain agnostic online semantic segmentation for multi-dimensional time series
abstract
The article Domain agnostic online semantic segmentation for multi-dimensional time series, written by Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn Keogh was originally published electronically on the publisher’s internet portal (currently SpringerLink) on 25 September 2018 without open access.
Shaghayegh Gharghabi, Chin-Chia Michael Yeh, Yifei Ding, Wei Ding 0003, Paul Hibbing, Samuel LaMunion, Andrew Kaplan, Scott E. Crouter, Eamonn J. Keogh
Data Min. Knowl. Discov.3
2018 Time series joins, motifs, discords and shapelets: a unifying view that exploits the matrix profile
Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Zachary Schall-Zimmerman, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh
Data Min. Knowl. Discov.5
2017 Matrix Profile VIII: Domain Agnostic Online Semantic Segmentation at Superhuman Performance Levels
abstract
Unsupervised semantic segmentation in the time series domain is a much-studied problem due to its potential to detect unexpected regularities and regimes in poorly understood data. However, the current techniques have several shortcomings, which have limited the adoption of time series semantic segmentation beyond academic settings for three primary reasons. First, most methods require setting/learning many parameters and thus may have problems generalizing to novel situations. Second, most methods implicitly assume that all the data is segmentable, and have difficulty when that assumption is unwarranted. Finally, most research efforts have been confined to the batch case, but online segmentation is clearly more useful and actionable. To address these issues, we present an algorithm which is domain agnostic, has only one easily determined parameter, and can handle data streaming at a high rate. In this context, we test our algorithm on the largest and most diverse collection of time series datasets ever considered, and demonstrate our algorithm's superiority over current solutions. Furthermore, we are the first to show that semantic segmentation may be possible at superhuman performance levels.
Shaghayegh Gharghabi, Yifei Ding, Chin-Chia Michael Yeh, Kaveh Kamgar, Liudmila Ulanova, Eamonn J. Keogh
ICDM2
2017 Query Suggestion to allow Intuitive Interactive Search in Multidimensional Time Series
abstract
In recent years, the research community, inspired by its success in dealing with single-dimensional time series, has turned its attention to dealing with multidimensional time series. There are now a plethora of techniques for indexing, classification, and clustering of multidimensional time series. However, we argue that the difficulty of exploratory search in large multidimensional time series remains underappreciated. In essence, the problem reduces to the "chicken-and-egg" paradox that it is difficult to produce a meaningful query without knowing the best subset of dimensions to use, but finding the best subset of dimensions is itself query dependent. In this work we propose a solution to this problem. We introduce an algorithm that runs in the background, observing the user's search interactions. When appropriate, our algorithm suggests to the user a dimension that could be added or deleted to improve the user's satisfaction with the query. These query dependent suggestions may be useful to the user, even if she does not act on them (by reissuing the query), as they can hint at unexpected relationships or redundancies between the dimensions of the data. We evaluate our algorithm on several real-world datasets in medical, human activity, and industrial domains, showing that it produces subjectively sensible and objectively superior results.
Yifei Ding, Eamonn J. Keogh
SSDBM1
2016 Matrix Profile I: All Pairs Similarity Joins for Time Series: A Unifying View That Includes Motifs, Discords and Shapelets
abstract
The all-pairs-similarity-search (or similarity join) problem has been extensively studied for text and a handful of other datatypes. However, surprisingly little progress has been made on similarity joins for time series subsequences. The lack of progress probably stems from the daunting nature of the problem. For even modest sized datasets the obvious nested-loop algorithm can take months, and the typical speed-up techniques in this domain (i.e., indexing, lower-bounding, triangular-inequality pruning and early abandoning) at best produce one or two orders of magnitude speedup. In this work we introduce a novel scalable algorithm for time series subsequence all-pairs-similarity-search. For exceptionally large datasets, the algorithm can be trivially cast as an anytime algorithm and produce high-quality approximate solutions in reasonable time. The exact similarity join algorithm computes the answer to the time series motif and time series discord problem as a side-effect, and our algorithm incidentally provides the fastest known algorithm for both these extensively-studied problems. We demonstrate the utility of our ideas for two time series data mining problems, including motif discovery and novelty discovery.
Chin-Chia Michael Yeh, Yan Zhu 0014, Liudmila Ulanova, Nurjahan Begum, Yifei Ding, Hoang Anh Dau, Diego Furtado Silva, Abdullah Mueen, Eamonn J. Keogh
ICDM5