Dan Li 0016

dblp:48/4185-16 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0002-3787-1673ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author
YearPublicationVenuePosition
2026 Lightweight Time Series Data Valuation on Time Series Foundation Models via In-Context Finetuning
Shunyu Wu, Tianyue Li, Yixuan Leng, Jingyi Suo, Jian Lou 0001, Dan Li 0016, See-Kiong Ng
DASFAA (3)6
2026 Large language models for explainable fault diagnosis of machines
Hamzah A. A. M. Qaid, Bo Zhang 0022, Shuai Su, Dan Li 0016, See-Kiong Ng, Wei Li 0019
Eng. Appl. Artif. Intell.4
2026 Restoring missing gaps and intervals via global consistency and local coherence
Richa Hu, Dan Li 0016, Jian Lou 0001, Ruibing Jin, Bingxin Lin, See-Kiong Ng, Zibin Zheng
Expert Syst. Appl.2
2025 Fostering Active Learning: A Study on Anonymous Q&A Software in Undergraduate Education
abstract
This research explores strategies for using one-way anonymous Q&A software to enhance class participation and teaching effectiveness of undergraduate students. The study finds that undergraduate students, influenced by cultural factors, ed-ucational systems, and upbringing environments, tend to display introverted psychological traits, resulting in insufficient class participation. One-way anonymous Q&A software significantly improves student participation enthusiasm through mechanisms such as breaking psychological barriers, promoting deep thinking, enhancing classroom interaction, and providing diverse feedback. Case analysis shows that by introducing the proposed anonymous Q&A software to teaching activities, students asked questions more frequently, and the quality of questions also improved accordingly with a wider adoption of the software. Besides, it also shows that the learning interests of students have increased after using the software. The research suggests that future studies should deepen technological innovation, pro-mote teaching model reform, and strengthen interdisciplinary integration to better meet the needs of a modern educational environment in universities.
Dan Li 0016, Zigui Jiang, Yuxin Su 0001, Zibin Zheng
SSE1
2025 LogExpertSolver: A Multi-Agent Framework for Domain-Specialized Log Parsing
abstract
Log parsing, as a process of extracting structured information from semi-structured raw log data, is a crucial step in log analysis workflows. Rule-based parsing methods often overlook the rich semantic information contained in logs. Recently, LLM-based methods face three major challenges: 1) limited understanding of domain-specific logs due to lack of professional domain knowledge; 2) significant parsing costs and potential data privacy risks associated with commercial LLMs like ChatGPT; and 3) the Group Accuracy (GA) is highly susceptible to individual anomalous data. To address these challenges, we propose LogExpertSolver, a multi-agent framework for domain-specialized log parsing. This framework effectively reduces manual intervention in log parsing through inter-agent collaboration mechanisms. LogExpertSolver employs locally deployed medium-scale language models to construct multiple agents with diverse domain expertise. By decomposing complex log parsing tasks into a series of fine-grained subtasks, which are handled by corresponding expert agents, the framework achieves efficient log parsing. Through evaluation on the LogHub2.0 dataset, LogExpertSolver achieves an average Parsing Accuracy (PA) of 0.915, surpassing state-of-the-art parsers (LibreLog and LILAC) by 6.2% and 7.3% respectively.
Chenxi Mao, Yuxin Su 0001, Dan Li 0016
APSEC3
2025 Integrating Time Series into LLMs via Multi-layer Steerable Embedding Fusion for Enhanced Forecasting
abstract
Time series (TS) data are ubiquitous across various application areas, rendering time series forecasting (TSF) a fundamental task. With the astounding advances in large language models (LLMs), a variety of methods have been developed to adapt LLMs for time series forecasting. Despite unlocking the potential of LLMs in comprehending TS data, existing methods are inherently constrained by their shallow integration of TS information, wherein LLMs typically access TS representations at shallow layers, primarily at the input layer. This causes the influence of TS representations to progressively fade in deeper layers and eventually leads to ineffective adaptation between textual embeddings and TS representations. In this paper, we propose the Multi-layer Steerable Embedding Fusion (MSEF), a novel framework that enables LLMs to directly access time series patterns at all depths, thereby mitigating the progressive loss of TS information in deeper layers. Specifically, MSEF leverages off-the-shelf time series foundation models to extract semantically rich embeddings, which are fused with intermediate text representations across LLM layers via layer-specific steering vectors. These steering vectors are designed to continuously optimize the alignment between time series and textual modalities and facilitate a layer-specific adaptation mechanism that ensures efficient few-shot learning capabilities. Experimental results on seven benchmarks demonstrate significant performance improvements by MSEF compared with baselines, with an average reduction of 31.8% in terms of MSE. The code is available at https://github.com/One1sAll/MSEF.
Zhuomin Chen, Dan Li 0016, Jiahui Zhou, Shunyu Wu, Haozheng Ye, Jian Lou 0001, See-Kiong Ng
CIKM2
2025 Soft label enhanced graph neural network under heterophily
Junyuan Fang, Jiajing Wu, Dan Li 0016, Zibin Zheng
Knowl. Based Syst.4
2024 eWAPA: An eBPF-based WASI Performance Analysis Framework for Web Assembly Runtimes
abstract
WebAssembly (Wasm) is a low-level bytecode format that can run in modern browsers. With the development of standalone runtimes and the improvement of the WebAssembly System Interface (WASI), Wasm has further provided a more complete sandboxed runtime experience for server-side applications, effectively expanding its application scenarios. However, the implementation of WASI varies across different runtimes, and suboptimal interface implementations can lead to performance degradation during interactions between the runtime and the operating system. Existing research mainly focuses on overall performance evaluation of runtimes, while studies on WASI implementations are relatively scarce. To tackle this problem, we propose an eBPF-based WASI performance analysis framework. It collects key performance metrics of the runtime under different I/O load conditions, such as total execution time, startup time, WASI execution time, and syscall time. We can comprehensively analyze the performance of the runtime's I/O interactions with the operating system. Additionally, we provide a detailed analysis of the causes behind two specific WASI performance anomalies. These analytical results will guide the optimization of standalone runtimes and WASI implementations, enhancing their efficiency.
Chenxi Mao, Yuxin Su 0001, Shiwen Shan, Dan Li 0016
SSE4
2024 GLA-DA: Global-Local Alignment Domain Adaptation for Multivariate Time Series
Gang Tu, Dan Li 0016, Bingxin Lin, Zibin Zheng, See-Kiong Ng
DASFAA (5)2
2024 Face It Yourselves: An LLM-Based Two-Stage Strategy to Localize Configuration Errors via Logs
abstract
Configurable software systems are prone to configuration errors, resulting in significant losses to companies. However, diagnosing these errors is challenging due to the vast and complex configuration space. These errors pose significant challenges for both experienced maintainers and new end-users, particularly those without access to the source code of the software systems. Given that logs are easily accessible to most end-users, we conduct a preliminary study to outline the challenges and opportunities of utilizing logs in localizing configuration errors. Based on the insights gained from the preliminary study, we propose an LLM-based two-stage strategy for end-users to localize the root-cause configuration properties based on logs. We further implement a tool, LogConfigLocalizer, aligned with the design of the aforementioned strategy, hoping to assist end-users in coping with configuration errors through log analysis.
Shiwen Shan, Yintong Huo, Yuxin Su 0001, Yichen Li 0003, Dan Li 0016, Zibin Zheng
ISSTA5
2023 CB-GAN: Generate Sensitive Data with a Convolutional Bidirectional Generative Adversarial Networks
Richa Hu, Dan Li 0016, See-Kiong Ng, Zibin Zheng
DASFAA (4)2
2022 MAD-SGCN: Multivariate Anomaly Detection with Self-learning Graph Convolutional Networks
abstract
Today's Cyber Physical Systems (CPSs) are large and complex data-intensive systems. Constant monitoring and analysis of the data generated by a multitude of interconnected sensors and actuators are required in order to detect anomalies due to possible intrusions or faults with high accuracy and timeliness. Recently, unsupervised anomaly detection techniques based on deep learning for multivariate time series have been proposed for detecting CPSs attacks with promising performance. However, the current methods are either limited by their representation learning methods in encoding the temporal and spatial information simultaneously and effectively, or cannot easily scale to other tasks without having explicit knowledge of the internal relationships between the different variables or sensors, which are both important for characterising CPSs data. In this paper, we propose a novel unsupervised anomaly detection method for multivariate time series MAD-SGCN which effectively captures the temporal and spatial correlations of the input sequences simultaneously using Long Short-Term Memory networks (LSTMs) and spectral-based Graph Convolutional Networks (GCNs). We design a self-supervised graph structure learning mechanism to minimize the usage of the prior knowledge about the network structures of the CPSs. Experiments on four CPS datasets demonstrate the superiority of the proposed method.
Panpan Qi, Dan Li 0016, See-Kiong Ng
ICDE2
2020 VC-GAN: Classifying Vessel Types by Maritime Trajectories using Generative Adversarial Networks
abstract
As maritime transport is the backbone of global trade, it is important to ensure the safety and security for sea transportation effectively. However, the dependence on experienced human operators for maritime surveillance does not scale in terms of coverage. While the ship information and trajectory data from the Automatic Identification System (AIS) can be used to automate maritime surveillance, the AIS data may be modified deliberately or accidentally, resulting in difficulties in the detection of illegal maritime activities. We have developed VC-GAN for Vessel Classification using Generative Adversarial Networks to identify vessel types based solely on the vessel trajectories in areas of interest. Our VC-GAN framework adversarially trains a multi-class classifier by learning from the labelled AIS data to classify the types of the vessels of interest, as well as to detect the out-of-class vessels. We evaluated the proposed VC-GAN method on two maritime datasets in Europe and Southeast Asia. The experimental results showed that VC-GAN significantly outperformed other vessel classification methods, especially in detecting out-of-class vessels.
Dan Li 0016, See-Kiong Ng
ICTAI1
2020 Handling Incomplete Sensor Measurements in Fault Detection and Diagnosis for Building HVAC Systems
abstract
Due to the development of sensor networks and information technology, data-driven fault detection and diagnosis (FDD) has been made possible with real-time multiple sensor measurements. However, due to inevitable sensor errors or communication failures, the raw data are usually incomplete with corrupted values, lost values, or undetected missing values. In practice, the incomplete data are usually dealt with by directly excluding incomplete measurements and abnormal spikes. In addition, some preprocessing methods, which naively impute data though averaging or smoothing, have also been widely applied. In this article, we address the building FDD problem with incomplete data by proposing a new approach, the adjacent information recovery (AIR) filter. The AIR filter is utilized to deal with the FDD for a typical air handling unit (AHU) system with incomplete data based on the ASHRAE Research Project 1312. Experimental results show that the proposed method improves FDD performance by recovering missing sensor measurements and outperforms the state-of-the-art methods.
Dan Li 0016, Yuxun Zhou, Guoqiang Hu 0001, Costas J. Spanos
IEEE Trans Autom. Sci. Eng.1
2019 MAD-GAN: Multivariate Anomaly Detection for Time Series Data with Generative Adversarial Networks
Dan Li 0016, Dacheng Chen, Baihong Jin, Jonathan Goh, See-Kiong Ng
ICANN (4)1
2019 Identifying Unseen Faults for Smart Buildings by Incorporating Expert Knowledge With Data
abstract
Thanks to the development of sensor networks and information technology, data-driven fault detection and diagnosis (FDD) is getting more and more popular with rich data. In the building FDD field, mature supervised learning algorithms and strategies have been applied to detect and diagnose known faults. However, it is out of the question to collect labeled training data for every possible fault. Thus, there is a necessity to study FDD when the training data for some faults are unavailable. To the authors' best knowledge, few works have reported how to identify “unseen faults.” In this paper, authors propose a novel expert knowledge-based unseen fault identification (EK-UFI) method to identify unseen faults by employing the similarities between known faults and unknown faults. The similarity is captured by incorporating essential expert knowledge that is encoded in the fault gene matrix. The fault gene is integrated with a latent incorporation matrix that transfers knowledge from known faults to unseen faults. With application to a real system, the proposed method is proven to be effective in identifying various building unknown faults with a high accuracy. Note to Practitioners- FDD is of great importance for saving energy and improving occupancy comfort levels and building safety levels. Identifying unseen faults in real application is challenging since: 1) building faults are complicated and confusing while well-labeled fault data is rare; 2) experimental fault data collected in laboratory test beds cannot be directly used as judgment criteria for real buildings; and 3) it is impossible to measure every possible fault ahead of time. Although supervised learning methods have been successfully applied in existing works to solve the building FDD, they could not attack the UFI problem. In this paper, a novel EK-UFI method is proposed to identify unseen faults by employing the similarities between known faults and unknown faults. Experimental results show that the proposed method is essential.
Dan Li 0016, Yuxun Zhou, Guoqiang Hu 0001, Costas J. Spanos
IEEE Trans Autom. Sci. Eng.1
2018 Design Automation for Smart Building Systems
abstract
Smart buildings today are aimed at providing safe, healthy, comfortable, affordable, and beautiful spaces in a carbon and energy-efficient way. They are emerging as complex cyber-physical systems with humans in the loop. Cost, the need to cope with increasing functional complexity, flexibility, fragmentation of the supply chain, and time-to-market pressure are rendering the traditional heuristic and ad hoc design paradigms inefficient and insufficient for the future. In this paper, we present a platform-based methodology for smart building design. Platform-based design (PBD) promotes the reuse of hardware and software on shared infrastructures, enables rapid prototyping of applications, and involves extensive exploration of the design space to optimize design performance. In this paper, we identify, abstract, and formalize components of smart buildings, and present a design flow that maps high-level specifications of desired building applications to their physical implementations under the PBD framework. A case study on the design of on-demand heating, ventilation, and air conditioning (HVAC) systems is presented to demonstrate the use of PBD.
Ruoxi Jia 0001, Baihong Jin, Ming Jin 0002, Yuxun Zhou, Ioannis C. Konstantakopoulos, Han Zou, Joyce Kim, Dan Li 0016, Weixi Gu, Reza Arghandeh, Pierluigi Nuzzo 0002, Stefano Schiavon, Alberto L. Sangiovanni-Vincentelli, Costas J. Spanos
Proc. IEEE8
2017 Optimal Sensor Configuration and Feature Selection for AHU Fault Detection and Diagnosis
abstract
Experiments show that operation efficiency and reliability of buildings can greatly benefit from rich and relevant datasets. More specifically, data can be analyzed to detect and diagnose system and component failures that undermine energy efficiency. Among the huge quantity of information, some features are more correlated with the failures than others. However, there has been little research to date focusing on determining the types of data that can optimally support fault detection and diagnosis (FDD). This paper presents a novel optimal feature selection method, named information greedy feature filter (IGFF), to select essential features that benefit building FDD. On one hand, the selection results can serve as reference for configuring sensors in the data collection stage, especially when the measurement resource is limited. On the other hand, with the most informative features selected by the IGFF, the performance of building FDD could be improved and theoretically justified. A case study on air-handling unit (AHU) is conducted based on the dataset of the ASHRAE Research Project 1312. Numerical results show that, compared with several baselines, the FDD performances of conventional classification methods are greatly enhanced by the IGFF.
Dan Li 0016, Yuxun Zhou, Guoqiang Hu 0001, Costas J. Spanos
IEEE Trans. Ind. Informatics1
2016 Optimal Training and Efficient Model Selection for Parameterized Large Margin Learning
Yuxun Zhou, Jae Yeon Baek, Dan Li 0016, Costas J. Spanos
PAKDD (1)3