VLDB 2026 Research / reviewers in the wild / expert
Zhijing Li 0007
dblp:149/3951-7
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-2066-9874ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MicLog: Towards Accurate and Efficient LLM-based Log Parsing via Progressive Meta In-Context LearningabstractLog parsing converts semi-structured logs into structured templates, forming a critical foundation for downstream analysis. Traditional syntax and semantic-based parsers often struggle with semantic variations in evolving logs and data scarcity stemming from their limited domain coverage. Recent large language model (LLM)-based parsers leverage in-context learning (ICL) to extract semantics from examples, demonstrating superior accuracy. However, LLM-based parsers face two main challenges: 1) underutilization of ICL capabilities, particularly in dynamic example selection and cross-domain generalization, leading to inconsistent performance; 2) time-consuming and costly LLM querying. To address these challenges, we present MicLog, the first progressive meta in-context learning (ProgMeta-ICL) log parsing framework that combines meta-learning with ICL on small open-source LLMs (i.e., Qwen-2.5-3B). Specifically, MicLog: i) enhances LLMs' ICL capability through a zero-shot to k-shot ProgMeta-ICL paradigm, employing weighted DBSCAN candidate sampling and enhanced BM25 demonstration selection; ii) accelerates parsing via a multi-level pre-query cache that dynamically matches and refines recently parsed templates. Evaluated on Loghub-2.0, MicLog achieves 10.3% higher parsing accuracy than the state-of-the-art parser while reducing parsing time by 42.4%. Jianbo Yu 0003, Junjielong Xu, Zhijing Li 0007, Pinjia He, Wanyuan Wang |
AAAI | 6 |
| 2026 | Enhanced Reasoning for Biomedical Document-Level Relation Extraction via a Novel Cascade Language Model FrameworkabstractHaohua Song, Wenhao Gu, Zhijing Li, Yunwenyu, Tiantian Zhu, Xiao Yang, Zexuan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haohua Song, Wenhao Gu, Zhijing Li 0007, Yunwen Yu, Xiao Yang 0019, Zexuan Zhu 0001 |
ACL (1) | 3 |
| 2026 | Surveying Root Cause Analysis Techniques: A Comprehensive Review of Aspects for Multi-Service ApplicationsabstractAs the landscape of industry and commerce continues to evolve, it presents increasing challenges for traditional operation and maintenance of services. These challenges arise from the need to adapt to dynamic market conditions, integrate complex systems and technologies, ensure uninterrupted service delivery, manage large volumes of data, and meet ever-growing customer expectations. Identifying the root cause of a failure is a crucial aspect of day-to-day operation and maintenance. With an accurate and prompt diagnosis, it becomes possible to take timely action and address the underlying issue at its core. Research on root cause analysis has been active in recent years, as it is recognized as a potential solution for effectively managing complex system states. This survey delves into the current state of research on root cause analysis, investigates recent research trends, and presents the commonly available public datasets. We organize studies along two orthogonal axes—application scenarios (cloud services, microservices, industrial systems) and input-data types (logs, traces, metrics, reports)—and synthesize algorithmic families, hybrid/LLM approaches, evaluation metrics, datasets, and tooling. To our knowledge, this is the first survey to classify RCA methods by both scenario and input-data perspectives while providing a consolidated inventory of datasets and tools, offering a roadmap for researchers Zhijing Li 0007, Jianbo Yu 0003, Zhijun Huang |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Research on Urban Data Visualization Based on Big Data: Transforming Insights into Action
Jibo Guo, Zhijing Li 0007, Yiping Jiang 0001, Qiuying Yu |
CogSci | 2 |
| 2024 | Relation Extraction in Biomedical Texts: A Cross-Sentence ApproachabstractRelation extraction, a crucial task in understanding the intricate relationships between entities in biomedical domains, has predominantly focused on binary relations within single sentences. However, in practical biomedical scenarios, relationships often extend across multiple sentences, leading to extraction errors with potential impacts on clinical decision-making and medical diagnosis. To overcome this limitation, we present a novel cross-sentence relation extraction framework that integrates and enhances coreference resolution and relation extraction models. Coreference resolution serves as the foundation, breaking sentence boundaries and linking entities across sentences. Our framework incorporates pre-trained deep language representations and leverages graph LSTMs to effectively model cross-sentence entity mentions. The use of a self-attentive Transformer architecture and external semantic information further enhances the modeling of intricate relationships. Comprehensive experiments conducted on two standard datasets, namely the BioNLP dataset and THYME dataset, demonstrate the state-of-the-art performance of our proposed approach. Zhijing Li 0007, Liwei Tian, Yiping Jiang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | Revisiting Log Parsing: The Present, the Future, and the UncertaintiesabstractIn the recent decade, the amount of software runtime logs has increased rapidly and spawned a line of automated log analysis research using machine learning or data mining algorithms. In the typical workflow of log analysis, log parsing, which aims to transform unstructured or semistructured logs into structured logs, is crucial to various downstream algorithms. While the state-of-the-art (SOTA) parsers achieve extremely high accuracy, recent research shows that these parsers are far from being useful under stricter evaluation metrics. Thus, researchers and practitioners are unclear about the current state of log parsing research and what might be important to explore in the future. To this end, we conduct an empirical study to revisit log parsing by running extensive experiments of SOTA parsers on 16 widely used log datasets under five evaluation metrics with different preprocessing settings. Our results show that the performance of log parsers varies significantly under different evaluation metrics. In addition, preprocessing plays an important role in the evaluation. In particular, preprocessing with common regular expressions can cause a 0.38 performance difference in group accuracy, highlighting the importance of reporting preprocessing details in parsing research. We also generalize the word-level regular expressions in preprocessing and try to use them to parse the whole logs, which leads to surprisingly decent accuracy. These results imply that formulating log parsing as a word-level classification task is a feasible future direction. Moreover, we find out that the most widely used dataset (i.e., LogHub) contains labeling errors. To address this issue, we make an extensive manual effort to fix the errors in the log dataset, providing a revised ground truth for future log parsing research. On the revised log dataset, our simple parser (word-level regular expression-based) achieves 0.97 precision-template accuracy on the Spark dataset and an average recall-template accuracy of 0.93 on 16 datasets, which outperforms all existing parsers. Zhijing Li 0007, Qiuai Fu, Zhijun Huang, Jianbo Yu 0003, Yiqian Li, Yuanhao Lai, Yuchi Ma |
IEEE Trans. Reliab. | 1 |
| 2023 | Hue: A User-Adaptive Parser for Hybrid LogsabstractLog parsing, which extracts log templates from semi-structured logs and produces structured logs, is the first and the most critical step in automated log analysis. While existing log parsers have achieved decent results, they suffer from two major limitations by design. First, they do not natively support hybrid logs that consist of both single-line logs and multi-line logs (Java Exception and Hadoop Counters). Second, they fall short in integrating domain knowledge in parsing, making it hard to identify ambiguous tokens in logs. This paper defines a new research problem, hybrid log parsing, as a superset of traditional log parsing tasks, and proposes Hue, the first attempt for hybrid log parsing via a user-adaptive manner. Specifically, Hue converts each log message to a sequence of special wildcards using a key casting table and determines the log types via line aggregating and pattern extracting. In addition, Hue can effectively utilize user feedback via a novel merge-reject strategy, making it possible to quickly adapt to complex and changing log templates. We evaluated Hue on three hybrid log datasets and sixteen widely-used single-line log datasets (Loghub). The results show that Hue achieves an average grouping accuracy of 0.845 on hybrid logs, which largely outperforms the best results (0.563 on average) obtained by existing parsers. Hue also exhibits SOTA performance on single-line log datasets. Junjielong Xu, Qiuai Fu, Zhouruixing Zhu, Yutong Cheng, Zhijing Li 0007, Yuchi Ma, Pinjia He |
ESEC/SIGSOFT FSE | 5 |
| 2023 | Log Parsing with Generalization Ability under New Log TypesabstractLog parsing, which converts semi-structured logs into structured logs, is the first step for automated log analysis. Existing parsers are still unsatisfactory in real-world systems due to new log types in new-coming logs. In practice, available logs collected during system runtime often do not contain all the possible log types of a system because log types related to infrequently activated system states are unlikely to be recorded and new log types are frequently introduced with system updates. Meanwhile, most existing parsers require preprocessing to extract variables in advance, but preprocessing is based on the operator’s prior knowledge of available logs and therefore may not work well on new log types. In addition, parser parameters set based on available logs are difficult to generalize to new log types. To support new log types, we propose a variable generation imitation strategy to craft a novel log parsing approach with generalization ability, called Log3T. Log3T employs a pre-trained transformer encoder-based model to extract log templates and can update parameters at parsing time to adapt to new log types by a modified test-time training. Experimental results on 16 benchmark datasets show that Log3T outperforms the state-of-the-art parsers in terms of parsing accuracy. In addition, Log3T can automatically adapt to new log types in new-coming logs. Siyu Yu, Yifan Wu 0002, Zhijing Li 0007, Pinjia He, Ningjiang Chen |
ESEC/SIGSOFT FSE | 3 |