Yifan Gan

dblp:228/0256 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 Textual and structural dual enhancement for knowledge graph completion with large language models
Yifan Gan, Yongfeng Dong
J. Intell. Inf. Syst.2
2024 On the Feasibility and Benefits of Extensive Evaluation
abstract
Benchmark and system parameters often have a significant impact on performance evaluation, which raises a long-lasting question about which settings we should use. This paper studies the feasibility and benefits of extensive evaluation. A full extensive evaluation, which tests all possible settings, is usually too expensive. This work investigates whether it is possible to sample a subset of the settings and, upon them, generate observations that match those from a full extensive evaluation. Towards this goal, we have explored the incremental sampling approach, which starts by measuring a small subset of random settings, builds a prediction model on these samples using the popular ANOVA approach, adds more samples if the model is not accurate enough, and terminates otherwise. To summarize our findings: 1) Enhancing a research prototype to support extensive evaluation mostly involves changing hard-coded configurations, which does not take much effort. 2) Some systems are highly predictable, which means that they can achieve accurate predictions with a low sampling rate, but some systems are less predictable. 3) We have not found a method that can consistently outperform random sampling + ANOVA. Based on these findings, we provide recommendations to improve artifact predictability and strategies for selecting parameter values during evaluation.
Yujie Hui, Miao Yu 0023, Hao Qi 0008, Yifan Gan, Tianxi Li, Yuke Li 0003, Xueyuan Ren, Sixiang Ma, Xiaoyi Lu 0001, Yang Wang 0009
Proc. ACM Manag. Data4
2022 IsoBugView: Interactively Debugging Isolation Bugs in Database Applications
abstract
Database applications frequently use weaker isolation levels, such as Read Committed, for better performance, which may lead to bugs that do not happen under Serializable. Although a number of works have proposed methods to identify such isolation-related bugs, the difficulty of analyzing reported bugs is often underestimated, since these bugs often involve multiple complicated transactions interleaved in a specific order and they often require users' feedback to improve the accuracy of bug analysis. This paper presents IsoBugView, a tool to visualize isolation bugs and incorporate users' feedback: to address the challenge that a complicated bug may include much information and thus is hard to present, IsoBugView displays a high-level overview of the bug first and displays further information of individual pieces if the developer needs further investigation. To incorporate users' feedback, IsoBugView embeds hook functions into the backend analysis tool to preprocess a dependency graph and postprocess a found cycle and further allows a user to apply predefined hook functions in its graphic user interface. Our experience shows that IsoBugView has greatly improved our productivity of analyzing isolation bugs.
Drew Ripberger, Yifan Gan, Xueyuan Ren, Spyros Blanas, Yang Wang 0009
Proc. VLDB Endow.2
2020 IsoDiff: Debugging Anomalies Caused by Weak Isolation
Yifan Gan, Xueyuan Ren, Drew Ripberger, Spyros Blanas, Yang Wang 0009
Proc. VLDB Endow.1
2019 Multi-objective optimization of energy consumption in crude oil pipeline transportation system operation based on exergy loss analysis
Yang Liu 0079, Qinglin Cheng, Yifan Gan, Zhidong Li
Neurocomputing3
2018 Evaluating Scalability Bottlenecks by Workload Extrapolation
abstract
Testing a scalability bottleneck requires a large system to generate sufficient load, which is usually not accessible to researchers. To address this problem, this paper extrapolates the workload to a bottleneck node. The key observation that motivates our approach is that systems at a large scale are often repeating their behaviors at small scales, by running a job more times, running more nodes of the same type, or running more iterations of the same loop. Following this observation, we record a node's workloads at small scales and extrapolate such workload at a large scale. Towards this goal, we have developed PatternMiner, a semi-automatic tool to identify how workload patterns change with scale. We have tested our method on HDFS NameNode and YARN's Resource Manager. Our evaluation shows that PatternMiner is able to predict 98% of the workloads for NameNode and 83% of the workloads for the Resource Manager. Furthermore, by utilizing the extrapolated workload, we are able to emulate a cluster of up to 60,000 nodes with only 8 physical machines to evaluate NameNode and Resource Manager.
Rong Shi, Yifan Gan, Yang Wang 0009
MASCOTS2
2018 wPerf: Generic Off-CPU Analysis to Identify Bottleneck Waiting Events
Yifan Gan, Sixiang Ma, Yang Wang 0009
OSDI2