VLDB 2026 Research / reviewers in the wild / expert
Kangjin Wang
dblp:207/3544
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementabstractThe surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cloud providers employ spot instances to reduce costs for low-priority (LP) tasks, existing schedulers still grapple with high eviction rates and lengthy queuing times. To address these limitations, we present GFS, a novel preemptive scheduling framework that enhances service-level objective (SLO) compliance for high-priority (HP) tasks while minimizing preemptions to LP tasks. Firstly, GFS utilizes a lightweight forecasting model that predicts GPU demand among different tenants, enabling proactive resource management. Secondly, GFS employs a dynamic allocation mechanism to adjust the spot quota for LP tasks with guaranteed durations. Lastly, GFS incorporates a preemptive scheduling policy that prioritizes HP tasks while minimizing the impact on LP tasks. We demonstrate the effectiveness of GFS through both real-world implementation and simulations. The results show that GFS reduces eviction rates by 33.0%, and cuts queuing delays by 44.1% for LP tasks. Furthermore, GFS enhances the GPU allocation rate by up to 22.8% in real production clusters. In a production cluster of more than 10,000 GPUs, GFS yields roughly $459,715 in monthly benefits. Jiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang, Kangjin Wang, Chenzhi Liao, Yinghao Yu, Qin Hua, Hanwen Hu, Dongqing Bao, Tianyu Lu, Jian Cao 0001, Guangtao Xue, Liping Zhang 0013, Gang Chen 0001 |
ASPLOS (1) | 5 |
| 2026 | When May We Eliminate Collective Bias? A GRA+ Perspectives on Fifty-Fifty CompromiseabstractIn the indivisible public goods allocation problem (IPGAP), it is always impossible to obtain a perfect allocation plan which can satisfy every group. To eliminate collective bias, decision-makers often lean toward applying the fifty-fifty compromise principle, which means compromising absolutely fairly among all groups. For this concern, a few relevant research investigates the effectiveness of utilizing this principle from a computational perspective due to the lack of quantitative analysis tools. Notably, the environment-classes, agents, roles, groups, and objects (E-CARGO) model, a mature and proven effective tool, demonstrates outstanding performance in the study of such social issues. With respect to the E-CARGO model and its submodel group role assignment (GRA), this article formalizes and explores the IPGAP. Based on the group role assignment with constraints (GRA+), this article provides novel insight into the effectiveness of fifty-fifty compromise, which can inspire decision-makers that unthinkingly conducting the fifty-fifty compromise may not always succeed in eliminating collective bias. Relevant large-scale simulation experiments are conducted in this article to explore when decision-makers may eliminate the bias between the two groups. This article reveals a social paradox: compromise sometimes may not eliminate group bias, and instead, both sides may be offended. Kangjin Wang, Haibin Zhu 0001, Dongning Liu |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?abstractLog-based software reliability maintenance systems are crucial for sustaining stable customer experience. However, existing deep learning-based methods represent a black box for service providers, making it impossible for providers to understand how these methods detect anomalies, thereby hindering trust and deployment in real production environments. To address this issue, this paper defines a trustworthiness metric—diagnostic faithfulness—for models to gain service providers’ trust, based on surveys of SREs at a major cloud provider. We design two evaluation tasks: attention-based root cause localization and event perturbation. Empirical studies demonstrate that existing methods perform poorly in diagnostic faithfulness. Consequently, we propose FaithLog, a faithful log-based anomaly detection system, which achieves faithfulness through a carefully designed causality-guided attention mechanism and adversarial consistency learning. Evaluation results on two public datasets and one industrial dataset demonstrate that the proposed method achieves state-of-the-art performance in diagnostic faithfulness. Minghua He, Chiming Duan, Pei Xiao 0005, Lingzhe Zhang, Kangjin Wang, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001 |
ASE | 6 |
| 2025 | Solve the Aquaculture Imbalance Problem between Supply and Demand via Extending GMRAabstractAquaculture, as a vital component of the fisheries industry, is assuming an increasingly significant role in meeting the growing global demand for aquatic products. However, the allocation of aquaculture resources has become increasingly complex. Overproduction of a single species can lead to market oversupply, resulting in sharp price declines and substantial profit losses for fishermen. The Environment - Classes, Agents, Roles, Groups, and Objects (E-CARGO) model has shown strong potential in addressing such socio-economic problems. This study extends the Group Multirole Allocation (GMRA) model to formalize and address the Aquaculture Imbalance Problem between Supply and Demand (AISDP). The objective is to maximize total profit while considering market demand, disaster risk, and the potential for oversupply. Extensive simulation experiments reveal that the maximum total profit does not occur at the threshold, but rather beyond it—meaning that even though the unit price decreases, total profit can still be increased by further increasing the stocking quantity. Furthermore, the results suggest that fishermen should select the regulatory parameter based on real-world market dynamics and their risk tolerance. This enables aquaculture enterprises to adopt optimal, diversified decision-making strategies aligned with resource availability and strategic development goals. Zigeng Huang, Kangjin Wang, Haibin Zhu 0001, Dongning Liu |
SMC | 2 |
| 2025 | Solving the Flexible Task Allocation Problem via Group Role Assignment in CrowdsourcingabstractFlexible employment remains a critical topic for crowdsourcing platforms. Decision-makers always focus their attention on maximizing operational efficiency, yet rarely prioritize the work experience of crowdsourcing employees. However, suboptimal working conditions contribute to user attrition, ultimately diminishing platform profitability. While the importance of worker experience in crowdsourcing platforms is recognized, the Flexible Task Allocation Problem (FTAP) has seen relatively little computational investigation due to the absence of powerful quantitative analytical tools. A notable exception is the Environment-Classes, Agents, Roles, Groups, and Objects (E-CARGO) model, a mature computational framework with demonstrated efficacy in resolving similar socio-technical challenges. Consequently, this study builds upon the E-CARGO model and its Group Role Assignment (GRA) sub-model to provide a formalization and systematic analysis of the FTAP. By incorporating adjustments for distance and role-switching, the enhanced GRA model can significantly optimize the work experience of crowdsourcing employees, albeit at a slight cost to overall performance. This improvement helps crowdsourcing platforms retain more users and expand their scale. Relevant simulation experiments are conducted in this study to rigorously evaluate the effectiveness of these optimizations. Kangjin Wang, Haibin Zhu 0001, Dongning Liu |
SMC | 1 |
| 2024 | Reducing Events to Augment Log-based Anomaly Detection Models: An Empirical StudyabstractAs software systems grow increasingly intricate, the precise detection of anomalies have become both essential and challenging. Current log-based anomaly detection methods depend heavily on vast amounts of log data leading to inefficient inference and potential misguidance by noise logs. However, the quantitative effects of log reduction on the effectiveness of anomaly detection remain unexplored. Therefore, we first conduct a comprehensive study on six distinct models spanning three datasets. Through the study, the impact of log quantity and their effectiveness in representing anomalies is qualifies, uncovering three distinctive log event types that differently influence model performance. Drawing from these insights, we propose LogCleaner: an efficient methodology for the automatic reduction of log events in the context of anomaly detection. Serving as middleware between software systems and models, LogCleaner continuously updates and filters anti-events and duplicative-events in the raw generated logs. Experimental outcomes highlight LogCleaner’s capability to reduce over 70% of log events in anomaly detection, accelerating the model’s inference speed by approximately 300%, and universally improving the performance of models for anomaly detection. Lingzhe Zhang, Kangjin Wang, Mengxi Jia, Yong Yang 0011, Ying Li 0012 |
ESEM | 3 |
| 2022 | Characterizing Job Microarchitectural Profiles at Scale: Dataset and AnalysisabstractUnderstanding the microarchitectural resource characteristics of datacenter jobs has become increasingly critical to guarantee the performance of jobs while improving resource utilization. Prior work studied the resource characteristics of datacenter jobs at the OS level, little reveals the deep and detailed characteristics at the microarchitecture level due to the lack of related open traces. In this paper, we provide a new open trace, AMTrace (Alibaba Microarchitecture Trace) 1, which is profiled from 8,577 high-end physical hosts from Alibaba’s datacenter by a hardware/software co-design monitoring method. AMTrace provides the microarchitectural metrics of 9.8 × 105 Linux containers with ”Per-Container-Per-Logic CPU” granularity. Different from existing open traces, AMTrace provides a new perspective to analyze the microarchitectural resource characteristics of datacenter jobs. Based on AMTrace, we first reveal the uneven resource usage of jobs among multiple logic CPUs. Then, we analyze the impact of resource contention of CPU and memory bandwidth on job performance. Finally, we analyze the job performance under different CPU provisioning modes from microarchitecture perspective. These analyses lead to constructive insights for datacenter resource management and optimization. Furthermore, we discuss possible research opportunities on AMTrace and we believe that AMTrace will inspire more exciting research on microarchitecture and resource management. Kangjin Wang, Ying Li 0012, Kingsum Chow, Yaoyong Dou, Guoyao Xu, Chuanjia Hou, Liping Zhang 0013 |
ICPP | 1 |