Kaiyuan Qi

dblp:34/6289 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
13since 2021 · last 2026
0009-0005-9043-106XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Discrepancy-aware contrastive learning with mixture of experts for cross-modal image-text semantic alignment
Gaigai Tang, Kaiyuan Qi, Guangfeng Su, Huiyun Zhang
Neurocomputing4
2026 GT-MARL: Graph- and Transformer-Enhanced Multiagent Reinforcement Learning for Cloud-Edge Collaborative Scheduling
abstract
Scheduling across cloud and edge systems must handle job heterogeneity, dynamic arrivals, and coupled resource constraints. To address the above issues, we present GT-MARL, a reinforcement learning framework with graph and Transformer enhancements for joint task selection and resource allocation across multiple clusters. At each decision epoch, GT-MARL encodes a heterogeneous graph over tasks and resources with Hierarchical Attention Network (HAN) to capture structural dependencies, applies a temporal Transformer to model backlog evolution and delayed interactions, and produces factorized decisions for each cluster through a Transformer-guided task selection head and discrete resource allocation with a Multi-Layer Perceptron (MLP) head. Training follows centralized training with decentralized execution (CTDE) using a centralized critic and masked action spaces. Experiments on both synthetic and trace-replay workloads show that GT-MARL consistently achieves the best p95 job completion time (JCT), while keeping mean JCT competitive with representative heuristic and learning baselines. Specifically, GT-MARL obtains 35.56 mean JCT and 70.76 p95 JCT on the synthetic workload, and 144.21 mean JCT and 264.47 p95 JCT on the trace-replay workload. On the synthetic workload, this corresponds to a 26.3% p95 JCT reduction relative to the best mean baseline with only a 3.0% increase in mean JCT. Additional queueing and service decomposition across workload intensities indicates that the tail-latency gain mainly comes from alleviating queue accumulation rather than shortening the inherent service time of tasks. GT-MARL provides a favorable tradeoff between mean latency and tail latency for scheduling under SLO constraints across cloud and edge systems.
Kaiyuan Qi, Li Pan 0001, Shijun Liu
IEEE Internet Things J.2
2025 Pre-tiering Matters: Proactive CXL Memory Tiering for Ephemeral Serverless Functions
Fengze Liu, Zhiyuan Su, Kaiyuan Qi, Laiping Zhao
ICA3PP (6)5
2025 Loop closure detection based on image feature matching and motion trajectory similarity for mobile robot
WeiLong Hao, Cui Ni, Wenjun Huangfu, Kaiyuan Qi
Appl. Intell.6
2025 Automatic generation of industrial internet attack graphs with graph neural networks and Bayesian models
Gaigai Tang, Kaiyuan Qi, Guangfeng Su, Huiyun Zhang
Comput. Networks4
2025 RehearMixup: Improving rehearsal-based continual learning
Yan Zhang 0145, Kaiyuan Qi, Guoqiang Wu, Yilong Yin
Neurocomputing2
2025 CodeQG: Automated Multiple Question Generation for Source Code Comprehension
abstract
During software maintenance and evolution, developers spend more than half of their time on code comprehension activities. In order to understand an unfamiliar code base, they would naturally ask different types of questions related to code snippets and try to find the answers. In this paper, we conduct an initial work to explore the possibility of automatic question generation for program comprehension. We construct a large-scale data set containing pairs of source code and questions that are automatically transformed from inline comments based on dependency analysis and semantic role labeling. We also build a comprehensive taxonomy of question types so as to generate questions concerning different aspects of code snippets, such as purpose, implementation details and so on. Then, we propose a deep learning-based prototype CodeQG to automatically generates multiple types of questions for code snippets. We evaluate CodeQG by using both typical performance metrics and manual evaluation. The results show that (1) we can achieve a value of 42.02 on BLEU4 and 60.81 on ROUGE-L for the generated questions; (2) overall, the questions are very correct in grammatical, semantic and format; (3) the questions are related to the corresponding code snippet and are helpful for developers in source code comprehension activities. Our work gives insights into automatically generating multiple types of questions for code comprehension. We expect this exploration will improve the applicability and generality of machine code comprehension.
Xiaowei Zhang 0018, Lin Chen 0015, Kaiyuan Qi, Weiqin Zou, Liye Pang, Lianfa Zhang, Peng Zhang 0083, Guanqun Xu
Int. J. Softw. Eng. Knowl. Eng.3
2025 An Online Algorithm for Inference Service Scheduling Using Combinations of Server-Based and Serverless Instances in Cloud Environments
abstract
With the continuous development in the field of machine learning, there is an increasing demand for the cloud-based machine learning inference services, which are latency-sensitive tasks, such as the service requests from the Internet of Things (IoT) devices. These inference services are generally accompanied by fluctuations and uncertainties, so they often require vastly varied numbers of servers at different time spots. As a result, how to dynamically and rationally schedule cloud servers for inference services has become an important issue. Alibaba Cloud currently provides its serverless instances called elastic container instance (ECI), and due to the advantages of pay-as-you-go billing and second-level elasticity, they are well-suited for handling bursty or fluctuating workloads. At the same time, Alibaba Cloud’s subscription-based elastic compute service (ECS) instances can be used for steady-state workloads. Our objective is to dynamically combine these two types of instances to deal with inference service requests. In this article, we propose a deterministic online algorithm that can rationally schedule these two types of instances to optimize costs without requiring knowledge of future workloads. We prove that the proposed online algorithm achieves a competitive ratio of no more than 2 compared to the optimal offline algorithm. Through simulation experiments, we demonstrate that our algorithm outperforms three benchmarks, which are all-reserved algorithm, all-on-demand algorithm, and traditional online algorithms that only use ECS instances. Our algorithm exhibits superiority across various workloads and can significantly reduces costs in most cases.
Li Pan 0001, Shijun Liu, Kaiyuan Qi
IEEE Internet Things J.4
2025 Cross-space topological contrastive learning for knowledge graph-aware issue recommendation
Leihong Zhang, Yuliang Shi, Kaiyuan Qi, Xinjun Wang 0003, Zhongmin Yan
Knowl. Inf. Syst.3
2025 MVSTD: Multi-view spatio-temporal graphs with external disturbance consideration for ride-hailing demand prediction
Xuanxuan Fan, Zihang Yin, Kaiyuan Qi, Zhijian Qu, Chongguang Ren
Knowl. Based Syst.3
2025 TA-ASTGCN: a trend-aware adaptive spatio-temporal graph convolutional network for traffic flow prediction
Chuang Cai, Huijie Guo, Tianfeng Dou, Kaiyuan Qi, Yuqin Bai, Chongguang Ren
Neural Comput. Appl.6
2024 MicroHFRCL: A History Faults Based Root Cause Localization Framework in Microservice Systems
abstract
At present, the microservice architecture is widely used in modern software development for its flexibility and scalability. However, the huge data scale and complex invocation relationships between services make the root cause localization of faults in microservice systems extremely difficult. Some of the current root cause localization methods based on metrics data use history fault data to effectively improve the localization results, but there are challenges such as not representing history faults effectively and limiting the localization results to repetitive faults that have occurred in history. In this paper, we propose an automatic root cause localization framework MicroHFRCL based on the history fault library to address the above issues. MicroHFRCL constructs an instance causal graph based on metric data for causal analysis. The instance causal graph is weighted by encoding the anomalous subgraph and calculating the similarity of history faults. The PageRank algorithm is used to locate the root cause of faults. Among them, MicroHFRCL learns the structure and feature information of fault anomalous subgraphs through GCN and Transformer models, achieving effective representation of history faults and fast calculation of similarity in history fault codes, improving the efficiency of repetitive fault localization, and effectively solving the problem of the limitation of using history faults for root cause localization results. We implemented MicroHFRCL and tested it on the fault dataset collected by a benchmark microservice test system. Compared with the latest baseline models, MicroHFRCL has significantly improved the localization accuracy, and can also achieve good results in the case of small-scale history faults.
Leyao Zhang, Yuliang Shi, Kaiyuan Qi, Xinjun Wang 0003, Zhongmin Yan
IJCNN3
2022 Balancing Supply and Demand for Mobile Crowdsourcing Services
Zhaoming Li, Wei He 0020, Ning Liu 0014, Li-Zhen Cui 0001, Kaiyuan Qi
ICSOC6
2015 RPECA-Rumor Propagation Based Eventual Consistency Assessment Algorithm
Zhiyuan Su, Kaiyuan Qi, Guomao Xin
APPT3
2012 MapReduce-Based Data Stream Processing over Large History Data
Kaiyuan Qi, Zhuofeng Zhao, Jun Fang 0006, Yanbo Han
ICSOC1
2010 A Model-Driven Approach for Business-Oriented Monitoring of Service Operation
abstract
The recent trend of "everything as a service", is fostering the Internet of services (IoS) and will promote the emergence of service operation. Service operation provides the business and technical base for advanced business models. Service monitoring is a crucial issue for the guaranteed service delivery in service operation. However, most service monitoring approaches are specific and focus on IT level. The challenge of how to monitor diverse business aspect of service simply and flexibly needs to be overcome. In this paper, we present a model-driven approach for service monitoring from business perspective. In the approach, a business-oriented service monitoring metamodel is put forward to define various monitoring models on demand. The model can flexibly specify the monitored information in both business level and IT level and the monitoring process. Also, a service monitor is implemented through model-driven way which brings the scalability to its implementation.
Zhuohao Wang, Zhuofeng Zhao, Kaiyuan Qi
ICSS3
2008 A Consistency-Preserving Mechanism for Web Services Response Caching
abstract
Web services are rapidly emerging as a popular standard technology for sharing data and functionality among heterogeneous systems. Service providers and consumers are loosely coupled and distributed across the network, either within an organization or across organizational boundaries, and therefore, performance becomes a major concern in such a distributed environment. Furthermore, XML is widely used as message format for service providers and consumers in Web services environment. XML message packaging and parsing brings extra overhead to both ends. Web services response latency, as well as throughput, is becoming a bottleneck problem. In this paper, We propose a consistency-preserving mechanism for Web services response caching, which reduces the volume of data transmitted without semantic interpretation of service requests or responses, and accelerates the services response finally. It achieves this reduction through the use of cryptographic hashing to detect similarities with previous results. Experiments with an initial prototype called SigsitAcclerator indicate that our mechanism can lead to significant performance improvement over more straightforward techniques.
Wubin Li, Zhuofeng Zhao, Kaiyuan Qi, Jun Fang 0006, Weilong Ding 0002
ICWS3