VLDB 2026 Research / reviewers in the wild / expert
Yujia Fu
dblp:214/5638
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLARE: Fine-Grained Length-Aware Routing for Resource-Efficient Heterogeneous LLM ServingabstractWith the rapid proliferation of large language models (LLMs), model pools have become increasingly heterogeneous in both capability and efficiency.Larger LLMs can improve quality but incur higher latency and cost, while smaller LLMs are the opposite, making perquery model selection crucial in practice.This has spawned LLM routers that dispatch each query to an appropriate model.Existing routers lack fine-grained resource awareness across deployment settings, which degrades efficiency metrics in real-world serving.To this end, we propose FLARE, a length-centric, resourceaware multi-LLM routing framework that uses length-based models to estimate per-query latency and cost.FLARE formulates routing as a discrete multi-objective optimization problem to achieve an efficient trade-off.Experiments show that FLARE reduces latency and cost by up to 68% and 75% while achieving sufficient accuracy, and can be easily applied to new datasets and LLMs. Yujia Fu, Heming Zhong, Dan Huang 0001, Yutong Lu |
ACL (1) | 1 |
| 2025 | IasRT: Interference-Aware and SLO-Driven GPU Scheduling for Real-Time DNN InferenceabstractDeep Neural Network (DNN) inference has become a cornerstone of latency-sensitive applications such as autonomous driving and augmented reality. While GPUs offer high throughput for DNN inference, they often suffer from underutilization due to coarse-grained scheduling and limited concurrency. Existing GPU-sharing methods either lack awareness of fine-grained kernel interference or fail to meet service-level objectives (SLOs) under multi-tenant, multi-priority workloads. In this paper, we propose IasRT, a runtime framework that enables interferenceaware and SLO-driven GPU sharing for real-time DNN inference. IasRT profiles kernel-level resource usage and interference sensitivity, and dynamically partitions GPU streaming multiprocessors (SMs) to collocate jobs with minimal performance degradation. Furthermore, it introduces a dynamic SLO controller to maintain latency targets for multiple latency-sensitive (LS) jobs simultaneously. Evaluations on real-world DNN workloads show that IasRT reduces the 99th percentile latency of LS jobs by up to 38% compared to the state-of-the-art GPU sharing methods, while maintaining similar overall throughput from multiple collocated workloads, demonstrating its effectiveness in high-concurrency environments. Heming Zhong, Jinhui Wei, Yujia Fu, Dan Huang 0001, Yutong Lu |
ICCD | 3 |
| 2025 | CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science MasteryabstractLarge language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round evaluation of the computer science field. To bridge this gap, we introduce CS-Bench, the first multilingual (English, Chinese, French, German) benchmark dedicated to evaluating the performance of LLMs in computer science. CS-Bench comprises approximately 10K meticulously curated test samples, covering 26 subfields across 4 key areas of computer science, encompassing various task forms and divisions of knowledge and reasoning. Utilizing CS-Bench, we conduct a comprehensive evaluation of over 30 mainstream LLMs, revealing the relationship between CS performance and model scales. We also quantitatively analyze the reasons for failures in existing LLMs and highlight directions for improvements, including knowledge supplementation and CS-specific reasoning. Further cross-capability experiments show a high correlation between LLMs' capabilities in computer science and their abilities in mathematics and coding. Moreover, expert LLMs specialized in mathematics and coding also demonstrate strong performances in several CS subfields. Looking ahead, we envision CS-Bench serving as a cornerstone for LLM applications in the CS field and paving new avenues in assessing LLMs' diverse reasoning capabilities. Our project homepage is available at https://csbench.github.io/. Xiaoshuai Song, Muxi Diao, Guanting Dong 0001, Yujia Fu, Runqi Qiao, Zhexu Wang, Dayuan Fu, Huangxuan Wu, Weihao Zeng 0003, Yejie Wang, Zhuoma Gongque, Jianing Yu 0001, Qiuna Tan, Weiran Xu |
ICLR | 5 |
| 2025 | An Alternating Guidance With Cross-View Teacher-Student Framework for Remote Sensing Semi-Supervised Semantic SegmentationabstractThe semantic segmentation of remote sensing images is crucial for Earth observation. The semi-supervised semantic segmentation method can effectively reduce the dependence of the training process on labeled data. Among them, the semi-supervised semantic segmentation method based on the teacher-student paradigm is currently one of the most mainstream methods. However, the issue of weight coupling has constrained further performance improvements. This article proposes an alternating guidance method that combines cross-view learning to improve the teacher-student paradigm and enhance the semantic segmentation performance of remote sensing images. The student model is designed by using two decoders with the same architecture but independently updated parameters. Two decoders process the input obtained after image and feature level perturbations. This allows the student model to generate unique feature representations and enhances its learning capability. The teacher model uses two decoders to construct an alternating supervision mechanism. The two decoders of the teacher model take turns outputting pseudo-labels to guide the training process of the student model. This alternating supervision strategy can provide richer supervision signals for student model while helping to alleviate weight coupling between teacher and student. The experiments on two remote sensing image datasets show that compared with the state-of-the-art (SOTA) semi-supervised semantic segmentation methods, the method proposed demonstrates excellent competitiveness. Yujia Fu, Mingyang Wang 0001, Gemine Vivone, Yunhong Ding |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Security Weaknesses of Copilot-Generated Code in GitHub Projects: An Empirical StudyabstractModern code generation tools utilizing AI models like Large Language Models have gained increased popularity due to their ability to produce functional code. However, their usage presents security challenges, often resulting in insecure code merging into the code base. Thus, evaluating the quality of generated code, especially its security, is crucial. While prior research explored various aspects of code generation, the focus on security has been limited, mostly examining code produced in controlled environments rather than open source development scenarios. To address this gap, we conducted an empirical study, analyzing code snippets generated by GitHub Copilot and two other AI code generation tools (i.e., CodeWhisperer and Codeium) from GitHub projects. Our analysis identified 733 snippets, revealing a high likelihood of security weaknesses, with 29.5% of Python and 24.2% of JavaScript snippets affected. These issues span 43 Common Weakness Enumeration (CWE) categories, including significant ones like CWE-330: Use of Insufficiently Random Values , CWE-94: Improper Control of Generation of Code , and CWE-79: Cross-site Scripting . Notably, eight of those CWEs are among the 2023 CWE Top-25, highlighting their severity. We further examined using Copilot Chat to fix security issues in Copilot-generated code by providing Copilot Chat with warning messages from the static analysis tools, and up to 55.5% of the security issues can be fixed. We finally provide the suggestions for mitigating security issues in generated code. Yujia Fu, Peng Liang 0001, Amjed Tahir, Zengyang Li, Mojtaba Shahin, Jinfu Chen 0006 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2024 | How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good DataabstractYejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu |
EMNLP | 8 |
| 2024 | Copilot-in-the-Loop: Fixing Code Smells in Copilot-Generated Python Code using CopilotabstractAs one of the most popular dynamic languages, Python experiences a decrease in readability and maintainability when code smells are present. Recent advancements in Large Language Models have sparked growing interest in AI-enabled tools for both code generation and refactoring. GitHub Copilot is one such tool that has gained widespread usage. Copilot Chat, released in September 2023, functions as an interactive tool aimed at facilitating natural language-powered coding. However, limited attention has been given to understanding code smells in Copilot-generated Python code and Copilot Chat's ability to fix the code smells. To this end, we built a dataset comprising 102 code smells in Copilot-generated Python code. Our aim is to first explore the occurrence of code smells in Copilot-generated Python code and then evaluate the effectiveness of Copilot Chat in fixing these code smells employing different prompts. The results show that 8 out of 10 types of code smells can be detected in Copilot-generated Python code, among which Multiply-Nested Container is the most common one. For these code smells, Copilot Chat achieves a highest fixing rate of 87.1%, showing promise in fixing Python code smells generated by Copilot itself. In addition, the effectiveness of Copilot Chat in fixing these smells can be improved by providing more detailed prompts. Beiqi Zhang, Peng Liang 0001, Qiong Feng, Yujia Fu, Zengyang Li |
ASE | 4 |