VLDB 2026 Research / reviewers in the wild / expert
Chenyu Yan
dblp:y/ChenyuYan
· DBLP profile ↗
23ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Theory of computation · 12 · 1 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical DatasetabstractJunhong Lai, Shuzhong Lai, Yanhao Yu, Wanlin Chen, Chenyu Yan, Haifeng Li, Lin Yao, Yueming Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhong Lai, Shuzhong Lai, Yanhao Yu, Wanlin Chen, Chenyu Yan, Lin Yao 0002, Yueming Wang 0001 |
ACL (1) | 5 |
| 2025 | Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text DetectionabstractLarge Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness. Xiaoye Zhu, Yiwen Yuan, Chak Tou Leong, Zuchao Li, Tang Long, Chenyu Yan, Guanghao Mei, Lefei Zhang |
AAAI | 11 |
| 2025 | Virtual Production: Global Collaboration and ChallengesabstractVirtual Production (VP) is a novel approach to filmmaking that blends real-time computer graphics, live-action footage, and cutting-edge technologies (such as LED volumes and game engines like Unreal Engine). As China’s higher education (HE) landscape continues to grow and diversify, Sino- foreign higher education institutions (SfHEIs) — partnerships between local Chinese universities and international institutions — have become increasingly prominent: They were created as an innovative solution to challenges such as insufficient domestic capacity, outdated curricula, regional imbalances, limited global engagement, and quality-assurance gaps, and have themselves become centers of pedagogical, research, and institutional innovation. This paper examines the experience of a team of students and faculty at an SfHEI who collaborated on a new approach to support VP, using in-camera virtual effects (ICVFX). The paper explores the background to the project, its core technical innovations, and the team dynamics. Parallel to the technical development work, an Open Educational Resource (OER) was also created. This OER contains not only the technical elements of the project, thus helping future VP/ICVFX workers, but also the team-development experiences. The paper will be of interest not only to the VP/ICVFX community, but also to SfHEI students and staff, and to the OER community. Juin Yang Lam, Changyu Li, Dave Towey, Lynne Chen, Levi Dean, Filippo Gilardi, Omar Zahran, Chenyu Yan |
COMPSAC | 10 |
| 2025 | A novel approach for understanding the parking demand for internal access roads in hub car parks
Qianyi Hu, Chenyu Yan, Jun Cheng 0005, Changyin Dong |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Just-in-time defect prediction for software hunksabstractAbstract Just‐in‐time defect prediction can remind software developers and managers to verify and fix bugs at the moment they appeared, thus improving the effectiveness and validity of bug fixing. Existing studies mainly focus on just‐in‐time prediction for software files (JIT‐F). JIT‐F is a binary classification problem, which classifies (hence predicts) a file change as buggy or clean. This article provides a detailed analysis of just‐in‐time defect prediction for software hunks (JIT‐H), which predicts bugs at a finer level of granularity, and hence further improves the efficiency of bug fixing. Classification is performed using the ensemble technique of bagging—aggregated combinations of random under sampling plus multiple classifiers (J48 and Random Forest). An empirical study with 10 open source projects was conducted to validate the effectiveness of JIT‐H. Experimental results show that JIT‐H is effective at predicting defects in software hunk changes. Compared with JIT‐F, JIT‐H is more cost effective. Additionally, analysis on the change features indicates that Text Vector features and hunk change level features are of more importance than features in other groups and levels. Xiaoyan Zhu 0003, Chenyu Yan, E. James Whitehead Jr., Binbin Niu, Lei Zhu 0011, Long Pan |
Softw. Pract. Exp. | 2 |
| 2013 | Speeding up distributed request-response workflowsabstractWe found that interactive services at Bing have highly variable datacenter-side processing latencies because their processing consists of many sequential stages, parallelization across 10s-1000s of servers and aggregation of responses across the network. To improve the tail latency of such services, we use a few building blocks: reissuing laggards elsewhere in the cluster, new policies to return incomplete results and speeding up laggards by giving them more resources. Combining these building blocks to reduce the overall latency is non-trivial because for the same amount of resource (e.g., number of reissues), different stages improve their latency by different amounts. We present Kwiken, a framework that takes an end-to-end view of latency improvements and costs. It decomposes the problem of minimizing latency over a general processing DAG into a manageable optimization over individual stages. Through simulations with production traces, we show sizable gains; the 99th percentile of latency improves by over 50% when just 0.1% of the responses are allowed to have partial results and by over 40% for 25% of the services when just 5% extra resources are used for reissues. Virajith Jalaparti, Peter Bodík, Srikanth Kandula, Ishai Menache, Mikhail Rybalkin, Chenyu Yan |
SIGCOMM | 6 |
| 2012 | Zeta: scheduling interactive services with partial executionabstractThis paper presents a scheduling model for a class of interactive services in which requests are time bounded and lower result quality can be traded for shorter execution time. These applications include web search engines, finance servers, and other interactive, on-line services. We develop an efficient scheduling algorithm, Zeta, that allocates processor time among service requests to maximize the quality and minimize the variance of the response. Yuxiong He, Sameh Elnikety, James R. Larus, Chenyu Yan |
SoCC | 4 |
| 2012 | Compact and low delay routing labeling scheme for Unit Disk Graphs
Chenyu Yan, Yang Xiang 0007, Feodor F. Dragan |
Comput. Geom. | 1 |
| 2010 | Collective Tree Spanners in Graphs with Bounded Parameters
Feodor F. Dragan, Chenyu Yan |
Algorithmica | 2 |
| 2010 | Network flow spannersabstractAbstract In this article, motivated by applications of ordinary (distance) spanners in communication networks and to address such issues as bandwidth constraints on network links, link failures, network survivability, etc., we introduce a new notion of flow spanner, where one seeks a spanning subgraph H = (V, E') of a graph G = (V, E) which provides a “good” approximation of the source‐sink flows in G. We formulate several variants of this problem and investigate their complexities. Special attention is given to the version where H is required to be a tree. © 2009 Wiley Periodicals, Inc. NETWORKS, 2010 Feodor F. Dragan, Chenyu Yan |
Networks | 2 |
| 2009 | Compact and Low Delay Routing Labeling Scheme for Unit Disk Graphs
Chenyu Yan, Yang Xiang 0007, Feodor F. Dragan |
WADS | 1 |
| 2008 | Single-level integrity and confidentiality protection for distributed shared memory multiprocessorsabstractMultiprocessor computer systems are currently widely used in commercial settings to run critical applications. These applications often operate on sensitive data such as customer records, credit card numbers, and financial data. As a result, these systems are the frequent targets of attacks because of the potentially significant gain an attacker could obtain from stealing or tampering with such data. This provides strong motivation to protect the confidentiality and integrity of data in commercial multiprocessor systems through architectural support. Architectural support is able to protect against software-based attacks, and is necessary to protect against hardware-based attacks. In this work, we propose architectural mechanisms to ensure data confidentiality and integrity in Distributed Shared Memory multiprocessors which utilize a point-to-point based interconnection network. Our approach improves upon previous work in this area, mainly in the fact that our approach reduces performance overheads by significantly reducing the amount of cryptographic operations required. Evaluation results show that our approach can protect data confidentiality and integrity in a 16-processor DSM system with an average overhead of 1.6% and a maximum of only 7% across all SPLASH-2 applications. Brian Rogers, Chenyu Yan, Siddhartha Chhabra, Milos Prvulovic, Yan Solihin |
HPCA | 2 |
| 2008 | Collective Additive Tree Spanners of Homogeneously Orderable Graphs
Feodor F. Dragan, Chenyu Yan, Yang Xiang 0007 |
LATIN | 2 |
| 2007 | Spanners for bounded tree-length graphs
Yon Dourisboure, Feodor F. Dragan, Cyril Gavoille, Chenyu Yan |
Theor. Comput. Sci. | 4 |
| 2006 | Distance Approximating Trees: Complexity and Algorithms
Feodor F. Dragan, Chenyu Yan |
CIAC | 2 |
| 2006 | Improving Cost, Performance, and Security of Memory Encryption and AuthenticationabstractProtection from hardware attacks such as snoopers and mod chips has been receiving increasing attention in computer architecture. This paper presents a new combined memory encryption/authentication scheme. Our new split counters for counter-mode encryption simultaneously eliminate counter overflow problems and reduce per-block counter size, and we also dramatically improve authentication performance and security by using the Galois/counter mode of operation (GCM), which leverages counter-mode encryption to reduce authentication latency and overlap it with memory accesses. Our results indicate that the split-counter scheme has a negligible overhead even with a small (32KB) counter cache and using only eight counter bits per data block. The combined encryption/authentication scheme has an IPC overhead of 5% on average across SPEC CPU 2000 benchmarks, which is a significant improvement over the 20% overhead of existing encryption/authentication schemes Chenyu Yan, Daniel Englender, Milos Prvulovic, Brian Rogers, Yan Solihin |
ISCA | 1 |
| 2006 | Network Flow Spanners
Feodor F. Dragan, Chenyu Yan |
LATIN | 2 |
| 2006 | Collective tree spanners of graphsabstractIn this paper we introduce a new notion of collective tree spanners. We say that a graph G=(V,E)admits a system of $\mu$ collective additive tree r-spanners if there is a system T(G) of at most $\mu$ spanning trees of G such that for any two vertices x,y of G a spanning tree T\in \cT(G) exists such that d_T(x,y)\leq d_G(x,y)+r. Among other results, we show that any chordal graph, chordal bipartite graph or cocomparability graph admits a system of at most log 2 n collective additive tree 2-spanners. These results are complemented by lower bounds, which say that any system of collective additive tree 1-spanners must have $\Omega(\sqrt{n})$ spanning trees for some chordal graphs and $\Omega(n)$ spanning trees for some chordal bipartite graphs and some cocomparability graphs. Furthermore, we show that any c-chordal graph admits a system of at most log 2 n collective additive tree (2\lfloor c/2\rfloor)-spanners, any circular-arc graph admits a system of two collective additive tree 2-spanners. Towards establishing these results, we present a general property for graphs, called (\al,r)$-decomposition, and show that any $(\al,r)$-decomposable graph G with n vertices admits a system of at most $\log_{1/\al} n$ collective additive tree $2r$-spanners. We discuss also an application of the collective tree spanners to the problem of designing compact and efficient routing schemes in graphs. For any graph on n vertices admitting a system of at most $\mu$ collective additive tree r-spanners, there is a routing scheme of deviation r with addresses and routing tables of size $O(\mu \log^2n/\log \log n)$ bits per vertex. This leads, for example, to a routing scheme of deviation $(2\lfloor c/2\rfloor)$ with addresses and routing tables of size $O(\log^3n/\log \log n)$ bits per vertex on the class of c-chordal graphs. Feodor F. Dragan, Chenyu Yan, Irina Lomonosov |
SIAM J. Discret. Math. | 2 |
| 2005 | Collective Tree Spanners in Graphs with Bounded Genus, Chordality, Tree-Width, or Clique-Width
Feodor F. Dragan, Chenyu Yan |
ISAAC | 2 |
| 2005 | Collective Tree 1-Spanners for Interval Graphs
Derek G. Corneil, Feodor F. Dragan, Ekkehard Köhler, Chenyu Yan |
WG | 4 |
| 2005 | Additive sparse spanners for graphs with bounded length of largest induced cycle
Victor Chepoi, Feodor F. Dragan, Chenyu Yan |
Theor. Comput. Sci. | 3 |
| 2004 | Collective Tree Spanners and Routing in AT-free Related Graphs
Feodor F. Dragan, Chenyu Yan, Derek G. Corneil |
WG | 2 |
| 2003 | Additive Spanners for k-Chordal Graphs
Victor Chepoi, Feodor F. Dragan, Chenyu Yan |
CIAC | 3 |