VLDB 2026 Research / reviewers in the wild / expert
Wenyuan Jiang
dblp:198/3262
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Test of Time: Rethinking Temporal Signal of Benchmark ContaminationabstractTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter, Punya Syon Pandey, Keenan Samway, Wenyuan Jiang, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Terry Jingchen Zhang, Gopal Dev, Max Obreiter, Wenyuan Jiang, Punya Syon Pandey, Keenan Samway, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin 0001 |
ACL (1) | 5 |
| 2026 | SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM SystemsabstractYuzhe Zhang, Feiran Liu, Yi Shan, Xinyi Huang, Xin Yang, Yueqi Zhu, Xuxin Cheng, Cao Liu, Ke Zeng, Terry Jingchen Zhang, Wenyuan Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Feiran Liu, Yi Shan 0001, Yueqi Zhu, Xuxin Cheng, Cao Liu, Terry Jingchen Zhang, Wenyuan Jiang |
ACL (1) | 11 |
| 2025 | Measuring and Augmenting Large Language Models for Solving Capture-the-Flag ChallengesabstractCapture-the-Flag (CTF) competitions are crucial for cybersecurity education and training. With the evolution of large language models (LLMs), there is growing interest in their ability to automate CTF challenge solving, with DARPA's AIxCC competition (since 2023) being a notable example. However,this demands a combination of multiple abilities of LLMs, from knowledge to reasoning and further to actions. In this paper, we highlight the importance of technical knowledge in solving CTF problems and deliberately construct a focused benchmark, CTFKnow, with 3,992 questions to measure LLMs' performance in this core aspect. Our study offers a focused and innovative measurement of LLMs' capability in understanding CTF knowledge and applying it to solve CTF challenges. Our key findings reveal that while LLMs possess substantial technical knowledge, they struggle to apply it accurately to specific scenarios and adapt based on feedback from CTF environments. Zimo Ji, Daoyuan Wu, Wenyuan Jiang, Pingchuan Ma 0004, Zongjie Li, Shuai Wang 0011 |
CCS | 3 |
| 2025 | Exploring the Jupyter Ecosystem: An Empirical Study of Bugs and VulnerabilitiesabstractBackground. Jupyter notebooks are one of the main tools used by data scientists. Notebooks include features (configuration scripts, markdown, images, etc.) that make them challenging to analyze compared to traditional software. As a result, existing software engineering models, tools, and studies do not capture the uniqueness of Notebook's behavior. Aims. This paper aims to provide a large-scale empirical study of bugs and vulnerabilities in the Notebook ecosystem. Method. We collected and analyzed a large dataset of Notebooks from two major platforms. Our methodology involved quantitative analyses of notebook characteristics (such as complexity metrics, contributor activity, and documentation) to identify factors correlated with bugs. Additionally, we conducted a qualitative study using grounded theory to categorize notebook bugs, resulting in a comprehensive bug taxonomy. Finally, we analyzed security-related commits and vulnerability reports to assess risks associated with Notebook deployment frameworks. Results. Our findings highlight that configuration issues are among the most common bugs in notebook documents, followed by incorrect API usage. Finally, we explore common vulnerabilities associated with popular deployment frameworks to better understand risks associated with Notebook development. Conclusions. This work highlights that notebooks are less wellsupported than traditional software, resulting in more complex code, misconfiguration, and poor maintenance. Wenyuan Jiang, Diany Pressato, Harsh Darji, Thibaud Lutellier |
ESEM | 1 |
| 2025 | Leveraging Runtime Information for LLM Quantization
Wenyuan Jiang |
PRICAI | 3 |
| 2025 | Can LLMs Write Fast System-Aware Numerical Computation Code?abstractWhile Large Language Models (LLMs) have demonstrated impressive capabilities in code generation and mathematical reasoning, their ability to produce correct and highly optimized numerical computation code remains largely unexplored. This paper presents a systematic evaluation framework for assessing LLMs’ performance in generating computationally efficient numerical code. We introduce a benchmark comprising carefully curated classical numerical computation problems where human experts typically achieve 3-10× speedups compared to straightforward implementations. These problems require a deep understanding of numerical computation and computer architecture, identification of performance bottlenecks, and application of multilevel optimization techniques. To address potential dataset contamination, we also develop a benchmark generator that creates novel variants of these optimization challenges. Our evaluation of mainstream LLMs reveals that even state-of-the-art models consistently fall short of human expert optimization levels. The benchmark is open-sourced to facilitate further research in this direction. Bintao Tang, Zimo Ji, Wenyuan Jiang |
SMC | 5 |
| 2025 | An Obfuscator for Securing Ring Confidential Transactions' Signing Keys of CryptocurrenciesabstractRing Confidential Transaction (RingCT) protocols are widely used in cryptocurrencies to protect user privacy. Consequently, a corresponding digital signature scheme, such as a ring signature scheme that hides the signers’ identities, is required. Accordingly, the security of a RingCT protocol depends on the confidentiality of the secret signing keys of the underlying ring signature scheme. However, existing solutions like hardware wallets, Trusted Execution Environments (TEEs), and threshold signature schemes have limitations such as specified expensive hardware, targeting attacks at CPUs on insufficiently secure hardware, and overheads caused by multiple parties. On the contrary, program obfuscation for signature schemes offers advantages over these existing approaches. Concretely, we propose a novel obfuscator that secures the secret keys of the concise linkable spontaneous anonymous group (CLSAG) signature scheme, which is the latest ring signature scheme used in Monero’s RingCT protocol. To achieve enhanced security, the proposed obfuscator leverages Paillier homomorphic encryption to transform secret keys into an obfuscated form resistant to attacks. The security of the proposed obfuscator has been formally proved. Computational efficiency has been both theoretically analyzed and experimentally evaluated with positive results on various testing platforms. Yang Shi 0002, Minyu Teng, Tianyuan Luo, Wenyuan Jiang, Jiayao Gao, Man Ho Au |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Killing Two Birds with One Stone: Cross-modal Reinforced Prompting for Graph and Language TasksabstractIn recent years, Graph Neural Networks (GNNs) and Large Language Models (LLMs) have exhibited remarkable capability in addressing different graph learning and natural language tasks, respectively. Motivated by this, integrating LLMs with GNNs has been increasingly studied to acquire transferable knowledge across modalities, which leads to improved empirical performance in language and graph domains. However, existing studies mainly focused on a single-domain scenario by designing complicated integration techniques to manage multimodal data effectively. Therefore, a concise and generic learning framework for multi-domain tasks, i.e., graph and language domains, is highly desired yet remains under-exploited due to two major challenges. First, the language corpus of downstream tasks differs significantly from graph data, making it hard to bridge the knowledge gap between modalities. Second, not all knowledge demonstrates immediate benefits for downstream tasks, potentially introducing disruptive noise to context-sensitive models like LLMs. To tackle these challenges, we propose a novel plug-and-play framework for incorporating a lightweight cross-domain prompting method into both language and graph learning tasks. Specifically, we first convert the textual input into a domain-scalable prompt, which not only preserves the semantic and logical contents of the textual input, but also highlights related graph information as external knowledge for different domains. Then, we develop a reinforcement learning-based method to learn the optimal edge selection strategy for useful knowledge extraction, which profoundly sharpens the multi-domain model capabilities. In addition, we introduce a joint multi-view optimization module to regularize agent-level collaborative learning across two domains. Finally, extensive empirical justifications over 23 public and synthetic datasets demonstrate that our approach can be applied to diverse multi-domain tasks more accurately, robustly, and reasonably, and improve the performances of the state-of-the-art graph and language models in different learning paradigms. Wenyuan Jiang, Wenwei Wu, Le Zhang 0010, Zixuan Yuan, Jingbo Zhou 0003, Hui Xiong 0001 |
KDD | 1 |