VLDB 2026 Research / reviewers in the wild / expert
Yushi Cao
dblp:274/2297
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-2886-4121ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding TasksabstractLarge Language Models (LLMs) are increasingly used not only to generate code, but also to judge it: comparing, ranking, or scoring competing solutions.However, their reliability in this evaluative role remains poorly understood.Inconsistent or flawed judgments can undermine benchmarks and distort training signals.This paper investigates the performance and robustness of LLMs when used as code judges.We introduce CodeJudgeBench, a benchmark explicitly designed to evaluate LLM-as-a-Judge models across three critical coding tasks: code generation, code repair, and unit test generation.We comprehensively benchmark the performance of 26 LLM-as-a-Judge models, encompassing general-purpose, code-tuned, and reasoning models.Our empirical findings reveal that relatively small reasoning models (e.g., Qwen3-8B) can outperform much larger non-reasoning models up to 70B.We further stress-test robustness by applying both general and code-specific perturbations.All models show significant instability and are sensitive to changes such as response ordering, variable naming, and misleading comments.These findings highlight serious concerns about the consistency and robustness of LLM-based judges for coding tasks. Hongchao Jiang, Yiming Chen 0010, Yushi Cao, Hung-yi Lee, Robby T. Tan |
ACL (1) | 3 |
| 2025 | Logic-Q: Improving Deep Reinforcement Learning-based Quantitative Trading via Program Sketch-based TuningabstractDeep reinforcement learning (DRL) has revolutionized quantitative trading (Q-trading) by achieving decent performance without significant human expert knowledge. Despite its achievements, we observe that the current state-of-the-art DRL models are still ineffective in identifying the market trends, causing them to miss good trading opportunity or suffer from large drawdowns when encountering market crashes. To address this limitation, a natural approach is to incorporate human expert knowledge in identifying market trends. Whereas, such knowledge is abstract and hard to be quantified. In order to effectively leverage abstract human expert knowledge, in this paper, we propose a universal logic-guided deep reinforcement learning framework for Q-trading, called Logic-Q. In particular, Logic-Q adopts the program synthesis by sketching paradigm and introduces a logic-guided model design that leverages a lightweight, plug-and-play market trend-aware program sketch to determine the market trend and correspondingly adjusts the DRL policy in a post-hoc manner. Extensive evaluations of two popular quantitative trading tasks demonstrate that Logic-Q can significantly improve the performance of previous state-of-the-art DRL trading strategies. Junzhe Jiang 0002, Yushi Cao, Aixin Cui, Bozhi Wu, Bo Li 0037, Yang Liu 0003, Danny Dongning Sun |
AAAI | 3 |
| 2024 | Unveiling Project-Specific Bias in Neural Code ModelsabstractDeep learning has introduced significant improvements in many software analysis tasks. Although the Large Language Models (LLMs) based neural code models demonstrate commendable performance when trained and tested within the intra-project independent and identically distributed (IID) setting, they often struggle to generalize effectively to real-world inter-project out-of-distribution (OOD) data. In this work, we show that this phenomenon is caused by the heavy reliance on project-specific shortcuts for prediction instead of ground-truth evidence. We propose a Cond-Idf measurement to interpret this behavior, which quantifies the relatedness of a token with a label and its project-specificness. The strong correlation between model behavior and the proposed measurement indicates that without proper regularization, models tend to leverage spurious statistical cues for prediction. Equipped with these observations, we propose a novel bias mitigation mechanism that regularizes the model’s learning behavior by leveraging latent logic relations among samples. Experimental results on two representative program analysis tasks indicate that our mitigation framework can improve both inter-project OOD generalization and adversarial robustness, while not sacrificing accuracy on intra-project IID data. Yanzhou Li, Tianlin Li, Mengnan Du, Bozhi Wu, Yushi Cao, Junzhe Jiang 0002, Yang Liu 0003 |
LREC/COLING | 6 |
| 2024 | Improving Neural Logic Machines via Failure ReflectionabstractReasoning is a fundamental ability towards artificial general intelligence (AGI). Fueled by the success of deep learning, the neural logic machines models (NLMs) have introduced novel neural-symbolic structures and demonstrate great performance and generalization on reasoning and decision-making tasks. However, the original training approaches of the NLMs are still far from perfect, the models would repeat similar mistakes during the training process which leads to sub-optimal performance. To mitigate this issue, we present a novel framework named Failure Reflection Guided Regularizer (FRGR). FRGR first dynamically identifies and summarizes the root cause if the model repeats similar mistakes during training. Then it penalizes the model if it makes similar mistakes in future training iterations. In this way, the model is expected to avoid repeating errors of similar root causes and converge faster to a better-performed optimum. Experimental results on multiple relational reasoning and decision-making tasks demonstrate the effectiveness of FRGR in improving performance, generalization, training efficiency, and data efficiency. Yushi Cao, Yan Zheng 0002, Xu Liu 0014, Bozhi Wu, Tianlin Li, Xiufeng Xu, Junzhe Jiang 0002, Yon Shin Teo, Shangwei Lin 0001, Yang Liu 0003 |
ICML | 2 |
| 2024 | Reinventing Node-centric Traffic Forecasting for Improved Accuracy and Efficiency
Xu Liu 0014, Yuxuan Liang 0002, Chao Huang 0001, Hengchang Hu, Yushi Cao, Bryan Hooi, Roger Zimmermann |
ECML/PKDD (3) | 5 |
| 2023 | An Automatic Test Plan Generation Approach for Automotive Software TestingabstractThe automotive industry is shifting from hardware-centric to software-centric with the emergence of various intelligent features powered by software. This poses a new challenge for software testers to ensure software reliability by designing test plans that satisfy the test objectives while abiding by the constraints like scope, time, as well as various automotive safety standards. This paper proposed an automatic test plan generation framework built on the evolutionary algorithm. A novel encoding mechanism is proposed to represent the multi-dimensional test plan, while a belief model is proposed to reveal the underlying correlations between the relevant test attributes. Experiments conducted on an actual automotive software in production environment developed by our industry partner show that our method can achieve around 50% improvements in finding defects and covering high-priority test cases as compared to typical evolutionary algorithms while abiding by multiple constraints such as the total run time and custom objectives set by users. Yushi Cao, Yanran Li, Yon Shin Teo, Yan Zheng 0002, Zhexin Liang, Shangwei Lin 0001 |
SoMeT | 1 |
| 2022 | GALOIS: Boosting Deep Reinforcement Learning via Generalizable Logic SynthesisabstractDespite achieving superior performance in human-level control problems, unlike humans, deep reinforcement learning (DRL) lacks high-order intelligence (e.g., logic deduction and reuse), thus it behaves ineffectively than humans regarding learning and generalization in complex problems. Previous works attempt to directly synthesize a white-box logic program as the DRL policy, manifesting logic-driven behaviors. However, most synthesis methods are built on imperative or declarative programming, and each has a distinct limitation, respectively. The former ignores the cause-effect logic during synthesis, resulting in low generalizability across tasks. The latter is strictly proof-based, thus failing to synthesize programs with complex hierarchical logic. In this paper, we combine the above two paradigms together and propose a novel Generalizable Logic Synthesis (GALOIS) framework to synthesize hierarchical and strict cause-effect logic programs. GALOIS leverages the program sketch and defines a new sketch-based hybrid program language for guiding the synthesis. Based on that, GALOIS proposes a sketch-based program synthesis method to automatically generate white-box programs with generalizable and interpretable cause-effect logic. Extensive evaluations on various decision-making tasks with complex logic demonstrate the superiority of GALOIS over mainstream baselines regarding the asymptotic performance, generalizability, and great knowledge reusability across different environments. Yushi Cao, Tianpei Yang, Hao Zhang 0004, Yan Zheng 0002, Yi Li 0008, Jianye Hao, Yang Liu 0003 |
NeurIPS | 1 |
| 2022 | A Holistic Automated Software Structure Exploration Framework for TestingabstractExploring the underlying structure of a Human-Machine Interface (HMI) product effectively while adhering to the pre-defined test conditions and methodology is critical for validating the quality of the software. We propose an reinforcement-learning powered Automated Software Structure Exploration Framework for Testing (ASSET), which is capable of interacting with and analyzing the HMI software under testing (SUT). The main challenge is to incorporate the human instructions into the ASSET phase by using the visual feedback such as the downloaded image sequence from the HMI, which could be difficult to analyze. Our framework combines both computer vision and natural language processing techniques to understand the semantic meanings of the visual feedback. Building on the semantic understanding, we develop a rules-guided software exploration algorithm via reinforcement learning and deterministic finite automaton (DFA). We conducted experiments on HMI software in actual production phase and demonstrate that the exploration coverage and efficiency of our framework outperforms current start-of-art methods. Yushi Cao, Yon Shin Teo, Yan Zheng 0002, Yuxuan Toh, Shangwei Lin 0001 |
SoMeT | 1 |
| 2022 | Fair and accurate age prediction using distribution aware data curation and augmentationabstractDeep learning-based facial recognition systems have experienced increased media attention due to exhibiting unfair behavior. Large enterprises, such as IBM, shut down their facial recognition and age prediction systems as a consequence. Age prediction is an especially difficult application with the issue of fairness remaining an open research problem (e.g., predicting age for different ethnicity equally accurate). One of the main causes of unfair behavior in age prediction methods lies in the distribution and diversity of the training data. In this work, we present two novel approaches for dataset curation and data augmentation in order to increase fairness through balanced feature curation and increase diversity through distribution aware augmentation. To achieve this, we introduce out-of-distribution detection to the facial recognition domain which is used to select the data most relevant to the deep neural network’s (DNN) task when balancing the data among age, ethnicity, and gender. Our approach shows promising results. Our best-trained DNN model outperformed all academic and industrial baselines in terms of fairness by up to 4.92 times and also enhanced the DNN’s ability to generalize outperforming Amazon AWS and Microsoft Azure public cloud systems by 31.88% and 10.95%, respectively. Yushi Cao, David Berend, Palina Tolmach, Guy Amit, Moshe Levy, Yang Liu 0003, Asaf Shabtai, Yuval Elovici |
WACV | 1 |
| 2021 | Automatic HMI Structure Exploration Via Curiosity-Based Reinforcement LearningabstractDiscovering the underlying structure of HMI software efficiently and sufficiently for the purpose of testing without any prior knowledge on the software logic remains a difficult problem. The key challenge lies in the complexity of the HMI software and the high variance in the coverage of current methods. In this paper, we introduce the PathFinder, an effective and automatic HMI software exploration framework. PathFinder adopts a curiosity-based reinforcement learning framework to choose actions that lead to the discovery of more unknown states. Additionally, PathFinder progressively builds a navigation model during the exploration to further improve state coverage. We have conducted experiments on both simulations and real-world HMI software testing environment, which comprise a full tool chain of automobile dashboard instrument cluster. The exploration coverage outperforms manual and fuzzing methods which are the current industrial standards. Yushi Cao, Yan Zheng 0002, Shangwei Lin 0001, Yang Liu 0003, Yon Shin Teo, Yuxuan Toh, Vinay Vishnumurthy Adiga |
ASE | 1 |