VLDB 2026 Research / reviewers in the wild / expert
Yifan Zhang 0019
dblp:57/4707-19
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0001-3030-9960ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Scores: Explainable Intelligent Assessment Strengthens Pre-service Teachers' Assessment LiteracyabstractAssessment literacy (AL) is essential for personalized education, yet difficult to cultivate in pre-service teachers. Conventional teacher preparation programs focus on theoretical knowledge, while digital assessment tools commonly provide opaque scores or parameters. These limitations hinder reflection and transfer, leaving AL underdeveloped. We propose XIA, an eXplainable Intelligent Assessment platform that extends statistics-informed support with visualized cognitive diagnostic reasoning, including contrastive and counterfactual explanations. In a pre-post controlled study with 21 pre-service teachers, we combined quantitative tasks and questionnaires with qualitative interviews. The findings offer preliminary evidence that XIA supported reflection, self-regulation, and assessment awareness, and helped reduce assessment errors. Interviews further showed a shift from score-based judgments toward evidence-based reasoning. This work contributes insights into the design of intelligent assessment tools, showing how explanatory scaffolding can bridge assessment theory and classroom practice and support the cultivation of AL in teacher education. Yuang Wei, Fei Wang 0063, Yifan Zhang 0019, Brian Y. Lim, Bo Jiang 0016 |
CHI | 3 |
| 2026 | Comparables XAI: Faithful Example-based AI Explanations with Counterfactual Trace AdjustmentsabstractExplaining with examples is an intuitive way to justify AI decisions. However, it is challenging to understand how a decision value should change relative to the examples with many features differing by large amounts. We draw from real estate valuation that uses Comparables—examples with known values for comparison. Estimates are made more accurate by hypothetically adjusting the attributes of each Comparable and correspondingly changing the value based on factors. We propose Comparables XAI for relatable example-based explanations of AI with Trace adjustments that trace counterfactual changes from each Comparable to the Subject, one attribute at a time, monotonically along the AI feature space. In modelling and user studies, Trace-adjusted Comparables achieved the highest XAI faithfulness and precision, user accuracy, and narrowest uncertainty bounds compared to linear regression, linearly adjusted Comparables, or unadjusted Comparables. This work contributes a new analytical basis for using example-based explanations to improve user understanding of AI decisions. Yifan Zhang 0019, Tianle Ren, Fei Wang 0063, Brian Y. Lim |
CHI | 1 |
| 2026 | Transferable XAI: Relating Understanding Across Domains with Explanation TransferabstractCurrent Explainable AI (XAI) focuses on explaining a single application, but when encountering related applications, users may rely on their prior understanding from previous explanations. This leads to either overgeneralization and AI overreliance, or burdensome independent memorization. Indeed, related decision tasks can share explanatory factors, but with some notable differences; e.g., body mass index (BMI) affects the risks for heart disease and diabetes at the same rate, but chest pain is more indicative of heart disease. Similarly, models using different attributes for the same task still share signals; e.g., temperature and pressure affect air pollution but in opposite directions due to the ideal gas law. Leveraging transfer of learning, we propose Transferable XAI to enable users to transfer understanding across related domains by explaining the relationship between domain explanations using a general affine transformation framework applied to linear factor explanations. The framework supports explanation transfer across various domain types: translation for data subspace (subsuming prior work on Incremental XAI), scaling for decision task, and mapping for attributes. Focusing on task and attributes domain types, in formative and summative user studies, we investigated how well participants could understand AI decisions from one domain to another. Compared to single-domain and domain-independent explanations, Transferable XAI was the most helpful for understanding the second domain, leading to the best decision faithfulness, factor recall, and ability to relate explanations between domains. This framework contributes to improving the reusability of explanations across related AI applications by explaining factor relationships between subspaces, tasks, and attributes. Fei Wang 0063, Yifan Zhang 0019, Brian Y. Lim |
IUI | 2 |
| 2025 | PAT-Agent: Autoformalization for Model CheckingabstractRecent advances in large language models (LLMs) offer promising potential for automating formal methods. However, applying them to formal verification remains challenging due to the complexity of specification languages, the risk of hallucinated output, and the semantic gap between natural language and formal logic. We introduce PAT-Agent, an end-to-end framework for natural language autoformalization and formal model repair that combines the generative capabilities of LLMs with the rigor of formal verification to automate the construction of verifiable formal models. In PAT-Agent, a Planning LLM first extracts key modeling elements and generates a detailed plan using semantic prompts, which then guides a Code Generation LLM to synthesize syntactically correct and semantically faithful formal models. The resulting code is verified using the Process Anal y sis Toolkit (PAT) model checker against user-specified properties, and when discrepancies occur, a Repair Loop is triggered to iteratively correct the model using counterexamples. To improve flexibility, we built a web-based interface that enables users, particularly non-FM-experts, to describe, customize, and verify system behaviors through user-LLM interactions. Experimental results on 40 systems show that PAT-Agent consistently outperforms baselines, achieving high verification success with superior efficiency. The ablation studies confirm the importance of both planning and repair components, and the user study demonstrates that our interface is accessible and supports effective formal modeling, even for users with limited formal methods experience. Xinyue Zuo, Yifan Zhang 0019, Hongshu Wang, Yufan Cai 0001, Jing Sun 0002, Jin Song Dong 0001 |
ASE | 2 |
| 2023 | On-the-Fly Adapting Code Summarization on Trainable Cost-Effective Language ModelsabstractDeep learning models are emerging to summarize source code to comment,
facilitating tasks of code documentation and program comprehension.
Scaled-up large language models trained on large open corpus have achieved good performance in such tasks.
However, in practice, the subject code in one certain project can be specific,
which may not align with the overall training corpus.
Some code samples from other projects may be contradictory and introduce inconsistencies when the models try to fit all the samples.
In this work, we introduce a novel approach, Adacom, to improve the performance of comment generators by on-the-fly model adaptation.
This research is motivated by the observation that deep comment generators
often need to strike a balance as they need to fit all the training samples.
Specifically, for one certain target code $c$,
some training samples $S_p$ could have made more contributions while other samples $S_o$ could have counter effects.
However, the traditional fine-tuned models need to fit both $S_p$ and $S_o$ from a global perspective,
leading to compromised performance for one certain target code $c$.
In this context, we design Adacom to
(1) detect whether the model might have a compromised performance on a target code $c$ and
(2) retrieve a few helpful training samples $S_p$ that have contradictory samples in the training dataset and,
(3) adapt the model on the fly by re-training the $S_p$ to strengthen the helpful samples and unlearn the harmful samples.
Our extensive experiments on 7 comment generators and 4 public datasets show that
(1) can significantly boost the performance of comment generation (BLEU4 score by on average 14.9\%, METEOR by 12.2\%, and ROUGE-L by 7.4\%),
(2) the adaptation on one code sample is cost-effective and acceptable as an on-the-fly solution, and
(3) can adapt well on out-of-distribution code samples. Yufan Cai 0001, Yun Lin 0001, Chenyan Liu, Jinglian Wu, Yifan Zhang 0019, Yeyun Gong, Jin Song Dong 0001 |
NeurIPS | 5 |
| 2023 | DeepDebugger: An Interactive Time-Travelling Debugging Approach for Deep ClassifiersabstractA deep classifier is usually trained to (i) learn the numeric representation vector of samples and (ii) classify sample representations with learned classification boundaries. Time-travelling visualization, as an explainable AI technique, is designed to transform the model training dynamics into an animation of canvas with colorful dots and territories. Despite that the training dynamics of the high-level concepts such as sample representations and classification boundaries are now observable, the model developers can still be overwhelmed by tens of thousands of moving dots across hundreds of training epochs (i.e., frames in the animation), which makes them miss important training events. Xianglin Yang, Yun Lin 0001, Yifan Zhang 0019, Linpeng Huang, Jin Song Dong 0001, Hong Mei 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2023 | Knowledge Expansion and Counterfactual Interaction for Reference-Based Phishing Detection
Yun Lin 0001, Yifan Zhang 0019, Penn Han Lee, Jin Song Dong 0001 |
USENIX Security Symposium | 3 |
| 2022 | RegMiner: mining replicable regression dataset from code repositoriesabstractIn this work, we introduce a tool, RegMiner, to automate the process of collecting replicable regression bugs from a set of Git repositories. In the code commit history, RegMiner searches for regressions where a test can pass a regression-fixing commit, fail a regressioninducing commit, and pass a previous working commit again. Technically, RegMiner (1) identifies potential regression-fixing commits from the code evolution history, (2) migrates the test and its code dependencies in the commit over the history, and (3) minimizes the compilation overhead during the regression search. Our experients show that RegMiner can successfully collect 1035 regressions over 147 projects in 8 weeks, creating the largest replicable regression dataset within the shortest period, to the best of our knowledge. In addition, our experiments further show that (1) RegMiner can construct the regression dataset with very high precision and acceptable recall, and (2) the constructed regression dataset is of high authenticity and diversity. The source code of RegMiner is available at https://github.com/SongXueZhi/RegMiner, the mined regression dataset is available at https://regminer.github.io/, and the demonstration video is available at https://youtu.be/yzcM9Y4unok. Xuezhi Song, Yun Lin 0001, Yijian Wu, Yifan Zhang 0019, Siang Hwee Ng, Xin Peng 0001, Jin Song Dong 0001, Hong Mei 0001 |
ESEC/SIGSOFT FSE | 4 |