VLDB 2026 Research / reviewers in the wild / expert
Haolin Jin
dblp:207/8891
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0007-0875-6813ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trust in One Round: Confidence Estimation for Large Language Models via Structural Signals
Pengyue Yang, Jiawen Wen, Haolin Jin, Linghan Huang, Huaming Chen, Ling Chen 0006 |
WWW | 3 |
| 2026 | Are LLMs reliable code reviewers? systematic overcorrection in requirement conformance judgementabstractAbstract Large language models (LLMs) have become essential tools in software development, widely used for requirements engineering, code generation and review tasks. Software engineers often rely on LLMs to verify if code implementation satisfy task requirements, thereby ensuring code robustness and accuracy. However, it remains unclear whether LLMs can reliably determine code against the given task descriptions, which is usually in a form of natural language specifications. In this paper, we uncover a systematic failure of LLMs in matching code to natural language requirements. Specifically, with widely adopted benchmarks and unified prompts design, we demonstrate that LLMs frequently misclassify correct code implementation as non-compliant or defective. Surprisingly, we find that more detailed prompt design, particularly with those requiring explanations and proposed corrections, leads to higher misjudgment rates, highlighting critical reliability issues for LLM-based code assistants. We further analyze the mechanisms driving these failures and evaluate the reliability of rationale-required judgments. Building on these findings, we propose a Fix-guided Verification Filter that treats the model proposed fix as executable counterfactual evidence, and validates the original and revised implementations using benchmark tests and spec-constrained augmented tests. Our results expose previously under-explored limitations in LLM-based code review capabilities, and provide practical guidance for integrating LLM-based reviewers with safeguards in automated review and development pipelines. Haolin Jin, Huaming Chen |
Autom. Softw. Eng. | 1 |
| 2025 | Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language SpecificationsabstractLarge language models (LLMs) have become essential tools in software development, widely used for requirements engineering, code generation and review tasks. Software engineers increasingly rely on LLMs to assess whether system code implementation satisfy task requirements, thereby enhancing code robustness and accuracy. However, it remains unclear whether LLMs can reliably determine whether the code complies fully with the given task descriptions, which is usually natural language specifications. In this paper, we uncover a systematic failure of LLMs in evaluating whether code aligns with natural language requirements. Specifically, using widely adopted benchmarks, we employ unified prompts to judge code correctness. Our results reveal that LLMs frequently misclassify correct code implementations as either "not satisfying requirements" or containing potential defects. Surprisingly, more complex prompting, especially when leveraging prompt engineering techniques involving explanations and proposed corrections, leads to higher misjudgment rate, which highlights the critical reliability issues in using LLMs as code review assistants. We further analyze the root causes of these misjudgments, and propose two improved prompting strategies for mitigation. For the first time, our findings reveals unrecognized limitations in LLMs to match code with requirements. We also offer novel insights and practical guidance for effective use of LLMs in automated code review and task-oriented agent scenarios. Haolin Jin, Huaming Chen |
ASE | 1 |
| 2020 | Discovering differential features: Adversarial learning for information credibility evaluation
Lianwei Wu, Yuan Rao 0004, Ambreen Nazir, Haolin Jin |
Inf. Sci. | 4 |
| 2019 | Different Absorption from the Same Sharing: Sifted Multi-task Learning for Fake News DetectionabstractLianwei Wu, Yuan Rao, Haolin Jin, Ambreen Nazir, Ling Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lianwei Wu, Yuan Rao 0004, Haolin Jin, Ambreen Nazir, Ling Sun 0004 |
EMNLP/IJCNLP (1) | 3 |