VLDB 2026 Research / reviewers in the wild / expert
Hanzhuo Tan
dblp:271/5760
· DBLP profile ↗
7ranked-venue papers
6as first author
6since 2021 · last 2026
0000-0001-5392-5435ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Security and privacy · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prompt-Based Code Completion via Multi-Retrieval Augmented GenerationabstractAutomated code completion, aiming at generating subsequent tokens from unfinished code, has significantly benefited from recent progress in pre-trained Large Language Models (LLMs). However, these models often suffer from coherence issues and hallucinations when dealing with complex code logic or extrapolating beyond their training data. Existing Retrieval Augmented Generation (RAG) techniques partially address these issues by retrieving relevant code with a separate encoding model where the retrieved snippet serves as contextual reference for code completion. However, their retrieval scope is subject to a singular perspective defined by the encoding model, which largely overlooks the complexity and diversity inherent in code semantics. To address this limitation, we propose ProCC, a code completion framework leveraging prompt engineering and the contextual multi-armed bandits algorithm to flexibly incorporate and adapt to multiple perspectives of code. ProCC first employs a prompt-based multi-retriever system which crafts prompt templates to elicit LLM knowledge to understand code semantics with multiple retrieval perspectives. Then, it adopts the adaptive retrieval selection algorithm to incorporate code similarity into the decision-making process to determine the most suitable retrieval perspective for the LLM to complete the code. Experimental results demonstrate that ProCC outperforms a widely studied code completion technique RepoCoder by 7.92% on the public benchmark CCEval, 3.19% in HumanEval-Infilling, 2.80% on our collected open-source benchmark suite, and 4.48% on the private-domain benchmark suite collected from Kuaishou Technology in terms of Exact Match. ProCC also allows augmenting fine-tuned techniques in a plug-and-play manner, yielding an averaged 6.5% improvement over the fine-tuned model. Hanzhuo Tan, Qi Luo 0001, Zizheng Zhan, Jing Li 0049, Haotian Zhang 0026, Yuqun Zhang |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2025 | Decompile-Bench: Million-Scale Binary-Source Function Pairs for Real-World Binary DecompilationabstractRecent advances in LLM-based decompilers have been shown effective to convert low-level binaries into human-readable source code. However, there still lacks a comprehensive benchmark that provides large-scale binary-source function pairs, which is critical for advancing the LLM decompilation technology. Creating accurate binary-source mappings incurs severe issues caused by complex compilation settings and widespread function inlining that obscure the correspondence between binaries and their original source code. Previous efforts have either relied on used contest‐style benchmarks, synthetic binary–source mappings that diverge significantly from the mappings in real world, or partially matched binaries with only code lines or variable names, compromising the effectiveness of analyzing the binary functionality. To alleviate these issues, we introduce Decompile-Bench, the first open-source dataset comprising two million binary-source function pairs condensed from 100 million collected function pairs, i.e., 450GB of binaries compiled from permissively licensed GitHub projects. For the evaluation purposes, we also developed a benchmark Decompile-Bench-Eval including manually crafted binaries from the well-established HumanEval and MBPP, alongside the compiled GitHub repositories released after 2025 to mitigate data leakage issues. We further explore commonly-used evaluation metrics to provide a thorough assessment of the studied LLM decompilers and find that fine-tuning with Decompile-Bench causes a 20% improvement over previous benchmarks in terms of the re-executability rate. Our code and data has been released in HuggingFace and Github. https://github.com/anonepo/LLM4Decompile Hanzhuo Tan, Xiaolong Tian, Hanrui Qi, Zuchen Gao, Qi Luo 0001, Yuqun Zhang |
NeurIPS | 1 |
| 2025 | HICL: Hashtag-Driven In-Context Learning for Social Media Natural Language UnderstandingabstractNatural language understanding (NLU) is integral to various social media applications. However, the existing NLU models rely heavily on context for semantic learning, resulting in compromised performance when faced with short and noisy social media content. To address this issue, we leverage in-context learning (ICL), wherein language models learn to make inferences by conditioning on a handful of demonstrations to enrich the context and propose a novel hashtag-driven ICL (HICL) framework. Concretely, we pretrain a model #Encoder, which employs #hashtags (user-annotated topic labels) to drive BERT-based pretraining through contrastive learning. Our objective here is to enable #Encoder to gain the ability to incorporate topic-related semantic information, which allows it to retrieve topic-related posts to enrich contexts and enhance social media NLU with noisy contexts. To further integrate the retrieved context with the source text, we employ a gradient-based method to identify trigger terms useful in fusing information from both sources. For empirical studies, we collected 45 M tweets to set up an in-context NLU benchmark, and the experimental results on seven downstream tasks show that HICL substantially advances the previous state-of-the-art results. Furthermore, we conducted an extensive analysis and found that the following hold: 1) combining source input with a top-retrieved post from #Encoder is more effective than using semantically similar posts and 2) trigger words can largely benefit in merging context from the source and retrieved posts. Hanzhuo Tan, Chunpu Xu, Jing Li 0049, Yuqun Zhang, Zeyang Fang, Baohua Lai |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | LLM4Decompile: Decompiling Binary Code with Large Language ModelsabstractDecompilation aims to convert binary code to high-level source code, but traditional tools like Ghidra often produce results that are difficult to read and execute.Motivated by the advancements in Large Language Models (LLMs), we propose LLM4Decompile, the first and largest open-source LLM series (1.3B to 33B) trained to decompile binary code.We optimize the LLM training process and introduce the LLM4Decompile-End models to decompile binary directly.The resulting models significantly outperform GPT-4o and Ghidra on the HumanEval and ExeBench benchmarks over 100% in terms of re-executability rate.Additionally, we improve the standard refinement approach to fine-tune the LLM4Decompile-Ref models, enabling them to effectively refine the decompiled code from Ghidra and achieve a further 16.2% improvement over the LLM4Decompile-End.LLM4Decompile 1 demonstrates the potential of LLMs to revolutionize binary code decompilation, delivering remarkable improvements in readability and executability while complementing conventional tools for optimal results. Hanzhuo Tan, Qi Luo 0001, Jing Li 0049, Yuqun Zhang |
EMNLP | 1 |
| 2022 | An extensive study on pre-trained models for program understanding and generationabstractAutomatic program understanding and generation techniques could significantly advance the productivity of programmers and have been widely studied by academia and industry. Recently, the advent of pre-trained paradigm enlightens researchers to develop general-purpose pre-trained models which can be applied for a broad range of program understanding and generation tasks. Such pre-trained models, derived by self-supervised objectives on large unlabelled corpora, can be fine-tuned in downstream tasks (such as code search and code generation) with minimal adaptations. Although these pre-trained models claim superiority over the prior techniques, they seldom follow equivalent evaluation protocols, e.g., they are hardly evaluated on the identical benchmarks, tasks, or settings. Consequently, there is a pressing need for a comprehensive study of the pre-trained models on their effectiveness, versatility as well as the limitations to provide implications and guidance for the future development in this area. To this end, we first perform an extensive study of eight open-access pre-trained models over a large benchmark on seven representative code tasks to assess their reproducibility. We further compare the pre-trained models and domain-specific state-of-the-art techniques for validating pre-trained effectiveness. At last, we investigate the robustness of the pre-trained models by inspecting their performance variations under adversarial attacks. Through the study, we find that while we can in general replicate the original performance of the pre-trained models on their evaluated tasks and adopted benchmarks, subtle performance fluctuations can refute the findings in their original papers. Moreover, none of the existing pre-trained models can dominate over all other models. We also find that the pre-trained models can significantly outperform non-pre-trained state-of-the-art techniques in program understanding tasks. Furthermore, we perform the first study for natural language-programming language pre-trained model robustness via adversarial attacks and find that a simple random attack approach can easily fool the state-of-the-art pre-trained models and thus incur security issues. At last, we also provide multiple practical guidelines for advancing future research on pre-trained models for program understanding and generation. Zhengran Zeng, Hanzhuo Tan, Haotian Zhang 0026, Jing Li 0049, Yuqun Zhang, Lingming Zhang 0001 |
ISSTA | 2 |
| 2021 | Minutiae Attention Network With Reciprocal Distance Loss for Contactless to Contact-Based Fingerprint IdentificationabstractInteroperability between contactless and conventional contact-based fingerprint recognition systems is fundamental for the success of emerging contactless fingerprint technologies which are highly sought, especially due to current pandemic. However, image formation differences and acquisition distortions between these two modalities pose significant challenges for such interoperability. In order to address these challenges, this paper presents a minutiae attention network with Siamese architecture and the reciprocal distance loss function to enable more accurate contactless to contact-based fingerprint identification. The proposed network contains two branches, a global-net branch to recover global features and a minutiae attention branch that focuses on the local minutiae areas. Attention mechanism is introduced to guide the minutiae attention branch to concentrate on distorted areas and recover minutiae/features correspondence for contactless and contact-based fingerprint images from the same fingers. Meanwhile, reciprocal distance loss is specifically designed to impose strong penalty towards contactless and contact-based fingerprint images from different fingers and guide the network to learn robust features for distinguishing identities. Experimental results on two publicly available databases illustrate significant performance improvements, over state-of-art methods in the literature, and validate the effectiveness of the proposed framework for the contactless to contact-based fingerprint identification. Hanzhuo Tan, Ajay Kumar 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Towards More Accurate Contactless Fingerprint Minutiae Extraction and Pose-Invariant MatchingabstractContactless fingerprint identification offers significantly higher user convenience, hygiene and has attracted increasing attention for the deployments. However, the presentation of fingers towards the contactless fingerprint sensors is hard to control and often results in unwanted pose changes that significantly degrade the contactless fingerprint matching accuracy. In order to address such problems and improve the fingerprint matching accuracy, this paper proposes a more precise minutiae extraction and pose-compensation approach. As compared with the conventional minutiae extraction approaches, our deep neural network-based approach does not require any image enhancement and is robust to spurious minutiae. All the minutiae extracted from our network are subjected to a three stage pose compensation framework: a) view angle estimation based on the location of core point, b) ellipsoid model formulation which simulates and compensate finger pose, c) intersection area estimation and alignment between different view angles. The proposed ellipsoid model is adaptive to both the silhouette of 2D contactless fingerprint image and the estimated view angle. The corresponding area between the different view angles can be theoretically estimated using this model and incorporated to align two contactless fingerprints for achieving superior matching accuracy. Our reproducible experimental results presented in this paper using public databases, and a database acquired during this work, validate the effectiveness of the proposed framework over the commercial software and earlier methods. Hanzhuo Tan, Ajay Kumar 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |