VLDB 2026 Research / reviewers in the wild / expert
Jinman Zhao
dblp:160/8761
· DBLP profile ↗
18ranked-venue papers
7as first author
13since 2021 · last 2026
0009-0001-1511-8875ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Theory of computation · 2Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Step Pruning: Information Theory Based Step-level Optimization for Self-Refining Large Language ModelsabstractLarge language models (LLMs) have shown impressive capabilities in natural language tasks, yet they continue to struggle with multi-step mathematical reasoning, where correctness depends on a precise chain of intermediate steps. Preference optimization methods such as Direct Preference Optimization (DPO) have improved answer-level alignment, but they often overlook the reasoning process itself, providing little supervision over intermediate steps that are critical for complex problem-solving. Existing fine-grained approaches typically rely on strong annotators or reward models to assess the quality of individual steps. However, reward models are vulnerable to reward hacking. To address this, we propose ISLA, a reward-model-free framework that constructs step-level preference data directly from SFT gold traces. ISLA also introduces a self-improving pruning mechanism that identifies informative steps based on two signals: their marginal contribution to final accuracy (relative accuracy) and the model’s uncertainty, inspired by the concept of information gain. Empirically, ISLA achieves better performance than DPO while using only 12% of the training tokens, demonstrating that careful step-level selection can significantly improve both reasoning accuracy and training efficiency. Jinman Zhao, Erxue Min, Ziheng Li 0003, Zexu Sun, Hengyi Cai, Shuaiqiang Wang, Xu Chen 0017, Gerald Penn |
AAAI | 1 |
| 2026 | CodeHacker: Automated Test Case Generation for Detecting Vulnerabilities in Competitive Programming SolutionsabstractThe evaluation of Large Language Models (LLMs) for code generation relies heavily on the quality and robustness of test cases.However, existing benchmarks often lack coverage for subtle corner cases, allowing incorrect solutions to pass.To bridge this gap, we propose CodeHacker, an automated agent framework dedicated to generating targeted adversarial test cases that expose latent vulnerabilities in program submissions.Mimicking the hack mechanism in competitive programming, CodeHacker employs a multistrategy approach, including stress testing, anti-hash attacks, and logic-specific targeting to break specific code submissions.To ensure the validity and reliability of these attacks, we introduce a Calibration Phase, where the agent iteratively refines its own Validator and Checker via self-generated adversarial probes before evaluating contestant code.Experiments demonstrate that CodeHacker significantly improves the True Negative Rate (TNR) of existing datasets, effectively filtering out incorrect solutions that were previously accepted.Furthermore, generated adversarial cases prove to be superior training data, boosting the performance of RL-trained models on benchmarks like LiveCodeBench.Our code is available at https://github.com/shi0712/ CodeHacker. Jingwei Shi, Xinxiang Yin, Shengyu Tao, Jinman Zhao |
ACL (1) | 5 |
| 2025 | PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR AccuracyabstractShuhao Guan, Moule Lin, Cheng Xu, Xinyi Liu, Jinman Zhao, Jiexin Fan, Qi Xu, Derek Greene. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuhao Guan, Moule Lin, Cheng Xu 0006, Jinman Zhao, Jiexin Fan, Derek Greene |
ACL (1) | 5 |
| 2025 | LongRecipe: Recipe for Efficient Long Context Generalization in Large Language ModelsabstractLarge language models (LLMs) face significant challenges in handling long-context tasks because of their limited effective context window size during pretraining, which restricts their ability to generalize over extended sequences. Meanwhile, extending the context window in LLMs through post-pretraining is highly resource-intensive.To address this, we introduce LongRecipe, an efficient training strategy for extending the context window of LLMs, including impactful token analysis, position index transformation, and training optimization strategies. It simulates long-sequence inputs while maintaining training efficiency and significantly improves the model’s understanding of long-range dependencies. Experiments on three types of LLMs show that LongRecipe can utilize long sequences while requiring only 30% of the target context window size, and reduces computational training resource over 85% compared to full sequence training. Furthermore, LongRecipe also preserves the original LLM’s capabilities in general tasks. Ultimately, we can extend effective context window of open-source LLMs from 8k to 128k, achieving performance close to GPT-4 with just one day of dedicated training using a single GPU with 80G memory.Our code is released at https://github.com/zhiyuanhubj/LongRecipe. Jinman Zhao, Suyuchen Wang, WangYan WangYan, Wei Shen 0005, Qing Gu 0001, Anh Tuan Luu, See-Kiong Ng, Zhiwei Jiang 0001, Bryan Hooi |
ACL (1) | 3 |
| 2025 | UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter Efficient Fine-Tuning of Large ModelsabstractThis paper introduces UoRA, a novel parameter-efficient fine-tuning (PEFT) approach for large language models (LLMs). UoRA achieves state-of-the-art efficiency by leveraging a low-rank approximation method that reduces the number of trainable parameters without compromising performance. Unlike existing methods such as LoRA and VeRA, UoRA employs a re-parametrization mechanism that eliminates the need to adapt frozen projection matrices while maintaining shared projection layers across the model. This results in halving the trainable parameters compared to LoRA and outperforming VeRA in computation and storage efficiency. Comprehensive experiments across various benchmarks demonstrate UoRA’s superiority in achieving competitive fine-tuning performance with minimal computational overhead. We demonstrate its performance on GLUE and E2E benchmarks and is effectiveness in instruction-tuning large language models and image classification models. Our contributions establish a new paradigm for scalable and resource-efficient fine-tuning of LLMs. Jinman Zhao, Zhifei Yang 0004, Yibo Zhong, Shuhao Guan, Linbo Cao |
ACL (1) | 2 |
| 2025 | Inside-Outside Algorithm for Probabilistic Product-Free Lambek Categorial GrammarabstractThe inside-outside algorithm is widely utilized in statistical models related to context-free grammars. It plays a key role in the EM estimation of probabilistic context-free grammars. In this work, we introduce an inside-outside algorithm for Probabilistic Lambek Categorical Grammar (PLCG) Jinman Zhao, Gerald Penn |
COLING | 1 |
| 2025 | Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-TuningabstractIn this work, we propose FoRA-UA, a novel method that, using only 1-5% of the standard LoRA's parameters, achieves state-ofthe-art performance across a wide range of tasks.Specifically, we explore scenarios with extremely limited parameter budgets and derive two key insights: (1) fix-sized sparse frequency representations approximate small matrices more accurately; and (2) with a fixed number of trainable parameters, introducing a smaller intermediate representation to approximate larger matrices results in lower construction error.These findings form the foundation of our FoRA-UA method.By inserting a small intermediate parameter set, we achieve greater model compression without sacrificing performance.We evaluate FoRA-UA across diverse tasks, including natural language understanding (NLU), natural language generation (NLG), instruction tuning, and image classification, demonstrating strong generalisation and robustness under extreme compression. 1 Jinman Zhao, Jiaru Li, Jingcheng Niu, Yulan Hu, Erxue Min, Gerald Penn |
EMNLP | 1 |
| 2025 | MSDet: Receptive Field Enhanced Multiscale Detection for Tiny Pulmonary NoduleabstractPulmonary nodules are critical for early lung cancer diagnosis, but traditional CT imaging methods suffer from low detection rates and poor localization. Small nodule detection is challenging due to subtle differences in density and issues like occlusion. Existing methods such as FPN, with its fixed feature fusion and limited receptive field, struggle to effectively overcome these issues. To address these challenges, our paper proposed three key contributions: Firstly, we proposed MSDet, a multiscale attention and receptive field network for detecting tiny pulmonary nodules. Secondly, we proposed the extended receptive domain (ERD) strategy to capture richer contextual information and reduce false positives caused by nodule occlusion. We also proposed the position channel attention mechanism (PCAM) to optimize feature learning and reduce multiscale detection errors, and designed the tiny object detection block (TODB) to enhance the detection of tiny nodules. Experiments on the LUNA16 dataset show an 8.8% improvement in mAP over YOLOv8, achieving state-of-the-art performance. The code is available at https://github.com/CaiGuoHui123/MSDet. Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang 0006, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhibin Liao, Yang Zhao 0019, Ying Cai 0002 |
ICME | 7 |
| 2025 | PolyMorphous: An MLIR-Based Polyhedral Compiler with Loop Transformation PrimitivesabstractWe present PolyMorphous, an MLIR-based polyhedral compiler that exposes a set of loop-based scheduling primitives, providing users with ample control over optimizations for input code. The primitives are expressed in a new Schedule dialect that is based on MLIR's Transform dialect. PolyMorphous' polyhedral engine collapses the primitives into a polyhedral schedule and checks for its legality. If necessary, and possible, the engine corrects an illegal schedule into a legal one. PolyMorphous is evaluated using the PolyBench suite. The evaluation validates PolyMorphous' approach in two ways. First, it shows that PolyMorphous enables users to explore the optimization space, and that it results in well-optimized code. Second, it shows that PolyMorphous' correction of illegal schedules enables users to optimize code with fewer primitives and improves their productivity by alleviating the need for manual correction. Specifically, the evaluation shows that PolyMorphous optimized code performs as well as code optimized by Pluto+ with its optimization flags tuned to achieve the best performance for each benchmark. Further, for several benchmarks, PolyMorphous optimized code performs better, confirming the value of providing users with control over optimizations. Averaged over all the benchmarks, PolyMorphous optimized code has a speedup of$1.24 \mathrm{x} / 1.25 \mathrm{x}$over that optimized by Pluto+ on Arm/X86 systems. PolyMorphous brings the approach of empowering users of polyhedral compilers with control over optimizations to a community-developed infrastructure, promoting the approach's adoptability. It does so while delivering performant code. Jinman Zhao, Seyed Aryan Vahabpour, Xingyu Yue, Kai-Ting Amy Wang, Tarek S. Abdelrahman |
IPDPS | 1 |
| 2024 | A Generative Model for Lambek Categorial SequentsabstractIn this work, we introduce a generative model, PLC+, for generating Lambek Categorial Grammar(LCG) sequents. We also introduce a simple method to numerically estimate the model’s parameters from an annotated corpus. Then we compare our model with probabilistic context-free grammars (PCFGs) and show that PLC+ simultaneously assigns a higher probability to a common corpus, and has greater coverage. Jinman Zhao, Gerald Penn |
LREC/COLING | 1 |
| 2023 | Better Context Makes Better Code Language Models: A Case Study on Function Call Argument CompletionabstractPretrained code language models have enabled great progress towards program synthesis. However, common approaches only consider in-file local context and thus miss information and constraints imposed by other parts of the codebase and its external dependencies. Existing code completion benchmarks also lack such context. To resolve these restrictions we curate a new dataset of permissively licensed Python packages that includes full projects and their dependencies and provide tools to extract non-local information with the help of program analyzers. We then focus on the task of function call argument completion which requires predicting the arguments to function calls. We show that existing code completion models do not yield good results on our completion task. To better solve this task, we query a program analyzer for information relevant to a given function call, and consider ways to provide the analyzer results to different code completion models during inference and training. Our experiments show that providing access to the function implementation and function usages greatly improves the argument completion performance. Our ablation study provides further insights on how different types of information available from the program analyzer and different ways of incorporating the information affect the model performance. Hengzhi Pei, Jinman Zhao, Leonard Lausen, Sheng Zha, George Karypis |
AAAI | 2 |
| 2023 | Large Language Models of Code Fail at Completing Code with Potential BugsabstractLarge language models of code (Code-LLMs) have recently brought tremendous advances to code completion, a fundamental feature of programming assistance and code intelligence. However, most existing works ignore the possible presence of bugs in the code context for generation, which are inevitable in software development. Therefore, we introduce and study the buggy-code completion problem, inspired by the realistic scenario of real-time code suggestion where the code context contains potential bugs – anti-patterns that can become bugs in the completed program. To systematically study the task, we introduce two datasets: one with synthetic bugs derived from semantics-altering operator changes (buggy-HumanEval) and one with realistic bugs derived from user submissions to coding problems (buggy-FixEval). We find that the presence of potential bugs significantly degrades the generation performance of the high-performing Code-LLMs. For instance, the passing rates of CODEGEN-2B-MONO on test cases of buggy-HumanEval drop more than 50% given a single potential bug in the context. Finally, we investigate several post-hoc methods for mitigating the adverse effect of potential bugs and find that there remains a large gap in post-mitigation performance. Tuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho, Leonard Lausen, Sheng Zha, George Karypis |
NeurIPS | 2 |
| 2021 | Code Prediction by Feeding Trees to TransformersabstractCode prediction, more specifically autocomplete, has become an essential feature in modern IDEs. Autocomplete is more effective when the desired next token is at (or close to) the top of the list of potential completions offered by the IDE at cursor position. This is where the strength of the underlying machine learning system that produces a ranked order of potential completions comes into play. We advance the state-of-the-art in the accuracy of code prediction (next token prediction) used in autocomplete systems. Our work uses Transformers as the base neural architecture. We show that by making the Transformer architecture aware of the syntactic structure of code, we increase the margin by which a Transformer-based system outperforms previous systems. With this, it outperforms the accuracy of several state-of-the-art next token prediction systems by margins ranging from 14% to 18%. We present in the paper several ways of communicating the code structure to the Transformer, which is fundamentally built for processing sequence data. We provide a comprehensive experimental evaluation of our proposal, along with alternative design choices, on a standard Python dataset, as well as on Facebook internal Python corpus. Our code and data preparation pipeline will be available in open source. Seohyun Kim 0001, Jinman Zhao, Yuchi Tian, Satish Chandra 0001 |
ICSE | 2 |
| 2019 | Counting hypergraph matchings up to uniqueness thresholdabstractWe study the problem of approximately counting matchings in hypergraphs of bounded maximum degree and maximum size of hyperedges. With an activity parameter λ , each matching M is assigned a weight λ | M | . The counting problem is formulated as computing a partition function that gives the sum of the weights of all matchings in a hypergraph. This problem unifies two extensively studied statistical physics models in approximate counting: the hardcore model (graph independent sets) and the monomer–dimer model (graph matchings). For this problem, the critical activity λ c = d d k ( d − 1 ) d + 1 is the threshold for the uniqueness of Gibbs measures on the infinite ( d + 1 ) -uniform ( k + 1 ) -regular hypertree. Consider hypergraphs of maximum degree at most k + 1 and maximum size of hyperedges at most d + 1 . We show that when λ < λ c , there is an FPTAS for computing the partition function; and when λ = λ c , there is a PTAS for computing the log-partition function. These algorithms are based on the decay of correlation (strong spatial mixing) property of Gibbs distributions. When λ > 2 λ c , there is no PRAS for the partition function or the log-partition function unless NP = RP. Towards obtaining a sharp transition of computational complexity of approximate counting, we study the local convergence from a sequence of finite hypergraphs to the infinite lattice with specified symmetry. We show a surprising connection between the local convergence and the reversibility of a natural random walk. This leads us to a barrier for the hardness result: The non-uniqueness of infinite Gibbs measure is not realizable by any finite gadgets. Renjie Song, Yitong Yin, Jinman Zhao |
Inf. Comput. | 3 |
| 2018 | Generalizing Word Embeddings using Bag of SubwordsabstractWe approach the problem of generalizing pretrained word embeddings beyond fixed-size vocabularies without using additional contextual information.We propose a subwordlevel word vector generation model that views words as bags of character n-grams.The model is simple, fast to train and provides good vectors for rare or unseen words.Experiments show that our model achieves stateof-the-art performances in English word similarity task and in joint prediction of part-ofspeech tag and morphosyntactic attributes in 23 languages, suggesting our model's ability in capturing the relationship between words' textual representations and their embeddings. Jinman Zhao, Sidharth Mudgal, Yingyu Liang |
EMNLP | 1 |
| 2018 | The Effect of Network Width on the Performance of Large-batch TrainingabstractDistributed implementations of mini-batch stochastic gradient descent (SGD) suffer from communication overheads, attributed to the high frequency of gradient updates inherent in small-batch training. Training with large batches can reduce these overheads; however it besets the convergence of the algorithm and the generalization performance. In this work, we take a first step towards analyzing how the structure (width and depth) of a neural network affects the performance of large-batch training. We present new theoretical results which suggest that--for a fixed number of parameters--wider networks are more amenable to fast large-batch training compared to deeper ones. We provide extensive experiments on residual and fully-connected neural networks which suggest that wider networks can be trained using larger batches without incurring a convergence slow-down, unlike their deeper variants. Lingjiao Chen, Hongyi Wang 0001, Jinman Zhao, Dimitris S. Papailiopoulos, Paraschos Koutris |
NeurIPS | 3 |
| 2018 | Neural-augmented static analysis of Android communicationabstractWe address the problem of discovering communication links between applications in the popular Android mobile operating system, an important problem for security and privacy in Android. Any scalable static analysis in this complex setting is bound to produce an excessive amount of false-positives, rendering it impractical. To improve precision, we propose to augment static analysis with a trained neural-network model that estimates the probability that a communication link truly exists. We describe a neural-network architecture that encodes abstractions of communicating objects in two applications and estimates the probability with which a link indeed exists. At the heart of our architecture are type-directed encoders (TDE), a general framework for elegantly constructing encoders of a compound data type by recursively composing encoders for its constituent types. We evaluate our approach on a large corpus of Android applications, and demonstrate that it achieves very high accuracy. Further, we conduct thorough interpretability studies to understand the internals of the learned neural networks. Jinman Zhao, Aws Albarghouthi, Vaibhav Rastogi, Somesh Jha, Damien Octeau |
ESEC/SIGSOFT FSE | 1 |
| 2016 | Counting Hypergraph Matchings up to Uniqueness ThresholdabstractWe study the problem of approximately counting matchings in hypergraphs of bounded maximum degree and maximum size of hyperedges. With an activity parameter lambda, each matching M is assigned a weight lambda^{|M|}. The counting problem is formulated as computing a partition function that gives the sum of the weights of all matchings in a hypergraph. This problem unifies two extensively studied statistical physics models in approximate counting: the hardcore model (graph independent sets) and the monomer-dimer model (graph matchings). For this model, the critical activity lambda_c= (d^d)/(k (d-1)^{d+1}) is the threshold for the uniqueness of Gibbs measures on the infinite (d+1)-uniform (k+1)-regular hypertree. Consider hypergraphs of maximum degree at most k+1 and maximum size of hyperedges at most d+1. We show that when lambda < lambda_c, there is an FPTAS for computing the partition function; and when lambda = lambda_c, there is a PTAS for computing the log-partition function. These algorithms are based on the decay of correlation (strong spatial mixing) property of Gibbs distributions. When lambda > 2lambda_c, there is no PRAS for the partition function or the log-partition function unless NP=RP. Towards obtaining a sharp transition of computational complexity of approximate counting, we study the local convergence from a sequence of finite hypergraphs to the infinite lattice with specified symmetry. We show a surprising connection between the local convergence and the reversibility of a natural random walk. This leads us to a barrier for the hardness result: The non-uniqueness of infinite Gibbs measure is not realizable by any finite gadgets. Renjie Song, Yitong Yin, Jinman Zhao |
APPROX-RANDOM | 3 |