VLDB 2026 Research / reviewers in the wild / expert
Hantian Ding
dblp:242/8095
· DBLP profile ↗
9ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 41% Efficient and distributed learning · 20% Information extraction and text analysis · 12% | |
| Software engineering, system software, and programming languages
4 papers |
Program synthesis and code generation · 84% Empirical software engineering · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program synthesis and code generation › code completion
fill-in-the-middle |
0.9 | 1 | 2025 | Planning-Aware Code Infilling via Horizon-Length Prediction · EMNLP 2025 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.8 | 1 | 2024 | Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
code representation learning |
0.8 | 1 | 2024 | Code Representation Learning at Scale · ICLR 2024 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.8 | 1 | 2024 | Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024 |
Machine learning › Efficient and distributed learning
KV cache management |
0.8 | 1 | 2024 | Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024 |
Natural language and speech › Language models and text generation
large language model training |
0.8 | 1 | 2024 | Fewer Truncations Improve Language Modeling · ICML 2024 |
Natural language and speech › Language models and text generation › large language model training › language model pretraining
pretraining data curation |
0.8 | 1 | 2024 | Fewer Truncations Improve Language Modeling · ICML 2024 |
Program synthesis and code generation
code language model |
0.8 | 1 | 2024 | Code Representation Learning at Scale · ICLR 2024 |
Natural language and speech › Language models and text generation
code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation
code completion |
0.7 | 1 | 2023 | CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion · NeurIPS 2023 |
Program synthesis and code generation
code generation evaluation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Program synthesis and code generation › code generation with language models
multilingual code generation |
0.7 | 1 | 2023 | Multi-lingual Evaluation of Code Generation Models · ICLR 2023 |
Bioinformatics and computational biology
protein engineering |
0.4 | 1 | 2020 | Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering · RECOMB 2020 |
Bioinformatics and computational biology › protein sequence analysis
protein sequence modeling |
0.4 | 1 | 2020 | Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering · RECOMB 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.4 | 1 | 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › lexical semantics › verb semantics
selectional preference |
0.4 | 1 | 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › lexical semantics › verb semantics
selectional preference acquisition |
0.4 | 1 | 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019 |
Natural language and speech › Language models and text generation
code language models |
0.3 | 1 | 2025 | Planning-Aware Code Infilling via Horizon-Length Prediction · EMNLP 2025 |
Information retrieval › document retrieval › domain-specific retrieval
code search |
0.2 | 1 | 2023 | CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.1 | 1 | 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
winograd schema challenge |
0.1 | 1 | 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019 |
Methods — techniques the papers use, named apart from their topics
next-token prediction · 1.7lookahead planning · 1.7horizon-length prediction · 1.7masked language modeling · 1.5contrastive learning · 1.5static analysis · 1.3large language model · 1.3language model prompting · 1.3multi-query attention · 0.8combinatorial optimization · 0.8best-fit packing · 0.8GEMM · 0.8deep sequence modeling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Planning-Aware Code Infilling via Horizon-Length PredictionabstractFill-in-the-Middle (FIM), or infilling, has become integral to code language models, enabling generation of missing code given both left and right contexts.However, the current FIM training paradigm which performs next-token prediction (NTP) over reordered sequence often leads to models struggling to generate content that aligns well with the surrounding context.We hypothesize that NTP alone is insufficient for models to learn effective planning conditioned on the distant right context, a critical factor for successful code infilling.To overcome this, we propose Horizon-Length Prediction (HLP), a novel training objective that teaches models to predict the number of remaining middle tokens at each step.HLP advances FIM with lookahead planning, enabling models to inherently learn infilling boundaries for arbitrary left and right contexts without relying on dataset-specific post-processing.Our evaluation across different model families and sizes shows that HLP significantly improves FIM performance by up to 24% relatively on diverse benchmarks, across file-level and repository-level.Furthermore, the enhanced planning capability gained through HLP boosts model performance on code reasoning.Importantly, HLP incurs negligible training overhead and no additional inference cost, ensuring its practicality for real-world scenarios. Hantian Ding, Shiqi Wang 0002, Qing Sun 0013, Zijian Wang 0002 |
EMNLP | 2 |
| 2024 | Code Representation Learning at ScaleabstractRecent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size. Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang |
ICLR | 4 |
| 2024 | Bifurcated Attention for Single-Context Large-Batch SamplingabstractIn our study, we present bifurcated attention, a method developed for language model inference in single-context batch sampling contexts. This approach aims to reduce redundant memory IO costs, a significant factor in latency for high batch sizes and long context lengths. Bifurcated attention achieves this by dividing the attention mechanism during incremental decoding into two distinct GEMM operations, focusing on the KV cache from prefill and the decoding process. This method ensures precise computation and maintains the usual computational load (FLOPs) of standard attention mechanisms, but with reduced memory IO. Bifurcated attention is also compatible with multi-query attention mechanism known for reduced memory IO for KV cache, further enabling higher batch size and context length. The resulting efficiency leads to lower latency, improving suitability for real-time applications, e.g., enabling massively-parallel answer generation without substantially increasing latency, enhancing performance when integrated with post-processing techniques such as reranking. Ben Athiwaratkun, Sujan K. Gonugondla, Sanjay Krishna Gouda, Haifeng Qian, Hantian Ding, Qing Sun 0013, Jun Wang 0022, Jiacheng Guo, Liangfu Chen, Parminder Bhatia, Ramesh Nallapati, Sudipta Sengupta, Bing Xiang |
ICML | 5 |
| 2024 | Fewer Truncations Improve Language ModelingabstractIn large language model training, input documents are typically concatenated together and then split into sequences of equal length to avoid padding tokens. Despite its efficiency, the concatenation approach compromises data integrity—it inevitably breaks many documents into incomplete pieces, leading to excessive truncations that hinder the model from learning to compose logically coherent and factually consistent content that is grounded on the complete context. To address the issue, we propose Best-fit Packing, a scalable and efficient method that packs documents into training sequences through length-aware combinatorial optimization. Our method completely eliminates unnecessary truncations while retaining the same training efficiency as concatenation. Empirical results from both text and code pre-training show that our method achieves superior performance (e.g., +4.7% on reading comprehension; +16.8% in context following; and +9.2% on program synthesis), and reduces closed-domain hallucination effectively by up to 58.3%. Hantian Ding, Zijian Wang 0002, Giovanni Paolini, Anoop Deoras, Dan Roth 0001, Stefano Soatto |
ICML | 1 |
| 2023 | Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati |
ICLR | 12 |
| 2023 | CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code CompletionabstractCode completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers. Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang |
NeurIPS | 4 |
| 2023 | Regularized multi-trait multi-locus linear mixed models for genome-wide association studies and genomic selection in cropsabstractBACKGROUND: We consider two key problems in genomics involving multiple traits: multi-trait genome wide association studies (GWAS), where the goal is to detect genetic variants associated with the traits; and multi-trait genomic selection (GS), where the emphasis is on accurately predicting trait values. Multi-trait linear mixed models build on the linear mixed model to jointly model multiple traits. Existing estimation methods, however, are limited to the joint analysis of a small number of genotypes; in fact, most approaches consider one SNP at a time. Estimating multi-dimensional genetic and environment effects also results in considerable computational burden. Efficient approaches that incorporate regularization into multi-trait linear models (no random effects) have been recently proposed to identify genomic loci associated with multiple traits (Yu et al. in Multitask learning using task clustering with applications to predictive modeling and GWAS of plant varieties. arXiv:1710.01788 , 2017; Yu et al in Front Big Data 2:27, 2019), but these ignore population structure and familial relatedness (Yu et al in Nat Genet 38:203-208, 2006). RESULTS: This work addresses this gap by proposing a novel class of regularized multi-trait linear mixed models along with scalable approaches for estimation in the presence of high-dimensional genotypes and a large number of traits. We evaluate the effectiveness of the proposed methods using datasets in maize and sorghum diversity panels, and demonstrate benefits in both achieving high prediction accuracy in GS and in identifying relevant marker-trait associations. CONCLUSIONS: The proposed regularized multivariate linear mixed models are relevant for both GWAS and GS. We hope that they will facilitate agronomy-related research in plant biology and crop breeding endeavors. Aurélie C. Lozano, Hantian Ding, Naoki Abe, Alexander E. Lipka |
BMC Bioinform. | 2 |
| 2020 | Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering
Yunan Luo, Lam Vo, Hantian Ding, Yufeng Su, Yang Liu 0097, Wesley Wei Qian, Huimin Zhao 0007, Jian Peng 0001 |
RECOMB | 3 |
| 2019 | SP-10K: A Large-scale Evaluation Set for Selectional Preference AcquisitionabstractSelectional Preference (SP) is a commonly observed language phenomenon and proved to be useful in many natural language processing tasks.To provide a better evaluation method for SP models, we introduce SP-10K, a largescale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.Three representative SP acquisition methods based on pseudo-disambiguation are evaluated with SP-10K.To demonstrate the importance of our dataset, we investigate the relationship between SP-10K and the commonsense knowledge in ConceptNet5 and show the potential of using SP to represent the commonsense knowledge.We also use the Winograd Schema Challenge to prove that the proposed new SP relations are essential for the hard pronoun coreference resolution problem. Hongming Zhang 0009, Hantian Ding, Yangqiu Song |
ACL (1) | 2 |