Hantian Ding

dblp:242/8095 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 41% Efficient and distributed learning · 20% Information extraction and text analysis · 12%
Software engineering, system software, and programming languages
4 papers
Program synthesis and code generation · 84% Empirical software engineering · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Program synthesis and code generation › code completion
fill-in-the-middle
0.912025
Planning-Aware Code Infilling via Horizon-Length Prediction · EMNLP 2025
Machine learning › Deep learning architectures and training
attention mechanism
0.812024
Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
code representation learning
0.812024
Code Representation Learning at Scale · ICLR 2024
Machine learning › Efficient and distributed learning
inference efficiency
0.812024
Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024
Machine learning › Efficient and distributed learning
KV cache management
0.812024
Bifurcated Attention for Single-Context Large-Batch Sampling · ICML 2024
Natural language and speech › Language models and text generation
large language model training
0.812024
Fewer Truncations Improve Language Modeling · ICML 2024
Natural language and speech › Language models and text generation › large language model training › language model pretraining
pretraining data curation
0.812024
Fewer Truncations Improve Language Modeling · ICML 2024
Program synthesis and code generation
code language model
0.812024
Code Representation Learning at Scale · ICLR 2024
Natural language and speech › Language models and text generation
code generation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Natural language and speech › Language models and text generation › evaluation of language models
multilingual evaluation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Program synthesis and code generation
code completion
0.712023
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion · NeurIPS 2023
Program synthesis and code generation
code generation evaluation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Program synthesis and code generation › code generation with language models
multilingual code generation
0.712023
Multi-lingual Evaluation of Code Generation Models · ICLR 2023
Bioinformatics and computational biology
protein engineering
0.412020
Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering · RECOMB 2020
Bioinformatics and computational biology › protein sequence analysis
protein sequence modeling
0.412020
Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering · RECOMB 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412019
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › lexical semantics › verb semantics
selectional preference
0.412019
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › lexical semantics › verb semantics
selectional preference acquisition
0.412019
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019
Natural language and speech › Language models and text generation
code language models
0.312025
Planning-Aware Code Infilling via Horizon-Length Prediction · EMNLP 2025
Information retrieval › document retrieval › domain-specific retrieval
code search
0.212023
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion · NeurIPS 2023
Natural language and speech › Information extraction and text analysis
coreference resolution
0.112019
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
winograd schema challenge
0.112019
SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

next-token prediction · 1.7lookahead planning · 1.7horizon-length prediction · 1.7masked language modeling · 1.5contrastive learning · 1.5static analysis · 1.3large language model · 1.3language model prompting · 1.3multi-query attention · 0.8combinatorial optimization · 0.8best-fit packing · 0.8GEMM · 0.8deep sequence modeling · 0.4
YearPublicationVenuePosition
2025 Planning-Aware Code Infilling via Horizon-Length Prediction
abstract
Fill-in-the-Middle (FIM), or infilling, has become integral to code language models, enabling generation of missing code given both left and right contexts.However, the current FIM training paradigm which performs next-token prediction (NTP) over reordered sequence often leads to models struggling to generate content that aligns well with the surrounding context.We hypothesize that NTP alone is insufficient for models to learn effective planning conditioned on the distant right context, a critical factor for successful code infilling.To overcome this, we propose Horizon-Length Prediction (HLP), a novel training objective that teaches models to predict the number of remaining middle tokens at each step.HLP advances FIM with lookahead planning, enabling models to inherently learn infilling boundaries for arbitrary left and right contexts without relying on dataset-specific post-processing.Our evaluation across different model families and sizes shows that HLP significantly improves FIM performance by up to 24% relatively on diverse benchmarks, across file-level and repository-level.Furthermore, the enhanced planning capability gained through HLP boosts model performance on code reasoning.Importantly, HLP incurs negligible training overhead and no additional inference cost, ensuring its practicality for real-world scenarios.
Hantian Ding, Shiqi Wang 0002, Qing Sun 0013, Zijian Wang 0002
EMNLP2
2024 Code Representation Learning at Scale
abstract
Recent studies have shown that code language model at scale demonstrate significant performance gains on downstream tasks, i.e., code generation. However, most of the existing works on code representation learning train models at a hundred million parameter scale using very limited pretraining corpora. In this work, we fuel code representation learning with a vast amount of code data via a two-stage pretraining scheme. We first train the encoders via a mix that leverages both randomness in masking language modeling and implicit structure and semantic aspects of programming language. We then enhance the representations via contrastive learning with hard negative and hard positive constructed in an unsupervised manner. We establish an off-the-shelf encoder model that persistently outperforms the existing models on a wide variety of downstream tasks by large margins. To comprehend the factors contributing to successful code representation learning, we conduct detailed ablations and share our findings on (i) a customized and effective token-level denoising scheme for source code; (ii) the importance of hard negatives and hard positives; (iii) how the proposed bimodal contrastive learning boost the cross-lingual semantic search performance; and (iv) how the pretraining schemes decide the downstream task performance scales with the model size.
Dejiao Zhang, Wasi Uddin Ahmad, Hantian Ding, Ramesh Nallapati, Dan Roth 0001, Xiaofei Ma 0001, Bing Xiang
ICLR4
2024 Bifurcated Attention for Single-Context Large-Batch Sampling
abstract
In our study, we present bifurcated attention, a method developed for language model inference in single-context batch sampling contexts. This approach aims to reduce redundant memory IO costs, a significant factor in latency for high batch sizes and long context lengths. Bifurcated attention achieves this by dividing the attention mechanism during incremental decoding into two distinct GEMM operations, focusing on the KV cache from prefill and the decoding process. This method ensures precise computation and maintains the usual computational load (FLOPs) of standard attention mechanisms, but with reduced memory IO. Bifurcated attention is also compatible with multi-query attention mechanism known for reduced memory IO for KV cache, further enabling higher batch size and context length. The resulting efficiency leads to lower latency, improving suitability for real-time applications, e.g., enabling massively-parallel answer generation without substantially increasing latency, enhancing performance when integrated with post-processing techniques such as reranking.
Ben Athiwaratkun, Sujan K. Gonugondla, Sanjay Krishna Gouda, Haifeng Qian, Hantian Ding, Qing Sun 0013, Jun Wang 0022, Jiacheng Guo, Liangfu Chen, Parminder Bhatia, Ramesh Nallapati, Sudipta Sengupta, Bing Xiang
ICML5
2024 Fewer Truncations Improve Language Modeling
abstract
In large language model training, input documents are typically concatenated together and then split into sequences of equal length to avoid padding tokens. Despite its efficiency, the concatenation approach compromises data integrity—it inevitably breaks many documents into incomplete pieces, leading to excessive truncations that hinder the model from learning to compose logically coherent and factually consistent content that is grounded on the complete context. To address the issue, we propose Best-fit Packing, a scalable and efficient method that packs documents into training sequences through length-aware combinatorial optimization. Our method completely eliminates unnecessary truncations while retaining the same training efficiency as concatenation. Empirical results from both text and code pre-training show that our method achieves superior performance (e.g., +4.7% on reading comprehension; +16.8% in context following; and +9.2% on program synthesis), and reduces closed-domain hallucination effectively by up to 58.3%.
Hantian Ding, Zijian Wang 0002, Giovanni Paolini, Anoop Deoras, Dan Roth 0001, Stefano Soatto
ICML1
2023 Multi-lingual Evaluation of Code Generation Models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang 0002, Xiaopeng Li 0002, Wasi Uddin Ahmad, Shiqi Wang 0002, Qing Sun 0013, Mingyue Shang, Sujan K. Gonugondla, Hantian Ding, Nathan Fulton, Arash Farahani, Siddhartha Jain 0001, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati
ICLR12
2023 CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion
abstract
Code completion models have made significant progress in recent years, yet current popular evaluation datasets, such as HumanEval and MBPP, predominantly focus on code completion tasks within a single file. This over-simplified setting falls short of representing the real-world software development scenario where repositories span multiple files with numerous cross-file dependencies, and accessing and understanding cross-file context is often required to complete the code correctly. To fill in this gap, we propose CrossCodeEval, a diverse and multilingual code completion benchmark that necessitates an in-depth cross-file contextual understanding to complete the code accurately. CrossCodeEval is built on a diverse set of real-world, open-sourced, permissively-licensed repositories in four popular programming languages: Python, Java, TypeScript, and C#. To create examples that strictly require cross-file context for accurate completion, we propose a straightforward yet efficient static-analysis-based approach to pinpoint the use of cross-file context within the current file. Extensive experiments on state-of-the-art code language models like CodeGen and StarCoder demonstrate that CrossCodeEval is extremely challenging when the relevant cross-file context is absent, and we see clear improvements when adding these context into the prompt. However, despite such improvements, the pinnacle of performance remains notably unattained even with the highest-performing model, indicating that CrossCodeEval is also capable of assessing model's capability in leveraging extensive context to make better code completion. Finally, we benchmarked various methods in retrieving cross-file context, and show that CrossCodeEval can also be used to measure the capability of code retrievers.
Yangruibo Ding, Zijian Wang 0002, Wasi Uddin Ahmad, Hantian Ding, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth 0001, Bing Xiang
NeurIPS4
2023 Regularized multi-trait multi-locus linear mixed models for genome-wide association studies and genomic selection in crops
abstract
BACKGROUND: We consider two key problems in genomics involving multiple traits: multi-trait genome wide association studies (GWAS), where the goal is to detect genetic variants associated with the traits; and multi-trait genomic selection (GS), where the emphasis is on accurately predicting trait values. Multi-trait linear mixed models build on the linear mixed model to jointly model multiple traits. Existing estimation methods, however, are limited to the joint analysis of a small number of genotypes; in fact, most approaches consider one SNP at a time. Estimating multi-dimensional genetic and environment effects also results in considerable computational burden. Efficient approaches that incorporate regularization into multi-trait linear models (no random effects) have been recently proposed to identify genomic loci associated with multiple traits (Yu et al. in Multitask learning using task clustering with applications to predictive modeling and GWAS of plant varieties. arXiv:1710.01788 , 2017; Yu et al in Front Big Data 2:27, 2019), but these ignore population structure and familial relatedness (Yu et al in Nat Genet 38:203-208, 2006). RESULTS: This work addresses this gap by proposing a novel class of regularized multi-trait linear mixed models along with scalable approaches for estimation in the presence of high-dimensional genotypes and a large number of traits. We evaluate the effectiveness of the proposed methods using datasets in maize and sorghum diversity panels, and demonstrate benefits in both achieving high prediction accuracy in GS and in identifying relevant marker-trait associations. CONCLUSIONS: The proposed regularized multivariate linear mixed models are relevant for both GWAS and GS. We hope that they will facilitate agronomy-related research in plant biology and crop breeding endeavors.
Aurélie C. Lozano, Hantian Ding, Naoki Abe, Alexander E. Lipka
BMC Bioinform.2
2020 Evolutionary Context-Integrated Deep Sequence Modeling for Protein Engineering
Yunan Luo, Lam Vo, Hantian Ding, Yufeng Su, Yang Liu 0097, Wesley Wei Qian, Huimin Zhao 0007, Jian Peng 0001
RECOMB3
2019 SP-10K: A Large-scale Evaluation Set for Selectional Preference Acquisition
abstract
Selectional Preference (SP) is a commonly observed language phenomenon and proved to be useful in many natural language processing tasks.To provide a better evaluation method for SP models, we introduce SP-10K, a largescale evaluation set that provides human ratings for the plausibility of 10,000 SP pairs over five SP relations, covering 2,500 most frequent verbs, nouns, and adjectives in American English.Three representative SP acquisition methods based on pseudo-disambiguation are evaluated with SP-10K.To demonstrate the importance of our dataset, we investigate the relationship between SP-10K and the commonsense knowledge in ConceptNet5 and show the potential of using SP to represent the commonsense knowledge.We also use the Winograd Schema Challenge to prove that the proposed new SP relations are essential for the hard pronoun coreference resolution problem.
Hongming Zhang 0009, Hantian Ding, Yangqiu Song
ACL (1)2