VLDB 2026 Research / reviewers in the wild / expert
Indraneil Paul
dblp:232/3376
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0001-8215-4764ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Program synthesis and code generation · 60% Software maintenance and evolution · 21% Compilers and program optimization · 18% | |
| Artificial intelligence
3 papers |
Language models and text generation · 88% Trustworthy machine learning · 12% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
code language models |
1.1 | 2 | 2025 | ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding · ICLR 2025 IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model training
pre-training objectives |
0.9 | 1 | 2025 | ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding · ICLR 2025 |
Software maintenance and evolution › program comprehension
code comprehension |
0.9 | 1 | 2025 | ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding · ICLR 2025 |
Program synthesis and code generation › code generation evaluation
code generation benchmark |
0.9 | 1 | 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions · ICLR 2025 |
Program synthesis and code generation
code generation with language models |
0.9 | 1 | 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions · ICLR 2025 |
Compilers and program optimization
intermediate representation |
0.8 | 1 | 2024 | IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators · ACL (1) 2024 |
Program synthesis and code generation › code generation with language models
multilingual code generation |
0.8 | 1 | 2024 | IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.3 | 1 | 2025 | Droid: A Resource Suite for AI-Generated Code Detection · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
obfuscation-based pre-training · 1.7de-obfuscation objectives · 1.7cross-lingual transfer · 1.5continued causal language modelling · 1.5uncertainty-based resampling · 0.9multi-task learning · 0.9metric learning · 0.9large language model · 0.9benchmark construction · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Droid: A Resource Suite for AI-Generated Code DetectionabstractWe present DroidCollection 1 2 , the most extensive open data suite for training and evaluating machine-generated code detectors, comprising over a million code samples, seven programming languages, outputs from 43 coding models, and three real-world coding domains.Alongside fully AI-generated examples, our collection includes human-AI co-authored code, as well as adversarial examples explicitly crafted to evade detection.Subsequently, we develop DroidDetect, a suite of encoderonly detectors trained using a multi-task objective over DroidCollection.Our experiments show that existing detectors' performance fails to generalise to diverse coding domains and programming languages outside of their narrow training data.We further demonstrate that while most detectors are easily compromised by humanising the output distributions using superficial prompting and alignment approaches, this problem can be easily amended by training on a small number of adversarial examples.Finally, we demonstrate the effectiveness of metric learning and uncertainty-based resampling as way to enhance detector training on possibly noisy distributions. Daniil Orel, Indraneil Paul, Iryna Gurevych, Preslav Nakov |
EMNLP | 2 |
| 2025 | ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation GroundingabstractLanguage models (LMs) have become a staple of the code-writing toolbox. Their pre-training recipe has, however, remained stagnant over recent years, barring the occasional changes in data sourcing and filtering strategies. In particular, research exploring modifications to Code-LMs' pre-training objectives, geared towards improving data efficiency and better disentangling between syntax and semantics, has been noticeably sparse, especially compared with corresponding efforts in natural language LMs. In this work, we examine grounding on obfuscated code as a means of helping Code-LMs look beyond the surface-form syntax and enhance their pre-training sample efficiency. To this end, we compile ObscuraX, a dataset of approximately 55M source and obfuscated code pairs in seven languages. Subsequently, we pre-train ObscuraCoder models, ranging in size from 255M to 2.8B parameters, on a 272B-token corpus that includes ObscuraX and demonstrate that our obfuscation-based pre-training recipe leads to consistent improvements in Code-LMs' abilities compared to both vanilla autoregressive pre-training as well as existing de-obfuscation (DOBF) objectives. ObscuraCoder demonstrates sizeable gains across multiple tests of syntactic and semantic code understanding, along with improved capabilities in multilingual code completion, multilingual code commit summarization, and multi-purpose library-oriented code generation. Indraneil Paul, Haoyi Yang, Goran Glavas, Kristian Kersting, Iryna Gurevych |
ICLR | 1 |
| 2025 | BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex InstructionsabstractTask automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing **diverse function calls as tools** to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately understanding **complex instructions**. Fulfilling both of these characteristics can pose a great challenge for LLMs. To assess how well LLMs can solve challenging and practical tasks via programs, we introduce BigCodeBench, a benchmark that challenges LLMs to invoke multiple function calls as tools from 139 libraries and 7 domains for 1,140 fine-grained tasks. To evaluate LLMs rigorously, each task encompasses 5.6 test cases with an average branch coverage of 99%. In addition, we propose a natural-language-oriented variant of BigCodeBench, BigCodeBench-Instruct, that automatically transforms the original docstrings into short instructions containing only essential information. Our extensive evaluation of 60 LLMs shows that **LLMs are not yet capable of following complex instructions to use function calls precisely, with scores up to 60%, significantly lower than the human performance of 97%**. The results underscore the need for further advancements in this area. Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu 0011, Wenhao Yu 0002, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong 0005, James Hoang, Armel Zebaze, Xiaoheng Hong, Wen-Ding Li, Jean Kaddour, Zhihan Zhang 0001, Prateek Yadav |
ICLR | 10 |
| 2024 | IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code GeneratorsabstractCode generation has fast become one of the most popular applications of language models (LMs).Nonetheless, research on multilingual aspects of Code-LMs, such as cross-lingual transfer between different programming languages, language-specific data augmentation, and post-hoc LM adaptation, alongside the exploitation of data sources other than the original textual content, has been much sparser than for their natural language counterparts.In particular, most mainstream Code-LMs have been pre-trained on source code files alone.In this work, we investigate the prospect of leveraging readily available compiler intermediate representations (IR)-shared across programming languages-to improve the multilingual capabilities of Code-LMs and facilitate crosslingual transfer.To this end, we first compile SLTrans, 1,2 a parallel dataset consisting of nearly 4M self-contained source code files coupled with their respective intermediate representations.Next, starting from various base Code-LMs (ranging from 1.1B to 7.3B parameters), we carry out continued causal language modelling training on SLTrans, forcing the Code-LMs to (1) learn the IR language and (2) align the IR constructs with respective constructs of various programming languages.Our resulting models, dubbed IRCoder, display sizeable and consistent gains across various code generation tasks and metrics, including prompt robustness, multilingual code completion, code understanding, and instruction following. Indraneil Paul, Goran Glavas, Iryna Gurevych |
ACL (1) | 1 |
| 2022 | Sub-Task Imputation via Self-Labelling to Train Image Moderation Models on Sparse Noisy DataabstractE-commerce marketplaces protect shopper experience and trust at scale by deploying deep learning models trained on human annotated moderation data, for the identification and removal of advert imagery that does not comply with moderation policies (a.k.a. defective images). However, human moderation labels can be hard to source for smaller advert programs that target specific device types with separate formats or for recently launched locales with unique moderation policies. Additionally, the sourced labels can be noisy due to annotator biases or policy rules clubbing multiple types of transgressions into a single category. Therefore, training advert image moderation models necessitates an approach that can effectively improve the sample efficiency of training, weed out noise and discover latent moderation sub-labels in one go. Indraneil Paul, Sumit Negi |
CIKM | 1 |