Petr Babkin

dblp:161/0027 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0004-2737-9820ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Perturb Your Data: Paraphrase-Guided Training Data Watermarking
abstract
Training data detection is critical for enforcing copyright and data licensing, as Large Language Models (LLM) are trained on massive text corpora scraped from the internet. We present SPECTRA, a watermarking approach that makes training data reliably detectable even when it comprises less than 0.001% of the training corpus. SPECTRA works by paraphrasing text using an LLM and assigning a score based on how likely each paraphrase is, according to a separate scoring model. A paraphrase is chosen so that its score closely matches that of the original text, to avoid introducing any distribution shifts. To test whether a suspect model has been trained on the watermarked data, we compare its token probabilities against those of the scoring model. We demonstrate that SPECTRA achieves a consistent p-value gap of over nine orders of magnitude when detecting data used for training versus data not used for training, which is greater than all baselines tested. SPECTRA equips data owners with a scalable, deploy‑before‑release watermark that survives even large‑scale LLM training.
Pranav Shetty, Mirazul Haque, Petr Babkin, Xiaomo Liu, Manuela M. Veloso
AAAI3
2025 Nova: Generative Language Models for Assembly Code with Hierarchical Attention and Contrastive Learning
abstract
Binary code analysis is the foundation of crucial tasks in the security domain; thus building effective binary analysis techniques is more important than ever. Large language models (LLMs) although have brought impressive improvement to source code tasks, do not directly generalize to assembly code due to the unique challenges of assembly: (1) the low information density of assembly and (2) the diverse optimizations in assembly code. To overcome these challenges, this work proposes a hierarchical attention mechanism that builds attention summaries to capture the semantics more effectively and designs contrastive learning objectives to train LLMs to learn assembly optimization. Equipped with these techniques, this work develops Nova, a generative LLM for assembly code. Nova outperforms existing techniques on binary code decompilation by up to 14.84 -- 21.58% higher Pass@1 and Pass@10, and outperforms the latest binary code similarity detection techniques by up to 6.17% Recall@1, showing promising abilities on both assembly generation and understanding tasks.
Nan Jiang 0012, Chengxiao Wang, Kevin Liu, Xiangzhe Xu, Lin Tan 0001, Xiangyu Zhang 0001, Petr Babkin
ICLR7
2024 DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding
abstract
Dongsheng Wang, Natraj Raman, Mathieu Sibue, Zhiqiang Ma, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Dongsheng Wang 0005, Natraj Raman, Mathieu Sibue, Petr Babkin, Simerjot Kaur, Yulong Pei, Armineh Nourbakhsh, Xiaomo Liu
ACL (1)5
2023 An Automated Code Update Tool For Python Packages
abstract
The adoption of libraries provides developers with pre-existing functionality that is both robust and easy to use. As a code base grows over time, it is natural that the libraries become stale and require updates in order to sustain innovation. Nonetheless, updating a library comes at a cost; it can potentially introduce breaking changes in the code. Thus it may require a large portion of developer time for maintenance. To alleviate this, we propose to use artificial intelligence to parse release notes documentation and automatically recommend code updates to become compatible with new versions. The solution comes in the form of an IDE plugin that can automatically detect deprecated library usages in live code bases and suggest the recommended fixes in a user-friendly way. The system architecture is comprised of three components: a web crawler that sources and pre-processes library deprecation texts, a deprecations parser that transforms those texts into a structured form, and an IDE plugin that finds and updates the parsed deprecations in a code base. In order to validate our approach, we collect and annotate a dataset of 426 API deprecations from 7 popular Python libraries and obtain an overall average weighted subtree overlap of 31.3 and 45.9 on method deprecations. Finally, we demo our tool in the form of a plugin for the industry leading Python IDE PyCharm and test it on 33 internal repositories at J.P Morgan Chase.
Nacho Navarro, Salwa Alamir, Petr Babkin, Sameena Shah
ICSME3
2023 How Effective Are Neural Networks for Fixing Security Vulnerabilities
abstract
Security vulnerability repair is a difficult task that is in dire need of automation. Two groups of techniques have shown promise: (1) large code language models (LLMs) that have been pre-trained on source code for tasks such as code completion, and (2) automated program repair (APR) techniques that use deep learning (DL) models to automatically fix software bugs.
Yi Wu 0023, Nan Jiang 0012, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan 0001, Petr Babkin, Sameena Shah
ISSTA7
2023 BizGraphQA: A Dataset for Image-based Inference over Graph-structured Diagrams from Business Domains
abstract
Graph-structured diagrams, such as enterprise ownership charts or management hierarchies, are a challenging medium for deep learning models as they not only require the capacity to model language and spatial relations but also the topology of links between entities and the varying semantics of what those links represent. Devising Question Answering models that automatically process and understand such diagrams have vast applications to many enterprise domains, and can move the state-of-the-art on multimodal document understanding to a new frontier. Curating real-world datasets to train these models can be difficult, due to scarcity and confidentiality of the documents where such diagrams are included. Recently released synthetic datasets are often prone to repetitive structures that can be memorized or tackled using heuristics. In this paper, we present a collection of 10,000 synthetic graphs that faithfully reflect properties of real graphs in four business domains, and are realistically rendered within a PDF document with varying styles and layouts. In addition, we have generated over 130,000 question instances that target complex graphical relationships specific to each domain. We hope this challenge will encourage the development of models capable of robust reasoning about graph structured images, which are ubiquitous in numerous sectors in business and across scientific disciplines.
Petr Babkin, William Watson, Lucas Cecchi, Natraj Raman, Armineh Nourbakhsh, Sameena Shah
SIGIR1
2015 Automatic Ellipsis Resolution: Recovering Covert Information from Text
abstract
Ellipsis is a linguistic process that makes certain aspects of text meaning not directly traceable to surface text elements and, therefore, inaccessible to most language processing technologies. However, detecting and resolving ellipsis is an indispensable capability for language-enabled intelligent agents. The key insight of the work presented here is that not all cases of ellipsis are equally difficult: some can be detected and resolved with high confidence even before we are able to build agents with full human-level semantic and pragmatic understanding of text. This paper describes a fully automatic, implemented and evaluated method of treating one class of ellipsis: elided scopes of modality. Our cognitively-inspired approach, which centrally leverages linguistic principles, has also been applied to overt referring expressions with equally promising results.
Marjorie McShane, Petr Babkin
AAAI2