EDBT 2026 Demo / reviewers in the wild / expert
Abhisek Dey
dblp:64/10510
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0001-5553-8279ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MoleculeMiner: Extracting and Linking Molecule Figures with Tabular MetadataabstractDespite an ongoing shift in automated chemical literature search methods, many are fairly limited in ability to find very specific relevant information about a drawn molecule and its associated property data. We aim to tackle the challenge of converting drawn molecules to a machine readable representation and co-reference any associated molecule data. MoleculeMiner is a system where a user can feed in their own patent or paper to obtain each drawn molecule along with any specific metadata (chemical name, chemical reactivity, yield, purity etc.) provided anywhere in the PDF in a tabular format, using an interactive user-friendly environment. We also present MolScribeV2, a molecular image parser which improved upon the original MolScribe by introducing pixel-based self attention positional embedding technique. Along with other changes, MolScribeV2 is robust to varied styles of compound drawings commonly found in patents and papers--scanned or born digital. Our extraction and user interactive system can be found at https://github.com/insitro/MoleculeMiner. Abhisek Dey, Nathaniel H. Stanley |
IJCAI | 1 |
| 2025 | Targeted Multi-Modal Passage Search for Molecules and their Synthesis PathwaysabstractWe present a chemical extraction and search pipeline intended to support information tasks related to drug discovery. Commonly used search tools for drug discovery such as Reaxys and SciFinder do not allow users to obtain retrieval results at the passage level. To address this, we present a passage retrieval tool for chemical patents that supports queries combining text and molecule diagrams expressed in SMILES. When SMILES is provided as a part of a query, the system refines text retrieval results through matching both textual names and drawn figures based on extracted SMILES representations. Molecule matches are obtained through substructure matching and structural similarity. This functionality was motivated by a chemist's need to find synthesis pathways for specific molecules containing a substructure of interest that binds and thus inhibits specific human genes. For this demonstration, we index a collection of 131 PDF patents categorized into 12 specific genes enabling a user to search on them. There are 32,301 document pages in the collection. Our user interface can be accessed at https://unichemfinder.gccis.rit.edu/. Our source code and data is available at https://gitlab.com/dprl/unichemfinder. Abhisek Dey, Nathaniel H. Stanley, Richard Zanibbi |
SIGIR | 1 |
| 2025 | Multimodal Search in Chemical Documents and ReactionsabstractWe present a multimodal search tool for retrieval of chemical reactions, molecular structures, and associated text from scientific literature.Queries may combine molecular diagrams, textual descriptions, and reaction data, allowing users to connect different chemical information representations.Indexing includes chemical diagram extraction and parsing, extraction of reaction data from text in tabular form, and cross-modal linking of diagrams with their mentions in text.We describe the system's architecture and retrieval features, along with expert assessments of the system.Our demo highlights the workflow and search components.Online demo: https://www.cs.rit.edu/ Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong 0005, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han 0001 |
SIGIR | 2 |
| 2024 | ChemScraper: leveraging PDF graphics instructions for molecular diagram parsing
Ayush Kumar Shah, Bryan Amador, Abhisek Dey, Ming Creekmore, Blake Ocampo, Scott E. Denmark, Richard Zanibbi |
Int. J. Document Anal. Recognit. | 3 |
| 2023 | Searching the ACL Anthology with Math Formulas and TextabstractMathematical notation is a key analytical resource for science and technology. Unfortunately, current math-aware search engines require LATEX or template palettes to construct formulas, which can be challenging for non-experts. Also, their indexed collections are primarily web pages where formulas are represented explicitly in machine-readable formats (e.g., LATEX, Presentation MathML). The new MathDeck system searches PDF documents in a portion of the ACL Anthology using both formulas and text, and shows matched words and formulas along with other extracted formulas in-context. In PDF, formulas are not demarcated: a new indexing module extracts formulas using PDF vector graphics information and computer vision techniques. For non-expert users and visual editing, a central design feature of MathDeck's interface is formula 'chips' usable in formula creation, search, reuse, and annotation with titles and descriptions in cards. For experts, LATEX is supported in the text query box and the visual formula editor. MathDeck is open-source, and our demo is available online. Bryan Amador, Matt Langsenkamp, Abhisek Dey, Ayush Kumar Shah, Richard Zanibbi |
SIGIR | 3 |
| 2021 | A Math Formula Extraction and Evaluation Framework for PDF Documents
Ayush Kumar Shah, Abhisek Dey, Richard Zanibbi |
ICDAR (2) | 2 |