Ayush Kumar Shah

dblp:301/3151 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0001-6158-7632ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Multimodal Search in Chemical Documents and Reactions
abstract
We present a multimodal search tool for retrieval of chemical reactions, molecular structures, and associated text from scientific literature.Queries may combine molecular diagrams, textual descriptions, and reaction data, allowing users to connect different chemical information representations.Indexing includes chemical diagram extraction and parsing, extraction of reaction data from text in tabular form, and cross-modal linking of diagrams with their mentions in text.We describe the system's architecture and retrieval features, along with expert assessments of the system.Our demo highlights the workflow and search components.Online demo: https://www.cs.rit.edu/
Ayush Kumar Shah, Abhisek Dey, Leo Luo, Bryan Amador, Patrick Philippy, Ming Zhong 0005, Siru Ouyang, David Mark Friday, David Bianchi, Nick Jackson, Richard Zanibbi, Jiawei Han 0001
SIGIR1
2024 ChemScraper: leveraging PDF graphics instructions for molecular diagram parsing
Ayush Kumar Shah, Bryan Amador, Abhisek Dey, Ming Creekmore, Blake Ocampo, Scott E. Denmark, Richard Zanibbi
Int. J. Document Anal. Recognit.1
2023 Line-of-Sight with Graph Attention Parser (LGAP) for Math Formulas
Ayush Kumar Shah, Richard Zanibbi
ICDAR (5)1
2023 Searching the ACL Anthology with Math Formulas and Text
abstract
Mathematical notation is a key analytical resource for science and technology. Unfortunately, current math-aware search engines require LATEX or template palettes to construct formulas, which can be challenging for non-experts. Also, their indexed collections are primarily web pages where formulas are represented explicitly in machine-readable formats (e.g., LATEX, Presentation MathML). The new MathDeck system searches PDF documents in a portion of the ACL Anthology using both formulas and text, and shows matched words and formulas along with other extracted formulas in-context. In PDF, formulas are not demarcated: a new indexing module extracts formulas using PDF vector graphics information and computer vision techniques. For non-expert users and visual editing, a central design feature of MathDeck's interface is formula 'chips' usable in formula creation, search, reuse, and annotation with titles and descriptions in cards. For experts, LATEX is supported in the text query box and the visual formula editor. MathDeck is open-source, and our demo is available online.
Bryan Amador, Matt Langsenkamp, Abhisek Dey, Ayush Kumar Shah, Richard Zanibbi
SIGIR4
2021 A Math Formula Extraction and Evaluation Framework for PDF Documents
Ayush Kumar Shah, Abhisek Dey, Richard Zanibbi
ICDAR (2)1