EDBT 2026 Demo / reviewers in the wild / expert
Amit Pathak
dblp:63/6545
· DBLP profile ↗
8ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0003-4006-5119ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Database system architecture and tuning · 43% Query processing and optimization · 26% Graph data management · 14% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 67% Language models and text generation · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 100% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › document understanding
document layout analysis |
0.9 | 1 | 2025 | Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments · ACL (1) 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis › document analysis › document information extraction
visually rich document information extraction |
0.9 | 1 | 2025 | Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments · ACL (1) 2025 |
Query processing and optimization
analytical query processing |
0.9 | 1 | 2025 | The HANA Native Query Engine for Lakehouse Systems · Proc. VLDB Endow. 2025 |
Storage systems › buffer management
buffer cache management |
0.5 | 2 | 2019 | Native Store Extension for SAP HANA · Proc. VLDB Endow. 2019 BTrim - Hybrid In-Memory Database Architecture for Extreme Transaction Processing in VLDBs · Proc. VLDB Endow. 2018 |
Transaction processing and concurrency control › OLTP
in-memory transaction processing |
0.3 | 1 | 2018 | BTrim - Hybrid In-Memory Database Architecture for Extreme Transaction Processing in VLDBs · Proc. VLDB Endow. 2018 |
Storage systems › data management › database storage
columnar storage |
0.3 | 1 | 2025 | The HANA Native Query Engine for Lakehouse Systems · Proc. VLDB Endow. 2025 |
Graph data management
personalized pagerank |
0.2 | 2 | 2008 | Fast algorithms for topk personalized pagerank queries · WWW 2008 Index Design for Dynamic Personalized PageRank · ICDE 2008 |
Graph data management
graph indexing |
0.1 | 1 | 2011 | Index design and query processing for graph conductance search · VLDB J. 2011 |
Indexing and storage engines
column store |
0.1 | 1 | 2019 | Native Store Extension for SAP HANA · Proc. VLDB Endow. 2019 |
Graph data management
graph query processing |
0.1 | 1 | 2008 | Fast algorithms for topk personalized pagerank queries · WWW 2008 |
Information retrieval › indexing
index design |
0.1 | 1 | 2008 | Index Design for Dynamic Personalized PageRank · ICDE 2008 |
Information retrieval
search engines |
0.1 | 1 | 2008 | Index Design for Dynamic Personalized PageRank · ICDE 2008 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2008 | Fast algorithms for topk personalized pagerank queries · WWW 2008 |
Graph algorithms and graph theory › centrality › pagerank
personalized pagerank |
0.0 | 1 | 2008 | Index Design for Dynamic Personalized PageRank · ICDE 2008 |
Methods — techniques the papers use, named apart from their topics
pushdown architecture · 1.7direct access architecture · 1.7caching · 1.7semantic segmentation of documents · 0.9large language model · 0.9run-length encoding · 0.8prefetching · 0.8dictionary encoding · 0.8in-memory row store · 0.7data tiering · 0.7cost-benefit optimization · 0.2query processing · 0.1performance modeling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual SegmentsabstractInformation extraction (IE) from Visually Rich Documents (VRDs) containing layout features along with text is a critical and well-studied task. Specialized non-LLM NLP-based solutions typically involve training models using both textual and geometric information to label sequences/tokens as named entities or answers to specific questions. However, these approaches lack reasoning, are not able to infer values not explicitly present in documents, and do not generalize well to new formats. Generative LLMs-based approaches proposed recently are capable of reasoning, but struggle to comprehend clues from document layout especially in previously unseen document formats, and do not show competitive performance in heterogeneous VRD benchmark datasets. In this paper, we propose BLOCKIE, a novel LLM-based approach that organizes VRDs into localized, reusable semantic textual segments called \textit{semantic blocks}, which are processed independently. Through focused and more generalizable reasoning,our approach outperforms the state-of-the-art on public VRD benchmarks by 1-3% in F1 scores, is resilient to document formats previously not encountered and shows abilities to correctly extract information not explicitly present in documents. Aniket Bhattacharyya, Anurag Tripathi, Ujjal Das, Archan Karmakar, Amit Pathak, Maneesh Gupta |
ACL (1) | 5 |
| 2025 | Fast yet force-effective mode of supracellular collective cell migration due to extracellular force transmissionabstractCell collectives, like other motile entities, generate and use forces to move forward. Here, we ask whether environmental configurations alter this proportional force-speed relationship, since aligned extracellular matrix fibers are known to cause directed migration. We show that aligned fibers serve as active conduits for spatial propagation of cellular mechanotransduction through matrix exoskeleton, leading to efficient directed collective cell migration. Epithelial (MCF10A) cell clusters adhered to soft substrates with aligned collagen fibers (AF) migrate faster with much lesser traction forces, compared to random fibers (RF). Fiber alignment causes higher motility waves and transmission of normal stresses deeper into cell monolayer while minimizing shear stresses and increased cell-division based fluidization. By contrast, fiber randomization induces cellular jamming due to breakage in motility waves, disrupted transmission of normal stresses, and heightened shear driven flow. Using a novel motor-clutch model, we explain that such 'force-effective' fast migration phenotype occurs due to rapid stabilization of contractile forces at the migrating front, enabled by higher frictional forces arising from simultaneous compressive loading of parallel fiber-substrate connections. We also model 'haptotaxis' to show that increasing ligand connectivity (but not continuity) increases migration efficiency. According to our model, increased rate of front stabilization via higher resistance to substrate deformation is sufficient to capture 'durotaxis'. Thus, our findings reveal a new paradigm wherein the rate of leading-edge stabilization determines the efficiency of supracellular collective cell migration. Amrit Bagchi, Bapi Sarker, Marcus Foston, Amit Pathak |
PLoS Comput. Biol. | 5 |
| 2025 | The HANA Native Query Engine for Lakehouse SystemsabstractModern enterprise applications and data warehouse systems move data into data lakes for economical and scalability reasons. Data is then stored in popular columnar file formats like Parquet which are optimized for writing using open table formats like Iceberg or Delta. This presents new challenges for existing database systems and their execution engines because excellent performance and scalability when accessing this data in complex analytical queries is expected while data is located in a remote data lake. In this work, we present how we adapted the HANA Cloud Database Engine for efficient processing of files in data lakes, which we call SQL-on-Files (SoF). We motivate this evolution by its relevance for Business Data Cloud, SAP's Lakehouse, we discuss the viability of general architecture choices like pushdown and direct access architectures, and give insights into our SoF design decisions towards scalable, analytical query processing around execution engine, optimizer and caching. Our evaluation of SoF shows benefits of direct access over pushdown architectures for a new warehouse benchmark with complex, analytical workloads. Daniel Ritter 0001, Mihnea Andrei, Sukhyeun Cho, Maik Goergens, Taehyung Lee 0002, Norman May, Amit Pathak, Paul R. Willems |
Proc. VLDB Endow. | 7 |
| 2019 | Native Store Extension for SAP HANAabstractWe present an overview of SAP HANA's Native Store Extension (NSE). This extension substantially increases database capacity, allowing to scale far beyond available system memory. NSE is based on a hybrid in-memory and paged column store architecture composed from data access primitives. These primitives enable the processing of hybrid columns using the same algorithms optimized for traditional HANA's in-memory columns. Using only three key primitives, we fabricated byte-compatible counterparts for complex memory resident data structures (e.g. dictionary and hash-index), compressed schemes (e.g. sparse and run-length encoding), and exotic data types (e.g. geo-spatial). We developed a new buffer cache which optimizes the management of paged resources by smart strategies sensitive to page type and access patterns. The buffer cache integrates with HANA's new execution engine that issues pipelined prefetch requests to improve disk access patterns. A novel load unit configuration, along with a unified persistence format, allows the hybrid column store to dynamically switch between in-memory and paged data access to balance performance and storage economy according to application demands while reducing Total Cost of Ownership (TCO). A new partitioning scheme supports load unit specification at table, partition, and column level. Finally, a new advisor recommends optimal load unit configurations. Our experiments illustrate the performance and memory footprint improvements on typical customer scenarios. Reza Sherkat, Colin Florendo, Mihnea Andrei, Rolando Blanco, Adrian Dragusanu, Amit Pathak, Pushkar Khadilkar, Neeraj Kulkarni, Christian Lemke, Sebastian Seifert, Sarika Iyer, Sasikanth Gottapu, Robert Schulze, Chaitanya Gottipati, Nirvik Basak, Vivek Kandiyanallur, Santosh Pendap, Dheren Gala, Rajesh Almeida, Prasanta Ghosh |
Proc. VLDB Endow. | 6 |
| 2018 | BTrim - Hybrid In-Memory Database Architecture for Extreme Transaction Processing in VLDBsabstractTo address the need for extreme OLTP performance on commodity multi-core hardware supporting large amounts of memory, SAP ASE is re-architected to tightly integrate an In-Memory Row Store (IMRS) within the existing database engine. The IMRS is both a store and a caching layer to host "hot" rows in-memory, in a row-oriented format. The IMRS is an extension to the traditional buffer-cache which deals with data in a page-oriented storage format (referred to as the page-store). Data in individual tables marked as IMRS-enabled can be fully memory-resident or can straddle the page store and the IMRS. Cold data in the IMRS is organically identified, harvested, and "packed" back to the page store. SQL statements and transactions can access data transparently from both stores for the same table. All Transact-SQL capabilities and language constructs are supported with no application or stored procedure code changes required. Full durability for in-memory data is provided, including support for backup and restore of database archives and periodic transaction dumps. The high-level system design supporting this architecture, along with experimental results and performance benefits is presented. Aditya Gurajada, Dheren Gala, Amit Pathak, Zhan-Feng Ma |
Proc. VLDB Endow. | 4 |
| 2011 | Index design and query processing for graph conductance search
Soumen Chakrabarti, Amit Pathak, Manish S. Gupta |
VLDB J. | 2 |
| 2008 | Index Design for Dynamic Personalized PageRankabstractPersonalized page rank, related to random walks with restarts and conductance in resistive networks, is a frequent search paradigm for graph-structured databases. While efficient batch algorithms exist for static whole-graph page rank, interactive query-time personalized page rank has proved more challenging. Here we describe how to select and build indices for a popular class of page rank algorithms, so as to provide real-time personalized page rank and smoothly trade off between index size, preprocessing time, and query speed. We achieve this by developing a precise, yet efficiently estimated performance model for personalized page rank query execution. We use this model in conjunction with a query workload in a cost-benefit type index optimizer. On millions of queries from CiteSeer and its data graphs with 74-320 thousand nodes, our algorithm runs 50-400 x faster than whole-graph page rank, the gap growing with graph size. Index size is 10-20% of a text index. Ranking accuracy is above 94%. Amit Pathak, Soumen Chakrabarti, Manish S. Gupta |
ICDE | 1 |
| 2008 | Fast algorithms for topk personalized pagerank queriesabstractIn entity-relation (ER) graphs (V,E), nodes V represent typed entities and edges E represent typed relations. For dynamic personalized PageRank queries, nodes are ranked by their steady-state probabilities obtained using the standard random surfer model. In this work, we propose a framework to answer top-k graph conductance queries. Our top-k ranking technique leads to a 4X speedup, and overall, our system executes queries 200-1600X faster than whole-graph PageRank. Some queries might contain hard predicates i.e. predicates that must be satisfied by the answer nodes. E.g. we may seek authoritative papers on public key cryptography, but only those written during 1997. We extend our system to handle hard predicates. Our system achieves these substantial query speedups while consuming only 10-20% of the space taken by a regular text index. Manish S. Gupta, Amit Pathak, Soumen Chakrabarti |
WWW | 2 |