VLDB 2026 Research / reviewers in the wild / expert
Arun Iyer
dblp:262/6555
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2023
0000-0001-7377-7599ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf GraphabstractMulti-document Summarization (MDS) characterizes compressing information from multiple source documents to its succinct summary.An ideal summary should encompass all topics and accurately model cross-document relations expounded upon in the source documents.However, existing systems either impose constraints on the length of tokens during the encoding or falter in capturing the intricate cross-document relationships.These limitations impel the systems to produce summaries that are non-factual and unfaithful, thereby imparting an unfair comprehension of the topic to the readers.To counter these limitations and promote the information equivalence between the source document and generated summary, we propose FABRIC, a novel encoder-decoder model that uses pre-trained BART to comprehensively analyze linguistic nuances, simplicial complex layer to apprehend inherent properties that transcend pairwise associations and sheaf graph attention to effectively capture the heterophilic properties.We benchmark FABRIC with eleven baselines over four widely-used MDS datasets -Multinews, CQASumm, DUC and Opinosis, and show that FABRIC achieves consistent performance improvement across all the evaluation metrics (syntactical, semantical and faithfulness).We corroborate these improvements further through qualitative human evaluation. Yash Kumar Atri, Arun Iyer, Tanmoy Chakraborty 0002, Vikram Goyal |
EMNLP | 2 |
| 2023 | FiGURe: Simple and Efficient Unsupervised Node Representations with Filter AugmentationsabstractUnsupervised node representations learnt using contrastive learning-based methods have shown good performance on downstream tasks. However, these methods rely on augmentations that mimic low-pass filters, limiting their performance on tasks requiring different eigen-spectrum parts. This paper presents a simple filter-based augmentation method to capture different parts of the eigen-spectrum. We show significant improvements using these augmentations. Further, we show that sharing the same weights across these different filter augmentations is possible, reducing the computational load. In addition, previous works have shown that good performance on downstream tasks requires high dimensional representations. Working with high dimensions increases the computations, especially when multiple augmentations are involved. We mitigate this problem and recover good performance through lower dimensional embeddings using simple random Fourier feature projections. Our method, FiGURe, achieves an average gain of up to 4.4\%, compared to the state-of-the-art unsupervised models, across all datasets in consideration, both homophilic and heterophilic. Our code can be found at: https://github.com/Microsoft/figure. Chanakya Ajit Ekbote, Ajinkya Pankaj Deshpande, Arun Iyer, Sundararajan Sellamanickam, Ramakrishna Bairi |
NeurIPS | 3 |
| 2022 | Jigsaw: Large Language Models meet Program SynthesisabstractLarge pre-trained language models such as GPT-3 [10], Codex [11], and Google's language model [7] are now capable of generating code from natural language specifications of programmer intent. We view these developments with a mixture of optimism and caution. On the optimistic side, such large language models have the potential to improve productivity by providing an automated AI pair programmer for every programmer in the world. On the cautionary side, since these large language models do not understand program semantics, they offer no guarantees about quality of the suggested code. In this paper, we present an approach to augment these large language models with post-processing steps based on program analysis and synthesis techniques, that understand the syntax and semantics of programs. Further, we show that such techniques can make use of user feedback and improve with usage. We present our experiences from building and evaluating such a tool Jigsaw, targeted at synthesizing code for using Python Pandas API using multi-modal inputs. Our experience suggests that as these large language models evolve for synthesizing code from intent, Jigsaw has an important role to play in improving the accuracy of the systems. Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy Iyengar, Sriram K. Rajamani, Rahul Sharma 0001 |
ICSE | 3 |
| 2022 | A Piece-Wise Polynomial Filtering Approach for Graph Neural Networks
Vijay Lingam, Manan Sharma, Chanakya Ajit Ekbote, Rahul Ragesh, Arun Iyer, Sundararajan Sellamanickam |
ECML/PKDD (2) | 5 |
| 2022 | Landmarks and regions: a robust approach to data extractionabstractWe propose a new approach to extracting data items or field values from semi-structured documents. Examples of such problems include extracting passenger name, departure time and departure airport from a travel itinerary, or extracting price of an item from a purchase receipt. Traditional approaches to data extraction use machine learning or program synthesis to process the whole document to extract the desired fields. Such approaches are not robust to for- mat changes in the document, and the extraction process typically fails even if changes are made to parts of the document that are unrelated to the desired fields of interest. We propose a new approach to data extraction based on the concepts of landmarks and regions. Humans routinely use landmarks in manual processing of documents to zoom in and focus their attention on small regions of interest in the document. Inspired by this human intuition, we use the notion of landmarks in program synthesis to automatically synthesize extraction programs that first extract a small region of interest, and then automatically extract the desired value from the region in a subsequent step. We have implemented our landmark based extraction approach in a tool LRSyn, and show extensive valuation on documents in HTML as well as scanned images of invoices and receipts. Our results show that the our approach is robust to various types of format changes that routinely happen in real-world settings Suresh Parthasarathy Iyengar, Lincy Pattanaik, Anirudh Khatry, Arun Iyer, Arjun Radhakrishna, Sriram K. Rajamani, Mohammad Raza |
PLDI | 4 |
| 2021 | HeteGCN: Heterogeneous Graph Convolutional Networks for Text ClassificationabstractWe consider the problem of learning efficient and inductive graph convolutional networks for text classification with a large number of examples and features. Existing state-of-the-art graph embedding based methods such as predictive text embedding (PTE) and TextGCN have shortcomings in terms of predictive performance, scalability and inductive capability. To address these limitations, we propose a heterogeneous graph convolutional network (HeteGCN) modeling approach that unites the best aspects of PTE and TextGCN together. The main idea is to learn feature embeddings and derive document embeddings using a HeteGCN architecture with different graphs used across layers. We simplify TextGCN by dissecting into several HeteGCN models which (a) helps to study the usefulness of individual models and (b) offers flexibility in fusing learned embeddings from different models. In effect, the number of model parameters is reduced significantly, enabling faster training and improving performance in small labeled training set scenario. Our detailed experimental studies demonstrate the efficacy of the proposed approach. Rahul Ragesh, Sundararajan Sellamanickam, Arun Iyer, Ramakrishna Bairi, Vijay Lingam |
WSDM | 3 |
| 2019 | Synthesis and machine learning for heterogeneous extractionabstractWe present a way to combine techniques from the program synthesis and machine learning communities to extract structured information from heterogeneous data. Such problems arise in several situations such as extracting attributes from web pages, machine-generated emails, or from data obtained from multiple sources. Our goal is to extract a set of structured attributes from such data. Arun Iyer, Manohar Jonnalagedda, Suresh Parthasarathy Iyengar, Arjun Radhakrishna, Sriram K. Rajamani |
PLDI | 1 |